A real-time dynamic simulation method and system for a soil remediation process based on digital twinning

CN122595797APending Publication Date: 2026-08-18开封市生态环境监测和安全中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610675285.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]为了解决现有技术无法对土壤修复进程进行准确预测、难以响应外部环境变化的技术问题,本申请提供一种基于数字孪生的土壤修复过程实时动态仿真方法,包括以下步骤:

Benefits of technology

附图说明

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122595797A_ABST
    Figure CN122595797A_ABST
Patent Text Reader

Abstract

This invention relates to the field of soil remediation technology, specifically to a real-time dynamic simulation method and system for soil remediation processes based on digital twins. The method includes: acquiring real-time environmental data and real-time soil state data from the previous moment, inputting them into a digital twin model to generate target substance prediction data for the current moment, including the predicted content of the target substance; acquiring the measured content of the target substance at the current moment, calculating the relative deviation between the measured content and the predicted content; determining whether the relative deviation exceeds an accuracy threshold; if so, calculating the fitness based on the predicted and measured contents of the target substance at the current moment, and using a genetic algorithm to iteratively optimize the digital twin model with fitness maximization as the optimization objective; if not, outputting the predicted content of the target substance. This application effectively solves the problem of delayed response to changes in the external environment during soil remediation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of soil remediation technology, specifically to a real-time dynamic simulation method and system for soil remediation processes based on digital twins. Background Technology

[0002] The purpose of soil remediation is to reduce the content of certain substances in the soil to meet specific targets. For example, in the treatment of heavy metal pollution, the concentration of heavy metals (such as cadmium, lead, and arsenic) in the soil needs to be reduced to below the national soil environmental quality standards to restore the normal function of the soil and ensure the safety of agricultural products and the health of the ecological environment. Throughout the remediation process, key parameters such as leaching time and remediation agent dosage need to be precisely controlled. However, changes in the external environment (such as precipitation, temperature, and sunlight) have a significant impact on the remediation process, especially in scenarios involving organic pollutants and microbial activity. Furthermore, the impact of external environmental changes on the soil remediation process is delayed; that is, it takes time for changes in the external environment to be reflected in changes of the components of interest in the soil. Traditional soil monitoring methods can only measure one or more specific indicators in real time, and cannot predict the impact of external environmental changes on the soil remediation process. This leads to delays in optimizing and adjusting remediation plans, low remediation efficiency, and may even result in resource waste or secondary pollution due to over-remediation or under-remediation.

[0003] Scholars have made considerable progress in understanding the mechanisms of action of various technologies in soil remediation, and have developed mathematical models for the mechanisms of chemical reactions and biodegradation during the remediation process. Some researchers have also attempted to use these mathematical models to predict the soil remediation process. However, due to the highly complex environments of the soil to be remediated, these mathematical models are based on ideal assumptions and often fail to provide satisfactory predictions in practical applications. Therefore, most of these models remain at the theoretical or experimental stage. Summary of the Invention

[0004] To address the technical problems of existing technologies being unable to accurately predict soil remediation processes and struggling to respond to changes in the external environment, this application provides a real-time dynamic simulation method for soil remediation processes based on digital twins, comprising the following steps:

[0005] The real-time environmental data and real-time soil condition data of the previous moment are acquired and input into the digital twin model to generate the target substance prediction data of the current moment, which includes the predicted content of the target substance.

[0006] Obtain the measured content of the target substance at the current moment, and calculate the relative deviation between the measured content of the target substance and the predicted content of the target substance at the current moment;

[0007] Determine if the relative deviation is greater than the accuracy threshold. If so, calculate the fitness based on the predicted content and the measured content of the target substance at the current moment, and use a genetic algorithm to iteratively optimize the digital twin model with fitness maximization as the optimization objective. If not, output the predicted content of the target substance.

[0008] Based on the above method, this application also provides a real-time dynamic simulation system for soil remediation processes based on digital twins, comprising:

[0009] The data acquisition module is used to collect real-time environmental status data, real-time soil status data, and measured content of target substances in the target application scenario.

[0010] The simulation prediction module is used to output target substance prediction data based on real-time environmental state data and real-time soil state data using a digital twin model. The target substance prediction data includes the predicted content of the target substance.

[0011] The data comparison module is used to calculate the relative deviation between the measured content of the target substance at the current moment and the predicted content of the target substance at the current moment.

[0012] The parameter optimization module is used to calculate the fitness based on the predicted content and the measured content of the target substance at the current moment when the relative deviation is greater than the accuracy threshold, and to use a genetic algorithm to iteratively optimize the digital twin model module with fitness maximization as the optimization objective.

[0013] Technical effects and advantages of the invention: Attached Figure Description

[0014] Figure 1 This is an overall flowchart of the method provided by the present invention.

[0015] Figure 2 This is a flowchart illustrating the process of obtaining a digital twin model in the method provided by this invention.

[0016] Figure 3 This is a schematic diagram of the overall structure of the system provided by the present invention. Detailed Implementation

[0017] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] In soil remediation, to achieve precise control and process optimization of remediation effects, it is necessary to monitor the concentration distribution, migration and transformation patterns of heavy metals in the soil, as well as the effectiveness of remediation agents in real time. Traditional soil remediation monitoring methods mostly rely on offline sampling and laboratory analysis, which suffers from problems such as long monitoring cycles, data lag, and insufficient spatial representativeness, making it difficult to meet the real-time control requirements of dynamic remediation processes. Furthermore, existing mathematical models often fail to provide satisfactory prediction results in complex real-world environments, and most remain at the theoretical or experimental stage, unable to accurately predict soil remediation progress and respond to changes in the external environment.

[0019] In this regard, refer to Figure 1 This application provides a real-time dynamic simulation method for soil remediation processes based on digital twins, comprising the following steps:

[0020] S1. Obtain the real-time environmental data and real-time soil state data of the previous moment, input them into the digital twin model, and generate the target substance prediction data of the current moment. The target substance prediction data includes the target substance prediction content.

[0021] Specifically, real-time environmental data and real-time soil condition data can be collected periodically at the soil remediation site using portable sensors or sampling tools through manual inspections, and then manually entered into the data processing system.

[0022] In soil remediation scenarios involving heavy metal pollution, considering that the heavy metal content in soil generally does not change rapidly, and that accurate determination of heavy metal content requires a complex processing and detection procedure, laboratory chemical analysis methods are typically employed. This involves digesting the soil sample and then using large, precision instruments (high-performance spectrometers or mass spectrometers) for measurement. The entire detection cycle is lengthy; therefore, the "real-time acquisition" step S1 can be performed at longer intervals (e.g., weekly or monthly sampling).

[0023] S2. Obtain the measured content of the target substance at the current moment, and calculate the relative deviation between the measured content of the target substance and the predicted content of the target substance at the current moment;

[0024] Specifically, at a point in time, either simultaneously with or slightly after data collection, representative soil samples can be selected from the soil remediation site and sent to a specialized laboratory for chemical analysis to obtain the actual content of the target substances. For example, for heavy metal pollutants, atomic absorption spectrometry can be used for detection. After obtaining the measured content, it is compared with the predicted content output by the digital twin model. The relative deviation can be calculated using simple mathematical operations; for example, dividing the difference between the measured and predicted content by the predicted content yields a percentage deviation value.

[0025] S3. Determine whether the relative deviation is greater than the accuracy threshold. If yes, calculate the fitness based on the predicted content and the measured content of the target substance at the current moment, and use a genetic algorithm to iteratively optimize the digital twin model with fitness maximization as the optimization objective. If no, output the predicted content of the target substance.

[0026] Specifically, a fixed accuracy threshold can be preset, such as 5% or 10%. The system compares the calculated relative deviation with this threshold. If the relative deviation exceeds the preset accuracy threshold, it indicates that the prediction result of the current digital twin model differs significantly from the actual situation and needs optimization. At this point, a fitness value is calculated based on the predicted content and the measured content of the target substance. For example, fitness can be defined as 1 divided by the square of the prediction error (the absolute difference between the measured content and the predicted content). Subsequently, a genetic algorithm is initiated to adjust the internal parameters of the digital twin model. The genetic algorithm generates a "population" of model parameters by simulating the biological evolution process, evaluates the fitness of each parameter set, and gradually evolves better model parameters through operations such as selection, crossover, and mutation to maximize fitness. This process is repeated until the model parameters converge or the preset number of optimizations is reached. If the relative deviation does not exceed the accuracy threshold, the prediction result of the current digital twin model is considered acceptable, and the predicted content is directly output.

[0027] This method effectively solves the problems of inaccurate prediction and delayed response to changes in the external environment in traditional soil remediation processes by acquiring environmental and soil condition data in real time and using a digital twin model for dynamic prediction. When there is a large deviation between the predicted results and the measured data, the system can adaptively use a genetic algorithm to iteratively optimize the digital twin model, thereby continuously improving the model's prediction accuracy and adaptability, ensuring that the soil remediation strategy can be adjusted in a timely and accurate manner, and thus improving the remediation effect.

[0028] Specifically, in some embodiments, reference is made to Figure 2 The steps to obtain a digital twin model include:

[0029] S11. Select the target application scenario and match it with the preset digital twin model library to determine whether there is a suitable digital twin model.

[0030] S12. If a suitable digital twin model exists, the digital twin model is called and its optimal parameter range is extracted as the initial parameter population. If no such model exists, the required basic mechanism model is selected based on the target application scenario, and the initial parameter population is set based on experience. The digital twin model of the target application scenario is constructed based on the initial parameter population.

[0031] S13. Obtain multiple sets of historical soil condition data, historical environmental condition data, and historical measured target substance content for the target application scenario. Input the historical soil condition data and historical environmental condition data into the digital twin model and output the predicted target substance content.

[0032] S14. Calculate the fitness based on the predicted content of the target substance and the historical measured content of the target substance, and use a genetic algorithm to iteratively optimize the digital twin model with fitness maximization as the optimization objective.

[0033] Specifically, in step S11, selecting the target application scenario refers to clarifying the specific site, pollution type, and remediation target for which real-time dynamic simulation of the soil remediation process is required. The pre-set digital twin model library is a collection of numerous developed, validated, and optimized digital twin models for different soil remediation scenarios. These models may cover the degradation mechanisms of different pollutants, the physicochemical properties of different soil types, and the application modes of different remediation technologies.

[0034] Specifically, the process of determining whether a suitable digital twin model exists can be implemented in several ways: one is based on vector operations.

[0035] S111. Establish a scene feature vector based on the real-time environmental state data and real-time soil state data of the target application scenario. The components of the scene feature vector correspond to a parameter value in the real-time environmental state data or the real-time soil state data, respectively.

[0036] In the process of constructing vectors, to eliminate the influence of different parameter units and numerical ranges, each parameter value can be normalized during vector construction. For example, Min-Max normalization or Z-score standardization can be used to ensure that the values ​​fall within a uniform numerical range. Furthermore, based on the selected parameters, feature engineering methods can be used to combine or transform the original parameters to generate more representative features. For example, the gradient, rate of change, or statistics of certain parameters can be calculated, and different weights can be assigned to each component of the scene feature vector according to the importance of different parameters to the soil remediation process, thus highlighting the influence of key parameters.

[0037] S112. Calculate the similarity between the scene feature vector of the target application scenario and the scene feature vector of each digital twin model in the digital twin model library. When the similarity is greater than the matching threshold, it is determined that there is a suitable digital twin model.

[0038] Vector similarity can be calculated using cosine similarity, Euclidean distance, or Manhattan distance, all of which are existing techniques and will not be elaborated further.

[0039] Another approach is to use an expert system or rule-based reasoning engine to logically match the descriptive information of the target application scenario (such as soil type, pollutant type, climate conditions, etc.) with the metadata of models in the model library to determine whether there are highly relevant models.

[0040] Specifically, in step S12, calling the adapted digital twin model means directly loading the existing model structure and preset parameters in the model library that are highly matched to the current scenario. Extracting its optimal parameter range refers to obtaining the parameter value range or probability distribution that performed best in historical applications from the past optimization records or metadata of the adapted model. These parameter ranges will serve as the initial search space for subsequent genetic algorithm optimization, thereby accelerating the optimization process and improving optimization efficiency.

[0041] Specifically, when a suitable digital twin model does not exist, there are several methods to establish a new digital twin model. One method is:

[0042] S121. Select the required basic mechanism model from the basic mechanism model library according to the target application scenario, and establish the coupling relationship between different basic mechanism models;

[0043] S122. Establish a discrete grid model based on the terrain data of the target application scenario, and assign values ​​to the discrete grid model according to the initial parameter population. This forms a solvable set of coupled equations with the coupled mathematical model, which is the digital twin model.

[0044] Specifically, in step S121, the required basic mechanism model is selected from the basic mechanism model library according to the target application scenario, in order to ensure that the constructed digital twin model can accurately and effectively reflect the actual situation of a specific soil remediation scenario.

[0045] For example, based on predefined rules or expert systems, appropriate mechanistic models can be automatically matched and selected according to the characteristics of the target application scenario (such as pollutant type, soil medium, remediation technology, etc.). Alternatively, feature extraction or semantic analysis of the description of the target application scenario can be performed, and similarity matching can be conducted with the metadata of each model in the basic mechanistic model library to select the most relevant model. In addition, a user interface can be provided to allow engineers to manually select the required basic mechanistic model based on their expertise in the application scenario.

[0046] After selecting the fundamental mechanistic model, it is necessary to establish the coupling relationships between different fundamental mechanistic models. This is to simulate the complex interactions and dependencies between various physical, chemical, and biological processes during soil remediation, thereby constructing a comprehensive model that more closely approximates reality. For example, a sequential coupling approach can be used, where the output of one model is used as the input of another, such as the output of a water flow model (groundwater velocity and direction) being used as the input of a contaminant migration model. An iterative coupling approach can also be used, where different models repeatedly exchange data at each time step until convergence is reached, to handle strongly coupled systems, such as hydrogeochemical coupling models. Alternatively, a unified simulation platform or interface can be used to achieve data sharing and synchronization between models, ensuring that each model can work collaboratively during the simulation.

[0047] Specifically, in step S122, a discrete grid model is established based on the terrain data of the target application scenario. The purpose is to spatially discretize the actual soil remediation area to facilitate numerical simulation in a computer and accurately reflect spatial heterogeneity such as terrain undulations and soil stratification. For example, the finite difference method can be used to divide the terrain data into regular grid cells (such as squares or rectangles), and model variables can be defined at each grid point. Alternatively, the finite element method can be used to divide complex terrain into irregular, more adaptable cells (such as triangles or tetrahedrons) to better fit irregular boundaries and complex geometries. The terrain data can originate from a digital elevation model (DEM), LiDAR scan data, or field survey data.

[0048] Subsequently, values ​​are assigned to the discrete grid model based on the initial parameter population. This aims to provide initial parameter values ​​with a certain range or distribution for each mechanistic model within each discrete grid cell in the digital twin model, laying the foundation for subsequent model optimization. For example, geostatistical methods (such as Kriging interpolation and inverse distance weighting) can be used to interpolate a limited number of measured soil parameters (such as permeability coefficient, porosity, and initial pollutant concentration) into all grid cells. Alternatively, based on the regional division of the application scenario (such as different soil type zones or different pollution zones), a set of empirical parameter values ​​can be assigned to grid cells within each region. For parameters with high uncertainty, values ​​can be randomly sampled from a pre-defined initial parameter population to cover the possible parameter space.

[0049] Finally, the discrete grid model is combined with the coupled mathematical model to form a solvable set of coupled equations. This process integrates all selected fundamental mechanism models, their coupling relationships, and the parameters assigned on the discrete grid into a complete mathematical system, ensuring that this system can be solved numerically, thereby achieving dynamic simulation of the soil remediation process. For example, when the fundamental mechanism model is mainly described by partial differential equations (PDEs), after discretization and coupling, it forms a large set of algebraic equations (linear or nonlinear), which can be solved using mature numerical solvers (such as the Newton-Raphson method and the conjugate gradient method).

[0050] The model established through the above modeling method incorporates terrain data of the entire target application scenario. Therefore, it can deduce the distribution of target substances in the entire target application scenario based on the content of target substances at multiple monitoring points, thereby realizing dynamic monitoring of the restoration process of the entire area.

[0051] The modeling method described above is relatively simple, with a straightforward process and easy implementation. However, since most mechanistic models are idealized, they have inherent limitations. The established digital twin model can only optimize a limited number of parameters, resulting in inaccurate initial predictions. This necessitates additional optimization steps, increasing the computational burden and reducing the efficiency of real-time dynamic simulation.

[0052] In response, this application proposes another method for establishing a new digital twin model:

[0053] S123. Select the required basic mechanism model from the basic mechanism model library according to the target application scenario, use the basic mechanism model as the loss function of the constrained embedded neural network, establish a digital twin model, and construct the uncertain parameters in the digital twin model as the initial parameter population.

[0054] By incorporating knowledge of the mechanisms of the physical world into data-driven neural network models, the physical plausibility and generalization ability of the models can be improved. The loss function is a metric used to measure the model's prediction error during neural network training. By adding the underlying mechanistic model as a constraint term to the loss function, the neural network can be guided to learn data features while also adhering to known physical laws. For example, the architecture of a Physical Information Neural Network (PINN) can be used, adding the partial differential equations (PDEs) or algebraic equations of the mechanistic model as regularization terms to the neural network's loss function. Alternatively, this can be achieved through soft or hard constraints. Soft constraints involve directly adding the residual terms of the mechanistic model to the loss function, while hard constraints can be implemented through network structure design (such as physical coding layers) or parameterization to ensure that the neural network's output naturally satisfies certain mechanistic constraints.

[0055] Specifically, in step S123, setting the initial parameter population based on experience means setting a reasonable initial value range or initial value set for the key parameters in the newly established digital twin model based on the knowledge of experts in the field, historical project data, literature, or preliminary experimental results, so as to ensure that the model can start searching from a meaningful starting point in subsequent optimization.

[0056] Specifically, in step S123, uncertain parameters refer to parameters in the digital twin model whose precise values ​​are difficult to determine due to data limitations, model simplification, or environmental complexity. Organizing these parameters into a population is the basis for the genetic algorithm to perform global search and optimization. For example, parameters in the digital twin model that have a significant impact on the prediction results and are uncertain can be identified, and a reasonable initial value range can be set for each uncertain parameter. Then, a set of parameter combinations can be randomly generated from these ranges to form the initial parameter population. Alternatively, key uncertain parameters in the model can be determined through sensitivity analysis or expert experience, and a probability distribution or interval can be defined for each parameter. Then, multiple parameter instances can be generated through sampling (such as Monte Carlo sampling) to constitute the initial parameter population.

[0057] By employing the aforementioned technical solution, the fundamental mechanism model is used as the loss function embedded in the neural network to constrain its learning. This allows the neural network to follow known physical laws while learning data features, significantly enhancing the model's physical reliability and generalization ability. It effectively reduces the risk of overfitting that may occur with purely data-driven models, thereby improving the predictive accuracy of the digital twin model. This mechanism-data fusion approach to establishing the digital twin model provides a more robust and reasonable starting point for subsequent iterative optimization. Furthermore, by constructing an initial parameter population from the uncertain parameters in the digital twin model, the uncertain parameters are directly transformed into the optimization starting point for the genetic algorithm. This enables the genetic algorithm to perform global search and optimization more efficiently, reducing the impact of initial setting errors on overall simulation accuracy. The digital twin model established in step S123 overcomes the limitations of traditional empirical initial parameter setting, reduces reliance on additional optimization steps, thereby reducing computational burden and improving the efficiency and accuracy of real-time dynamic simulation.

[0058] Specifically, in steps S1 and S13, soil state data and environmental state data (real-time environmental state data, real-time soil state data, historical soil state data, and historical environmental state data) refer to real data of the target application scenario obtained from sensor monitoring or laboratory analysis. This data includes soil state data such as soil moisture content, pH value, bulk density, and porosity, as well as environmental state data such as ambient temperature, ambient humidity, rainfall, and atmospheric pressure. These historical data are input into the currently constructed or invoked digital twin model. The model will simulate and output the predicted content of the target substance at the corresponding time point based on its internal mechanisms and parameters. The data types included in the soil state data and environmental state data need to be adjusted according to different digital twin models.

[0059] Specifically, in step S14, calculating fitness refers to quantifying the model's accuracy by comparing the predicted content of the target substance output by the model with the historically measured content of the target substance. For example, the reciprocal of the mean squared error (MSE), root mean square error (RMSE), or mean absolute error (MAE) can be used as the fitness function; the smaller the error, the higher the fitness. Iterative optimization using a genetic algorithm involves encoding the key parameters of the digital twin model as chromosomes. By simulating selection, crossover, and mutation in biological evolution, new parameter combinations (population) are continuously generated. Based on the fitness of each parameter combination, the best and worst are eliminated, gradually searching for the optimal set of parameters that maximizes fitness, thereby enabling the digital twin model to more accurately predict the content of the target substance.

[0060] In some of the above implementation methods, model optimization based on the deviation between prediction and measurement is proposed to improve prediction accuracy. However, in the process of implementation, the prediction data itself may have uncertainty, which may lead to the optimization process being unstable or unreliable.

[0061] In this regard, this application further proposes that the target substance prediction data also include a confidence level corresponding to the predicted content of the target substance. This confidence level is used to quantify the reliability or uncertainty of the predicted content of the target substance. For example, the confidence level can be expressed as the probability that the predicted value falls within a certain interval, or it can be reflected by the variance of the prediction results of different models in an ensemble learning model (such as random forest or gradient boosting tree). Alternatively, the predicted value and its confidence interval can be directly output using methods such as Bayesian neural networks to obtain the confidence level. By introducing the confidence level, the system can self-assess the reliability of the prediction results.

[0062] Furthermore, the method also includes the following steps: when the confidence level is less than a preset confidence threshold, the system will trigger a specific optimization mechanism, such as using a genetic algorithm to perform bi-objective optimization on the digital twin model with fitness maximization and confidence level greater than the confidence threshold as optimization objectives.

[0063] The confidence threshold is a preset value used to define the reliability level of the prediction result. For example, the confidence threshold can be set to 0.8 or 0.9 based on historical data analysis and expert experience, indicating that when the confidence level of the prediction result is lower than 80% or 90%, the reliability of the prediction result is considered insufficient. When insufficient confidence level of the prediction result is detected, the system will calculate the fitness based on the predicted content and the measured content of the target substance at the current moment.

[0064] Genetic algorithms are optimization algorithms that simulate natural selection and genetic mechanisms. They search for optimal solutions in the parameter space through operations such as selection, crossover, and mutation. Bi-objective optimization means considering two related or potentially conflicting objectives simultaneously during the optimization process: maximizing the model's predictive accuracy (reflected by fitness) and ensuring the model's predictive reliability (reflected by a confidence score greater than a confidence threshold). For example, multi-objective optimization algorithms such as NSGA-II (Non-dominated Sorting Genetic Algorithm II) can be used. In each iteration, the algorithm generates a set of Pareto optimal solutions, where each solution represents a trade-off between accuracy and reliability. In this way, the system can improve the reliability of prediction results while maintaining prediction accuracy.

[0065] By introducing confidence levels into the target substance prediction data, the system can assess the reliability of the prediction results in real time. When the confidence level is lower than a preset confidence threshold, it indicates that the current prediction result may have significant uncertainty. At this point, the system will not blindly perform single-objective optimization, but will trigger a more robust dual-objective optimization mechanism. This mechanism utilizes a genetic algorithm to iteratively optimize the digital twin model, simultaneously aiming to maximize prediction accuracy (reflected by fitness) and ensure prediction reliability (reflected by confidence levels exceeding the confidence threshold). This allows the digital twin model to effectively adjust and learn even under low confidence conditions, thereby avoiding optimization bias or model degradation caused by prediction uncertainty. This not only improves the prediction accuracy of the digital twin model for soil remediation processes, but more importantly, significantly enhances the reliability of the prediction results and the robustness of the model.

[0066] Based on the above method, this application also proposes a real-time dynamic simulation system for soil remediation processes based on digital twins, referencing... Figure 3 It includes: a data acquisition module, a simulation prediction module, a data comparison module, and a parameter optimization module.

[0067] The data acquisition module continuously acquires real-time environmental status data such as ambient temperature, humidity, and rainfall, as well as real-time soil status data such as soil moisture content and pH value, through a sensor network deployed at the target application scenario (soil remediation site). At the same time, it simultaneously collects the measured content of the target substance.

[0068] The simulation prediction module drives the digital twin model based on the real-time data (real-time soil condition data and real-time environmental condition data). This model integrates soil remediation mechanism and dynamic process simulation, and outputs the predicted content of target substances, overcoming the limitations of traditional mathematical models in complex environments.

[0069] The data comparison module is used to calculate the relative deviation between the measured content of the target substance at the current moment and the predicted content of the target substance at the current moment, quantify the degree of difference between the prediction results and the measured data, and provide an objective basis for model optimization.

[0070] The parameter optimization module is used to initiate the genetic algorithm optimization process when the relative deviation is greater than the accuracy threshold. It calculates the fitness based on the predicted content and the measured content of the target substance at the current moment, and uses the genetic algorithm to iteratively optimize the digital twin model module with fitness maximization as the optimization objective.

[0071] The advantage of the aforementioned system lies in its dynamic integration of data acquisition, simulation prediction, data comparison, and parameter optimization modules through a closed-loop feedback mechanism. This allows for real-time detection of prediction deviations and adaptive optimization of the digital twin model when external environmental changes affect the soil remediation process. This avoids the prediction failure issues caused by static models in traditional methods, achieving accurate prediction of the soil remediation process and supporting timely adjustments to remediation strategies. For example, after a rainfall event, the system uses real-time collected soil moisture content change data to drive the digital twin model to update its predictions. Simultaneously, the data comparison module identifies deviations in pollutant degradation rates, and the parameter optimization module triggers a genetic algorithm for optimization, ensuring timely adjustments to the remediation strategy after environmental changes. This system effectively solves the technical problems of unpredictable external environmental changes and insufficient model accuracy during soil remediation, achieving dynamic simulation and precise control of the remediation process, thereby significantly improving soil remediation effectiveness.

[0072] Specifically, in some embodiments, the simulation prediction module includes a digital twin model library and a mechanism model library. The digital twin model library is used to store optimized digital twin models, and the mechanism model library is used to store preset basic mechanism models.

[0073] The simulation prediction module is also used to match the target application scenario, determine whether there is a digital twin model in the digital twin model library that is compatible with the target application scenario, if there is, call the digital twin model and extract its optimal parameter range as the initial parameter population, if there is no, select the required basic mechanism model based on the target application scenario, set the initial parameter population based on experience, and construct the digital twin model of the target application scenario based on the initial parameter population.

[0074] The simulation prediction module can be implemented as a standalone software service module, such as a microservice deployed on a server, or as a subsystem integrated into a larger simulation platform. The digital twin model library is a data storage unit specifically designed to store trained, optimized, and validated digital twin models. These models are typically parameter-tuned for specific soil remediation scenarios or problem types and can be directly invoked to improve simulation efficiency and accuracy. Its implementation can be a relational database storing model metadata and paths to model files, or a distributed file system directly storing serialized model objects. The mechanistic model library is similar to the digital twin model library, implemented in the same way, and stores fundamental mechanistic models describing various physical, chemical, and biological processes in soil remediation. These models are typically scientifically validated sets of mathematical equations or algorithms, such as hydrological models, solute transport models, and biodegradation kinetic models.

[0075] Specifically, in some embodiments, the data acquisition module includes multiple soil state sensors and environmental state sensors. Multiple monitoring points are preset in the target application scenario (remediation area), and each monitoring point is equipped with a soil state sensor and an environmental state sensor. The soil state sensors are used to collect real-time soil state data and measured content of target substances in the target application scenario in real time. The environmental state sensors are used to collect real-time environmental state data of the target application scenario in real time. The simulation prediction module is also used to establish a discrete grid model based on the terrain data of the target application scenario, and assign values ​​to the discrete grid model according to the initial parameter population. It forms a solvable coupled equation set with the coupled mathematical model as a digital twin model, so as to deduce the predicted content distribution of target substances in the entire target application scenario based on the predicted content of target substances at multiple monitoring points.

[0076] Specifically, soil condition sensors can employ multi-parameter soil sensor integrated modules. These modules can simultaneously measure conventional parameters such as soil moisture content, pH value, and electrical conductivity, and integrate specific ion-selective electrodes or spectral analysis modules for direct or indirect measurement of the actual content of target substances (such as heavy metal ions and organic pollutant concentrations). Another approach combines in-situ sensors with portable rapid detection equipment. For example, in-situ sensors continuously monitor soil moisture content and pH value, while the actual content of target substances is obtained through periodic sampling at monitoring points and rapid on-site detection using portable XRF (X-ray fluorescence spectrometer) or rapid colorimetric methods.

[0077] Specifically, environmental status sensors can be small weather stations or integrated environmental sensors, capable of real-time monitoring of various environmental parameters such as ambient temperature, humidity, rainfall, atmospheric pressure, wind speed and direction, and solar radiation. Alternatively, data interfaces from existing meteorological monitoring networks can be utilized, combined with data from regional weather stations near the target application scenario, and supplemented and calibrated using locally deployed temperature and humidity sensors to obtain more refined real-time environmental status data.

[0078] Specifically, in the process of establishing a discrete grid model based on the terrain data of the target application scenario, Geographic Information System (GIS) software can be used to import the elevation data (DEM, Digital Elevation Model) of the target application scenario and divide it into regular two-dimensional or three-dimensional grids (such as rectangular grids or triangular grids) using GIS tools, which will serve as the basis for the subsequent establishment of finite difference method or discrete element method models.

[0079] The digital twin model established above can utilize the predicted content at monitoring points as boundary conditions or calibration points. Combined with simulations of physical, chemical, and biological processes within the model, the predicted content of the target substance in each cell of the entire discrete grid can be calculated, thus obtaining a predicted content distribution map for the entire application scenario. For example, during the digital twin model solution process, the predicted content at monitoring points can be used as the model's output verification points. The model calculates the predicted content of all grid cells through numerical simulation, and then uses visualization techniques (such as heatmaps and contour maps) to render these discrete grid data into a continuous and intuitive distribution map. Alternatively, data assimilation techniques can be used to fuse the predicted content at real-time monitoring points with the simulation results of the digital twin model. Through methods such as Kalman filtering or ensemble Kalman filtering, the model state can be continuously corrected, making the predicted content distribution of the entire scenario output by the model closer to the actual situation and dynamically reflecting the remediation process.

[0080] The system provided in this application overcomes the limitations of traditional monitoring methods, which can only acquire discrete point data and cannot comprehensively grasp the dynamics of the entire remediation scenario. Through a high-density sensor network and refined digital twin modeling, it achieves accurate derivation from local point data to the distribution of target materials throughout the application scenario, greatly improving the accuracy and comprehensiveness of real-time dynamic simulation of the soil remediation process. Engineers can intuitively understand the spatiotemporal variation trends of target materials within the entire remediation area, enabling them to adjust remediation strategies more promptly and accurately, optimize remediation effects, and effectively address the impact of external environmental changes on the remediation process.

[0081] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A real-time dynamic simulation method for soil remediation processes based on digital twins, characterized in that, Includes the following steps: The real-time environmental data and real-time soil condition data of the previous moment are acquired and input into the digital twin model to generate the target substance prediction data of the current moment, which includes the predicted content of the target substance. Obtain the measured content of the target substance at the current moment, and calculate the relative deviation between the measured content of the target substance and the predicted content of the target substance at the current moment; Determine if the relative deviation is greater than the accuracy threshold. If so, calculate the fitness based on the predicted content and the measured content of the target substance at the current moment, and use a genetic algorithm to iteratively optimize the digital twin model with fitness maximization as the optimization objective. If not, output the predicted content of the target substance.

2. The method according to claim 1, characterized in that, Obtain a digital twin model by following these steps: Select the target application scenario and match it with the preset digital twin model library to determine whether there is a suitable digital twin model; If a suitable digital twin model exists, the model is invoked and its optimal parameter range is extracted as the initial parameter population. If no suitable model exists, the required basic mechanism model is selected based on the target application scenario, and the initial parameter population is set based on experience. The digital twin model of the current scenario is then constructed based on the initial parameter population. Acquire multiple sets of historical soil condition data, historical environmental condition data, and historical measured target substance content for the target application scenario. Input the historical soil condition data and historical environmental condition data into the digital twin model and output the predicted target substance content. Fitness is calculated using the predicted content of the target substance and the historical measured content of the target substance. A genetic algorithm is then used to iteratively optimize the digital twin model with fitness maximization as the optimization objective.

3. The method according to claim 1, characterized in that, The following steps determine whether a suitable digital twin model exists: A scenario feature vector is established based on the real-time environmental state data and real-time soil state data of the target application scenario. The components of the scenario feature vector correspond to a parameter value in the real-time environmental state data or the real-time soil state data, respectively. Calculate the similarity between the scene feature vector of the target application scenario and the scene feature vector of each digital twin model in the digital twin model library. When the similarity is greater than the matching threshold, it is determined that there is a suitable digital twin model.

4. The method according to claim 1, characterized in that, The new digital twin model will be established through the following steps: Based on the target application scenario, select the required basic mechanism model from the basic mechanism model library and establish the coupling relationship between different basic mechanism models; A discrete grid model is established based on the terrain data of the target application scenario, and the discrete grid model is assigned values ​​according to the initial parameter population. This forms a solvable set of coupled equations with the coupled mathematical model, which is the digital twin model.

5. The method according to claim 2, characterized in that, The new digital twin model will be established through the following steps: Based on the target application scenario, select the required basic mechanism model from the basic mechanism model library, use the basic mechanism model as the loss function of the constrained embedded neural network, establish a digital twin model, and construct the uncertain parameters in the digital twin model as the initial parameter population.

6. The method according to claim 1, characterized in that, The real-time environmental status data includes: ambient temperature, ambient humidity, rainfall, and atmospheric pressure; The real-time soil condition data includes: soil moisture content, pH value, bulk density, and porosity.

7. The method according to claim 1, characterized in that, The target substance prediction data also includes the confidence level corresponding to the predicted content of the target substance; The method further includes the following steps: When the confidence level is less than the confidence threshold, the fitness is calculated based on the predicted content and the measured content of the target substance at the current moment. Then, a genetic algorithm is used to perform bi-objective optimization on the digital twin model with the optimization objectives of maximizing the fitness and ensuring that the confidence level is greater than the confidence threshold.

8. A real-time dynamic simulation system for soil remediation processes based on digital twins, used to implement the method described in any one of claims 1-7, characterized in that, include: The data acquisition module is used to collect real-time environmental status data, real-time soil status data, and measured content of target substances in the target application scenario. The simulation prediction module is used to output target substance prediction data based on real-time environmental state data and real-time soil state data using a digital twin model. The target substance prediction data includes the predicted content of the target substance. The data comparison module is used to calculate the relative deviation between the measured content of the target substance at the current moment and the predicted content of the target substance at the current moment. The parameter optimization module is used to calculate the fitness based on the predicted content and the measured content of the target substance at the current moment when the relative deviation is greater than the accuracy threshold, and to use a genetic algorithm to iteratively optimize the digital twin model module with fitness maximization as the optimization objective.

9. The system according to claim 8, characterized in that, The simulation prediction module includes a digital twin model library and a mechanism model library; The digital twin model library is used to store optimized digital twin models; The mechanism model library is used to store preset basic mechanism models; The simulation prediction module is also used to match the target application scenario, determine whether there is a digital twin model in the digital twin model library that is compatible with the target application scenario, if there is, call the digital twin model and extract its optimal parameter range as the initial parameter population, if there is no, select the required basic mechanism model based on the target application scenario, set the initial parameter population based on experience, and construct the digital twin model of the target application scenario based on the initial parameter population.

10. The system according to claim 9, characterized in that, The data acquisition module includes multiple soil condition sensors and environmental condition sensors. Multiple monitoring points are preset in the target application scenario, and each monitoring point is equipped with a soil condition sensor and an environmental condition sensor. The soil state sensor is used to collect real-time soil state data and measured content of target substances in the target application scenario, and the environmental state sensor is used to collect real-time environmental state data in the target application scenario. The simulation prediction module is also used to establish a discrete grid model based on the terrain data of the target application scenario, and assign values ​​to the discrete grid model according to the initial parameter population. It forms a solvable set of coupled equations with the coupled mathematical model as a digital twin model, so as to deduce the target material prediction content distribution of the entire target application scenario based on the target material prediction content of multiple monitoring points.