A Method for Decoupling and Predicting Multi-Source Signals in Termites Based on eDNA and Machine Learning
By constructing a transport-response mapping mechanism and an improved machine learning regression framework, combined with sparse source inversion, the difficulty of decoupling multi-source signals in multi-nest inversion in dams was solved, enabling accurate location and scale calculation of termite nests and improving the scientific nature of dam safety monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING HYDRAULIC RES INST
- Filing Date
- 2026-02-11
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies face difficulties in decoupling multi-source signals and lack physical transport mechanisms when dealing with the multi-nest inversion problem in complex dam environments, leading to fuzzy termite nest location and biased biomass estimation.
A transport-response mapping mechanism is constructed, which combines machine learning and physical models. Through sparse source inversion and an improved machine learning regression framework, the precise location and scale calculation of termite nests are achieved. A spatiotemporally coupled transport mechanism is used to handle multi-source signal aliasing.
It has enabled precise location and scale calculation of hidden termite nests inside dams, solved the problems of difficult decoupling of multi-source signals and low accuracy of non-steady-state transport inversion, and provided a scientific basis for dam safety management.
Smart Images

Figure CN121682782B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of dam safety monitoring technology, and in particular to a method for decoupling and predicting multi-source signals of termites based on eDNA and machine learning. Background Technology
[0002] The detection and assessment of termite hazards within dams and other water conservancy projects is a key technical challenge in the field of geotechnical engineering safety monitoring. Termites construct complex nest systems and tunnel networks within the dam structure, leading to reduced soil shear strength and abnormal permeability, which in turn can cause piping and even dam failure risks. Environmental DNA (eDNA) technology, as a molecular tracing method, can non-destructively track underground hidden biological sources by detecting residual biomolecule fragments in the soil medium, providing new physical and biological evidence for the precise location of termite nests and the quantitative inversion of biomass scale within dams.
[0003] Existing termite detection techniques for dams primarily rely on geophysical methods such as high-density electrical resistivity tomography (EDT) and ground-penetrating radar (GPR), as well as manual inspections based on surface signs. In studies incorporating eDNA technology, current mainstream approaches typically employ static gridded sampling, combined with general machine learning regression algorithms (such as random forests and support vector machines) or simplified steady-state diffusion models (such as Gaussian plume models) for data processing. These methods attempt to establish a direct mapping between soil eDNA concentration and potential nest coordinates, inferring nest locations through statistical fitting or simple spatial interpolation.
[0004] However, the aforementioned existing technologies have limitations in handling the multi-nest inversion problem in complex dam environments, mainly manifested in the difficulty of decoupling multi-source signals and the lack of physical transport mechanisms. Specifically, existing machine learning regression methods typically treat the source-field relationship as a black box, making it difficult to mathematically separate the superimposed signals generated by multiple adjacent nests. This results in semantic ambiguity in localization in multi-nest coexistence scenarios and an inability to resolve individual size. Furthermore, the use of a steady-state Gaussian model ignores the reaction-diffusion kinetics of eDNA as a biomolecule in porous media, failing to consider the degradation and decay over time and the hysteresis effect caused by soil adsorption. This lack of a non-steady-state physical mechanism leads to the inability to distinguish between nearby old traces and distant new traces in the inversion results, resulting in significant biases in nest activity determination and biomass estimation. Summary of the Invention
[0005] The purpose of this invention is to provide a method for decoupling and predicting multi-source signals in termites based on eDNA and machine learning, in order to solve one of the problems mentioned above in the existing technology.
[0006] Technical solution: A method for decoupling and predicting multi-source signals in termites based on eDNA and machine learning, comprising:
[0007] Soil samples were collected from the target dam area and quantitatively analyzed to obtain soil environmental DNA concentration distribution data;
[0008] A transport-response mapping mechanism was constructed to characterize the spatiotemporal correlation between termite nest source terms and soil environmental DNA concentration distribution data.
[0009] Based on soil environmental DNA concentration distribution data, a multi-source inversion analysis was performed using a transport-response mapping mechanism to calculate the termite nest occurrence state parameters, which include the nest spatial coordinates and biomass scale.
[0010] Output the termite nest status parameters.
[0011] Beneficial effects: This invention constructs a physically constrained sparse source inversion or improved machine learning regression framework. Specifically: when data is sufficient and meets the training conditions for the machine learning model, a machine learning method is used; when data is insufficient to support machine learning model training, a physically constrained inversion method is used. For dense sampling points, a random forest model is used; for sparse sampling points with significant anisotropy, physically constrained inversion is employed. Furthermore, a spatiotemporally coupled transport mechanism is introduced to achieve accurate location and scale calculation of hidden termite nests inside dams, solving the problems of difficult decoupling of multi-source signal aliasing and low accuracy of non-steady-state transport inversion. Attached Figure Description
[0012] Figure 1 This is a flowchart of a termite multi-source signal decoupling and prediction method based on eDNA and machine learning in an embodiment of this application.
[0013] Figure 2 The embodiments of this application generate virtual termite nest scene diagrams containing different numbers and locations based on a preset diffusion model.
[0014] Figure 3 This is a flowchart of the machine learning model in an embodiment of this application.
[0015] Figure 4 This is a field drilling sampling diagram from an embodiment of this application.
[0016] Figure 5 This is a flowchart illustrating the steps of the sparse source inversion strategy with physical constraints in the multi-source inversion analysis of this application embodiment.
[0017] Figure 6 This is a graph showing the composition of the optimized objective function in the embodiments of this application.
[0018] Figure 7 This is a flowchart illustrating the steps involved in obtaining soil environmental DNA concentration distribution data in an embodiment of this application. Detailed Implementation
[0019] Example 1 describes the general logical closed loop of a termite multi-source signal decoupling and prediction method based on eDNA and machine learning, from field data acquisition to final inversion result output, such as... Figure 1 As shown, this provides a unified framework for subsequent implementations based on machine learning or physical models.
[0020] Step 101: Collect soil samples from the target dam area and perform quantitative analysis to obtain soil environmental DNA concentration distribution data.
[0021] In this embodiment, data acquisition is fundamental to the entire method. Collecting soil samples from the target dam area typically involves systematic physical sampling at key locations within the dam. Specifically, the dam can be divided into multiple sampling units, for example, by arranging drilling points at a predetermined grid spacing (e.g., 5m x 5m or 10m x 10m). The sampling area should primarily cover the lower part of the upstream slope (e.g., 0 to 2 meters above the waterline), the middle of the downstream slope (e.g., near the overflow point of the seepage line), and the dam crest, as these areas are often high-risk for termite nesting. Soil sample collection is typically accomplished using small, portable drilling equipment, such as a 5cm diameter soil drill. To obtain vertical distribution information, soil samples need to be collected at specific depth intervals (e.g., every 1 meter) at each drilling point until a predetermined depth (e.g., 5 meters) is reached. The collected soil samples must be immediately placed in sterile containers and stored at low temperatures to prevent DNA degradation. Subsequently, soil samples undergo quantitative analysis, which typically includes extraction of total soil DNA and real-time quantitative PCR (qPCR) amplification of termite-specific gene fragments (such as COI or 16S rRNA genes). The raw signals obtained (such as the cycle threshold Ct) are converted into the number of eDNA copies per unit mass of soil after standard curve calculation. The summation of all sampling points at different spatial locations (three-dimensional coordinates) and their corresponding eDNA concentrations constitutes the soil environmental DNA concentration distribution data. This dataset is formally represented as D. in ={(x i ,y i ,z i C i )|i=1...N}, where N is the total number of samples.
[0022] In some alternative implementations, to more intuitively display the sampling results, spatial interpolation processing of soil environmental DNA concentration distribution data can be performed using a geographic information system (such as ArcGIS software). For example, Kriging interpolation can be used to extrapolate discrete sampling point data into a continuous concentration field distribution map, aiding in the identification of high-concentration anomaly areas. This type of visualized distribution feature can serve as a priori reference for subsequent inversion analysis, but is not directly used as quantitative inversion input.
[0023] Step 102: Construct a transport-response mapping mechanism to characterize the spatiotemporal correlation between termite nest source terms and soil environmental DNA concentration distribution data.
[0024] In this embodiment, the transport-response mapping mechanism serves as a bridge connecting the unknown source (termite nest) and the known observation (eDNA concentration). It describes how a potential termite nest (source) generates an observable eDNA concentration field (response) in the dam medium. This type of mapping relationship can be a data-driven model based on statistical laws or an analytical model based on physical diffusion laws. Specifically, if it is a data-driven approach, the mechanism is embodied in a pre-trained machine learning model, such as... Figure 3 As shown, it learns a large number of source-field sample pairs to master nonlinear rules mapping spatial location and concentration distribution characteristics to nest parameters. If based on a physical analysis approach, this mechanism manifests as a set of mathematical equations or observation matrices that explicitly describe the physical processes of eDNA molecule decay with distance, degradation over time, and adsorption hindrance in porous media. Regardless of the form, the mechanism must be able to quantitatively reflect how changes in source terms (such as increased biomass or location shift) lead to corresponding changes in concentration distribution data.
[0025] Step 103: Based on the soil environmental DNA concentration distribution data, multi-source inversion analysis is performed using the transport-response mapping mechanism to calculate the termite nest occurrence state parameters, which include the nest spatial coordinates and biomass scale.
[0026] In this embodiment, multi-source inversion analysis is a step of inversely deducing source term attributes using observed data. Given that multiple termite nests may exist simultaneously inside the dam, and their signals spatially superimposed, multi-source analysis is necessary. When the transport-response mapping mechanism is a machine learning model, multi-source inversion analysis specifically involves inputting the concentration distribution data collected on-site (usually processed by feature engineering) into the model, which then directly predicts the parameters of the output nest. When the mechanism is a physical model, the analysis process involves solving an inverse problem, i.e., finding the optimal source term distribution that minimizes the error between the theoretical concentration field generated by the source term through forward mapping and the actually observed concentration distribution data. The calculated termite nest occurrence state parameters are key indicators describing the degree of termite damage. Among them, the nest spatial coordinates (x, y, z) clarify the three-dimensional location of the hazard, guiding subsequent excavation or grouting treatment; the biomass scale (e.g., biomass M in grams or volume V in cubic meters) calculates the size of the nest and its potential destructive power. The output results are usually presented in list form, such as P out ={(x k ,y k ,z k M k V k )|k=1...K}, where K is the number of nests obtained by inversion.
[0027] Step 104: Output the termite nest status parameters.
[0028] In this embodiment, the final output step transforms the inverted parameters into user-usable information. This can be presented through a computer terminal display interface or by generating an electronic document containing a list of nest locations, size assessments, and risk levels. Preferably, the predicted nest locations can be overlaid with a three-dimensional terrain model of the dam to generate an intuitive distribution map of termite damage on the dam, marking hazardous areas requiring special attention, and providing a scientific basis for dam safety management and termite control projects.
[0029] Example 2 describes how to use machine learning techniques to construct a transport-response mapping mechanism and solve the problems of scarce field samples and insufficient model generalization ability through virtual data generation and transfer learning fine-tuning.
[0030] Step 201: The multi-source inversion analysis adopts a pre-trained random forest multi-output regression model. Based on the soil environmental DNA concentration distribution data, the multi-source inversion analysis is performed using the transport-response mapping mechanism. Specifically, this includes: extracting the spatial coordinates of sampling points and the corresponding concentration values from the soil environmental DNA concentration distribution data, and constructing a joint input feature vector; inputting the joint input feature vector into the pre-trained random forest multi-output regression model, and outputting termite nest occurrence state parameters, which include the nest location coordinates, biomass, and nest volume around each sampling point.
[0031] In this embodiment, a random forest multi-output regression model is chosen as the inversion tool because this model can simultaneously predict multiple related target variables (such as coordinates, biomass, and volume). Specifically, when performing field predictions, the discrete soil environmental DNA concentration distribution data needs to be organized into a format acceptable to the model. For each sampling point, its three-dimensional coordinates (x, y, z) are combined with the measured eDNA concentration value C at that point, or the data from multiple adjacent sampling points within a sampling unit are combined to construct a joint input feature vector. For example, the feature vector can be represented as V in =[x,y,z,C,ΔC xy ,ΔC z This vector contains information about spatial location, absolute concentration, and horizontal and vertical gradients of concentration. After inputting this vector into a pre-trained model, numerous decision trees within the model perform inferences, and output the existence state parameters of the termite nest to which the sampling point belongs or which is a neighboring termite nest through ensemble voting or averaging.
[0032] In some alternative implementations, to improve the interpretability of the model, the Shapley Additive Explanations (SHAP) method can be introduced to analyze the model's prediction results. By calculating the Shapley value for each input feature (such as concentration values at different depths), the contribution of that feature to the final nest location or size prediction can be calculated.
[0033] Step 202: The pre-trained random forest multi-output regression model is trained based on a virtual training dataset. The construction process of the virtual training dataset includes: generating virtual termite nest scenes with different numbers and locations based on a pre-defined diffusion model, such as... Figure 2 As shown, the mixed environment DNA concentration at virtual sampling points in the virtual termite nest scene is calculated. The mixed environment DNA concentration is obtained by linearly superimposing the independent concentration contributions of each virtual termite nest and applying a concentration change rate correction. The mixed environment DNA concentration and the coordinates of the virtual sampling points are used as input features, and the corresponding virtual termite nest scene parameters are used as output labels to generate a virtual training dataset.
[0034] In this embodiment, constructing a virtual training dataset is a key strategy to address the lack of training data for machine learning models. Because excavating and confirming termite nests in person is costly and destructive, it is difficult to obtain massive amounts of concentration-nest measured label data. Therefore, using computer simulation to generate virtual data becomes an inevitable choice. The specific process is as follows: In a virtual three-dimensional space equivalent to the actual dam size, 1 to 10 termite nests are randomly selected. The location (x0, y0, z0) and biomass M of each nest are randomly selected within a reasonable range, forming a virtual scene. Based on a pre-set diffusion model (such as a Gaussian diffusion model or a spatiotemporal coupling model), the independent concentration contribution of each nest at any point in space is calculated. For a virtual sampling point p, its mixed concentration C... mix (p) Initially obtained by linearly summing the contributions of each nest. Real biological and physical processes are often nonlinear, so a concentration change rate correction is needed to simulate the interaction between multiple nests. The generated virtual dataset contains tens of thousands of samples, each consisting of known nest parameters (labels) and simulated mixed concentrations (features), sufficient to support the training of the deep random forest model.
[0035] In some alternative implementations, besides generating data based on physical models, conditional generative adversarial networks (cGANs) can be used to enhance the realism of the virtual data. Specifically, a generator G is constructed to generate a simulated concentration distribution, and a discriminator D is constructed to distinguish the simulated data from a small amount of measured data. Through adversarial training, the virtual data generated by the generator approximates the distribution characteristics of real data. To ensure generation quality, the Wasserstein distance can be introduced as an evaluation metric to calculate the difference between the generated data distribution and the measured data distribution. High-quality virtual samples with a Wasserstein distance less than a preset threshold are selected and added to the training set to improve the robustness of the model.
[0036] Step 203, the specific calculation process of concentration change rate correction includes: based on the preset diffusion model, simulating and calculating the single-nest concentration value of the virtual termite nest when it exists independently and the multi-nest concentration value when multiple nests coexist; based on the relative difference between the multi-nest concentration value and the single-nest concentration value, calculating the concentration change rate index; using the concentration change rate index to perform weighted correction on the linear superposition results of each virtual termite nest to obtain the mixed environment DNA concentration, characterizing the nonlinear interaction between multiple nests.
[0037] In this embodiment, the concentration change rate correction aims to eliminate the error caused by simple linear superposition. When constructing virtual data, the measured concentration (multi-nest concentration value C) when two or more nests coexist under specific distance and medium conditions can be obtained through simulation or by consulting experimental databases. multi The sum of their theoretical concentrations when they exist individually (linear superposition value C) and their theoretical concentrations when they exist individually.lin The deviation between (p and concentration change rate) can be defined as R(p) = C. multi -C lin / C lin Where R(p) is the concentration change rate index, which is dimensionless; C multi The actual concentration values of multiple blocks during the coexistence time; C lin This is the linear superposition prediction of the concentration when each individual block exists independently. If R(p) > 0, it indicates a synergistic enhancement effect; if R(p) < 0, it indicates an inhibitory or competitive effect. This index is used for correction when calculating the mixed concentration of virtual sampling points, and the formula can be expressed as C. mix (p)=C lin (p)·(1+R(p)). Where, C mix (p) represents the corrected mixed environment DNA concentration at sampling point p; p is the three-dimensional spatial coordinate of the sampling point; C lin R(p) represents the linearly superimposed predicted concentration at sampling point p; R(p) is the concentration change rate index. This type of correction makes the virtual training data closer to the real soil environment, improving the prediction accuracy of the model in practical applications.
[0038] Step 204: Before applying the pre-trained random forest multi-output regression model to field prediction, it is also fine-tuned using a small field sample dataset. The fine-tuning process includes: selecting some representative field drilling sampling point data, freezing the parameters of the model's preceding feature extraction layer, and iteratively updating only the parameters of the model's final output layer until the model's prediction error on the field validation set is less than a preset threshold.
[0039] In this embodiment, fine-tuning through transfer learning is a crucial step in enabling the model to transition from the virtual world to the real world. Although virtual data is abundant and comprehensive, it cannot fully simulate all the micro-environmental characteristics of a specific dam site (such as interference from specific soil chemical compositions). Therefore, a model trained directly on purely virtual data may not be suitable for the actual field. To address this issue, after model pre-training, a small amount of data with real labels collected on-site (e.g., dozens of sets of data obtained through small-scale excavation verification) is used to fine-tune the model. Specifically, the pre-structures used for feature extraction and nonlinear transformation in the random forest (such as split nodes and thresholds) are kept unchanged, i.e., these parameters are frozen, while only the output layer weights used for final numerical regression are retrained. This strategy preserves the general patterns learned by the model from massive amounts of virtual data while adapting it to the specific environment of the current dam. The training process typically employs a small learning rate and includes an early stopping mechanism. That is, when the model's prediction error (e.g., root mean square error RMSE) on the field validation set no longer decreases or falls below a preset threshold (e.g., a positional error of 2.5 meters), iteration is stopped to prevent overfitting.
[0040] Step 205, mapping the sum of source strength values within each termite cluster to the biomass of the termite nest, is based on a pre-defined nonlinear power function relationship; the nonlinear power function relationship is represented by B. k =α·(Q k ) β B k For biomass, Q k The sum of source strengths within the cluster is denoted as α, and β are the mapping coefficients and exponents obtained from indoor control experiments.
[0041] In this embodiment, although it primarily serves the physical inversion branch, the same mapping logic is often used as a preferred post-processing step or embedded in the model's target output in the machine learning branch. To determine the quantitative relationship between biomass and eDNA signal intensity, indoor controlled experiments are required. In these experiments, termite nest models with different biomass gradients can be constructed, and the concentration or source strength of the eDNA produced can be measured. Based on the experimental data, a nonlinear power function relationship B is fitted. k =α·(Q k ) β In some preferred embodiments, taking into account the influence of environmental factors on the mapping relationship, the formula can be further extended to include an environmental correction term, for example, M = α·C β +γ·K T ·K pH ·K soilWhere M is the predicted termite nest biomass (e.g., in grams); α is the basic mapping coefficient; C is the representative concentration or source strength; β is the nonlinear mapping exponent; γ is the environmental impact weighting coefficient; K T K pH K soil These are correction factors for temperature, pH, and soil type, respectively. By incorporating these environmental factors, biomass prediction results can maintain greater stability under different seasons and geological conditions.
[0042] Example 3 describes a crucial nonlinear correction step in acquiring soil environmental DNA concentration distribution data. Given the complexity of dam soil environments, raw detection signals are often distorted by both detection link suppression and soil adsorption, and nonlinear superposition effects exist when multiple nests coexist. This example provides a systematic correction method that can improve the accuracy and signal-to-noise ratio of the input data.
[0043] Step 301, the process of obtaining soil environmental DNA concentration distribution data, specifically includes, for example... Figure 7 As shown: The detection signal of the collected samples is converted into apparent concentration; an inhibition index characterizing the degree of biochemical inhibition of the detection link is obtained by adding an internal standard to the collected samples; an adsorption index characterizing the degree of physical adsorption of environmental DNA by the soil medium is obtained by measuring the physicochemical properties of the collected samples; the apparent concentration is compensated and corrected by combining the inhibition index and the adsorption index of the samples to obtain the true concentration. The inhibition index characterizes the degree of biochemical inhibition of the detection link, and the adsorption index characterizes the degree of physical adsorption of environmental DNA by the soil medium; the true concentration is corrected by using a preset interaction residual coefficient to eliminate the nonlinear superposition bias when multiple nests coexist, and the soil environmental DNA concentration distribution data is obtained.
[0044] In this embodiment, the data processing flow begins with the cycle threshold (Ct value) obtained from real-time quantitative PCR (qPCR). Since the Ct value has a linear relationship with the logarithm of DNA concentration, the Ct value is converted to the apparent concentration N using a pre-established standard curve. The standard curve equation is typically expressed as Ct = ab * log10(N), where a is the intercept and b is the slope. Therefore, the apparent copy number N can be calculated using the formula N = 10^2 / 10^2. (Ci-a) / b The calculation yielded the result. Where N is the converted epigenetic eDNA copy number; a is the intercept of the qPCR standard curve; C... i 1 is the cycle threshold measured by real-time quantitative PCR; b is the cycle threshold measured by real-time quantitative PCR. However, the apparent concentration N is often lower than the actual concentration present in the soil because humic acid in the soil inhibits the PCR reaction, and soil particles adsorb DNA, making it difficult to extract.
[0045] To restore the true concentration, this embodiment introduces inhibition index I and adsorption index H for compensation correction. Specifically, inhibition index I can be determined by adding a known concentration of internal standard (such as non-target DNA) to the sample and measuring the offset of its Ct value. If the internal standard Ct value is delayed, it indicates the presence of inhibition. Adsorption index H can be obtained by measuring the physicochemical properties of the soil sample, such as parameters like clay content, organic matter content, or humic acid concentration. The corrected true concentration C i It can be calculated using empirical formulas, such as C. i =N i *exp(γ0+γ1I i +γ2H i ), where C i N represents the corrected true concentration for the i-th sample; i γi represents the apparent copy number (or apparent concentration) of the i-th sample; i is the sample index; γ0 is the baseline coefficient; γ1 and γ2 are the correction weight coefficients for inhibition and adsorption, respectively; Ii i H is the suppression index for the i-th sample (e.g., the offset of the internal standard Ct value); i Let be the adsorption index (e.g., soil clay or organic matter content) for the i-th sample. These coefficients can be pre-calibrated through indoor addition recovery experiments. This step resolves the problem of false negatives or measurement bias caused by differences in soil texture.
[0046] Building upon this, to address the nonlinearity issue in the superposition of multi-nest signals, an interaction residual coefficient R'(p) is further introduced. In a multi-source diffusion field, the actual observed concentration at a certain point is often not equal to the simple linear sum of the concentrations of each source individually. This may be due to competitive adsorption or co-release by organisms. The interaction residual coefficient is defined as R'(p) = C mix (p)-C lin (p) / C lin (p)+ε, where R'(p) is the interaction residual coefficient at sampling point p; C mix For mixed observation concentrations; C lin The concentration is predicted by linear superposition; ε is a very small positive number. In practical processing, the residual coefficient R at the sampling point can be estimated using a pre-defined parameterized model based on the distance between the sampling point and the center of the potential nest, and the medium parameters s (such as moisture content and pH). For example, an exponential decay model can be used: R = ρ0 * exp(-κd) + ρ1 T *s. Where R is the estimated interaction residual coefficient; ρ0 is the fundamental coefficient of residual strength; κ is the rate constant of residual decay with distance; d is the distance between the sampling point and the nearest potential nest center; ρ1 T's' is the transpose of the coefficient vector representing the influence of the environmental medium; 's' is the vector of environmental medium parameters (such as water content, porosity, etc.). This coefficient is used to correct the true concentration, yielding high-confidence soil environmental DNA concentration distribution data, which serves as input for subsequent inversion.
[0047] Example 4 describes a physically constrained multi-nest sparse source inversion method (PI-MSI). This example constructs a physical mechanism-based white-box inversion framework and utilizes compressed sensing theory and sparse optimization algorithms to achieve accurate location and scale analysis of underground hidden termite nests with limited sampling points.
[0048] Step 401, the multi-source inversion analysis adopts a physically constrained sparse source inversion strategy, such as... Figure 5 As shown, the specific steps include: discretizing the target dam area into a set of candidate source points containing multiple grid points, and defining the source strength vector of the candidate source points corresponding to the set of candidate source points; concretizing the transport-response mapping mechanism into an observation matrix, and constructing a linear superposition observation model that describes the relationship between soil environmental DNA concentration distribution data and the source strength vector of the candidate source points.
[0049] In this embodiment, the inversion process is modeled as a linear inverse problem. The three-dimensional space inside the dam is discretized into a set of candidate source points Ω={r} containing M grid points. j |j=1...M}, each r j This represents a possible location of the nest center. Meanwhile, the positions of the N sampling points are denoted as p. i The source strength vector q is an M-dimensional column vector, and its j-th element q j Let represent the release intensity of the j-th candidate source point. The linear superposition observation model can be expressed as y = Aq + e. Here, y is an N-dimensional observation data vector; A is an N-row, M-column observation matrix; q is an M-dimensional source intensity vector (representing the release intensity of all candidate source points); and e is a noise vector. The physical meaning of this model is that the observation value y at all sampling points is the superposition of all potential source points q after propagation through some transport mechanism A.
[0050] Step 402: The element values in the observation matrix are determined by the anisotropic transport kernel function between the sampling point and the candidate source point. The construction process of the anisotropic transport kernel function includes: calculating the coordinate component distance between the sampling point and the candidate source point in different spatial dimensions; and calculating the concentration contribution weight of the candidate source point to the sampling point by combining the preset anisotropic propagation scale parameter and the comprehensive attenuation coefficient. The anisotropic propagation scale parameter is used to describe the differences in medium transport in different directions.
[0051] In this embodiment, the accuracy of constructing the observation matrix A directly determines the accuracy of the inversion. Since dams are typically composed of compacted soil layers, their horizontal permeability and diffusivity are usually greater than their vertical permeability, exhibiting anisotropy. Therefore, an isotropic Gaussian kernel cannot be simply used. The elements A of the observation matrix... i j (i.e., the contribution of the j-th source to the i-th sampling point) is determined by the anisotropic transport kernel function G. i j is calculated. The specific formula is: G ij =exp(-((x i -x j ) 2 / l x 2 +(y i -y j ) 2 / l y 2 +(z i -z j ) 2 / l z 2 ))*exp(-λd ij ), of which G ij The anisotropic transport kernel function value (i.e., matrix element A) of the j-th candidate source point to the i-th sampling point. ij ); (x i ,y i ,z i ) and (x j ,y j ,z j ) represent the coordinates of the sampling point and the candidate source point, respectively; l x ,l y ,l z These are the anisotropic propagation scale parameters in the x, y, and z directions, which are pre-set or determined based on the compaction degree and layering structure of the dam soil; λ is the comprehensive attenuation coefficient; d ij The Euclidean distance between the two is given. By introducing these parameters, the kernel function can accurately characterize the non-uniform attenuation of the eDNA signal in the layered dam medium.
[0052] Step 403, multi-source inversion analysis specifically includes: constructing an optimization objective function, which includes residual terms for fitting the observed data and sparse constraint terms for constraining the distribution characteristics of the source strength vector; the optimization objective function is specifically composed of the following three weighted parts, such as... Figure 6As shown: the data fidelity term is used to characterize the Euclidean distance error between the predicted values of the linear superposition observation model and the soil environmental DNA concentration distribution data; the sparsity regularization term is constructed based on the first norm of the source strength vector and is used to constrain the sparsity of termite nests in spatial distribution; the spatial consistency regularization term is constructed based on the spatial difference or total variation of the source strength vector and is used to suppress isolated noise and maintain the spatial continuity of nest source strength.
[0053] In this embodiment, since the number of candidate source points M is much larger than the number of sampling points N, directly solving the linear equation system y=Aq is an ill-posed problem with infinitely many solutions. To select a unique solution that conforms to physical facts, prior constraints must be introduced. Based on the fact that the number of termite nests inside the dam is very small, the source strength vector q should be sparse (most elements are 0); based on the fact that the nests have a certain volume, the source strength should be spatially continuous, rather than isolated noise points. Therefore, the following optimization objective function is constructed: q*=argmin q≥0 {||Wy-Aq||2 2 +λ1||q||1+λ2||Dq||1}. Where q* is the optimal source strength vector obtained from the solution; argmin represents the variable value that minimizes the objective function; q≥0 indicates a non-negative source strength constraint; W is the weight matrix, which can be set according to the confidence level of each sampling point; y is the observation data vector; A is the observation matrix; ||·||2 2 λ is the square of the L2 norm (data fitting term); λ1 is the sparse weight coefficient; ||q||1 is the L1 norm of the source strength vector (sparse constraint term); λ2 is the smoothing weight coefficient; D is the three-dimensional difference operator, and ||Dq||1 is the total variation (TV) regularization term, which promotes the spatial contiguous distribution of non-zero source strength regions. This optimization problem is a convex optimization problem, which can be solved efficiently by algorithms such as the Fast Iterative Shrinking Threshold Algorithm (FISTA) or the Alternating Direction Multiplier Method (ADMM).
[0054] In this embodiment, the selection of the sparse weighting coefficient λ1 and the smoothing weighting coefficient λ2 requires a trade-off between sparsity and spatial continuity. Empirical evidence suggests that λ1 is typically 0.1 to 1 times the observation noise variance, and λ2 is typically 0.5 to 2 times λ1. In one specific implementation, λ1 = 0.01 and λ2 = 0.015 are set. When λ1 increases, the sparsity of the inversion results is enhanced, which is beneficial for distinguishing adjacent nests, but weak signal sources may be lost; when λ2 increases, spatial continuity is enhanced, which is beneficial for suppressing noise, but may lead to the merging of adjacent nests.
[0055] Step 404, multi-source inversion analysis specifically includes: solving the source strength vector by minimizing the optimization objective function, and extracting termite nest occurrence state parameters from the solved source strength vector; extracting termite nest occurrence state parameters from the solved source strength vector specifically includes: performing local maximum detection on the solved source strength vector to identify the source strength peak point; using the source strength peak point as the center, performing connected component clustering using a preset distance threshold to obtain several independent nest clusters; calculating the geometric center of each nest cluster as the spatial coordinates of the termite nest, and calculating the sum of the source strength values within each nest cluster, mapping it to the biomass scale of the termite nest; mapping the sum of the source strength values within each nest cluster to the biomass scale of the termite nest is based on a pre-calibrated nonlinear power function relationship; the nonlinear power function relationship is expressed as B k =α*(Q k ) β Among them, B k For biomass, Q k The sum of source strengths within the cluster is denoted as α, and β are the mapping coefficients and exponents obtained from indoor control experiments.
[0056] In this embodiment, the obtained optimal source strength vector q* represents the potential source strength of each grid point in space. To convert it into engineering-usable nest parameters, post-processing is required. This involves finding local maxima in q*, i.e., points where the source strength is greater than a preset threshold τ. q Points with source strength greater than that of all points in their neighborhood are identified as potential nest center candidates. Connected component clustering is performed, grouping points whose spatial distance is less than a preset distance threshold τ. d All non-zero source strength points (e.g., 0.5 meters) are grouped into nest clusters. For the k-th nest cluster, the weighted average of the coordinates of all its constituent grid points is calculated as the final spatial coordinates (x, y) of that nest. k ,y k ,z k The total equivalent source strength Q of the nest is obtained by summing the source strength values of all grid points within the cluster. k Using a pre-calibrated nonlinear power function B k =α*(Q k ) β The dimensionless or concentration-dimensional source strength Q k Converted into actual biomass B k (e.g., grams). Here, α and β are empirical constants, determined by fitting and determining them through pre-planting nests of known biomass in indoor Rubik's Cube experiments and inverting their source strength. For example, B can be set... k =0.023*(Q k ) 1.2 As a conversion formula, the final output parameter set P out It includes precise coordinates and biomass with clear physical meaning.
[0057] Suppose termite detection is conducted within a dam area measuring 50 meters long, 10 meters wide, and 5 meters deep. A grid with 2.5-meter intervals is used, resulting in 21 × 5 × 3 = 315 candidate source points. 25 sampling points are then deployed within this area for eDNA detection, yielding concentration values ranging from 50 to 2000 copies / gram.
[0058] After constructing a 315×25 observation matrix A, the FISTA optimization algorithm was executed. λ1=0.01, λ2=0.015, the maximum number of iterations was 500, and the convergence threshold was 1e-6. Convergence was achieved after 127 iterations.
[0059] Local maxima detection is performed on the obtained source strength vector, with a threshold τ. q =0.1 identified three peak points with coordinates (12.5, 5.0, 2.0), (30.0, 7.5, 3.0), and (42.5, 2.5, 1.5).
[0060] With distance threshold τ d Clustering of connected components at a distance of 0.5 meters yielded three independent nest clusters. The geometric center of each cluster was calculated as the nest coordinates, and the sums of the source strengths within the clusters were Q1=0.85, Q2=1.23, and Q3=0.56, respectively.
[0061] Substitute into mapping formula B k =0.023×(Q k ) 1.2 The calculated biomasses were B1=19.8g, B2=29.7g, and B3=12.1g.
[0062] The final predicted results for the three nests are as follows: Nest 1 is located at (12.5, 5.0, 2.0) meters with a biomass of approximately 20 grams; Nest 2 is located at (30.0, 7.5, 3.0) meters with a biomass of approximately 30 grams; and Nest 3 is located at (42.5, 2.5, 1.5) meters with a biomass of approximately 12 grams.
[0063] Example 5 describes the spatiotemporal decay-diffusion coupling model of eDNA and its key parameter calibration method. This example theoretically corrects the physical defects of the existing steady-state Gaussian model, introduces a reaction-diffusion mechanism, and provides a standard experimental procedure for obtaining model parameters (such as the effective diffusion coefficient and adsorption hysteresis factor), ensuring the full disclosure and reproducibility of the technical solution.
[0064] Step 501: The transport-response mapping mechanism is constructed based on the spatiotemporal decay-diffusion coupling model. The spatiotemporal decay-diffusion coupling model is described by the reaction-diffusion equation. It is characterized by including a time decay term for characterizing the dynamic degradation of environmental DNA over time, and an adsorption hysteresis factor for characterizing the blocking effect of soil medium, which describes the anisotropic transport process of termite nest source terms under non-steady-state conditions.
[0065] In this embodiment, a spatiotemporal coupling model based on the reaction-diffusion equation is used to more realistically describe the behavior of eDNA in soil. Unlike the Gaussian model that assumes that eDNA instantaneously reaches a steady-state distribution, this model explicitly considers the time dimension and the adsorption hindrance of the medium. Its governing equation is in the form: ΨC / Ψt=D eff ▽ 2 C-λ total C+S(t). Where ΨC / Ψt is the rate of change of concentration over time, and Ψ is the partial derivative; D eff The effective diffusion coefficient; ▽ 2 λ is the Laplace operator, representing spatial diffusion; total The time decay term represents the degradation and disappearance of eDNA; C is the environmental DNA concentration; and S(t) is the source term function. Furthermore, considering that the adsorption of DNA molecules by soil particles slows their migration, an adsorption lag factor R is introduced into the model. f The modified apparent diffusion equation can be expressed as: R f *(ΨC / Ψt)=D eff ▽ 2 C-λ total C+S(t). Where R f The adsorption hysteresis factor (dimensionless); the wavefront propagation speed of eDNA concentration is slower than the diffusion speed in pure water by R. f The particular solution (steady-state solution) of this model can be expressed as C under steady-state conditions (i.e., time t approaches infinity). ss (r)=(S0 / 4πD eff *r)exp(-r*sqrt(λ total / D eff Among them, C ss (r) represents the steady-state concentration at a distance r from the source point; S0 is the constant source strength; π is pi; D eff λ is the effective diffusion coefficient; r is the radial distance from the source point; λ total Let be the total decay rate constant; exp be the natural exponential function. This formula reveals that the concentration decays exponentially with distance r rather than Gaussianly, which is more consistent with the actual situation in porous media.
[0066] Based on the steady-state solution of the spatiotemporal decay-diffusion coupling model, the effective detection radius of environmental DNA under current soil conditions is calculated. In this embodiment, the effective detection radius refers to the maximum detection distance corresponding to the time when the concentration of environmental DNA released from the termite nest decreases to the detection limit after diffusion and decay in the soil. This parameter determines the maximum allowable spacing between sampling points.
[0067] The effective detection radius is calculated based on the steady-state solution formula. By setting the concentration term in the steady-state solution formula as the detection limit concentration, the implicit equation satisfied by the effective detection radius can be obtained. The specific form of this implicit equation is:
[0068] r eff =L×ln[S0 / (4×π×D eff ×C LOD ×r eff )];
[0069] Where, r eff The effective detection radius is in meters (m); L is the characteristic attenuation length in meters (m); ln is the natural logarithm function; S0 is the environmental DNA release source intensity of the termite nest in copies / h; π is pi, with a value of 3.14159; D eff The effective diffusion coefficient is expressed in m. 2 / h;C LOD The detection limit concentration for real-time quantitative PCR is given in copies / g.
[0070] The characteristic attenuation length L is calculated by the following formula:
[0071] L=(D eff / λ total ) 0.5 ;
[0072] Where, λ total The total decay rate constant is expressed in h. (-1) .
[0073] Because r eff When both sides of an equation appear simultaneously, the equation is a transcendental equation and has no analytical solution, requiring a numerical iterative method to solve it. This embodiment uses the Newton-Raphson iterative method for solution. The objective function is defined as follows:
[0074] f(r) = rL × ln[S0 / (4 × π × D)] eff ×C LOD ×r)];
[0075] Where f(r) is the objective function value; r is the detection radius variable, in meters.
[0076] Calculate the derivative of the objective function with respect to r:
[0077] f'(r) = 1 + L / r;
[0078] Here, f'(r) is the first derivative of the objective function f(r) with respect to r.
[0079] The iterative formula is:
[0080] r (n+1) =r n -f(r n ) / f'(r n );
[0081] Where n is the iteration number index, which takes the value of a non-negative integer such as 0, 1, 2, 3, etc.; r n r is the estimated detection radius for the nth iteration, in meters. (n+1) f(r) is the estimated detection radius for the (n+1)th iteration, in meters; n ) for r n The function value obtained by substituting it into the objective function; f'(r n ) for r n The derivative value is obtained by substituting into the derivative function.
[0082] The specific steps of the iterative process are as follows: First, set an initial value r0 based on engineering experience, typically between 2m and 5m. Second, substitute the r0 value into the iterative formula to calculate r1. Third, determine the convergence condition.
[0083] |r (n+1) -r n | / r n <τ;
[0084] Where |·| represents the absolute value operation; τ is the preset convergence threshold, usually set to 0.001. If the above convergence condition is met, the iteration stops, and r is output. (n+1) This serves as the effective detection radius; if it is not satisfied, return to the second step and continue iterating.
[0085] As a specific calculation example, for typical dam soil conditions, the following parameters are set: soil temperature T = 25℃, pH = 7, and volumetric water content θ = 20%. λ is calculated based on the multi-factor model. total =0.03h (-1) Effective diffusion coefficient D eff =5×10 (-10) m 2 / s=1.8×10 (-6) m 2 / h. Termite nest source strength S0=10 6 copies / h corresponds to a medium-sized nest with approximately 100g of biomass. Detection limit concentration C LOD =100 copies / g.
[0086] Substituting the above parameters, we first calculate the characteristic attenuation length:
[0087] L=(1.8×10 (-6) / 0.03) 0.5 =0.0077m;
[0088] With an initial value of r0 = 2m, the system converges after 5 iterations, yielding the effective detection radius r. eff ≈2.1m.
[0089] Based on the above calculations, the design of the field sampling grid density should ensure that the spacing between sampling points does not exceed twice the effective detection radius. Therefore, for the typical conditions described above, it is recommended that the sampling spacing not exceed 4m. If the soil temperature is high or the moisture content is high, leading to an increased attenuation rate, the effective detection radius will decrease accordingly, requiring a denser sampling grid.
[0090] Step 502: The spatiotemporal decay-diffusion coupling model includes a total decay rate constant, which is constructed as a function of multiple environmental factors. These environmental factors include soil temperature, pH, and water content. Soil temperature is associated with the microbial degradation rate through the Arrhenius equation, and pH is associated with the hydrolysis rate through a nonlinear sensitivity coefficient.
[0091] In this embodiment, the total decay rate constant λ total It is not a fixed value, but a dynamic parameter determined by multiple mechanisms such as microbial degradation, chemical hydrolysis, and soil adsorption. To adapt to the environmental differences in different areas of the dam, it is modeled as a function of multiple environmental factors: λ total =λ bio +λ chem +λ ads Among them, the microbial degradation rate λ bio It is the dominant factor, influenced by temperature, and described by the Arrhenius equation: λ total =λ bio_ref *exp(E a / R*(1 / T ref -1 / T))*f pH *f θ Wherein, λ total λ is the total decay rate constant; bio_ref E represents the microbial degradation rate constant under reference conditions. a is the activation energy; R is the gas constant; T ref Thermodynamic temperature; T is the actual soil thermodynamic temperature; f pH and f θ These are correction factors for pH and moisture content, respectively. For example, f pHIt can be constructed as a Gaussian function of pH, reflecting the optimal activity of microorganisms at a specific pH value. Through multi-factor parameterized modeling, the model can automatically adapt to the effects of seasonal changes (temperature changes) and rainfall events (water content changes).
[0092] Step 503: The effective diffusion coefficient and adsorption hysteresis factor in the spatiotemporal decay-diffusion coupling model are determined through a two-stage experiment: First stage: The partition coefficient of environmental DNA in soil is determined by batch adsorption experiment, and the adsorption hysteresis factor is calculated accordingly; Second stage: The apparent diffusion coefficient of environmental DNA is determined by soil column diffusion experiment, and the effective diffusion coefficient is obtained by combining the adsorption hysteresis factor.
[0093] In this embodiment, a rigorous two-stage indoor experiment was designed to obtain the key physical parameters in the model. In the first stage, the batch adsorption experiment, typical soil samples from the target dam were taken, air-dried, sieved, and mixed with eDNA solutions of different concentrations. Adsorption equilibrium was reached in a constant-temperature shaker. After centrifugation, the concentration of the supernatant was measured, adsorption isotherms were plotted, and the linear partition coefficient K was calculated. d (Unit: mL / g). The adsorption hysteresis factor R is calculated based on this. f =1+(ρ b / θ)*K d Among them, R f ρ is the adsorption hysteresis factor; b θ is the soil bulk density; K is the volumetric water content; d The apparent diffusion coefficient D is the linear distribution coefficient. In the second-stage soil column diffusion experiment, a cylindrical soil column device was constructed, with eDNA tracer injected into one end. After standing for different times, the soil column was sliced and sampled to obtain the concentration distribution curve along the axial direction. The experimental data were fitted with the analytical solution of the one-dimensional transient diffusion equation to obtain the apparent diffusion coefficient D. app Finally, using formula D... eff =D app *R f The true effective diffusion coefficient is obtained by conversion. Where, D eff D represents the effective diffusion coefficient (true porosity diffusion capacity). app R is the apparent diffusion coefficient (obtained from soil column experiments); f This is the adsorption hysteresis factor. For example, for typical clay, R is measured... f The value could be as high as 100, indicating extremely slow apparent diffusion. This conversion is necessary to accurately represent the true porosity diffusion capacity. This experimental procedure ensures that the model parameters are derived from objective measurements, rather than subjective assumptions.
[0094] Step 504: Based on the time difference between the sampling time and the termite activity time, perform time correction on the environmental DNA concentration obtained from the on-site detection to eliminate systematic bias caused by inconsistent sampling times.
[0095] In this embodiment, since on-site sampling is usually not possible simultaneously with termite activity, a termite activity detector is used to record the peak time to determine the moment of termite activity. There is a time interval between the sampling time and the moment the termites last engaged in activity. During this time interval, the environmental DNA released by the termites continues to degrade and decay, resulting in a detected concentration lower than the actual release concentration during termite activity. If this time difference is not corrected, samples collected at different times will not be comparable, affecting the accuracy of subsequent nest location.
[0096] Time correction is based on the first-order decay kinetics of environmental DNA. Let the concentration detected at the sampling time be C. obs If the time difference between the sampling time and the termite activity time is Δt, then the corrected concentration C corr Calculated by the following formula:
[0097] C corr =C obs ×exp(λ total ×Δt);
[0098] Among them, C corr This represents the time-corrected concentration of environmental DNA, expressed in copies / g; C obs λ represents the concentration of environmental DNA obtained from on-site sampling and testing, expressed in copies / g; exp is an exponential function with base e; λ total The total decay rate constant is expressed in h. (-1) Δt is the time difference between the sampling time and the time of termite activity, in hours.
[0099] The physical meaning of the above formula is: to compensate the detected concentration value back to the time of termite activity, that is, to uniformly calculate the concentration data of all sampling points to the same reference time in the time dimension. In other words, the concentrations of samples collected at different times are comparable, and an accurate concentration gradient field can be constructed in space.
[0100] The method for determining the time difference Δt is as follows. The timing of termite activity can be estimated based on the termite's activity patterns. Termites are mainly active at night, with peak activity typically occurring 2 to 4 hours after sunset. Therefore, if the sampling time is 9:00 AM the following morning, Δt can be estimated to be 12 to 14 hours. In practical applications, a more accurate time difference estimate can also be obtained by deploying automatic monitoring equipment to record the timestamps of termite activity signs.
[0101] As a concrete calculation example, suppose a soil sample is collected at a sampling point at 10:00 AM, and the concentration C is measured. obs=500 copies / g. Based on termite activity patterns, the estimated time of termite activity is 10 PM the previous night, therefore the time difference Δt = 12 hours. λ under soil conditions total =0.03h (-1) Substitute into the time correction formula:
[0102] C corr =500×exp(0.03×12)=500×exp(0.36)=500×1.433=716.5copies / g;
[0103] The corrected concentration was about 43% higher than the observed concentration, and this correction range affected the accuracy of nest location.
[0104] The time correction method addresses the fact that environmental DNA detection methods typically only focus on the presence or relative level of concentration, without considering the impact of time on the absolute value of concentration. In contrast, this invention establishes a spatiotemporal coupling model, incorporating the time dimension into the concentration analysis framework, thereby achieving standardized processing of concentration data at different sampling time points and improving the accuracy and reliability of nest location.
[0105] Example 6 describes an uncertainty-driven adaptive sampling closed-loop method (UDAS). Addressing the engineering constraints of high drilling costs and low permissible damage levels in dam drilling, this example changes the blindness of traditional one-time fixed grid sampling by introducing a dynamic loop mechanism of prediction-evaluation-decision. It utilizes the uncertainty of the inversion results to guide the next round of drilling, achieving precise positioning with the minimum number of sampling points.
[0106] Step 601, the method further includes an adaptive sampling closed-loop step, specifically including: based on the current multi-source inversion analysis results, evaluating the variance distribution of the inversion results through a statistical resampling method to generate a spatial uncertainty field; determining whether the spatial uncertainty field meets a preset stopping criterion; if not, generating a recommended sampling point set based on the spatial uncertainty field and obtaining incremental environmental DNA concentration data at the recommended sampling point set; updating the incremental environmental DNA concentration data to the soil environmental DNA concentration distribution data, and returning to execute multi-source inversion analysis until the stopping criterion is met.
[0107] In this embodiment, the adaptive sampling closure lies in calculating the reliability of the inversion results. Specifically, the Bootstrap statistical resampling method is used to evaluate the stability of the inversion results. Assuming the existing observation dataset is D, D is randomly resampled T times with replacement, generating T sets of resampled datasets. For each dataset, sparse source inversion is performed to obtain T sets of source strength vector solutions {q*(t)|t=1...T}. For each candidate source point location r... j Calculate the statistical variance σ of its source strength in T inversions. j2 =(1 / T-1*Σ t=1 T (q j (t)-μ j ) 2 Among them, σ j 2 q represents the source strength variance (i.e., uncertainty measure) of the j-th candidate source location; T represents the number of Bootstrap resampling attempts; j (t) represents the source strength value at the j-th point obtained from the t-th resampling inversion; μ j Let be the mean source strength of the j-th point in T inversions. The spatial uncertainty field U(r) is composed of the variances of all locations, i.e., U(r) j )=σ j 2 Intuitively, the larger the U(r) value, the weaker the constraint of the current data on that region, and the less reliable the inversion results. The stopping criterion can be defined as: when the maximum value of the spatial uncertainty field max(U(r)) is less than a preset uncertainty threshold τ... U Alternatively, the nest position offset obtained from the two rounds of iterative inversion is less than the preset stability threshold τ. r If the criteria are met, sampling stops and the final result is output. If not, the next round of incremental sampling begins, new data is collected and merged into the original dataset, and the entire inversion process is repeated.
[0108] Step 602 generates a recommended set of sampling points based on the spatial uncertainty field, specifically including: constructing the observation kernel vector of the candidate sampling points for the potential source terms; calculating the information gain score of each candidate sampling point based on the observation kernel vector and the spatial uncertainty field, the information gain score characterizing the expected contribution of the sampling point to reducing the inversion uncertainty; selecting several candidate sampling points with the highest information gain scores, and forming a recommended set of sampling points after applying the minimum distance constraint.
[0109] In this embodiment, the selection of the next round of sampling points is crucial. This embodiment employs a prediction variance maximization strategy to calculate the information gain score of candidate points. The candidate sampling point set Π = {p c |c=1...C}, these points must satisfy the passage and avoidance conditions of the dam project. For each candidate point p c Construct its observation kernel vector g(p) c )=[G(p c ,r1),G(p c ,r2),...,G(p c ,r M )] T Among them, g(p) c ) represents candidate sampling point p cThe observed kernel vector; G(p c ,r M ) is a candidate source point r M For candidate sampling point p c The transport kernel function value; T denotes the vector transpose. G is the anisotropic transport kernel function, which describes the physical connection between the sampling point and all potential source points. Construct the source strength uncertainty covariance matrix Σ. q For simplicity, it can be represented as a diagonal matrix with U(r) as its diagonal elements. Calculate the information gain score Score(p). c )=g(p c ) T *Σ q *g(p c Among them, Score(p) c ) represents candidate sampling point p c Information gain score; g(p) c ) represents candidate sampling point p c The observed kernel vector; Σ q The source strength uncertainty covariance matrix (usually composed of variance σ) j 2 (Composition). The physical meaning of this formula is: if a candidate point has a strong physical connection (large G) with a source point that currently has high uncertainty (large U(r)), then sampling at this point will yield the maximum information gain. After calculating the scores of all candidate points, sort them from high to low, and select the top L points (e.g., 3). To avoid overly dense sampling points and waste, a minimum distance constraint is applied: the distance between a newly selected point and existing sampling points and other newly selected points must be greater than a preset threshold τ. p (For example, 2 meters). The selected point set is the recommended sampling point set, which guides on-site personnel to conduct precise supplementary surveys.
[0110] Example 7 describes a method for inferring the time window of termite activity. This example fully utilizes the decay rate constant, a byproduct of the spatiotemporal coupled physical model, to extend the detection dimension from spatial positioning to temporal backtracking. It can distinguish between active nests and abandoned nests that still retain eDNA, providing a deeper basis for prevention and control decisions.
[0111] Step 701, the method also includes a termite activity time window inference step: based on the time decay term or total decay rate constant in the spatiotemporal decay-diffusion coupling model, calculate the half-life of environmental DNA under the current environmental conditions; combine the concentration peak and detection limit in the soil environmental DNA concentration distribution data, use the half-life to back-calculate the time interval when the termite source term stops releasing environmental DNA, obtain the termite activity time window, and determine the active state of the nest.
[0112] In this embodiment, the degradation of eDNA in soil follows first-order kinetics. This is based on the multi-factor total decay rate constant λ. total It can calculate the half-life t of eDNA under the current soil environment (specific temperature, humidity, pH). 1 / 2 =ln2 / λ total Among them, t 1 / 2 λ represents the half-life of environmental DNA under current environmental conditions; ln2 is the natural logarithm of 2 (approximately 0.693); λ total This represents the total decay rate constant. For example, under conditions of 25 degrees Celsius and pH=7, λ was measured. total =0.0268h -1 The half-life is approximately 25.9 hours. When inversion reveals a high concentration of eDNA signal peak C at a certain location... measured At the same time, it can be combined with the lowest detection limit C of qPCR. LOD (e.g., 100 copies / gram) to infer the time when the source termite might cease activity. Assuming the time between the termite ceasing activity (no longer releasing eDNA) and the sampling time is Δt, then according to the decay formula C... measured =C0*exp(-λ total *Δt) can be used to derive the maximum possible time interval t. activity_max =(1 / λ total )*ln(C measured / C LOD ), where t activity_max λ represents the maximum possible time interval between the termites ceasing activity and the sampling time. total C is the total decay rate constant; measured C represents the peak concentration of environmental DNA measured at the sampling point. LOD This is the lowest detection limit concentration for the detection method. If the calculated Δt is very short (e.g., less than a few times the half-life), and C... measured If the concentration is significantly higher than the background value, the nest is considered to be active. Conversely, if the measured concentration is low and the calculated Δt is long, it may be a remnant old nest. This termite activity time window information is of great reference value for assessing the timeliness of control measures (such as grouting poisoning).
[0113] Example 8: According to one aspect of this application, the method includes: constructing a quantitative model of termite eDNA concentration and biomass through indoor experiments.
[0114] Step S1.1: Constructing the experimental model: A 90cm×90cm×90cm model box was constructed using plexiglass and filled with silty clay taken from a certain dike project. The soil moisture content was controlled at 20%, the temperature at 25℃, and the humidity at 75% to simulate the on-site environment of the dike. A termite nest (black-winged subterranean termite) was placed in the center of the model box with initial biomass of 10g, 50g, 100g, 200g, 500g, and 1000g, with three replicates for each biomass gradient. Subsequently, 2-5 nests were placed in the model box, with a spacing of 10-30cm between nests and an initial biomass of 50-500g, to construct a multi-nest experimental scenario.
[0115] Step S1.2: Establishing a standard curve: Termite eDNA standards were prepared with a concentration gradient of 10⁻¹⁰ copies / μL. The Ct values of each concentration standard were detected using qPCR technology, and the data are shown in Table 1.
[0116] Table 1. Relationship between concentration gradient of termite eDNA standard and corresponding Ct value
[0117]
[0118] Linear regression was performed with lgN as the abscissa and Ct as the ordinate, and the standard curve equation was obtained as: Ct=-3.3·lgN+41.6 (R²=0.998), where a=-3.3 and b=41.6.
[0119] Step S1.3: eDNA Concentration Determination: 81 sampling points were set up at the 9×9×9 grid nodes of the model box to collect soil samples from a depth of 0-10 cm. After extracting eDNA, the Ct value was measured, and the eDNA concentration was calculated using the standard curve equation. For example, if the Ct value of a certain sampling point is 26.3, substituting it into the formula, with k=1, the calculated eDNA concentration is C=10. (26.3-41.6) / (-3.3) =10 (4.64) =4.37×10 4 copies / μL.
[0120] Step S1.5: Constructing a quantitative model: Calculate the representative concentration C of eDNA in the model chamber under each biomass gradient. Using C as the independent variable and biomass M as the dependent variable, perform nonlinear fitting to obtain the single-cell eDNA-biomass quantitative model: M = 0.018·C 1.15 (R²=0.98).
[0121] Step S1.6: Multi-House Interaction Law: In the two-house experiment, based on the concentration change rate correction logic, the nonlinear superposition deviation when multiple houses coexist is calculated:
[0122] When the nest spacing is 10cm, the concentration change rate R = -12.3% (inhibitory effect), i.e., the mixed concentration C mix =C lin •(1+R)=C lin ×0.877;
[0123] When the spacing is 30cm, R=2.1% (no significant interaction), i.e., C mix =C lin ×1.021;
[0124] The validation results show that the eDNA interaction in multi-nest coexistence is closely related to the nest spacing. The corrected mixing concentration can be used as input for subsequent machine learning models or sparse source inversion to ensure consistent data processing logic.
[0125] According to another aspect of this application, the method further includes: on-site drilling sampling and prediction of termite infestation scale.
[0126] Step S8: Study area. A homogeneous earth dam (1000m in length, 8m in height, and 5m in top width) was selected as the study object. The dam was divided into 40 sampling units (25m×25m).
[0127] Step S9: Drilling and sampling. Within each sampling unit, 25 drilling sampling points are arranged in a 5m × 5m grid. Small drilling equipment is used for drilling, and soil samples are collected at 1m depth intervals from 0 to 5m depth, for a total of 40 × 25 × 5 = 5000 soil samples. Figure 4 As shown.
[0128] Step S10: eDNA detection. eDNA was extracted from the samples and the Ct value was measured. The eDNA concentration was calculated using a standard curve. High eDNA concentrations (1.2 × 10⁻⁶) were found at three sampling points on the lower part of the water-facing slope (1-2 m depth). 5 -3.5×10 5 (copies / μL).
[0129] Step S11: Model prediction. Input the coordinates of the high-concentration sampling points (e.g., X=250m, Y=5m, Z=1.5m) and eDNA concentration data into the random forest model. The prediction results show that there is one termite nest in the corresponding area, located at (251m, 6m, 1.8m), with a biomass of approximately 120g and a volume of approximately 0.05m³.
[0130] Step S13: On-site verification. Manual excavation was carried out at the predicted location, and a complete termite nest was found. The actual location was (251.2m, 5.8m, 1.9m), with a biomass of 132g and a volume of 0.056m³. The prediction errors were 1.7% (location), 9.1% (biomass), and 10.7% (volume), all less than 15%, indicating that the prediction results are reliable.
[0131] To address the difficulty in decoupling multi-source signal aliasing, this scheme constructs a physically constrained sparse source inversion strategy (PI-MSI). By establishing a linear superposition observation model and introducing L1 norm sparsity regularization and total variational space consistency constraints, the mixed observation signals are mathematically forced to separate into sparsely distributed independent source terms. This overcomes the limitation of traditional machine learning black-box regression in failing to clearly analyze the multi-nest superposition effect, and achieves accurate localization and scale calculation of multi-source targets.
[0132] To address the issue of missing physical transport mechanisms (distortion of the steady-state Gaussian model), this scheme abandons the idealized steady-state assumption and establishes a spatiotemporal coupling mapping mechanism based on the reaction-diffusion equation. By constructing an anisotropic transport kernel function and introducing an adsorption hysteresis factor, the actual transport behavior of eDNA in heterogeneous and anisotropic media such as dam compacted soil layers is accurately characterized, thus improving the physical fidelity of the inversion model.
[0133] To address the challenge of distinguishing between nearby old traces and distant new traces, this solution introduces a dynamic time decay term and a parameterized model of multiple environmental factors (temperature, pH, and moisture content) to calculate the degradation pattern of eDNA signals over time. By combining half-life calculation and activity time window inference techniques, the temporal dimension of nest activity status is successfully determined, effectively eliminating interference from residual signals in abandoned nests. Furthermore, the uncertainty-driven adaptive sampling closed loop further solves the problem of incomplete information acquisition under static sampling.
Claims
1. A method for decoupling and predicting multi-source signals in termites based on eDNA and machine learning, characterized in that, include: Soil samples were collected from the target dam area and quantitatively analyzed to obtain soil environmental DNA concentration distribution data; A transport-response mapping mechanism was constructed to characterize the spatiotemporal correlation between termite nest source terms and soil environmental DNA concentration distribution data. Based on soil environmental DNA concentration distribution data, a multi-source inversion analysis was performed using a transport-response mapping mechanism to calculate the termite nest occurrence state parameters, which include the nest spatial coordinates and biomass scale. Output the termite nest configuration status parameters; The multi-source inversion analysis employs a physically constrained sparse source inversion strategy, specifically including: The target dam area is discretized into a set of candidate source points containing multiple grid points, and a source strength vector of the candidate source points corresponding to the set of candidate source points is defined. The transport-response mapping mechanism is concretized into an observation matrix, and a linear superposition observation model is constructed to describe the relationship between soil environmental DNA concentration distribution data and candidate source intensity vectors. Construct an optimization objective function, which includes a residual term for fitting the observed data and a sparse constraint term for constraining the distribution characteristics of the source strength vector; The source strength vector is solved by minimizing the optimization objective function, and the termite nest occurrence state parameters are extracted from the solved source strength vector. The process of obtaining soil environmental DNA concentration distribution data specifically includes: The detection signal of the collected sample is converted into apparent concentration; By adding internal standards to the collected samples, inhibition indicators characterizing the degree of biochemical inhibition of the detection chain were obtained; by measuring the physicochemical properties of the collected soil samples, adsorption indicators characterizing the degree of physical adsorption of environmental DNA by the soil medium were obtained. By combining the inhibition and adsorption indices of the samples, the apparent concentration was compensated and corrected to obtain the true concentration. The inhibition index characterizes the degree of biochemical inhibition of the detection link, and the adsorption index characterizes the degree of physical adsorption of environmental DNA by the soil medium. The true concentration is corrected by using the preset interaction residual coefficient to eliminate the nonlinear superposition bias when multiple nests coexist, and the soil environmental DNA concentration distribution data is obtained. It also includes an adaptive sampling closed-loop step, which is as follows: based on the current multi-source inversion analysis results, the variance distribution of the inversion results is evaluated by statistical resampling method to generate a spatial uncertainty field; it is determined whether the spatial uncertainty field meets the preset stopping criterion; if not, a recommended sampling point set is generated based on the spatial uncertainty field, and the incremental environmental DNA concentration data at the recommended sampling point set is obtained; the incremental environmental DNA concentration data is updated to the soil environmental DNA concentration distribution data, and the multi-source inversion analysis is returned to be executed until the stopping criterion is met.
2. The method according to claim 1, characterized in that, include: The multi-source inversion analysis uses a pre-trained random forest multi-output regression model; Based on soil environmental DNA concentration distribution data, a multi-source inversion analysis was performed using a transport-response mapping mechanism, specifically including: Extract the spatial coordinates of sampling points and the corresponding concentration values from soil environmental DNA concentration distribution data, and construct a joint input feature vector; The joint input feature vector is input into a pre-trained random forest multi-output regression model, which outputs termite nest occurrence state parameters, including the nest location coordinates, biomass, and nest volume around each sampling point.
3. The method according to claim 2, characterized in that, The pre-trained random forest multi-output regression model is obtained by training on a virtual training dataset. The process of constructing the virtual training dataset includes: Virtual termite nest scenes with varying numbers and locations are generated based on a pre-defined diffusion model. The mixed environment DNA concentration at virtual sampling points in the virtual termite nest scene is calculated. The mixed environment DNA concentration is obtained by linearly superimposing the independent concentration contributions of each virtual termite nest and applying a concentration change rate correction. Using the mixed environment DNA concentration and virtual sampling point coordinates as input features, and the corresponding virtual termite nest scene parameters as output labels, a virtual training dataset is generated.
4. The method according to claim 3, characterized in that, The specific calculation process for the concentration change rate correction includes: Based on a pre-set diffusion model, the concentration values of a single virtual termite nest when it exists independently and the concentration values of multiple nests when they coexist are simulated and calculated respectively. The concentration change rate index is calculated based on the relative difference between the concentration values of multiple nests and the concentration values of a single nest. By using the concentration change rate index to weight and correct the linear superposition results of each virtual termite nest, the DNA concentration in the mixed environment is obtained, which characterizes the nonlinear interaction between multiple nests.
5. The method according to claim 1, characterized in that, The objective function is specifically composed of the following three weighted parts: The data fidelity term is used to characterize the Euclidean distance error between the predicted values of the linear superposition observation model and the soil environmental DNA concentration distribution data. The sparsity regularization term, constructed based on the first norm of the source strength vector, is used to constrain the sparsity of the spatial distribution of termite nests. The spatial consistency regularization term, constructed based on the spatial difference or total variation of the source strength vector, is used to suppress isolated noise and maintain the spatial continuity of the nest source strength.
6. The method according to claim 1, characterized in that, The numerical values of the elements in the observation matrix are determined by the anisotropic transport kernel function between the sampling points and the candidate source points; the construction process of the anisotropic transport kernel function includes: Calculate the distance between the sampling point and the candidate source point in different spatial dimensions using their coordinate components. By combining the preset anisotropic propagation scale parameters and the comprehensive attenuation coefficient, the concentration contribution weight of the candidate source point to the sampling point is calculated, where the anisotropic propagation scale parameters are used to describe the differences in medium transport in different directions.
7. The method according to claim 1, characterized in that, The termite nest occurrence state parameters are extracted from the solved source strength vector, specifically including: Local maxima detection is performed on the obtained source intensity vector to identify the source intensity peak point; Centered on the peak point of the source strength, connected component clustering is performed using a preset distance threshold to obtain several independent nest clusters; The geometric center of each nest cluster is calculated as the spatial coordinates of the termite nest, and the sum of the source strength values within each nest cluster is calculated and mapped to the biomass scale of the termite nest.
8. The method according to claim 1, characterized in that, include: The transport-response mapping mechanism is constructed based on a spatiotemporal decay-diffusion coupling model; The spatiotemporal decay-diffusion coupling model is described by the reaction-diffusion equation, which includes a time decay term to characterize the dynamic degradation of environmental DNA over time, and an adsorption hysteresis factor to characterize the blocking effect of soil media, describing the anisotropic transport process of termite nest source terms under unsteady conditions.
Citation Information
Patent Citations
Acidity and pressure dual-mode termite nest positioning method and equipment
CN121186881A
Enhancement and release seedling resource class evaluation method based on environmental DNA polymerization analysis
CN121331216A