Hydrodynamic tracing method that couples microbial succession with multidimensional evidence

By combining microbial community data, environmental DNA data, and chemical tracer data, a multidimensional evidence fusion model was constructed, which solved the problem of difficulty in fusing multi-source evidence in existing technologies, and achieved accurate estimation of water mass duration and improved tracer reliability.

CN121415857BActive Publication Date: 2026-03-06TAIHU BASIN ADMINISTRATION BUREAU OF HYDROLOGY (INFORMATION CENT) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511993139.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-03-06
Estimated Expiration
2045-12-26

AI Technical Summary

Technical Problem

Existing hydrodynamic tracing technologies struggle to effectively integrate multiple pieces of evidence in complex hydrological environments, leading to misjudgments of water mass transport duration and impact status, and lacking robustness and accuracy.

Method used

By acquiring microbial community data, environmental DNA data, and chemical tracer data, a community succession index and an eDNA decay index are constructed. Combined with the travel time field of the hydrodynamic model, Bayesian fusion is performed to generate a posterior probability grid of incoming water and determine the area affected by incoming water.

Benefits of technology

It improves the accuracy of quantitative estimation of water mass duration and the reliability of tracing, solves the problem of multi-source evidence fusion, and provides more robust hydrodynamic tracing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415857B_ABST
    Figure CN121415857B_ABST
Patent Text Reader

Abstract

This invention discloses a hydrodynamic tracing method coupling microbial succession and multidimensional evidence, comprising: acquiring microbial community data, environmental DNA data, and chemical tracer data of the water body to be tested; constructing a community succession index based on the microbial community data and an eDNA decay index based on the environmental DNA data, and jointly inverting them to obtain a biological age raster characterizing the water mass duration; comparing the biological age raster with the travel time field of the hydrodynamic model to construct an age consistency likelihood; and, within a Bayesian framework, adaptively weighting and fusing the age consistency likelihood, the chemical likelihood constructed based on the chemical tracer data, and the source contribution likelihood constructed based on the microbial community data to obtain the final posterior probability raster of the incoming water, and determining the incoming water influence zone accordingly. Through the above method, this invention provides a quantitative estimate of the water mass duration, improving the accuracy and reliability of tracing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to hydrodynamic tracing methods, and more particularly to a hydrodynamic tracing method that couples microbial succession with multidimensional evidence. Background Technology

[0002] Inter-basin water transfer is a key engineering measure to ensure regional water supply security and improve the aquatic ecological environment. During the water transfer process, accurately and dynamically tracking the transport path, mixing range, and residence time of incoming water is of significant scientific importance and practical value for scientifically assessing the macro-hydraulic and ecological benefits of water transfer projects, optimizing scheduling plans, and providing early warnings of potential water quality risks. Therefore, developing a high-precision, highly robust hydrodynamic tracing technology system is an important technical support for modern refined water resource management.

[0003] Currently, the mainstream techniques for hydrodynamic tracing mainly include chemical tracing and biological tracing. Chemical tracing typically utilizes naturally occurring conserved ions (such as chloride ions, Cl-) between source and sink water bodies. - Concentration differences are used to infer the physical mixing and diffusion processes of water masses by monitoring the spatiotemporal changes in ion concentration in the receiving water body. Biotracing methods, particularly microbial source apportionment techniques based on high-throughput sequencing, such as the SourceTracker model, quantitatively estimate the contribution of external water bodies to the target area by comparing the similarities and differences in microbial community structures between different water bodies. Furthermore, environmental DNA (eDNA) technology is also used to track the distribution of certain indicator species as an auxiliary tracing method.

[0004] However, existing tracing technologies face technical bottlenecks when dealing with dynamic processes in complex hydrological environments, including limited information dimensions and difficulties in evidence fusion. Specifically, existing biological tracing methods ignore the inherent temporal dynamics of microbial signals. When a foreign water body enters a new environment, the microbial community it carries undergoes rapid succession, causing its source characteristic signals to decay over time. Traditional source apportionment models can only provide static contribution ratios and cannot distinguish between newly arrived water masses and those that have long been stagnant and undergoing succession, leading to misjudgments of the transport duration and impact state of water masses. Furthermore, there is a lack of effective fusion mechanisms between different tracer evidence (such as chemical, microbial community, and eDNA). Various tracers have their own applicable conditions and limitations; single evidence is insufficient to accurately reflect the whole picture, while multi-source evidence often yields inconsistent or even contradictory results. Existing technologies lack a unified framework that can comprehensively assess the uncertainties of various pieces of evidence and adaptively weight and fuse them based on physical and biological constraints, resulting in the inability to form robust and reliable final judgments. Summary of the Invention

[0005] The purpose of this invention is to provide a hydrodynamic tracing method that couples microbial succession with multidimensional evidence, so as to solve the above-mentioned problems existing in the prior art.

[0006] According to one aspect of this application, a hydrodynamic tracing method coupling microbial succession with multidimensional evidence includes:

[0007] Acquire microbial community data, environmental DNA data, and chemical tracer data of the water body to be tested;

[0008] A community succession index was constructed based on microbial community data, and an eDNA decay index was constructed based on environmental DNA data. The combined inversion yielded a biological age raster characterizing the duration of water masses.

[0009] By comparing the biological age grid with the travel time field of the hydrodynamic model, an age consistency likelihood is constructed.

[0010] Bayesian fusion was performed on the age consistency likelihood, the chemical likelihood constructed based on chemical tracer data, and the source contribution likelihood constructed based on microbial community data to obtain the posterior probability grid of incoming water.

[0011] The area affected by incoming water is determined based on the posterior probability grid of the incoming water.

[0012] Beneficial effects: Through the above technical solutions, this invention solves the problems of unknown temporal dynamics of biological signals and difficulty in fusing multi-source evidence in existing methods, provides quantitative estimation of water mass duration, and improves the accuracy and reliability of tracing. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the overall process of a hydrodynamic tracing method that couples microbial succession with multidimensional evidence.

[0014] Figure 2 This is a schematic diagram of the process of obtaining a biological age raster through joint inversion.

[0015] Figure 3 This is a schematic diagram illustrating the construction process of the community succession index.

[0016] Figure 4 This is a schematic diagram of the process for constructing age-consistent likelihood. Detailed Implementation

[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1 describes the overall framework of a hydrodynamic tracing method that couples microbial succession with multidimensional evidence.

[0019] In this embodiment, as Figure 1 As shown, the method may specifically include the following steps:

[0020] Step S101: Obtain microbial community data, environmental DNA data, and chemical tracer data for the water body to be tested. The water body to be tested can be a lake, reservoir, river, or estuary, etc. Microbial community data typically refers to a table of operational taxonomic units or amplicon sequence variants reflecting the community structure obtained by high-throughput sequencing of marker genes such as the 16S rRNA gene of microorganisms in the water sample. Environmental DNA (eDNA) data refers to the relative or absolute abundance data of a specific indicator species in the water body obtained by quantitative PCR or high-throughput sequencing of DNA fragments of that species. Chemical tracer data refers to the concentration data of specific chemical components in the water body, such as chloride ions (Cl). - , sulfate SO4 2- Sodium ions (Na) + Concentration of isoconservative ions.

[0021] Specifically, the acquisition of the above data is preferably achieved through an event-driven intelligent sampling method, which will be detailed in Example 5. This method can utilize high-frequency sensors to monitor water quality changes in real time. When hydrodynamic disturbance events (such as gate opening or sudden flow surges) are detected, high-density sampling is automatically triggered to capture sample data during key dynamic processes.

[0022] After acquiring the raw data, standardization processing is required. This process includes: mapping data from different sources onto a unified geospatial grid using spatial interpolation methods (such as inverse distance weighting, ordinary kriging, or preferred flow field constraint interpolation); aligning the data to a unified time step using temporal interpolation (such as linear interpolation); and performing rigorous quality control, such as screening by comparison with blank and negative control samples, correcting outliers in the data, and performing median or bandpass filtering to denoise high-frequency time series signals.

[0023] Step S102 involves constructing a community succession index based on microbial community data and an eDNA decay index based on environmental DNA data, and jointly inverting these to obtain a biological age raster characterizing the duration of the water mass. Biological age refers to the time it takes for a water mass to travel from its source (such as an inlet) to the current sampling point. It is not dependent on traditional physical hydrodynamic models but is inferred from the dynamic changes in biological signals within the water body over time. The biological age raster is a two-dimensional or three-dimensional gridded map, where the value of each grid cell is an estimated biological age of the water body at that location. The community succession index quantifies the degree of change in the current aquatic microbial community compared to the source community; the eDNA decay index quantifies the degree of decay of the DNA signal of a certain indicator species after its release from the source over time. By jointly optimizing the model and simultaneously fitting biological signals at two different lifecycle scales, a more robust biological age can be obtained. The specific method of this joint inversion will be elaborated in Example 2.

[0024] Step S103 involves comparing the biological age grid with the travel time field of the hydrodynamic model to construct an age consistency likelihood. The travel time field of the hydrodynamic model is the most probable physical travel time required for water flow to reach each grid location from its source, calculated by a mature hydrodynamic model (such as the Delft3D 3D water environment simulation system, the MIKE series models, or a simplified convection dispersion model). The age consistency likelihood is a probability grid where the value of each grid cell represents the probability that the biological age at that location, derived by the biological method of this invention, matches the physical travel time calculated by the hydrodynamic model. In this invention, the comparison is not a simple matter of determining equality, but rather the construction of a probabilistic model to assess the probability that the estimated biological age falls within a tolerable range around the physical travel time. This step represents the first-ever fusion of independent evidence from biological tracing with prior knowledge from the physical model, the specific construction of which will be detailed in Example 4.

[0025] Step S104 involves Bayesian fusion of the age-consistency likelihood, the chemical likelihood constructed based on chemical tracer data, and the source contribution likelihood constructed based on microbial community data to obtain the posterior probability grid of the incoming water. The chemical likelihood is based on chemical tracers (such as Cl-). - The probability of an impact from incoming water is determined by changes in the concentration of Cl in a certain area. For example, when Cl is detected in a certain area... -When the concentration is significantly lower than the background value, the chemical likelihood value of the area will be relatively high. Source contribution likelihood, based on the analysis results of microbial source tracing models (such as SourceTracker), determines the proportion of contribution from a specific source (such as the Yangtze River) in the current aquatic microbial community and converts this proportion into a probability. Bayesian fusion, within a unified probabilistic framework, fuses three independent pieces of evidence—age consistency likelihood, chemical likelihood, and source contribution likelihood—with hydrodynamic priors (e.g., the potential impact range of incoming water constructed based on scheduled flow information) to obtain the final posterior probability grid of the incoming water. This posterior probability grid integrates all information; the higher its value, the greater the certainty that the location is affected by incoming water. Various preferred implementations of this fusion process will be described in detail in Examples 4 and 6.

[0026] Step S105: Determine the water influx impact zone based on the posterior probability grid of the incoming water. This is done by setting a probability threshold p. _Star For example, p _Star =0.7, the posterior probability grid with values ​​greater than p _Star The identified area is the water influx area. To obtain a more regular boundary that better reflects physical reality, a series of image processing operations can be performed on the binarized influence area mask, such as morphological cleansing (opening or closing operations) to eliminate isolated noise points, and connected component analysis to extract the main influence patches. The boundary of the influence area can be vectorized to generate the leading edge of the water influx. This determination and extraction process will be refined in Example 9.

[0027] Example 2 provides an implementation method for retrieving the age of water masses from a two-layer biological clock, detailing how a biological age grid is obtained through joint inversion, such as... Figure 2 As shown.

[0028] Specifically, the joint inversion to obtain the biological age raster includes: step S201, constructing a community succession model characterizing the dynamic changes of the community succession index; step S202, constructing an eDNA decay model characterizing the dynamic changes of the eDNA decay index; step S203, establishing a joint objective function that couples the deviation between the community succession index and the community succession model, as well as the deviation between the eDNA decay index and the eDNA decay model; and step S204, minimizing the joint objective function to obtain the biological age raster.

[0029] Before constructing a community succession model, it is necessary to construct a community succession index. In this embodiment, such as... Figure 3 As shown, the construction of the community succession index includes: obtaining source-sink community data representing the source; calculating the Aitchison distance or Bray-Curtis distance between the microbial community data and the source-sink community data; and normalizing the calculated distance values ​​to generate the community succession index.

[0030] Source-reservoir community data refers to multiple sample data collected from the source water body (such as the Wangyu River inlet in the Yangtze River-Taihu Lake Diversion Project) before the incoming water enters the water body to be tested, which can represent the initial microbial community composition of the incoming water. Before performing distance calculations, it is usually necessary to dilute the microbial community data, that is, to randomly sample to a uniform sequencing depth, for example, 10,000 sequences per sample, to eliminate the influence of sequencing depth differences on community structure comparisons.

[0031] Preferably, before calculating the Aitchison distance, the method further includes performing a centered logarithmic transformation (clr) on the microbial community data to eliminate its inherent compositional effects and generate a centered logarithmic transformation matrix. The calculation of the Aitchison distance is based on the centered logarithmic transformation matrix.

[0032] Compositional effects refer to the fact that microbial community data, as relative abundance data, have a constant sum of all components of 1, leading to spurious negative correlations between components. Directly using conventional measures such as Euclidean distance will introduce bias. The clr transformation is achieved through the following formula: clr(x) = [log(x...] _1 / g(x)),log(x _2 / g(x)),...,log(x _D / g(x))], where x=[x _1 ,x _2 ,...,x _D [x] is the relative abundance vector of D species in a sample, and g(x) is the geometric mean of x. The Aitchison distance D _A (x,y) represents the Euclidean distance between the vectors of two samples after the CLR transformation. As an alternative, the Bray-Curtis distance can be used, which is calculated based on the ratio of the sum of differences in abundance to the sum of abundance for each species, and is suitable for raw relative abundance data. The calculated distance values ​​are normalized to the [0,1] interval using linear mapping or other methods to obtain the observed community succession index raster SI. _obs The closer the value is to 1, the greater the difference from the source community, and the longer the succession time may be.

[0033] Next, we construct the community succession model SI. _model (t). This model describes the theoretical variation of the community succession index with time t. An exemplary model is the exponential model: SI _model(t) = 1 - exp(-λ*t), where λ is the community succession rate constant, characterizing the rate of community change. In another preferred embodiment, a piecewise logistic model can be used to better describe the lag phase that may exist in the early stage of community succession and the plateau phase in the later stage. Model parameters such as the community succession rate constant λ need to be calibrated through indoor controlled experiments or small-scale in-situ experiments. For example, after mixing source water and lake water in different proportions, the mixture can be continuously cultured under controlled conditions (such as isothermal) and sampled and sequenced at regular intervals. The observed succession index change curve over time can be fitted using the least squares method or Bayesian method to determine the value of λ.

[0034] Furthermore, an eDNA decay model Ce was constructed. _model (t). In this embodiment, constructing the eDNA decay model includes: establishing the model as an exponential decay function; the exponential decay function describes the decay process of environmental DNA over time using the initial eDNA concentration and the decay rate constant. Its mathematical form can be: Ce _model (t)=Ce _0 *exp(-k*t), where t is time, Ce _0 To indicate the initial eDNA concentration of a species at the source, k is the decay rate constant, Ce _model (t) represents the theoretical eDNA concentration at time t.

[0035] To improve model accuracy, the decay rate constant k was calculated using the following method: controlled-parameter decay experimental data were obtained; environmental factor data for water temperature, pH, and suspended solids concentration were acquired; based on the controlled-parameter decay experimental data and environmental factor data, the coupling function was calibrated; the coupling function was used to calculate the decay rate constant and characterize its multiple dependencies on temperature, pH, and suspended solids concentration. eDNA degradation is influenced by a complex array of environmental factors; a preferred coupling function form is a log-linear model.

[0036] ln(k)=β _0 +β _T’ *T'+β _pH *pH+β _S *SPM;

[0037] Where T' represents temperature, SPM represents suspended solids concentration, and the β-series parameters are regression coefficients obtained by measuring the eDNA degradation rate under different combinations of temperature, pH, and suspended solids in a laboratory setting, followed by multiple linear regression. Optionally, when temperature is the dominant factor, a modified Arrhenius equation can also be used to describe the relationship between k and temperature.

[0038] The coupling function for the attenuation rate constant k is preferably calibrated using a family of log-linear functions. A complete and specific example is as follows: ln(k) = β _0+β _T’ *T'+β _pH *pH+β _S *SPM+β _T’S *T'*SPM+β _T’2 *T' 2 .

[0039] Where: ln(k) is the natural logarithm of the decay rate constant k; T' is the water temperature; pH is the water pH value; SPM is the suspended solids concentration in the water; β _0 β is the intercept term of the model, or the logarithm of the fundamental decay rate; _T’ β _pH β _S These are the first-order (linear) effect coefficients for temperature, pH, and suspended solids concentration, respectively; β _T’S β is the coefficient of the interaction term between temperature and suspended solids concentration, characterizing, for example, the nonlinear relationship that high temperatures amplify the adsorption / protective effect of suspended solids; _T’2 The coefficients of the second-order (squared) term for temperature characterize the nonlinear relationship between the degradation rate and temperature (e.g., the existence of an optimal degradation temperature). It is understood that the form of this family of functions is not uniquely limited, and those skilled in the art can add or remove other environmental factors (such as light intensity) or higher-order interaction terms (such as T*pH) based on controlled-parameter experimental data to obtain the best fit.

[0040] Based on this, steps S203 and S204 are executed. For each grid cell (grid) g, a joint objective function is established:

[0041] J _g (t)=w1*[SI _model (t)-SI _obs (g)] 2 +w2*[Ce _model (t;g)-Ce _obs (g)] 2 ;

[0042] Among them, SI _obs (g) and Ce _obs (g) represent the community succession index and eDNA concentration observed in this grid cell, respectively; the community succession model SI _model (t) and Ce _model (t;g) is the predicted value from the theoretical model, where Ce _model The dependence on g is due to the different environmental factors at various locations, leading to different values ​​of k; w1 and w2 are weighting coefficients, usually taken as the reciprocals of their respective observation error variances, used to balance the contributions of the two biological signals. Numerical optimization algorithms (such as the golden section method or Newton's method) are used to search the one-dimensional search space t∈[0,T] _maxInside, search for J _g (t) The smallest t value, which is the estimated biological age t of that grid cell. _hat (g) Repeat this process for all grid cells to obtain the final biological age raster T. _hat .

[0043] Preferably, the solved biological age grid can be quality controlled. For example, physical constraints can be imposed to ensure that all age values ​​are non-negative and do not exceed the maximum possible travel time given by the hydrodynamic model. For outliers that significantly deviate from their neighborhood mean, methods such as neighborhood weighted averaging can be used for smoothing correction to improve the spatial continuity and rationality of the age field.

[0044] According to one aspect of this application, when diluting microbial community data, a uniform sequencing depth q can be determined based on the overall sequencing quality and species richness of the samples. For example, each sample can be diluted to 10,000 valid reads, or to the depth corresponding to the 90th percentile of the lowest number of valid reads across all samples, to achieve a balance between preserving the majority of samples and eliminating sequencing depth bias.

[0045] According to one aspect of this application, joint inversion specifically refers to the process of performing a unified optimization objective function (such as J...). _g In the equation (t), two biological signals with different timescale characteristics are simultaneously coupled: the community succession index, representing the dynamic changes of the living community, and the eDNA decay index, representing the residual biological signals of death. The unique biological age t is solved by minimizing this joint objective function. _hat This age value is the optimal solution that simultaneously satisfies the constraints of two biological clock models, overcoming the limitations of a single signal source.

[0046] According to one aspect of this application, in the joint objective function J _g In (t), the weights w1 and w2 are preferably determined using the inverse variance weighting method. Specifically, w1 is proportional to the observed community succession index SI. _obs The reciprocal of the variance, w², is proportional to the observed eDNA concentration Ce. _obs The reciprocal of the variance. The variance mentioned above can be estimated through repeated sampling, model fitting residuals, or uncertainty propagation. For example, if SI... _obs The estimated standard deviation is 0.05, Ce _obs The estimated standard deviation (after normalization) is 0.1, so the ratio of w1 to w2 can be set as (1 / 0.05). 2 ):(1 / 0.1 2 The ratio is 4:1, to ensure that the more accurate signal is dominant in the inversion.

[0047] Example 3: A preferred embodiment of age field spatial regularization is provided, when the biological age grid T _hat The steps in this embodiment can be enabled when there is significant spatial heterogeneity or noise due to sparse sampling points, measurement noise, or model inversion uncertainty.

[0048] Specifically, after obtaining the biological age raster through joint inversion, and before constructing the age-consistent likelihood, the method also includes:

[0049] Step S301: Apply spatial regularization to the biological age raster to smooth noise and enhance spatial continuity. Spatial regularization is a post-processing technique based on the assumption that physically adjacent water masses should have similar water mass ages. By applying smoothing constraints in space, isolated outliers in the original inversion results are corrected, and missing data regions are filled.

[0050] Step S302: Spatial regularization is performed using a Markov random field or total variational regularization method to generate a regularized age raster. As a preferred implementation, this embodiment employs the total variational TV regularization method. This method achieves this by minimizing the energy function L(T), which includes a data fidelity term and a regularization term. Specifically, a regularized age raster T is sought that minimizes the following objective function L(T):

[0051] L(T) = Σ _g [(T(g)-T _hat (g)) 2 / σ _T (g) 2 ]+λ _TV *Σ _〈g,h〉 |T(g)-T(h)|;

[0052] Where: g and h are the indices of the raster, and <g,h> represents all spatially adjacent raster pairs; T(g) is the age value of the regularized raster g to be solved; T _hat (g) represents the original biological age estimate for grid g; σ _T (g) represents the biological age uncertainty (e.g., variance) of grid g, obtained through the uncertainty propagation process; |T(g)-T(h)| is the absolute value of the age difference between adjacent grids g and h, i.e., the total variation term; λ _TV is the regularization coefficient, a non-negative weighting parameter used to balance data fidelity and spatial smoothness.

[0053] In the objective function L(T), the first term Σ _g [(T(g)-T _hat (g)) 2 / σ _T (g) 2[This is a data fidelity item, requiring the new age value T(g) to be close to the original estimated value T] _hat (g), and the degree of closeness is subject to uncertainty σ. _T Modulation of (g): When the uncertainty σ of the original estimate _T When (g) is very small, i.e., the estimation is very reliable, the penalty for this term is large, causing T(g) to be pulled towards T. _hat (g); conversely, if the uncertainty is large, then a large deviation of T(g) is allowed. The second term λ _TV *Σ _〈g,h〉 |T(g)-T(h)| is a regularization term that penalizes excessive age jumps between adjacent grid cells, forcing spatial smoothness across the entire grid. λ _TV The larger the value of , the smoother the resulting age grid T will be. The final regularized age grid T can be obtained by solving this minimization problem using a numerical optimization algorithm (such as gradient descent or convex optimization). _hat_reg .

[0054] Alternatively, a Markov random field (MRF) model can be used. In this model, the age value of each grid cell is treated as a random variable, and its probability distribution depends only on the age values ​​of its neighboring grid cells. An energy function or Gibbs distribution is defined, which incorporates a penalty and the observed data (biological age grid T). _hat The data terms that deviate and the smoothing terms that penalize the differences between adjacent grids can also be used to solve for the optimal age field by minimizing this energy function.

[0055] Step S303: The age consistency likelihood is constructed based on a regularized age raster. After performing the regularization step, the resulting regularized age raster T... _hat_reg and its corresponding uncertainty σ _T_reg Will replace the original T _hat and σ _T , as input data for constructing age-consistent likelihood.

[0056] According to one aspect of this application, in the total variational regularization objective function L(T), the regularization coefficient λ _TV The value of λ balances data fidelity and spatial smoothness. The optimal value of this parameter can be determined using techniques known in the art, such as the L-curve method or K-fold cross-validation. For example, a series of candidate values ​​for λ can be selected. _TV The values ​​are then regularized, and the residual norm and half-norm of the solution are calculated for each result. The optimal value is obtained at the inflection point of the L-curve; alternatively, cross-validation is used to select the λ value that minimizes the prediction error. _TV value.

[0057] According to one aspect of this application, the constructed total variational regularized objective function L(T) is a convex optimization problem (when the biological age uncertainty σ _T (g) When known). This minimization problem can be solved using a variety of mature numerical optimization algorithms, such as the alternating direction multiplier method, the primal-dual mixed gradient method, the split Bregman iteration algorithm, or by using existing convex optimization toolkits (such as the CVX toolkit).

[0058] Example 4 provides an implementation of multi-evidence Bayesian fusion, describing how to effectively and adaptively fuse biological age with other evidence such as chemical and microbial origins.

[0059] The fusion process is based on Bayes' theorem: P(Incoming Water | Evidence) ∝ P(Evidence | Incoming Water) * P(Incoming Water). Here, P(Incoming Water) is the hydrodynamic prior; P(Evidence | Incoming Water) is the likelihood product of each piece of evidence.

[0060] The various likelihood functions and priors required to construct the fusion:

[0061] Constructing chemical likelihood L _ion Based on the gridded ion concentration field C _ion Background ion field C _bg and source-end ion field C _Src Calculate the mixed fraction f _ion (g) For example, using a two-terminal hybrid model:

[0062] f _ion (g)=(C _obs (g)-C _bg (g)) / (C _Src (g)-C _bg (g)). The physical mixture fraction f _ion (g) The chemical likelihood L is obtained by mapping it to a probability using a nonlinear function (such as the Logistic function). _ion Preferably, the shape of the mapping function (e.g., the slope of the logistic function) should be subject to the measurement uncertainty σ. _ion Modulation: When the measurement uncertainty is high, the function curve should be flatter, indicating a more conservative attitude towards the judgment of the mixed fraction.

[0063] Construct source contribution likelihood L _Src The posterior probability P for microbial source apportionment (such as SourceTracker) _Src Processing can be performed. Time-varying regularization and hydrodynamic constraints can be applied. After processing, the contribution ratio q of the target source (such as the Yangtze River) can be calculated. _reg Mapped to source contribution likelihood raster L _Src .

[0064] Constructing hydrodynamic priors _hydro This prior represents the probability of incoming water based solely on hydrodynamic knowledge without considering on-site observations. One implementation method is to base it on historical scheduled flow rates U. _gate The prior probability is obtained by convolving the prior probability with the hydrodynamic response kernel K(x,y,Δt). Another simplified implementation is to directly utilize the travel time field t output by the hydrodynamic model. _HD For example, let the prior probability be t _HD The reciprocal of the probability (i.e., the closer the distance and the shorter the travel time, the higher the prior probability) is set to zero for physically unreachable regions.

[0065] Constructing age consistency likelihood L _Age ,like Figure 4 Specifically, a biological age raster or a regularized age raster, and a travel time field are obtained. A tolerance threshold τ representing age consistency is set. For example, τ can be set to 12 hours, meaning that biological age and physical age are considered consistent if the difference is within 12 hours. This threshold τ can be determined using cross-validation. The age difference raster between the biological age raster or the regularized age raster and the travel time field is calculated. Specifically, for each raster g, d(g) = T is calculated. _hat (g)-t _HD (g) Assess the probability that the absolute value of the age difference grid falls within the tolerance threshold, and establish this probability as the age consistency likelihood. In practice, it is preferable to consider the uncertainties of both ages. Calculate the combined uncertainty standard deviation σ. _D (g)=sqrt(σ _T (g) 2 +σ _HD (g) 2 ), where σ _T (g) is the uncertainty of biological age, σ _HD (g) represents the uncertainty of the travel time in the hydrodynamic model (usually given based on model validation experience). In this case, the age difference d(g) is considered to follow a mean of T. _hat (g)-t _HD (g) Standard deviation is σ _D (g) follows a normal distribution. It can be understood that the probability of consistency is the probability that the true value of this distribution falls within the interval [-τ, τ]. Age consistency likelihood L _Age (g) is calculated as follows:

[0066] L _Age (g)=Φ((τ-d(g)) / σ _D (g))-Φ((-τ-d(g)) / σ _D(g)); where Φ is the cumulative distribution function of the standard normal distribution. This formula calculates the observed difference d(g) and uncertainty σ. _D Under condition (g), the probability that the actual age difference is less than τ.

[0067] After the above four steps are completed, perform Bayesian fusion:

[0068] Acquire the biological age uncertainty generated during the joint inversion process. When acquiring microbial community data, environmental DNA data, and chemical tracer data, their respective measurement variances and missing measurement masks are explicitly preserved. During the joint inversion process, the measurement variance is propagated to estimate the generated biological age uncertainty σ. _T (g)

[0069] Based on biological age uncertainty σ _T (g) Adaptively determine the age likelihood weight w for age consistency likelihood. _Age Preferably, the weights of all likelihoods are adaptively determined based on their uncertainty (variance). For example, an inverse variance weighting strategy can be used: w _i =(1 / σ _i 2 ) / Σ _j (1 / σ _j 2 ), where i,j∈{ion,src,age}σ _i 2 This represents the uncertainty (variance) of the i-th piece of evidence (ion, source contribution, age). For example, σ _Age 2 This refers to the combined uncertainty σ used in the construction of age-consistent likelihood. _D (g) 2 This method assigns greater weight to evidence with lower uncertainty (more credible) in the fusion process.

[0070] In the logarithmic domain, age-consistent likelihood, chemical likelihood, and source contribution likelihood are weighted and fused using age-likelihood weights. To prevent floating-point underflow caused by multiplying multiple probability values ​​between [0,1], the fusion calculation is performed in the logarithmic domain.

[0071] The weighted fusion results are combined with hydrodynamic priors to generate a posterior probability grid of incoming water. An example of the fusion calculation formula is as follows:

[0072] log(P _Tilde (g))=w _ion (g)*log(L _ion (g))+w _Src (g)*log(L _Src (g))+w _Age(g)*log(L _Age (g))+log(Prior _hydro (g)); where P _Tilde (g) is the unnormalized posterior probability after fusion, w _ion w _Src w _Age These are the weights corresponding to each piece of evidence. Let P... _Tilde The raster is processed through exponential transformation and normalization, such as the Softmax function, to obtain the final posterior probability raster P of the incoming water with probability values ​​in the range [0,1]. _posterior .

[0073] Preferably, the method further includes employing any of the following mechanisms to correct for the correlation or interaction between likelihoods:

[0074] Option A (Nonlinear Interaction Modulation): This mechanism generates a nonlinear interaction term based on age-consistent likelihood and uses the interaction term to dynamically modulate the fusion weights of chemical likelihood and source contribution likelihood. For example, define the interaction term φ. _Age =exp(-|d(g)| / τ), this value is close to 1 when there is high consistency in age and close to 0 when there is poor consistency. The weights of other evidence are modulated to w. _ion_new =w _ion_base +α*φ _Age Among them, w _ion_new For the updated ion evidence weights, w _ion_base The base weights are α, where α is the modulation intensity coefficient. When the biological age is highly consistent with the physical age, the credibility of the entire biological tracing system (including source contributions) is enhanced, thus nonlinearly amplifying the weights of other relevant evidence and reflecting the effect of evidence synergy.

[0075] Option B (Gaussian Copula Correction): This mechanism estimates the correlation matrix R between each likelihood and applies a Gaussian Copula function to perform correlation correction on the fusion result, thus avoiding redundant calculations of evidence. This method mathematically handles the evidence (e.g., source contribution likelihood L) more rigorously. _Src Consistency L with age _Age The correlation between the data (all derived from microbial data) is used to explicitly remove redundant information through the R matrix, avoiding overconfidence in the posterior probability caused by repeatedly calculating the same evidence.

[0076] Option C (Evidence Discount Correction): This mechanism calculates the mutual information of evidence (MI) between likelihoods. _ij To determine the correlation discount coefficient δ _i =1 / (1+Σ _j MI _ij The uncertainty of each piece of evidence (such as the coefficient of variation, cv) should be considered. _iCalculate the confidence discount γ _i =1 / (1+cv _i 2 The weighted fusion is then adjusted. The final fusion weight w _i Proportional to δ _i *γ _i This approach is more easily implemented in engineering, while penalizing correlation with other evidence (δ). _i and its own uncertainty γ _i It is highly interpretable and computationally stable.

[0077] According to one aspect of this application, in alternative scheme B (Gaussian Copula correction), the correlation matrix R is used to describe the different likelihoods (L... _ion L _Src L _Age The statistical correlation between the two sequences is calculated. This matrix R can be obtained empirically by: obtaining the likelihood value sequences at multiple different spatiotemporal sampling points (e.g., including affected and unaffected regions); applying the inverse function of the cumulative distribution function of the standard normal distribution (i.e., the probit transformation) to the likelihood value sequences to transform them into sequences that follow a Gaussian distribution; and calculating the Pearson correlation coefficient matrix between the transformed sequences, which is the empirical estimate of R.

[0078] Example 5 provides an implementation method for event-driven intelligent sampling, which elaborates on the data acquisition process. It identifies hydrodynamic disturbance events through high-frequency real-time sensing and triggers high-density automatic sampling based on these events, thus solving the problem that traditional timed and fixed-point sampling is difficult to capture key instantaneous processes.

[0079] The acquisition of microbial community data, environmental DNA data, and chemical tracer data for the water body under test is achieved through an event-driven intelligent sampling method, which includes:

[0080] Step S501: Deploy high-frequency sensors to monitor water quality in real time and obtain high-frequency sensor data. In this embodiment, the system hardware is deployed at key monitoring sections, such as near the water inlet to the lake. The deployed hardware includes: one or more multi-parameter water quality probes, such as the YSI EXO2 series, which can simultaneously measure parameters such as conductivity, temperature, and pH at frequencies up to 4 Hz; an acoustic Doppler current profiler for acquiring water flow velocity profile data; a programmable automatic sampler, such as the ISCO 6712, which contains several sterile sampling bottles and can collect water samples according to external commands; and an edge computing device, such as a waterproof control box equipped with a Raspberry Pi or similar microprocessor, for receiving sensor data and executing real-time decision-making algorithms. The above devices constitute a collaborative monitoring and sampling unit. Its system architecture is as follows: the high-frequency sensor transmits the data stream to the edge computing device in real time, the device performs real-time analysis and judgment, and if the triggering conditions are met, it immediately issues a sampling control command to the automatic sampler.

[0081] It is understandable that key hydrodynamic phenomena, such as the concentration shock wave generated at the moment a gate opens or the concentration pulsation in the turbulent mixing zone, have timescales ranging from seconds to minutes, far less than the frequency of conventional manual sampling (hours or days). By deploying high-frequency sensors, high-resolution time-series data necessary to capture transient phenomena can be obtained.

[0082] Step S502: A change point detection algorithm is applied to the high-frequency sensor data to identify hydrodynamic disturbance events in real time. The change point detection algorithm runs on the edge computing device, continuously analyzing real-time data streams from the high-frequency sensor, such as conductivity time-series data. Preferably, a cumulative sum control chart algorithm is used. Specifically, under steady-flow conditions, the baseline mean μ and standard deviation of the conductivity data are determined. The cumulative sum S is calculated in real time. _n :

[0083] S _n =max(0,S _n-1 +(x _n -μ-k'));

[0084] Among them, S _n S is the cumulative sum at time point n; _n-1 x is the cumulative sum at the previous time point; _n is the sensor reading at the current time n; μ is the historical baseline mean; k' is the relaxation parameter or sensitivity parameter, usually set to half the baseline standard deviation, to allow data to fluctuate normally within a certain range without triggering an alarm.

[0085] When the calculated cumulative sum S _nWhen the threshold value h' is exceeded (for example, h' can be set to 5 times the baseline standard deviation), a hydrodynamic disturbance event is determined to have occurred.

[0086] Alternatively, a more complex Bayesian online change point detection algorithm can be used. This algorithm can not only detect change points, but also provide the posterior probability of a change point occurring at each time point, providing richer information for decision-making, but its computational complexity is relatively high.

[0087] Step S503: When a disturbance event is identified, a trigger signal is generated. Once the change point detection algorithm in the edge computing device determines S... _n Upon receiving the signal >h', the device immediately generates a digital or analog trigger signal through its I / O interface. This could be a high-level TTL pulse or a command code sent via an RS-232 serial port.

[0088] Step S504 involves using a trigger signal to drive an automated sampler to perform high-density sampling to acquire microbial community data, environmental DNA data, and chemical tracer data. The trigger signal is sent to the control port of the automated sampler. The sampler is pre-programmed with one or more high-density sampling procedures, for example, collecting a 500 mL water sample every 5 minutes for 2 hours after receiving the trigger signal. Upon receiving the trigger signal, the sampler immediately starts the corresponding program, automatically completing a series of high-frequency water sample collections and storing the samples in separate sterile sampling bottles. Furthermore, the samples precisely captured during the event are used for laboratory analysis to obtain microbial community, environmental DNA, and chemical tracer data that can finely characterize the hydrodynamic disturbance process. Through this method, the present invention can acquire key dynamic data that are unavailable through traditional methods in a relatively intelligent and efficient manner, providing relatively high-quality data for the accuracy of subsequent biological age inversion and multi-evidence fusion.

[0089] Example 6 provides a preferred embodiment for likelihood construction, describing and optimizing the construction process of chemical likelihood and source contribution likelihood to improve the robustness and physical authenticity of traditional tracer evidence.

[0090] Chemical likelihood is constructed based on chemical tracer data, organizing the data into multi-ion vectors containing various ions. Traditional chemical tracers mainly rely on a single chloride ion, but they are susceptible to collinearity effects caused by factors such as rainfall dilution. To address this issue, this embodiment employs a multi-parameter integrated tracer approach. At each sampling point, a set of conservative or semi-conservative ions, such as chloride ions (Cl), are measured. - , sulfate SO4 2- Sodium ions (Na) + Potassium ions K + And so on, and organize their concentration values ​​into a multidimensional vector X=[C _Cl C_SO4 C _na C _K ,...].

[0091] Furthermore, latent factor analysis is applied to the multi-ion vectors to extract latent factors that characterize the differences between source and background water bodies, reducing the dimensionality of multidimensional ion information and extracting the information most sensitive to source-sink mixing processes. Preferably, latent factor analysis methods such as nonnegative matrix factorization (NMF) or partial least squares (PLS) can be used. Taking NMF as an example, the multi-ion vectors of all samples are used to construct an observation matrix V, which is decomposed by NMF into the product of two low-rank nonnegative matrices W and H (V≈W*H). The rows of H can be interpreted as several latent factors, each a linear combination of the original multiple ions, possibly representing a certain geochemical process or source characteristic. Through analysis, the latent factor Z that maximizes the distinction between source and background water bodies (e.g., a latent factor with a high score at the source and a low score at the background) is selected.

[0092] Furthermore, a multi-ion latent factor mixing index is constructed based on latent factors to comprehensively reflect the mixing state of the water body. After extracting the key latent factor Z, mixing calculations are no longer performed on the concentration of a single ion, but rather on a two-endmember mixing calculation in the space of latent factor Z. That is: ILMI(g) = (Z _obs (g)-Z _bg ) / (Z _Src -Z _bg ), where ILMI is the multi-ion potential factor mixing index; Z _obs (g), Z _bg and Z _Src These represent the scores of the observation point, background water body, and source water body on the potential factor Z, respectively.

[0093] Based on this, the chemical likelihood is generated according to the multi-ion latent factor mixing index. The ILMI(g) is then converted to the chemical likelihood L using a method similar to that in Example 4, such as a logistic mapping. _ion Because the latent factor Z integrates information from multiple ions and eliminates collinearity interference, the chemical likelihood ratio generated based on ILMI is higher than that based on a single Cl- ion. - The likelihood is more robust and reliable.

[0094] Based on microbial community data, a source contribution likelihood is constructed. The source apportionment process for this data is optimized using at least one of the following constraints to generate a constrained source contribution vector. Traditional microbial source apportionment models (such as SourceTracker) are based on statistics and do not consider the physical processes of water flow, potentially leading to physically unreasonable source tracing results. This embodiment optimizes the model by introducing hydrodynamic and temporal continuity constraints.

[0095] Method A1: Obtain the travel time field of the hydrodynamic model, define the hydrodynamic reachability region, and use the hydrodynamic reachability region as a hard constraint in source resolution, ensuring that the source contribution of physically unreachable regions is zero. Specifically, before running the source resolution model, obtain the travel time field t from the hydrodynamic model. _HD Based on the duration T of the water diversion event _event We can define the hydrodynamic reachable region, that is, all regions that satisfy t _HD (g)<=T _event The set of grids g. In the optimization solution of the source resolution model, a hard constraint is applied: for all sampling points not within this reachable domain, the contribution ratio q from the target source (such as the Yangtze River) is... _Src It must be forced to be set to 0, directly integrating prior knowledge of hydrodynamics into the microbial model and excluding physically impossible solutions.

[0096] Method B1: Obtain the source contribution vector from the previous time step and introduce a time-varying regularization term into the optimization objective of the source apportionment model to smooth the changes in source contributions between adjacent time steps. The movement and mixing of water masses is a continuous process; therefore, the proportion of source contributions between adjacent time steps should not experience drastic, discontinuous jumps. To reflect this physical characteristic, a time-varying regularization term, such as γ*||q, can be added to the optimization objective function of the source apportionment model. _T -q _t-1 || _2 2 Among them, q _T and q _t-1 These are the source contribution vectors for the current and previous time steps, respectively, and γ is a regularization parameter that controls the smoothing intensity. This regularization term penalizes drastic changes in source contributions between adjacent time steps, making the output smoother and more reasonable in the time dimension.

[0097] Preferably, methods A1 and B1 can be used simultaneously to form a comprehensive hydrodynamically constrained source analysis HBCSD model. Its optimization objective is to minimize a function that includes the original likelihood term, the time-varying regularization term, and an optional spatial regularization term (such as the total variational TV regularization term), while satisfying the hard constraints of the hydrodynamic reachability domain.

[0098] Based on this, the source contribution likelihood is generated using the constrained source contribution vector. The constrained source contribution vector q obtained through the above optimization solution is... _reg Its physical realism and spatiotemporal continuity have been significantly improved. Taking the contribution proportion of the target source and mapping it to a probability value of [0,1] yields a more reliable source contribution likelihood L. _Src , used for final fusion.

[0099] Example 7 provides an example of an adaptive update closed loop, detailing the operation, maintenance, and self-optimization mechanism of the entire method. It describes a closed-loop feedback mechanism that triggers adaptive updates of model parameters through performance evaluation after the method runs.

[0100] The method further includes an adaptive update step: step S701, comparing the result of determining the influent influence zone with the evaluation benchmark to generate a performance evaluation report. In this embodiment, the result of determining the influent influence zone is the influent influence zone mask Region. _mask and the front line _vec The evaluation benchmark can be: (1) ground truth data obtained through high-density manual sampling or remote sensing image interpretation; (2) the judgment result of another known baseline method (e.g., using only chloride ion tracer) for relative performance comparison.

[0101] The performance evaluation report is a quantitative report, and preferably, it includes, but is not limited to, the following key performance indicators (KPIs):

[0102] Spatial accuracy metrics, such as accuracy, precision, recall, and F1 score, are used to evaluate the correctness of grid cells identified as affected areas.

[0103] Comprehensive performance metrics: such as the area under the receiver operating characteristic curve (AUC), which comprehensively measures the model's classification ability across all possible thresholds.

[0104] Timeliness indicators: For example, detection lag, which is the time taken from when the water flow actually reaches a certain point to when this method determines that the point is affected.

[0105] Positioning accuracy metrics: such as positioning error, i.e., the extracted front line. _vec The average Hausdorff distance between the ground truth front and the ground truth front.

[0106] Step S702: When the metrics in the performance evaluation report meet the preset update conditions, parameter updates are triggered. The preset update conditions are logical judgment rules used to initiate adaptive adjustments, and may specifically include:

[0107] Performance degradation trigger (preferred): For example, if the average AUC index drops by more than 5% over three consecutive water adjustment cycles, or if the time lag is found to be more than 12 hours longer than the baseline.

[0108] Periodic triggering: For example, regardless of performance, perform an update every 6 months to adapt to seasonal changes in the water body's background environment.

[0109] Event triggers: For example, when a significant change in the hydrological situation upstream (such as the Yangtze River) is detected, or when a significant drift occurs in the source-reservoir microbial community on which this method is based.

[0110] Step S703, parameter update includes: updating the circadian rhythm model parameters used in the joint inversion, and updating the threshold for determining the influx area. Specifically, the update includes:

[0111] Parameters of the biological clock model:

[0112] By re-executing small-scale in-situ experiments or utilizing the latest accumulated data, the community succession model SI can be re-evaluated. _model The community succession rate λ parameter in (t) was refitted. By supplementing with new parameter-controlled attenuation experiments, the regression coefficients (such as β) in the coupling function of the eDNA attenuation coefficient function k(T,pH,SPM) were refitted. _T’ ,β _pH Update (etc.) to reflect changes in the environmental medium.

[0113] Fusion and Judgment Thresholds:

[0114] Using the latest accumulated decision results and benchmark data pairs, the optimal posterior probability threshold p that maximizes the F1 score or balances specificity and sensitivity is re-searched and determined through K-fold cross-validation or time-series split cross-validation. _Star Similarly, through cross-validation, the tolerance threshold τ used to construct the age consistency likelihood was re-optimized.

[0115] Step S704: The updated biological clock model parameters and decision thresholds are applied to subsequent tracer method execution cycles to achieve adaptive loop closure of the model. The updated parameter set (e.g., the updated model parameter set Params) _vNext The updated parameter set is saved and used to replace the old one. When the next water diversion event occurs, the system will automatically call the updated parameter set to execute the entire process, ensuring that the method can continuously adapt to environmental changes and maintain its accuracy and reliability in long-term operation.

[0116] Example 8 provides a specific numerical calculation case based on the Yangtze River Diversion Project to Taiyuan, which fully demonstrates the entire calculation process of the present invention through a specific and reproducible numerical case.

[0117] Target location (g): Gonghu No. 7 monitoring point (GH7).

[0118] Target time (t): September 3.

[0119] Background values ​​(data from August 22nd before the Yangtze River diversion):

[0120] Background Cl -Concentration C _bg 37.66 mg / L.

[0121] Source reservoir (Wangyu River) microbial community S _lib : Acquired.

[0122] Source library eDNA indicator species g _BIyi10 Initial concentration Ce0 _Source Set to 10000 Reads (relative unit).

[0123] Source value (Wangyu River Estuary data):

[0124] Source Cl⁻ concentration C _Src 15.24 mg / L (the lowest value in the Wangyu River Estuary area on September 3).

[0125] Observations on September 3 (at GH7):

[0126] Observation Cl - Concentration C _obs (g): 31.8 mg / L (example value, below background).

[0127] Observation of microbial sample M _obs (g): Collected.

[0128] Observe the concentration of eDNA Ce _obs (g): g _BIyi10 50 Reads (example value).

[0129] Model parameters:

[0130] Succession rate constant λ: 0.04hr -1 .

[0131] eDNA decay function k(T',pH,SPM):

[0132] Assuming the environmental conditions at GH7 (T'=28°C, pH=8.2, SPM=30mg / L), k is calculated to be 0.05hr. -1 .

[0133] Hydrodynamic model output (at GH7 point):

[0134] Travel Time Field t _HD (g): 60hr.

[0135] Travel time uncertainty σ _HD (g): 10hr.

[0136] fusion parameters:

[0137] Age tolerance threshold τ: 12hr.

[0138] The probability threshold p _Star : 0.7.

[0139] Step A': Perform a two-layer biological clock age inversion.

[0140] Calculate SI _obs (g): Microbial sample M _obs (g) with source library S _lib Perform Aitchison distance calculation (after CLR transformation), and obtain SI after normalization. _obs (g) = 0.90 (indicating a significant difference from the source library).

[0141] Calculate Ce _obs (g): The observation value is 50 Reads.

[0142] Construct the objective function J _g (t): SI _model (t) = 1 - exp(-0.04*t); Ce _model (t) = 10000 * exp(-0.05 * t); To simplify the calculation, assume the weights w1 = w2 = 1, and for Ce... _obs Normalization is performed, assuming Ce _obs (g) The normalized value is 0.05;

[0143] J _g (t)=[(1-exp(-0.04*t))-0.90] 2 +[(10000*exp(-0.05*t)) / 10000-0.05] 2 ;

[0144] J _g (t)=[0.1-exp(-0.04*t)] 2 +[exp(-0.05*t)-0.05] 2 ;

[0145] Solve for t _hat (g): Solving arg min(J) through numerical optimization _g (t)).

[0146] If t=60, J _g (60) = [0.1 - 0.09] 2 +[0.05-0.05] 2 ≈0.0001;

[0147] If t=58, J _g (58) = [0.1 - 0.10] 2+[0.055-0.05] 2 ≈0.000025; ...;

[0149] The biological age t was obtained by solving the problem. _hat (g)≈58.5hr. Meanwhile, the biological age uncertainty σ is estimated using uncertainty propagation (e.g., the Delta method). _T (g)≈8hr.

[0150] Step B': Construct multiple evidence likelihood.

[0151] Chemical Likelihood L _ion (g): Use the multi-ion latent factor mixing index ILMI (or simplified to Cl) - ):

[0152] f _ion (g) = (31.8 - 37.66) / (15.24 - 37.66) = -5.86 / -22.42 ≈ 0.26. Applying the mixed fraction 0.26 to a Logistic regression and considering the uncertainty, we obtain L... _ion (g)≈0.35.

[0153] Source contribution likelihood L _Src (g): Analyze microbial samples M using the HBCSD method (or simplified to SourceTracker). _obs (g) The proportion obtained from the Yangtze River is q _reg (g) = 0.55 (example value, higher than 15.5% of the background). Mapped to likelihood, L _Src (g)≈0.60.

[0154] Age consistency likelihood L _Age (g):

[0155] t _hat (g) = 58.5hr,t _HD (g)=60hr,τ=12hr. d(g)=58.5-60=-1.5hr.

[0156] σ _D (g)=sqrt(σ _T (g) 2 +σ _HD (g) 2 =sqrt(8) 2 +10 2 )≈12.8hr.

[0157] L _Age(g)=Φ((12-(-1.5)) / 12.8)-Φ((-12-(-1.5)) / 12.8);

[0158] L _Age (g)=Φ(13.5 / 12.8)-Φ(-10.5 / 12.8)=Φ(1.05)-Φ(-0.82)≈0.853-0.206=0.647.

[0159] This high value indicates a good consistency between biological age and physical age.

[0160] Hydrodynamic Prior _hydro (g):

[0161] Because GH7 is located within the hydrodynamically accessible area and is relatively close, Prior _hydro (g)≈0.8.

[0162] Step C': Perform Bayesian fusion.

[0163] Uncertainty weighting: Assume the normalized weights of each piece of evidence are calculated as: w _ion =0.2,w _Src =0.4,w _Age =0.4.

[0164] Logarithmic field fusion:

[0165] log(P _Tilde (g))=0.2*log(0.35)+0.4*log(0.60)+0.4*log(0.647)+log(0.8)log(P _Tilde (g))≈0.2*(-1.05)+0.4*(-0.51)+0.4*(-0.43)+(-0.22)log(P _Tilde (g))≈-0.21-0.204-0.172-0.22=-0.806;

[0166] Normalization: P _Tilde (g) = exp(-0.806) ≈ 0.447. The fusion result here needs to be normalized to the unaffected probability. For simplicity, assume the normalized posterior probability P... _posterior (g)=0.75.

[0167] Judgment: P _posterior (g) = 0.75. Probability threshold: p _Star =0.7.

[0168] Since 0.75 > 0.7, it is determined that the Gonghu GH7 point was affected by the Yangtze River water flow on September 3.

[0169] This embodiment demonstrates in a relatively complete way how to obtain a quantifiable and interpretable probability of water inflow impact by starting from multidimensional observation data, through biological clock inversion and multi-evidence fusion.

[0170] Example 9 provides an implementation method for determining the inflow influence zone and extracting the leading edge, detailing how to process the posterior probability grid P of the inflow. _posterior The data is processed to generate a final, deliverable vector map of the water inflow impact zone and its leading edge.

[0171] The specific process of determining the impact zone of incoming water based on the posterior probability grid can include the following steps:

[0172] Step S901: Thresholding is performed on the posterior probability grid of the incoming water. In this step, the posterior probability grid P of each grid with values ​​between [0,1] is thresholded. _posterior Compared with the preset probability threshold p _Star (This threshold has been determined through cross-validation, e.g., p) _Star =0.7) for comparison. For grid g, if its posterior probability P _posterior (g) is greater than p _Star If the value is true, the raster is marked as 1 (belonging to the affected area); otherwise, it is marked as 0 (not belonging to the affected area). Through this operation, a binarized affected area mask raster (Region) can be obtained. _mask .

[0173] Step S902 involves morphological cleansing of the influence zone mask grid. Due to measurement noise or local model errors, the binarized mask generated in step S901 may contain isolated noise points (i.e., a single 1 pixel surrounded by 0 pixels) or internal voids (i.e., small 0 regions within a 1 region). To obtain a physically more continuous and reasonable water influx influence zone, the influence zone mask grid (Region) needs to be cleaned. _mask Morphological cleansing is performed. Preferably, opening and closing operations are performed sequentially. The opening operation performs erosion and dilation operations on the image, which can effectively remove isolated bright spots (noise) in the mask and smooth the boundaries of objects. The closing operation performs dilation and erosion operations sequentially, which can fill small holes inside objects and connect adjacent areas. By combining the two operations, a cleaned mask with more regular spatial structure and less noise can be obtained.

[0174] Step S903 involves performing connected component analysis and extraction on the purified mask. Within the purified mask, there may be multiple unconnected patches identified as influencing areas. Connected component analysis algorithms (e.g., eight-neighbor or four-neighbor algorithms) are used to identify and label all independent connected regions. Based on the area of ​​each connected region, i.e., the number of grid cells it contains, the main influencing areas of the incoming water can be filtered out. For example, an area threshold can be set to retain only connected regions with an area greater than that threshold, filtering out small patches that may be generated by model artifacts and have no practical significance.

[0175] Step S904: Vectorize the finally determined influence areas to generate the front line. For the selected main influence area raster, a contour extraction algorithm (e.g., Marching Squares algorithm) is used to identify its outer boundary. The raster boundary is then converted into a vector line segment composed of a series of geographic coordinate points, i.e., the water inflow influence front line. _vec Simultaneously, the entire affected area patch can also be converted into a vector polygon region. _vec Vectorized outputs (such as Shapefile or GeoJSON format) are the final deliverables and can be used for subsequent water quality early warning, water quantity calculation, and scheduling decision support.

[0176] According to one aspect of this application, in some application scenarios, single chloride ion tracing may be limited by the complexity of the background water or the presence of multiple pollution sources. As an alternative or supplement to the multi-ion potential factor mixing index scheme, the present invention can also employ differential ion ratio tracing technology. Specifically, this technology selects two conservative ion pairs with stable and significantly different ratios at the source and background ends, such as chloride / bromine ion (Cl... - / Br - According to literature and measured data, the Cl in the Yangtze River water... - / Br - The ratio (approximately 288±15) was significantly lower than that of Taihu Lake (approximately 456±28). When the water body is diluted by events such as rainfall, although the Cl... - and Br - The absolute concentrations of the bromide ions decrease proportionally, but their ratio theoretically remains unchanged, enhancing the robustness of the tracer under complex hydrological conditions. In practice, high-precision ion chromatography or inductively coupled plasma mass spectrometry is required to accurately measure low concentrations of bromide ions (typically at the μg / L level). This is achieved by calculating the Cl concentration at the sampling points. - / Br - The ratio can be used to inversely deduce the mixing ratio of the incoming water using a two-terminal mixing model. Alternatively, other ion pairs, such as sodium / potassium ions (Na+ / Na+), can also be used. + / K + ) or calcium / magnesium ions (Ca2+ / Mg 2+ This is on the premise that these ion pairs have stable and distinguishable ratio characteristics between source and sink bodies.

[0177] According to one aspect of this application, in areas such as the water inlet to the lake, there is intense turbulent mixing, causing tracer concentrations to fluctuate dramatically at a small spatial scale (centimeter to meter level), posing a significant challenge to traditional point sampling. This invention effectively addresses this problem by combining event-driven intelligent sampling with a predetermined monitoring strategy. Specifically, in the turbulent mixing zone, monitoring and sampling units are deployed in an array. The intelligent sampling system is triggered when the upstream gate opens (this information can be obtained in advance through an early warning mechanism) or when the high-frequency conductivity sensor array detects a dramatic change in the indicated concentration shock wave. At this time, not only a single sampler, but the entire sampler array is activated, performing high-density sampling synchronously or according to a preset time sequence at a small spatial scale. Through a turbulent vortex-capturing intelligent sampling strategy, samples with high spatiotemporal resolution within the turbulent structure can be obtained, providing more refined data support for accurately understanding the rapid mixing process in the early stages of lake entry, calibrating model parameters, and verifying tracer results.

[0178] The proposed dual-layer biological clock inversion scheme no longer treats microbial signals as static fingerprints, but instead constructs a community succession model SI. _model (t) and eDNA decay model Ce _model (t), which quantitatively describes the evolution and degradation of biological signals over time. Through joint inversion, a new dimension—biological age t—was obtained. _hat This study is the first to achieve an effective distinction between newly arrived and long-standing water masses, solving the problem of neglecting the temporal dynamics of signals in existing biological tracing methods.

[0179] This invention elaborates on a unified Bayesian fusion framework, introducing age-consistent likelihood L... _Age Using physical model t _HD biological age t _hat Constraints and verifications were implemented to resolve conflicts between pieces of evidence. The framework also addresses uncertainties propagated throughout the entire process (such as the biological age uncertainty σ). _T Adaptive weighted fusion is used to ensure that highly credible evidence has a higher weight in decision-making. Optimization of likelihood construction (such as HBCSD) and optional mechanisms for evidence relevance (such as discount correction) jointly construct a robust, self-consistent, and physically meaningful fusion system, solving the problem of the lack of an effective fusion mechanism for multi-source evidence.

[0180] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details of the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. A hydrodynamic tracer method that couples microbial succession with multidimensional evidence, characterized in that, The method comprises the following steps: obtaining microbial community data, environmental DNA data and chemical tracer data of a water body to be tested; constructing a community succession index based on the microbial community data, constructing an eDNA decay index based on the environmental DNA data, and jointly inverting a biological age grid representing the duration of a water mass; comparing the biological age grid with a travel time field of a hydrodynamic model to construct an age consistency likelihood; performing Bayesian fusion on the age consistency likelihood, a chemical likelihood constructed based on the chemical tracer data, and a source contribution likelihood constructed based on the microbial community data to obtain a water inflow posterior probability grid; determining a water inflow affected area according to the water inflow posterior probability grid; The travel time field of the hydrodynamic model is the most likely physical travel time of water flow from a source to each grid position calculated by a mature hydrodynamic model. The biological age grid representing the duration of a water mass is obtained by jointly inverting the following: constructing a community succession model representing the dynamic changes of the community succession index; constructing an eDNA decay model representing the dynamic changes of the eDNA decay index; establishing a joint objective function coupling the deviations between the community succession index and the community succession model and the deviations between the eDNA decay index and the eDNA decay model; minimizing the joint objective function to obtain the biological age grid.

2. The method of claim 1, wherein, The construction of the community succession index based on the microbial community data comprises the following steps: obtaining source pool community data representing a source; calculating the Aitchison distance or Bray-Curtis distance between the microbial community data and the source pool community data; normalizing the calculated distance values to generate the community succession index.

3. The method of claim 2, wherein, Before calculating the Aitchison distance, the following steps are further included: performing a centered log transformation on the microbial community data to generate a centered log transformation matrix; The Aitchison distance is calculated based on the centered log transformation matrix.

4. The method of claim 1, wherein, The construction of the eDNA decay model representing the dynamic changes of the eDNA decay index comprises the following steps: establishing the model as an exponential decay function; The exponential decay function describes the decay process of environmental DNA over time through an initial eDNA concentration and a decay rate constant.

5. The method of claim 4, wherein, The decay rate constant is calculated by the following method: obtaining controlled parameter decay experiment data; obtaining environmental factor data of the water body, including temperature, pH and suspended solids concentration; calibrating a coupling function based on the controlled parameter decay experiment data and the environmental factor data; The coupling function is used to calculate the decay rate constant and represents its multiple dependence on temperature, pH and suspended solids concentration.

6. The method of claim 1, wherein, After obtaining the biological age grid by jointly inverting, and before constructing the age consistency likelihood, the following steps are further included: applying spatial regularization processing to the biological age grid; The spatial regularization processing adopts a Markov random field or a total variation regularization method to generate a regularized age grid; The age consistency likelihood is constructed based on the regularized age grid.

7. The method according to claim 1 or 6, characterized in that, The construction of the age consistency likelihood comprises the following steps: obtaining the biological age grid or the regularized age grid, and the travel time field; setting a tolerance threshold representing age consistency; calculating an age difference grid between the biological age grid or the regularized age grid and the travel time field; The probability that the absolute value of the age difference grid falls within the tolerance threshold is evaluated and established as the age consistency likelihood.

8. The method of claim 1, wherein, The Bayesian fusion includes: Obtaining the biological age uncertainty generated in the joint inversion process; Adaptively determining the age likelihood weight of the age consistency likelihood based on the biological age uncertainty; In the logarithmic domain, the age consistency likelihood, the chemical likelihood, and the source contribution likelihood are weighted and fused by using the age likelihood weight; The weighted fusion result is combined with the hydrodynamic prior to generate the inflow posterior probability grid.

9. The method of claim 8, wherein, The weighted fusion further includes correcting the correlation or interaction between the likelihoods by using any of the following mechanisms: Mechanism A: generating a nonlinear interaction term based on the age consistency likelihood, and dynamically modulating the fusion weight of the chemical likelihood and the source contribution likelihood by using the interaction term; Mechanism B: estimating the correlation matrix between the likelihoods, and applying a Gaussian Copula function to perform correlation correction on the fusion result; Mechanism C: calculating the evidence mutual information between the likelihoods to determine the correlation discount coefficient, and calculating the confidence discount by combining the uncertainty of each evidence to correct the weighted fusion.

Citation Information

Patent Citations

  • Sewage treatment system and sewage treatment method

    CN117699958A

  • Method for quality assurance in the supply of drinking water and / or ground water

    EP3115464A1