Agricultural non-point source pollution monitoring data fusion and quality control method for heterogeneous sensors

By performing spatiotemporal alignment and feature extraction on agricultural non-point source pollution monitoring data, a watershed mechanism proxy model and a dynamic cross-domain event map are constructed. This solves the problem of insufficient utilization of cross-domain causal relationships in existing technologies, and realizes efficient consistency verification and repair of heterogeneous sensor data, thereby improving the accuracy and reliability of the data.

CN121706010APending Publication Date: 2026-03-20HEBEI AGRICULTURAL UNIV.
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511910840.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing agricultural non-point source pollution monitoring methods cannot effectively utilize cross-domain causal chains and time delay effects, leading to confusion in the judgment of data anomalies, reducing the credibility of monitoring data and the practicality of early warning systems. Furthermore, existing remediation methods cannot meet the pollution transmission patterns, resulting in distorted data.

Method used

By performing spatiotemporal alignment and feature extraction on heterogeneous sensor data, a watershed mechanism proxy model and a dynamic cross-domain event map are constructed. Machine learning models are used to trace the source in reverse and perform cross-domain logical consistency verification and optimization repair to generate self-consistent repair values ​​that conform to the pollution transmission law.

Benefits of technology

It significantly improves the accuracy of identifying complex pollution events and the reliability of sensor fault diagnosis, ensures the physical rationality of abnormal data repair results, and enhances data credibility and repair quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121706010A_ABST
    Figure CN121706010A_ABST
Patent Text Reader

Abstract

The invention relates to a heterogeneous sensor-oriented agricultural non-point source pollution monitoring data fusion and quality control method. According to the method, firstly, heterogeneous monitoring data is subjected to standardization processing, a drainage basin agent model fusing a physical mechanism is constructed, and then a dynamic event graph describing a cross-domain causal relationship is established according to the drainage basin agent model; when it is detected that data is abnormal, the system performs reverse traceability through the atlas and drives the agent model to generate a theoretical expected value, and logic consistency verification is performed through comparison so as to distinguish a real pollution event from a sensor fault; for the data judged to be faulty, an optimization model fused with physical constraints is further established and solved, and a self-consistent repair value conforming to a pollution transmission rule is generated, so that quality control normal form transformation from isolated threshold judgment to cross-domain logic verification is realized; the accuracy of complex pollution event identification, the reliability of sensor fault diagnosis and the physical rationality of an abnormal data restoration result are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of environmental monitoring technology, and in particular relates to a method for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors. Background Technology

[0002] In the current field of agricultural non-point source pollution monitoring, a wide range of heterogeneous sensor networks, including weather stations, soil sensors, and online water quality analyzers, have been deployed, accumulating massive amounts of time-series monitoring data. However, the control of these data quality (data quality control) still mainly relies on traditional single-point or single-sequence analysis methods, such as setting static thresholds to determine whether the data exceeds limits, or using statistical methods to detect abrupt changes within a single data sequence.

[0003] The aforementioned conventional methods suffer from a fundamental flaw: they treat sensor data from different physical domains, which should be closely related, as isolated signals, completely ignoring the inherent, physically governed, cross-domain causal chains and time-delay effects in agricultural non-point source pollution events. For example, a real pollution event typically follows a transmission path of "rainfall driving changes in soil moisture, which in turn triggers surface runoff, ultimately leading to increased turbidity and pollutant concentrations in downstream water bodies," with specific time lags between each stage. Existing quality control technologies cannot model and utilize this logical relationship across multiple domains, including meteorology, soil, hydrology, and water quality, leading to a core dilemma: when a system alarm indicates a sudden spike in a water quality parameter (such as ammonia nitrogen concentration), current technologies cannot reliably distinguish whether this is a real pollution event requiring immediate action or simply due to drift, blockage, or electronic malfunction of the water quality sensor itself. This confusion between "true" and "false" anomalies severely reduces the reliability of monitoring data and the practicality of early warning systems, making subsequent data analysis, model simulation, and management decisions based on uncertainty. Furthermore, for data that is simply judged as abnormal, existing repair methods (such as linear interpolation and mean filling) only pursue the smoothness and completeness of the data sequence mathematically. The repair results often violate basic physical laws such as mass conservation and energy conservation. Although the generated data looks "good", it is "distorted" and cannot be used for accurate pollution load accounting and process mechanism research.

[0004] Therefore, the industry urgently needs an innovative data fusion and quality control method that can break through the limitations of isolated data point analysis and instead start from the overall physical logic of pollution transmission to perform consistency verification and intelligent repair of cross-domain heterogeneous sensor data. Summary of the Invention

[0005] Therefore, it is necessary to provide a data fusion and quality control method for agricultural non-point source pollution monitoring based on heterogeneous sensors to address the above-mentioned technical problems.

[0006] Firstly, this application provides a method for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors, including:

[0007] S1. Perform spatiotemporal alignment and feature extraction processing on the heterogeneous raw monitoring data in the agricultural non-point source pollution monitoring network to generate a standardized time-series feature dataset for each watershed sub-unit;

[0008] S2. Construct a simplified mechanism module based on the physical mechanism of agricultural non-point source pollution transmission. Based on the simplified mechanism module, train a machine learning model using historical monitoring data to obtain a watershed mechanism proxy model that describes the watershed mechanism from upstream rainfall and soil characteristics to downstream water quality parameter response.

[0009] S3. Based on the driving relationship between input and output variables described by the watershed mechanism proxy model, and combined with the sensor network spatial topology of the agricultural non-point source pollution monitoring network, a dynamic cross-domain event graph is constructed. In the dynamic cross-domain event graph, the nodes represent sensor or hydrological events, and the edges represent the causal relationship between nodes and the expected time delay.

[0010] S4. When any sensor data is detected as abnormal, reverse tracing is performed along the causal influence edge in the dynamic cross-domain event graph to obtain the upstream node data; the upstream node data is used to drive the watershed mechanism proxy model to calculate the theoretical expected value range of the current abnormal data.

[0011] S5. By comparing the actual observed values ​​with the theoretical expected values, cross-domain logical consistency verification and fault conflict location are performed to obtain abnormal sensor data that are judged to be logically conflicted and the conflict is located as sensor fault.

[0012] S6. Based on abnormal sensor data, establish an optimized repair model that integrates physical law constraints and the mechanism relationship described by the watershed mechanism proxy model; generate self-consistent repair values ​​that conform to the pollution transmission law by solving the optimized repair model to replace the abnormal sensor data.

[0013] Secondly, this application also provides a data fusion and quality control device for agricultural non-point source pollution monitoring based on heterogeneous sensors, used to implement the method described in the first aspect, the device comprising:

[0014] The multi-source data preprocessing module is used to perform spatiotemporal alignment and feature extraction on heterogeneous raw monitoring data in the agricultural non-point source pollution monitoring network, generating a standardized time-series feature dataset for each watershed sub-unit;

[0015] The pollution transmission modeling module is used to construct a simplified mechanism module based on the physical mechanism of agricultural non-point source pollution transmission. Based on the simplified mechanism module, a machine learning model is trained using historical monitoring data to obtain a watershed mechanism proxy model that describes the watershed mechanism from upstream rainfall and soil characteristics to downstream water quality parameter response.

[0016] The dynamic relationship graph construction module is used to construct a dynamic cross-domain event graph based on the driving relationship between input and output variables described by the watershed mechanism proxy model and the sensor network spatial topology of the agricultural non-point source pollution monitoring network. In the dynamic cross-domain event graph, the nodes represent sensor or hydrological events, and the edges represent the causal relationship between nodes and the expected time delay.

[0017] The anomaly tracing and localization module is used to trace the upstream node data by following the causal influence edge in the dynamic cross-domain event graph when any sensor data is detected as an anomaly; and to use the upstream node data to drive the watershed mechanism proxy model to calculate the theoretical expected value range of the current anomaly data.

[0018] The data consistency verification module is used to perform cross-domain logical consistency verification and fault conflict location by comparing the actual observed values ​​with the theoretical expected values, and to obtain abnormal sensor data that is judged to be logically conflicted and the conflict is located as a sensor fault.

[0019] The self-consistent data repair module is used to establish an optimized repair model based on abnormal sensor data, which integrates the mechanistic relationship described by the physical law constraints and the watershed mechanism proxy model. By solving the optimized repair model, self-consistent repair values ​​that conform to the pollution transmission law are generated to replace the abnormal sensor data.

[0020] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a method for data fusion and quality control of agricultural non-point source pollution monitoring oriented towards heterogeneous sensors as described in the first aspect.

[0021] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for data fusion and quality control of agricultural non-point source pollution monitoring oriented towards heterogeneous sensors, as described in the first aspect.

[0022] The aforementioned method for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors first standardizes the heterogeneous monitoring data and constructs a watershed proxy model that integrates physical mechanisms. Based on this, a dynamic event graph characterizing cross-domain causal relationships is established. When data anomalies are detected, the system uses this graph to trace the source and drives the proxy model to generate theoretical expected values. Logical consistency verification is performed through comparison to distinguish between real pollution events and sensor malfunctions. For data determined to be faulty, an optimization model integrating physical constraints is further established and solved to generate self-consistent repair values ​​that conform to the pollution transmission law. This achieves a paradigm shift in quality control from isolated threshold judgment to cross-domain logical verification, significantly improving the accuracy of identifying complex pollution events, the reliability of sensor fault diagnosis, and the physical rationality of abnormal data repair results. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating a method for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors provided by the present invention;

[0025] Figure 2 This is a schematic diagram of the process of constructing a watershed mechanism proxy model in one optional embodiment of the present invention;

[0026] Figure 3 This is a schematic diagram of the structure of an agricultural non-point source pollution monitoring data fusion and quality control device for heterogeneous sensors provided by the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0028] refer to Figure 1 The document presents a flowchart illustrating a data fusion and quality control method for agricultural non-point source pollution monitoring based on heterogeneous sensors, as provided in this application. The method includes the following steps:

[0029] S1. Perform spatiotemporal alignment and feature extraction processing on the heterogeneous raw monitoring data in the agricultural non-point source pollution monitoring network to generate a standardized time-series feature dataset for each watershed sub-unit.

[0030] Specifically, the core of this step is to address the inconsistencies in the spatiotemporal dimensions of heterogeneous sensor data, and to provide unified basic data for subsequent modeling through standardization. First, the types of heterogeneous sensors included in the monitoring network are identified. Different sensors have different sampling frequencies due to differences in the monitored targets. Spatiotemporal alignment is processed from both the temporal and spatial dimensions.

[0031] Time alignment employs the principle of "high-frequency interpolation, low-frequency resampling," using the highest sampling frequency among all sensors as the baseline time granularity to interpolate and supplement low-sampling-frequency data. Specifically, for sensor data with low sampling frequencies and smooth parameter variations, cubic spline interpolation is used, with the following formula:

[0032]

[0033] In the formula, For any time The corresponding interpolation results, Given the time points, Time node The corresponding observed values, and ; , , These are the coefficients of the first, second, and third terms in cubic spline interpolation. They are obtained by solving for the continuity of the first derivative at adjacent time points and the natural boundary condition that the second derivative at the boundary is zero, ensuring the smoothness and trend fit of the interpolation results. For sensor data with stable parameter changes over a short period, linear interpolation is used, with the following formula:

[0034]

[0035] In the formula, For the next adjacent time node For the corresponding observations, this method uses a linear relationship to supplement the numerical values ​​between two known observations, meeting the interpolation accuracy requirements for stationary parameters. After time alignment, all data needs to be timestamped to unify the time format and remove records with duplicate or missing timestamps.

[0036] Spatial alignment, based on the hydrological segmentation results of the Digital Elevation Model (DEM), divides the entire monitoring watershed into several watershed sub-units. The segmentation process first uses a priority flood algorithm to fill depressions in the DEM data, eliminating the influence of local depressions on the determination of flow direction. Then, the D8 algorithm is used to determine the flow direction of each grid (i.e., the adjacent grid with the largest flow direction elevation difference). Based on the flow direction, a confluence accumulation matrix is ​​constructed, and the boundaries of the watershed sub-units are delineated using a confluence accumulation threshold. For each sub-unit, inverse distance weighted interpolation is used to spatially interpolate the data from the distributed sensors. The formula is:

[0037]

[0038] In the formula, For any target location within the sub-unit The interpolation results, For the first Observations from each sensor For the first Each sensor monitors the target location. The weight, For target location With the The spatial linear distance between the sensors, and ( The distance attenuation coefficient is determined by cross-validation to ensure interpolation accuracy. This method enables continuous spatial coverage of data within a sub-cell.

[0039] The feature extraction process covers statistical features, trend features, and physical correlation features. Statistical features include the mean, variance, maximum, minimum, and median of each parameter, with the mean and variance calculated using a sliding window. Trend features are represented by the linear regression slope and the sliding window trend coefficient. The linear regression slope is obtained by fitting a univariate linear regression equation to the time-series data, with the formula:

[0040]

[0041] In the formula, The slope of the linear regression. This represents the number of samples in the time series data. For the first Time points for each sample For the first The slope of the observed values ​​of each sample reflects the long-term trend of the parameter; the sliding window trend coefficient uses the slope of a linear regression with a fixed window size to reflect the short-term trend of the parameter. Physical correlation features are constructed based on the pollution transmission mechanism, including correlation coefficients and lag correlation coefficients between different parameters. The lag correlation coefficient is determined by calculating the Pearson correlation coefficient between two parameters at different time lags, and the lag time corresponding to the maximum correlation coefficient is taken as the feature value. All extracted features need to be Z-score standardized, using the formula: In the formula These are the standardized eigenvalues. These are the original eigenvalues. The population mean is a characteristic. The overall standard deviation of the feature is used to eliminate the difference in dimensions through standardization, and finally a standardized time series feature dataset for each watershed subunit is generated.

[0042] S2. Construct a simplified mechanism module based on the physical mechanism of agricultural non-point source pollution transmission. Based on the simplified mechanism module, train a machine learning model using historical monitoring data to obtain a watershed mechanism proxy model that describes the watershed mechanism from upstream rainfall and soil characteristics to downstream water quality parameter response.

[0043] Specifically, this step constructs a watershed mechanism proxy model that balances rationality and computational efficiency by combining "physical mechanism simplification + data-driven modeling". Agricultural non-point source pollution transport follows the physical mechanism of "rainfall-runoff generation-convergence-pollutant migration and transformation". Due to the complexity of actual watershed conditions, the core processes and key parameters are extracted by constructing a simplified mechanism module.

[0044] The simplified mechanism module comprises three core sub-processes: rainfall runoff generation, pollutant migration, and water quality response. The rainfall runoff generation sub-module can use the SCS-CN model to describe the relationship between rainfall and surface runoff, as shown in the formula: In the formula Surface runoff, For rainfall, The initial loss is the rainfall amount that does not generate runoff after the rainfall begins; the relationship between the initial loss and the curve number (CN value, reflecting the watershed's runoff potential) can be: In the formula, 25.4 is the unit conversion factor (which can be adjusted according to actual conditions); the CN value needs to be dynamically adjusted based on soil moisture and land use type, and the adjustment formula is as follows: In the formula This is the actual CN value. This represents the minimum CN value when the soil is saturated. This represents the maximum CN value under soil drought conditions. To measure soil moisture, This refers to the residual soil moisture content. This represents the saturated water content of the soil.

[0045] The pollutant migration submodule focuses on nitrogen and phosphorus pollutants, simplifying their migration processes in soil-runoff-water bodies, including the migration of dissolved and adsorbed pollutants. The formula for the migration flux of dissolved pollutants is: In the formula For the migration flux of dissolved pollutants, This refers to the concentration of dissolved pollutants in the soil. The pollutant leaching coefficient (reflecting the ability of pollutants to leach from soil to runoff); the formula for the migration flux of adsorbed pollutants is: In the formula For the migration flux of adsorbed pollutants, For soil bulk density, This represents the concentration of adsorbed pollutants in the soil. The pollutant partition coefficient (reflects the proportion of pollutants in the solid and liquid phases of the soil). It can be calculated from the soil organic matter content, using the following formula: In the formula This is the organic carbon adsorption coefficient (reflecting the ability of pollutants to adsorb organic carbon in the soil). This represents the soil organic matter content. Simultaneously, considering the loss of pollutants during migration, the formula for the pollutant flux reaching downstream water bodies is: In the formula For downstream water pollutant flux, It is the pollutant attenuation coefficient (reflecting the rate of loss through biodegradation, chemical transformation, etc.). The migration time of pollutants from upstream soil to downstream water quality monitoring points (calculated from surface runoff velocity and migration distance).

[0046] The water quality response submodule describes the relationship between pollutant flux and downstream water quality parameters, using the following formula: In the formula For downstream water quality parameter concentration, The water volume at the downstream monitoring section. This is the water quality response coefficient (reflecting the degree of uniformity of pollutant mixing in water). Calculations based on the hydrological characteristics of monitoring sections are used to ensure the rationality of the correlation between water quality parameters and pollutant flux.

[0047] Based on a simplified mechanism module, a Long Short-Term Memory (LSTM) network was selected as the machine learning model to train the watershed mechanism proxy model. The LSTM model can handle long-term temporal dependencies, which aligns with the time-lag effect of pollution transport. Historical monitoring data used for training must cover different hydrological seasons, and the data integrity must meet the modeling requirements, with known outlier data segments removed. Upstream parameters (rainfall, soil characteristics, etc.) from the standardized time-series feature dataset were used as model inputs, and downstream water quality parameters were used as model outputs. Input time windows and output prediction step sizes were set to learn the cumulative impact of previous inputs on current water quality.

[0048] Before model training, the data is normalized using Min-Max, with the following formula: In the formula For the normalized data, The minimum value of the parameter. The maximum value of the parameters is used to eliminate the impact of data range differences on training. The dataset is proportionally divided into training, validation, and test sets. The training set is used for parameter learning, the validation set is used to monitor overfitting and adjust hyperparameters, and the test set is used to evaluate model performance. The model is trained using the Adam optimizer, with a decaying learning rate and root mean square error loss function, as shown in the formula: In the formula The loss value. For the sample size, These are actual observed values. The model predicts the values, and through iterative training, the loss values ​​converge, ultimately yielding a proxy model of the watershed mechanism.

[0049] S3. Based on the driving relationship between input and output variables described by the watershed mechanism proxy model, and combined with the sensor network spatial topology of the agricultural non-point source pollution monitoring network, a dynamic cross-domain event graph is constructed. In the dynamic cross-domain event graph, the nodes represent sensor or hydrological events, and the edges represent the causal relationship between nodes and the expected time delay.

[0050] Specifically, this step aims to visually present the relationships between multi-domain data through a dynamic cross-domain event graph, providing a logical framework for subsequent anomaly tracing and verification. First, the definitions of nodes and edges in the graph are clarified. Nodes include sensor nodes and hydrological event nodes. Sensor nodes correspond to various heterogeneous sensors in the monitoring network, recording the sensor's monitoring parameter types and installation location information. Hydrological event nodes correspond to key events in agricultural non-point source pollution transmission (such as runoff events, confluence events, etc.), determining the occurrence and duration of events based on monitoring data and a simplified mechanism module.

[0051] The edges of the graph need to reflect the causal relationships and expected time delays between nodes. The construction process combines the driving relationships of the watershed mechanism surrogate model with the spatial topology of the sensor network. Based on the watershed mechanism surrogate model, the driving relationship between input variables (upstream parameters) and output variables (downstream parameters) can be determined. That is, changes in upstream parameters will drive downstream parameters to respond. This driving relationship corresponds to the direction of causal influence between nodes in the graph. At the same time, combined with the spatial topology of the sensor network, the spatial distribution of sensors and the association with watershed sub-units are analyzed to determine the spatial association strength between nodes. The closer the spatial distance between nodes and the more closely related they are to their respective watershed sub-units, the greater the weight of the causal relationship.

[0052] To quantify the expected time lag between nodes, lag correlation analysis is used to calculate the correlation coefficient of the time series data of corresponding parameters of two nodes at different time lags. The time lag corresponding to the maximum correlation coefficient is taken as the expected time lag between nodes. The formula is as follows:

[0053]

[0054] In the formula, For time lag The correlation coefficient below, For upstream node parameters in time The observed values, For downstream node parameters in time The observed values, The mean of the parameters of the upstream nodes. This represents the average value of the parameters of the downstream nodes. The number of samples in the time series data; by traversing different Value, find the The largest This refers to the expected time lag.

[0055] The dynamic cross-domain event map has the ability to update in real time. As new monitoring data is generated, the weights of causal influence relationships and expected time delays between nodes are recalculated periodically, and the edge attributes of the map are adjusted. At the same time, when a new hydrological event is detected, the corresponding hydrological event node is added in a timely manner, and its association with other nodes is determined based on the mechanism relationship and spatial topology, so as to ensure that the map can reflect the pollution transmission association status within the watershed in real time.

[0056] S4. When any sensor data is detected as abnormal, reverse tracing is performed along the causal influence edge in the dynamic cross-domain event graph to obtain the upstream node data; the upstream node data is used to drive the watershed mechanism proxy model to calculate the theoretical expected value range of the current abnormal data.

[0057] Specifically, this step provides a basis for subsequent consistency verification by tracing the source of anomalies and calculating theoretical expected values. First, an anomaly detection mechanism for sensor data is established, using a comprehensive detection method based on statistical methods and time-series data trend analysis. When sensor data exceeds the normal statistical range or its trend deviates significantly from historical patterns, the data is determined to be abnormal.

[0058] Upon detecting abnormal data, reverse tracing is performed in the dynamic cross-domain event graph. Reverse tracing starts from the sensor node corresponding to the abnormal data and searches for related nodes in the reverse direction of the causal influence edges in the graph (i.e., tracing from downstream nodes to upstream nodes). The tracing process must follow the principle of "hierarchical tracing + validity screening." Hierarchical tracing refers to tracing according to the weight of the causal influence relationship from high to low, prioritizing the acquisition of upstream node data with a greater impact on the abnormal node, and avoiding interference from irrelevant node data. Validity screening involves verifying the completeness and rationality of the traced upstream node data, removing upstream node data with a missing data rate exceeding a set proportion or that has been marked as abnormal, and retaining only upstream node data that is in normal condition and has complete data as valid input. The tracing terminates when: tracing reaches a basic driving node with no upstream related nodes (such as a meteorological station rainfall monitoring node, which has no upstream driving nodes), or tracing has reached the preset maximum tracing level (determined based on the length of the watershed pollution transmission path).

[0059] After obtaining valid upstream node data, it undergoes the same standardization process as in step S1 to ensure that the data format matches the input requirements of the watershed mechanism proxy model. The standardization process uses the Z-score standardization method, with the following formula: In the formula For standardized upstream node data, The original data of the upstream node, This represents the overall mean of upstream node data of this type. This represents the overall standard deviation of upstream node data for this type. The standardized upstream node data is then concatenated according to the input time window requirements of the watershed mechanism proxy model to form a continuous input data sequence, thus covering the cumulative impact of upstream parameters on downstream anomalous node data during pollution transmission.

[0060] The processed upstream node data is used to drive a watershed mechanism proxy model to calculate the theoretical expected value of the current abnormal data. To reflect the uncertainty of the data and the model prediction error, the range of the theoretical expected value can be further calculated. The range boundary is determined in the form of "predicted value ± confidence interval", expressed as: ,in The theoretical expected value output by the surrogate model of watershed mechanism. This represents the quantiles corresponding to the confidence level under the standard normal distribution (the confidence level is determined based on the prediction accuracy requirements during model training). Let be the standard deviation of the model's predictions under the current input conditions. The standard deviation of predictions is obtained through statistical analysis of the prediction error on the validation set during model training. Specifically, it involves calculating the standard deviation of multiple predictions under the same input features in the validation set, forming a prediction standard deviation lookup table. During actual calculation, the corresponding standard deviation is queried based on the data features of the current upstream node. This ensures that the range of theoretically expected values ​​can cover reasonable fluctuations in model predictions.

[0061] S5. By comparing the actual observed values ​​with the theoretical expected values, cross-domain logical consistency verification and fault conflict location are performed to obtain abnormal sensor data that are judged to be logically conflicting and the conflict is located as a sensor fault.

[0062] Specifically, this step distinguishes the nature of abnormal data (real pollution events or sensor failures) through cross-domain logical verification and accurately locates the source of the failure. The core is to verify the consistency of data based on the physical logic of watershed pollution transmission.

[0063] Cross-domain logical consistency verification uses the degree of matching between actual observed values ​​and theoretically expected values ​​as the core criterion. In practice, it first calculates the deviation parameter between the actual observed values ​​and the theoretically expected values, and then uses a combination of two parameters—deviation rate and exceedance magnitude—for judgment. The deviation rate is calculated using the following formula: In the formula The deviation rate, The actual observed values ​​of the sensor's abnormal data. This is the theoretical expected value. The formula for calculating the excess range is:

[0064]

[0065] In the formula, To exceed the expected range, when the actual observed value is within the theoretical expected value range, When the actual observed value exceeds the range, It is the sum of the absolute values ​​of the differences between the actual observed values ​​and the range boundary.

[0066] Based on deviation rate and exceeding the range Set consistency determination rules: If (If the actual observed value is within the theoretical expected value range), it is judged as "logically consistent," and the abnormal data likely corresponds to a real pollution event (further confirmation is needed based on upstream node data, such as multiple upstream nodes showing pollution-driven characteristics); if and If the deviation threshold determined based on historical model errors is exceeded, it is judged as a "logical conflict", indicating that there is a contradiction between the actual observation value and the expected result based on the physical mechanism, and further fault conflict localization is required.

[0067] The fault conflict localization adopts a three-step method: "upstream verification + intra-regional comparison + time-delay verification". Step 1: Upstream verification. Check for anomalies in the upstream node data obtained through reverse tracing. If the upstream node data is normal and the theoretical expected value calculation process is error-free, the anomaly caused by upstream driving factors is ruled out, and the current sensor is initially identified as faulty. If the upstream node data is abnormal, it is necessary to trace back to even more upstream nodes to determine if it is a chain reaction of upstream anomalies (e.g., anomalies in upstream soil sensors causing deviations in downstream water quality predictions). In this case, the upstream abnormal data must be repaired first, the theoretical expected value range recalculated, and then consistency verification performed. Step 2: Intra-regional comparison. Select data from other sensors (if they exist) belonging to the same watershed sub-unit as the abnormal sensor and monitoring the same parameters, and compare their data trends with those of the abnormal sensor data. If other sensor data in the same region are normal and the trend is stable, and only the current sensor data is abnormal, this further corroborates the sensor fault. If multiple sensor data in the same region are abnormal, it is necessary to combine dynamic cross-regional event maps to check for common upstream driving anomalies (e.g., uniform rainfall anomalies within the watershed sub-unit) to rule out regional real pollution events. The third step is time delay verification: Based on the expected time delay between nodes in the dynamic cross-domain event graph, check whether the occurrence time of abnormal data matches the driving time of upstream node data. If the occurrence time of abnormal data is earlier than the reasonable time delay range of upstream driving data (e.g., water quality abnormality is earlier than rainfall runoff event), it violates the time delay logic of pollution transmission and is judged as sensor failure; if the time matches and conforms to the time delay relationship, the rationality of the theoretical expected value range needs to be re-evaluated (e.g., whether special pollution transmission paths are omitted).

[0068] Through the above three steps of localization, abnormal data that is "logically conflicting and without upstream driver anomalies, with normal data in the same domain and time delay mismatch" is finally filtered out. Such data is judged as sensor abnormal data with "conflict localization as sensor failure", providing a clear target for subsequent repair.

[0069] S6. Based on abnormal sensor data, establish an optimized repair model that integrates physical law constraints and the mechanism relationship described by the watershed mechanism proxy model; generate self-consistent repair values ​​that conform to the pollution transmission law by solving the optimized repair model to replace the abnormal sensor data.

[0070] Specifically, the core of this step is to construct an optimized repair model that takes into account both physical rationality and mechanistic consistency, ensuring that the repaired data can not only meet the basic laws of pollution transmission, but also maintain logical consistency with upstream and downstream data, thus avoiding data distortion caused by traditional repair methods.

[0071] The objective function of the optimized repair model focuses on the consistency between the repair value and the expected mechanism, while also considering the continuity between the repair value and the time series. The formula can be:

[0072]

[0073] In the formula, The objective function value (which needs to be minimized). This is a repair value for abnormal sensor data. The theoretical expected value output by the surrogate model of watershed mechanism. , , These are the mechanistic consistency weight, the forward temporal continuity weight, and the backward temporal continuity weight (the sum of these three is 1, determined based on the verification effect of historical repair data, which can...). (Highest weight, ensuring mechanism priority) The time series window containing the abnormal data (including the normal data time points adjacent to the abnormal data). , These are the normal observations at the time points preceding and following the anomalous data, respectively. In the objective function, the first term ensures that the restored value closely matches the expected mechanism, while the second and third terms ensure a smooth transition between the restored value and the normal data before and after the time series, avoiding abrupt changes.

[0074] Constraints are divided into two categories: physical law constraints and mechanistic logic constraints. Physical law constraints are based on the fundamental physical laws of agricultural non-point source pollution and mainly include: 1) Concentration non-negativity constraints: 1) Because parameters such as pollutant concentration and soil moisture cannot be negative; 2) Mass conservation constraint: If the abnormal data is a pollutant flux parameter, it must satisfy the following conditions. ( For runoff volume, For upstream pollutant input flux, (For downstream pollutant output flux), ensuring that the flux after remediation conforms to the law of mass conservation; 3) Rate of change constraint: and , This is the maximum reasonable rate of change for this type of parameter (obtained based on the statistical analysis of the temporal variation of historical normal data, reflecting the gradual nature of pollutant migration and avoiding abrupt changes in the repair value).

[0075] Mechanistic logic constraints are set based on the watershed mechanism proxy model and dynamic cross-domain event graph, mainly including: 1) Upstream and downstream consistency constraints: The repair value... As input to downstream nodes, the expected values ​​for downstream nodes are calculated using the watershed mechanism surrogate model. This requires that the actual observed values ​​(if normal) of the downstream nodes fall within this expected range. The formula is as follows: ( This is the expected value for downstream nodes calculated based on the repair value. These are normal observations for downstream nodes. 1) The predicted standard deviation of the expected value of the downstream node); 2) Time lag matching constraint: The time point corresponding to the repair value must satisfy the expected time lag in the dynamic cross-domain event graph with the driving time of the upstream node data, that is For the repair value time point, For the upstream node data time point, (with expected time lag), ensuring that the repair value conforms to the time logic of pollution transmission.

[0076] The solution to the optimized repair model is obtained using a gradient descent-based numerical optimization algorithm. The specific steps are as follows: 1) Initialize the repair values. (Initial value set as theoretical expected value) ); 2) Calculate the objective function value corresponding to the current repair value. 3) Update the repair value along the gradient descent direction of the objective function, with the following formula: (The formula is missing from the original text.) ( For the number of iterations, For learning rate, For the objective function in 4) Check if the updated repair value satisfies all constraints. If not, map it to the constrained feasible region using the projection method; 5) Repeat steps 2-4 until the objective function value converges (the gradient at the point of intersection); 6) Repeat steps 2-4 until the objective function value converges (the gradient of two adjacent iterations). The result is obtained when the difference is less than a set threshold or when the maximum number of iterations is reached. This is the final self-consistent repair value.

[0077] Finally, the effectiveness of the repaired values ​​is verified by substituting them into the dynamic cross-domain event graph and watershed mechanism proxy model to check whether they can form a complete logical closed loop with upstream and downstream data (e.g., the repaired water quality data can be reasonably driven by upstream rainfall and soil data, and can also reasonably drive downstream data). After verification, the original abnormal data is replaced to complete the data repair.

[0078] The aforementioned method for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors first standardizes the heterogeneous monitoring data and constructs a watershed proxy model that integrates physical mechanisms. Based on this, a dynamic event graph characterizing cross-domain causal relationships is established. When data anomalies are detected, the system uses this graph to trace the source and drives the proxy model to generate theoretical expected values. Logical consistency verification is performed through comparison to distinguish between real pollution events and sensor malfunctions. For data determined to be faulty, an optimization model integrating physical constraints is further established and solved to generate self-consistent repair values ​​that conform to the pollution transmission law. This achieves a paradigm shift in quality control from isolated threshold judgment to cross-domain logical verification, significantly improving the accuracy of identifying complex pollution events, the reliability of sensor fault diagnosis, and the physical rationality of abnormal data repair results.

[0079] refer to Figure 2 In one optional embodiment, a simplified mechanism module is constructed based on the physical mechanism of agricultural non-point source pollution transmission. Based on the simplified mechanism module, a machine learning model is trained using historical monitoring data to obtain a watershed mechanism proxy model describing the response of water quality parameters from upstream rainfall and soil characteristics to downstream water quality parameters. This includes the following steps:

[0080] S11. Based on the principle of full storage and runoff generation, a simplified hydrological response module is constructed to describe the transformation process from rainfall to soil water and then to surface runoff. The simplified hydrological response module uses a water tank model to simulate the change in soil water storage and integrates a simplified form of the SCS curve number method to calculate surface runoff.

[0081] Specifically, this step constructs a simplified hydrological response module based on the principle of full water storage and runoff generation. The principle of full water storage and runoff generation refers to the fact that after the soil moisture content reaches the field capacity, all subsequent rainfall is converted into surface runoff or interflow. This module integrates the water tank model with the SCS curve number method to realize the physical description of the transformation process of rainfall-soil water-surface runoff, while taking into account both computational efficiency and mechanism rationality.

[0082] The water tank model is used to simulate the dynamic changes in soil water storage. A single-layer water tank structure simplifies the vertical movement of soil moisture. Its core formula is:

[0083]

[0084] In the formula, Let be the soil water storage at time t. Let be the soil water storage at time t-1. Let be the rainfall at time t. Let be the amount of infiltration at time t. Evapotranspiration at time t. Infiltration. A phased calculation method is adopted, which, combined with the characteristics of full-storage runoff, divides the process into a non-full-storage stage and a full-storage stage: when ( When the soil's maximum water storage capacity (corresponding to field capacity) is reached, the infiltration rate follows a simplified form of the Horton infiltration curve, as shown in the formula: In the formula To stabilize the infiltration rate (close to the soil saturation hydraulic conductivity). Let k be the initial infiltration rate, and k be the infiltration attenuation coefficient; when At that time, the infiltration rate is only the stable infiltration rate. Excess water is the primary source of surface runoff, which is generated through evapotranspiration. The calculation is performed using the Thornthwaite simplified formula, which correlates air temperature with soil moisture correction factors. The formula is as follows: In the formula Potential evapotranspiration This is the soil moisture correction factor, reflecting the limiting effect of soil moisture content on evapotranspiration.

[0085] The simplified form of the SCS curve number method directly correlates soil water storage with surface runoff. Its core is the dynamic adjustment of the curve number (CN value) to reflect the impact of soil moisture changes on runoff potential. In the traditional SCS model, the CN value is a fixed value. This module calculates soil water storage (deficit) using the soil water storage output from the tank model. The formula is as follows: In the formula, S represents the maximum possible retention capacity of the watershed, and CN represents the curve number. To achieve integration with the water tank model, the CN value is dynamically updated according to the soil water storage capacity, as shown in the formula: In the formula Let CN be the dynamic value at time t. This represents the minimum CN value under conditions of soil drought. This represents the maximum CN value when the soil is saturated. Given the current soil water storage deficit, substituting the dynamic CN value into the SCS runoff calculation formula yields a simplified form, for example:

[0086]

[0087] In the formula, Let be the surface runoff at time t. This represents the maximum possible retention capacity calculated based on the CN value at time t. This integrated approach retains the ability of the water tank model to simulate dynamic changes in soil moisture while achieving rapid runoff calculation through a simplified form of the SCS curve number method. It meets the dual requirements of real-time performance and mechanistic integrity in agricultural non-point source pollution monitoring. 0.2 and 0.8 are empirical coefficients that can be adjusted according to actual conditions.

[0088] S12. Based on empirical relationships of pollutant migration, a simplified pollution migration module is constructed to describe the migration and transformation process of pollutant concentration from surface runoff to river channels. The simplified pollution migration module uses a nonlinear function with time delay parameters to simulate the flushing and decay process of pollutant concentration after the runoff peak.

[0089] Specifically, this step constructs a simplified pollution migration module based on the empirical law of "scouring-transportation-attenuation" of pollutants migrating with surface runoff. The core of this module is to quantify the dynamic relationship between runoff drive and pollutant concentration response through a nonlinear function with time delay parameters, focusing on characterizing the changes in pollutant concentration after the runoff peak.

[0090] During pollutant migration, the scouring intensity of surface runoff on soil pollutants is positively correlated with the runoff volume. However, after entering the river channel, the pollutant concentration gradually decreases due to dilution and degradation. Furthermore, there is a time lag between the peak runoff concentration and the peak pollutant concentration. Therefore, the nonlinear function needs to integrate three core elements: scouring intensity, decay rate, and time lag. The function form adopts a piecewise construction method, covering both the runoff rise and decay stages. The overall formula is:

[0091]

[0092] In the formula, Let be the concentration of pollutants in the river channel at time t. The time of peak runoff This is a time lag parameter (the time difference between peak runoff and peak pollutant concentration). For lag Surface runoff over time The scour coefficient (reflects the ability of runoff to carry pollutants, and is related to soil texture and pollutant adsorption characteristics). This is the runoff-scour nonlinear coefficient (reflecting the nonlinear relationship between scour intensity and runoff volume). It is the pollutant attenuation coefficient (reflecting the rate of loss due to biodegradation, chemical transformation, etc.).

[0093] For this function, when time At that time, the pollutants had not yet been washed into the river or the concentration was below the detection limit, with a concentration value of 0; when At that time, the first part of the function In simulating the scouring process, the larger the runoff and the higher the scouring coefficient, the faster the pollutant concentration rises, and the nonlinear coefficient... This demonstrates the accelerating effect of scouring intensity on increasing runoff; Part Two The simulated decay process shows that the concentration decreases exponentially over time, and the decay coefficient... The larger the value, the faster the pollutant is lost. Time delay parameter The introduction of this parameter represents a key modification to the physical processes of pollution migration. Its magnitude is related to watershed slope and runoff velocity; the steeper the slope and the faster the velocity, the better. The smaller the value, the more consistent it is with the time lag pattern of pollutant migration with runoff. This nonlinear function can accurately capture the typical change curve of pollutant concentration after the runoff peak, which is "rapidly rising to the peak value and then slowly decaying to the background value", providing a physical constraint framework for pollutant migration for subsequent mechanism proxy models.

[0094] S13. Extract multi-source features from the standardized time-series feature dataset, including rainfall event pulse features, soil moisture steady-state features, and water quality parameter steady-state features from the historical sequence. Construct a training sample set based on the multi-source features. The input features of the training sample set are multi-source features, and the output label of the training sample set is the water quality feature at a specific time lag in the future.

[0095] Specifically, this step transforms the raw time-series data into structured features that the model can recognize through multi-source feature extraction. The core is to screen feature variables that can reflect the key driving factors and response patterns of pollution transmission, providing high-quality input-output data pairs for machine learning model training.

[0096] The extraction of rainfall event pulse characteristics focuses on the driving effect of rainfall on runoff generation. First, a sliding window method is used to identify independent rainfall events (where the interval between adjacent rainfall events exceeds a preset no-rainfall threshold). Then, features are extracted for individual rainfall events: 1) Rainfall pulse peak characteristics, including peak rainfall intensity (maximum instantaneous rainfall intensity within the event) and peak duration (the length of time the peak rainfall intensity is maintained); 2) Pulse morphology characteristics, including pulse rise rate (the average rate of change from the initial rainfall intensity to the peak), pulse fall rate (the average rate of change from the peak to the final rainfall intensity), and pulse asymmetry coefficient (the ratio of the rise phase duration to the fall phase duration); 3) Cumulative effect characteristics, including total rainfall, cumulative duration of rainfall intensity exceeding a certain threshold, and rainfall kinetic energy (calculated through the empirical relationship between rainfall intensity and raindrop spectrum). These features can quantify the impact of the intensity, duration, and morphology of rainfall events on soil water transformation and pollutant scouring.

[0097] Soil moisture steady-state characteristics are extracted based on the principle of water storage and runoff generation. For soil moisture time-series data before and after rainfall events, a "rate of change threshold method" is used to determine steady state: when the absolute value of the soil moisture change rate at multiple consecutive time points is less than a set threshold, steady state is considered to have been reached. The extracted features include: 1) pre-rainfall steady-state value (mean of the last steady-state period before the event); 2) post-rainfall steady-state value (mean of the first steady-state period after the event); 3) water storage response characteristics, including the time to reach field capacity (from the start of rainfall to soil moisture reaching...). The characteristics of steady-state conditions include: 1) duration of rainfall; 2) magnitude of humidity increase (difference between steady-state value before and after rainfall); 3) steady-state maintenance characteristics, including the duration of steady-state before rainfall and the duration of steady-state after rainfall. These characteristics directly reflect the response of soil moisture to rainfall and are key intermediate variables connecting rainfall and runoff.

[0098] The extraction logic for steady-state characteristics of water quality parameters is consistent with that for soil moisture, focusing on changes in water quality parameters before and after a pollution event. This includes: 1) background steady-state values ​​(the steady-state mean of water quality parameters before the pollution event); 2) peak response characteristics, including peak pollutant concentrations, peak occurrence time, and lag time of the peak relative to the peak runoff; 3) recovery steady-state characteristics, including the time to recover from the peak concentration to the background steady-state value and the average decay rate during the recovery phase; and 4) fluctuation characteristics, including the standard deviation and coefficient of variation of water quality parameters during the event (the ratio of the standard deviation to the background steady-state value). These characteristics comprehensively characterize the response patterns and recovery capabilities of water quality parameters to upstream driving factors.

[0099] The construction of the training sample set requires precise matching between input features and output labels: the input features are a set of the aforementioned rainfall event pulse features and soil moisture steady-state features, which are concatenated to form a high-dimensional feature vector; the output label is the water quality feature at a specific time lag in the future. The specific time lag is determined through lag correlation analysis, i.e., calculating the Pearson correlation coefficient of the water quality parameter corresponding to the input feature at different time lags, and taking the lag time corresponding to the largest correlation coefficient as the specific time lag. The output label is specifically the water quality parameter concentration value or concentration peak at that time lag. During the construction process, the samples need to be screened for validity, removing samples with incomplete rainfall events or where soil moisture or water quality parameter steady-state cannot be identified. The final training sample set is stored in a structured format of "input feature vector - output label" to provide data support for subsequent model training.

[0100] S14. The gradient boosting decision tree algorithm is adopted to jointly train and fit the undetermined parameters in the simplified hydrological response module and the simplified pollution migration module with the training sample set, so as to obtain a machine learning model for predicting the water quality characteristics of the downstream based on the multi-source characteristics of the upstream. The machine learning model is used as a proxy model for watershed mechanism.

[0101] Specifically, this step uses the Gradient Boosting Decision Tree (GBDT) algorithm to integrate "mechanism module parameter optimization + data-driven modeling". The core is to use machine learning algorithms to jointly fit the undetermined parameters in the simplified hydrological response module and the pollution migration module, taking into account both physical mechanism constraints and data fitting accuracy, and finally forming a watershed mechanism proxy model.

[0102] The core principle of the GBDT algorithm is to iteratively construct multiple decision trees, with each new tree fitting the residuals of the preceding model. The loss function is minimized via gradient descent, thus achieving a non-linear mapping from input features to output labels. This algorithm has advantages in handling high-dimensional features and capturing feature interaction effects. Furthermore, by setting parameter constraints, it can incorporate physical mechanisms, making it suitable as the core algorithm for mechanism proxy models. Undetermined parameters in the simplified module include: In the simplified hydrological response module, maximum soil water storage capacity is included. Initial infiltration rate Infiltration attenuation coefficient Dynamic CN value upper and lower limits and The simplified pollution migration module includes a scour coefficient. Runoff-scour nonlinear coefficient Pollutant attenuation coefficient Time delay parameters These parameters, as the parameters to be optimized in the GBDT model, must have values ​​that conform to physical meaning (e.g., , ).

[0103] The joint training process is divided into three stages: parameter initialization, model iterative training, and hyperparameter tuning. In the parameter initialization stage, based on fundamental data such as watershed soil texture and land use types, and referring to empirical parameter values ​​from similar studies, initial values ​​for undetermined parameters are set. For example, different soil types are determined according to soil texture handbooks. Initial values ​​are set according to the type of pollutant. The initial range. During the model iterative training phase, the input feature vector of the training sample set is input into the GBDT model. The output label (water quality characteristics at a specific future time lag) is the target, and joint parameter fitting is achieved through the following steps: 1) Run the simplified hydrological response module and pollution migration module with the initial parameter values ​​to obtain the mechanism prediction value; 2) Calculate the residual between the mechanism prediction value and the sample output label, and use the residual as the fitting target for the first decision tree of GBDT; 3) Construct the decision tree fitting residual and update the model prediction value (initial mechanism prediction value + prediction value of the first tree); 4) Fix the GBDT model parameters, and adjust the undetermined parameters in the simplified module based on the new prediction value (e.g., use the least squares method to make the module output closer to the GBDT prediction value); 5) Repeat steps 2-4, iteratively adding decision trees and adjusting the mechanism module parameters until the loss function converges. The loss function uses the root mean square error (RMSE), and the formula is... In the formula, Loss is the loss value, and N is the number of samples. Let i be the output label of the i-th sample. These are the model predictions (outputs of the fusion mechanism module and GBDT).

[0104] During the hyperparameter tuning phase, K-fold cross-validation is used to adjust key hyperparameters of GBDT, including the number of decision trees, tree depth, learning rate, and minimum number of samples per leaf node. The specific process is as follows: the training sample set is divided into K subsets, and K-1 subsets are selected sequentially as the training set and one subset as the validation set. The model is trained under different hyperparameter combinations, and the validation set loss is calculated. The hyperparameter combination with the minimum validation set loss is selected as the optimal setting. Simultaneously, parameter constraints are added during the tuning process, limiting the range of values ​​for undetermined parameters through regularization terms, such as adding an L2 regularization term to the loss function. In the formula The loss value is the one with regularization. The regularization coefficient is . To simplify the module's undetermined parameter set, Indicates traversing a collection Each undetermined parameter Ensure that the parameter values ​​are within a reasonable range.

[0105] After training, the resulting model integrates the physical constraints of the simplified mechanism module with the data fitting capability of the GBDT algorithm. It can accurately predict downstream water quality characteristics at specific time lags based on the pulse characteristics of upstream rainfall events and the steady-state characteristics of soil moisture. This model is the watershed mechanism proxy model. Finally, the model performance needs to be evaluated using an independent test set, with evaluation metrics including the coefficient of determination (COP). The model incorporates methods such as mean absolute error (MAE) to ensure good generalization ability and prediction accuracy, which can be used for subsequent verification and repair of abnormal data.

[0106] In one optional embodiment, a dynamic cross-domain event graph is constructed based on the driving relationship between input and output variables described by the watershed mechanism proxy model, combined with the sensor network spatial topology of the agricultural non-point source pollution monitoring network, including the following steps:

[0107] S21. Define each type of sensor at each monitoring point of the agricultural non-point source pollution monitoring network as a sensor node, and define each rainfall event exceeding the threshold as an independent event node. The set of nodes in the dynamic cross-domain event map is composed of all sensor nodes and event nodes.

[0108] Specifically, sensor nodes are defined by a unique combination of "monitoring point - sensor type". Each monitoring point corresponds to a fixed monitoring location within the watershed (such as a mountaintop meteorological monitoring point or a watershed outlet water quality monitoring point). Each type of sensor corresponds to a specific monitoring parameter (such as a soil moisture sensor or an ammonia nitrogen concentration sensor). Each sensor node needs to carry basic attribute information, including monitoring point number, sensor type, monitoring parameter name, installation latitude and longitude, and sampling frequency. These attributes are used for subsequent node association and data traceability.

[0109] For each event node, a threshold for the rainfall event is first determined. This threshold is set based on the critical conditions for runoff generation in the watershed, typically taking the minimum rainfall amount (or rainfall intensity) of the multi-year average runoff generation per rainfall event in the watershed. This ensures that rainfall events exceeding this threshold can effectively drive subsequent soil water, runoff, and pollutant migration processes. Each rainfall event that meets the threshold condition is defined as an independent event node. The attributes of an event node include event number, rainfall start time, rainfall end time, total rainfall amount, and maximum rainfall intensity. These attributes are used to characterize the intensity and duration of the rainfall event, providing a basis for subsequent causal association analysis.

[0110] Finally, all sensor nodes and event nodes are organized into a node set in the format of "unique identifier-attribute set". Each node is assigned a unique identifier code (e.g., sensor node code is "monitoring point number-sensor type code", event node code is "event occurrence date-serial number") to ensure the uniqueness and identifiability of nodes in subsequent map construction.

[0111] S22. Based on the sub-basin confluence relationship determined by the watershed digital elevation model, establish spatial topological edges representing the water flow transmission path between sensor nodes in the upstream and downstream sub-basins.

[0112] Specifically, this step determines the spatial association of sensor nodes based on the watershed topographic features. The core is to extract the sub-watershed confluence relationships based on the digital elevation model (DEM) and then construct spatial topological edges that reflect the water flow transmission path. First, the DEM data is preprocessed, including filling depressions (eliminating local concave areas to avoid errors in water flow direction judgment) and peak smoothing (removing isolated elevation anomalies) to ensure that the DEM can accurately reflect the watershed topographic undulations.

[0113] Based on the preprocessed DEM, the flow direction of each grid is determined by a flow direction algorithm (such as the D8 algorithm). Then, the flow accumulation of each grid is calculated by a flow accumulation algorithm. A flow accumulation threshold is set (determined based on the watershed area and river development characteristics). Grids with flow accumulation exceeding the threshold are divided into rivers, thereby segmenting multiple independent sub-watersheds and obtaining the flow relationship between the sub-watersheds. That is, the water flow of the upstream sub-watershed flows into the downstream sub-watershed through the river, forming a hierarchical flow structure of "upstream → midstream → downstream".

[0114] Based on the confluence relationship between sub-basins, spatial topological edges are established between sensor nodes in different sub-basins: if the water flow in the sub-basin where sensor node A is located eventually flows into the sub-basin where sensor node B is located, then a spatial topological edge from A to B is established between A and B. The attributes of the edge include the transmission path length (the shortest distance between the two sub-basins based on DEM calculation) and the average slope (the average slope of the terrain within the path range). These attributes are used to help determine the time and speed of water flow transmission. The establishment of spatial topological edges must ensure that all sensor nodes in sub-basins with confluence relationships are covered, and fully reflect the spatial correlation driven by water flow.

[0115] S23. Based on the physical driving relationship described by the watershed mechanism proxy model, establish causal influence edges between sensor nodes and event nodes that have causal relationships; wherein each causal influence edge has an initial time delay parameter representing the time required for the cause to lead to the result.

[0116] Specifically, the core of this step is to determine the causal relationships between nodes based on physical mechanisms and assign initial time delay parameters to reflect the time lag characteristics of pollution transmission. First, physical driving relationships are extracted from the watershed mechanism proxy model. These relationships clarify the causal logic of "input variable → output variable". For example, rainfall events (input) drive changes in soil moisture (output), changes in soil moisture (input) drive changes in surface runoff (output), and changes in surface runoff (input) drive changes in water quality parameters (output).

[0117] Based on the above physical driving relationship, causal influence edges are established between the corresponding nodes: between the rainfall event node (cause node) and the soil moisture sensor node (result node), between the soil moisture sensor node (cause node) and the hydrological sensor node (result node), and between the hydrological sensor node (cause node) and the water quality sensor node (result node), causal influence edges are established from the cause node to the result node.

[0118] Each causal influence edge is assigned an initial time delay parameter, which is determined based on the prediction results of the watershed mechanism surrogate model. This parameter is calculated by averaging the time interval between changes in the causal node parameter and the response of the result node parameter, using the average of multiple simulation results. For example, the average time from the occurrence of a simulated rainfall event to the start of significant changes in soil moisture is used as the initial time delay parameter for the causal influence edge between the rainfall event node and the soil moisture sensor node; similarly, the average time from the occurrence of changes in surface runoff to the start of changes in water quality parameters is used as the initial time delay parameter for the causal influence edge between the hydrological sensor node and the water quality sensor node. The unit of the initial time delay parameter is consistent with the sensor sampling frequency (e.g., minutes, hours).

[0119] S24. During the online operation of the agricultural non-point source pollution monitoring network, the data of the nodes at both ends of the causal influence edge are continuously monitored. By analyzing the sequential relationship of the node state changes, the initial time delay parameters of each causal influence edge are dynamically adjusted and updated to obtain the actual time delay parameters.

[0120] Specifically, this step optimizes the time lag parameters by monitoring data in real time to ensure that the causal influence edge accurately reflects the time lag pattern of actual pollution transmission. First, a continuous monitoring cycle is set (synchronized with the sensor sampling frequency, such as every 15 minutes or every hour). Within each monitoring cycle, real-time data of the nodes at both ends of the causal influence edge is collected, with a focus on the "state change moment" of the node parameters, that is, the time point when the parameters start to change significantly from a stable state (the change exceeds a preset threshold, such as 10% of the average).

[0121] For each causal influence edge, record the state change time of the cause node (denoted as t1) and the state change time of the result node (denoted as t2). Calculate the single time difference Δt = t2 - t1. If Δt is positive (consistent with the logic that cause precedes result), it is included in the time delay statistics sample. If Δt is negative (possibly due to sensor failure or data transmission delay), the abnormal sample is removed to avoid interfering with parameter adjustment.

[0122] The statistical samples are processed periodically (e.g., daily, weekly, determined based on the amount of accumulated data). The time-delay parameters are updated using either the moving average method (taking the average of the most recent N valid Δt values) or the median method (taking the median of the N valid Δt values). The updated parameters are then used as the actual time-delay parameters. For example, for the causal relationship between rainfall event nodes and soil moisture sensor nodes, all valid Δt values ​​within that week are collected weekly, and the average value is calculated as the new actual time-delay parameter, replacing the original initial time-delay parameter. This ensures that the time-delay parameters can be dynamically adjusted according to seasonal changes (such as time-delay changes caused by differences in soil infiltration capacity between rainy and dry seasons) and changes in the watershed underlying surface (such as changes in vegetation cover), thus conforming to the actual monitoring scenario.

[0123] S25. By combining sensor nodes, event nodes, spatial topology edges, causal influence edges, and the actual time delay parameters corresponding to each causal influence edge, a dynamic cross-domain event graph is constructed.

[0124] Specifically, these elements are organized according to a graph structure and stored using a graph database (such as Neo4j). Nodes are stored in the form of "unique identifier-attribute", and edges are stored in the form of "start point identifier-end point identifier-edge type-attribute" (edge ​​types are divided into "spatial topology edge" and "causal influence edge", and the attributes correspond to the path length and average slope of the spatial topology edge, and the actual time delay parameter of the causal influence edge, respectively).

[0125] Simultaneously, a real-time update mechanism for the event map is established. When a new sensor node is added to the monitoring network, the corresponding node is automatically added to the map, and spatial topological edges and causal influence edges are supplemented based on sub-basin confluence relationships and physical driving relationships. When a new rainfall event occurs, an event node is automatically created, and causal influence edges with downstream sensor nodes are established. When the actual time-delay parameters are updated, the attributes of the corresponding causal influence edges are modified synchronously. Through element integration and real-time updates, the final constructed dynamic cross-domain event map can completely and in real-time reflect the spatial association, causal logic, and time-delay characteristics of multi-domain nodes in the agricultural non-point source pollution monitoring network, providing a basic framework for subsequent anomaly tracing and consistency verification.

[0126] In one optional embodiment, during the online operation of the agricultural non-point source pollution monitoring network, the data of the nodes at both ends of the causal influence edge are continuously monitored. By analyzing the temporal relationship of the node state changes, the initial time delay parameters of each causal influence edge are dynamically adjusted and updated to obtain the actual time delay parameters. This includes the following steps:

[0127] S31. For a causal edge pointing from cause node A to result node B, continuously monitor the data of cause node A. When a change in the state of cause node A exceeding a first preset range is detected, record the start time of the corresponding change event. .

[0128] Specifically, first, clarify the data type of the monitoring node A, such as rainfall for the rainfall node and soil moisture for the soil sensor node. Then, determine the first preset range based on the historical normal operation data of this node. Statistical methods can be used to calculate the mean and standard deviation of the historical data of node A. The mean ± k × standard deviation can be used as the first preset range, which represents the fluctuation range of the normal operation status of node A. The value of k is set according to the stability of the node data.

[0129] During monitoring, a sliding time window is used to calculate the rolling average or instantaneous value of node A's data in real time and compare it with a first preset range. The size of the sliding time window is set according to the node's sampling frequency. When the data value exceeds the first preset range within a consecutive window period, or when a single data value exceeds the range and subsequent sampled data does not return to the range, it is determined that the state of node A has undergone a "change exceeding the first preset range." At this time, the timestamp of the first occurrence of this change is recorded, which is the start time of the change event. The core of this operation is to eliminate the interference of random fluctuations in node data, ensuring that the captured state changes are those with real causal driving significance, rather than spurious changes caused by random noise.

[0130] S32, at the start time The next maximum possible time delay window Internally, continuously monitor the data of result node B, and when result node B is detected... When the state changes beyond the second preset range, record the start time of the corresponding response event. This forms a pair of causal event observation pairs. ;in, This represents the maximum possible time delay.

[0131] Specifically, this step establishes a time correspondence between causal events and response events by setting a time window to lock the response period of the result node. Maximum possible time delay. The determination needs to combine the physical characteristics of the watershed and the spatial location of the nodes, based on the sub-watershed confluence path length and average surface runoff velocity extracted from the watershed digital elevation model, through... The base value is calculated as (confluence path length / average runoff velocity), then corrected by combining the longest response times of cause node A and result node B from historical monitoring data, to finally determine the value. The value of is chosen to ensure that the window range covers all possible response periods while avoiding the introduction of interference data from irrelevant time intervals.

[0132] The setting logic for the second preset range is consistent with that of the first preset range. It is obtained based on the historical normal data statistics of result node B, and the k value is adjusted according to the monitoring parameter type of node B. For example, the k value can be appropriately increased when the water quality parameter stability is low. Within the window, the same sliding window monitoring method as S31 is used to track the data changes of node B in real time. When the data of node B first exceeds the second preset range and meets the judgment conditions of "continuous window exceedance" or "no regression after single exceedance", the start time of the response event is recorded. Under the same causal influence edge and Pairing to form causal event observation pairs This observation directly reflects the time interval between cause and effect in a single causal event, providing a basic data unit for subsequent time delay parameter calculation.

[0133] S33. Collect a series of causal event observation pairs from history to form a causal event observation pair set; calculate different time delays based on the causal event observation pair set. The statistical association strength is obtained by calculating the association strength sequence; where mutual information is used as the statistical association strength, and the formula for calculating mutual information is:

[0134]

[0135] in, Indicates time delay Mutual information under the following Let A represent the set of all possible discrete states of cause node A. Let B represent the set of all possible discrete states of the result node B. Indicates that cause node A is at time [time]. The discrete state values, This indicates that the result node B is at time [time]. The discrete state values; express The marginal probability distribution, express The marginal probability distribution, express and The joint probability distribution.

[0136] Specifically, this step quantifies the correlation strength between causal and result nodes under different time lags using mutual information. The core is to transform discrete causal event observation pairs into statistically significant correlation indicators. First, the node states need to be discretized. Based on the physical meaning of the monitoring parameters of nodes A and B, their continuous data are divided into a finite number of discrete states. For example, rainfall is divided into "no rain," "light rain," "moderate rain," and "heavy rain," and water quality concentration is divided into "normal," "slightly exceeding the standard," and "severely exceeding the standard." Each discrete state corresponds to a value range, thus forming... and .

[0137] Collect causal event observation pairs covering different hydrological seasons to form an observation pair set. For each time delay to be calculated... , The value range is 0 to The step size is set according to the node sampling frequency, and "cause node A" is extracted from the set of observation pairs. status "and result node B in status "Statistics of all" The number of times a state combination occurs is used to calculate the marginal probability and joint probability.

[0138] Substituting the above probability values ​​into the mutual information formula for calculation ,when A larger value indicates a time lag. Below, the state of cause node A is strongly correlated with the state of result node B, meaning that the time delay better reflects the actual causal response relationship between the two; when A smaller value indicates a weaker correlation between the two, and the time lag does not reflect reality. This can be determined by iterating through all... Value, calculated Follow The changing correlation strength sequence provides data support for subsequently determining the optimal time delay.

[0139] S34. Finding mutual information Reaching maximum time delay Time lag This serves as the actual time delay parameter for the corresponding causal influence edge.

[0140] Specifically, the core of this step is to select the optimal time delay that best conforms to the causal response law from the correlation strength sequence, and use it as the actual time delay parameter of the causal influence edge. This is done after obtaining the correlation strength sequence (i.e., different...). corresponding After finding the value, an extreme value search method is used. To find the maximum value, first iterate through all... value corresponding Determine the maximum Values ​​and their corresponding values Value; if multiple values ​​exist The maximum value corresponding to the same Then calculate these The arithmetic mean of the values, as the final... .

[0141] Will The physical significance of determining the actual time delay parameter as the causal influence edge lies in the fact that the state change of the cause node A and the state change of the result node B are most strongly correlated under this time delay, and best reflect the true causal response time difference between the two. This is in contrast to the initial time delay parameter (set based on physical experience). Based on historical observation data, this data more closely reflects the actual operation of the monitoring network and the real pollution transport patterns in the watershed. In practical applications, steps S31-S34 are repeated periodically to update the set of causal event observation pairs and the correlation strength sequence, dynamically adjusting the data. This ensures that the time delay parameters can adapt to the effects of watershed hydrological conditions (such as runoff velocity changes caused by seasonal variations) or sensor status changes, and maintains the accuracy of dynamic cross-domain event maps.

[0142] In one optional embodiment, based on abnormal sensor data, an optimized remediation model is established that integrates physical constraints and the mechanistic relationships described by a watershed mechanism proxy model; by solving the optimized remediation model, a self-consistent remediation value conforming to the pollution transport law is generated to replace the abnormal sensor data, including the following steps:

[0143] S41. Identify the sensor corresponding to the abnormal data of the sensor to be repaired, and designate it as the abnormal sensor; determine the time of the fault occurrence and the abnormal value to be repaired corresponding to the abnormal data of the sensor to be repaired; based on the dynamic cross-domain event graph, query all other sensors that are directly connected to the abnormal sensor through causal influence edges at the time of the corresponding fault occurrence and whose data status is reliable, and designate them as target sensors; construct an associated observation vector based on the observation values ​​of the target sensors.

[0144] Specifically, firstly, key information is extracted from the sensor anomaly detection results, and the sensor that generates abnormal data is defined as an abnormal sensor, and its device identifier (such as sensor number and installation location) is directly associated with it. Then, the time of the fault occurrence (i.e., the timestamp of the first appearance of the abnormal value) and the abnormal value to be repaired (i.e., the specific data value that exceeds the normal range or has logical conflicts) are determined from the abnormal data record to ensure that the object to be repaired is accurately located.

[0145] The query process for dynamic cross-domain event graphs focuses on two core conditions: "directly connected causal influence edges" and "reliable data status." "Directly connected causal influence edges" means that there is a direct causal influence edge (not an indirect relationship) between the target sensor and the anomalous sensor in the graph. For example, if the anomalous sensor is a downstream water quality sensor, the upstream soil sensor and runoff sensor directly connected to it are potential target sensors. The observed values ​​of these sensors directly drive the theoretical values ​​of the anomalous sensor. The determination of "reliable data status" is based on the results of previous cross-domain logical consistency checks. That is, the data from the target sensor at the time of the fault must be determined to be "logically consistent" (the actual observed value is within the theoretical expected value range) and not marked as anomalous or faulty, ensuring the reliability of the associated data.

[0146] When constructing the associated observation vector, the observations from the target sensors are integrated according to preset rules. Based on the causal correlation strength between the target sensor and the faulty sensor, from high to low, the observation value of each target sensor at the moment of the fault occurs is used as one dimension of the vector, forming an ordered multidimensional vector. For example, if the target sensors include an upstream soil moisture sensor and a surface runoff sensor, their observed values ​​are respectively... (Soil moisture) (Runoff), then the associated observation vector is This vector fully carries the reliable observation information that directly affects the values ​​of the anomalous sensors, providing an input basis for subsequent mechanism calculations.

[0147] S42. Based on the associated observation vectors and the mechanism function derived from the watershed mechanism proxy model, an optimization problem model for data restoration is established; the objective function of the optimization problem model aims to minimize the deviation between the restored value and the initial estimate, as well as the deviation between the restored value and the theoretical value calculated from the mechanism function; the objective function is:

[0148]

[0149] in, The repair value to be solved. This is the initial estimate obtained by interpolation of outlier values. For associated observation vectors, For mechanism function, For the corresponding theoretical value, and These are the preset weighting coefficients.

[0150] Specifically, The repair value to be solved is the correction value of the abnormal sensor data that needs to be determined. The interpolation method is used to obtain the initial estimate of the outlier value. The interpolation method can be selected based on the normal observations before and after the outlier value (such as linear interpolation, time-series smoothing interpolation, etc.) to provide an initial reference benchmark for the repair value. This is the associated observation vector, which is the set of reliable observations of the target sensor constructed in S41; This is a mechanistic function derived from a watershed surrogate model. This function is constructed based on the "input-output" driving relationship within the surrogate model. For example, if the surrogate model takes upstream soil and runoff data as input and downstream water quality data as output, then... This is the mathematical expression of the input-output mapping relationship; This is the theoretical value calculated through the mechanism function, that is, the theoretical data value that the abnormal sensor should have, derived from the observation value of the target sensor. and The preset weighting coefficients, Used to control the degree of influence of the initial estimate on the repaired value. The constraint strength between the theoretical value of the mechanism function and the repair value is used to control the constraint strength. Both are set according to the priority of "temporal continuity" and "mechanistic rationality" in the repair scenario. For example, when more emphasis is placed on the mechanism, the constraint strength can be increased. When prioritizing time-series smoothing, it can be increased .

[0151] The design logic of the objective function lies in balancing "time series data continuity" and "physical mechanism consistency". (First term) Ensure a smooth transition between the repaired values ​​and the normal data before and after the abnormal values, avoiding abrupt changes in the time series curve; the second item To ensure that the remediation values ​​conform to the mechanisms of watershed pollution transport, it is crucial to avoid "distorted" data that violates physical logic. By minimizing this objective function, preliminary candidate remediation values ​​that balance both requirements can be obtained.

[0152] S43. Add constraints derived from physical laws to the optimization problem model to obtain a constrained optimization repair model; among which, the constraints include: non-negativity constraint of repair value, upper limit constraint of repair value concentration, and constraint based on the conservation of pollutant mass.

[0153] Specifically, this step further limits the reasonable range of the repair value by adding physical law constraints, ensuring that the repair results conform to the basic physical laws of pollutant transport in the actual environment, and avoiding the optimization model from solving only from a mathematical perspective while ignoring the physical meaning.

[0154] The non-negativity constraint for remediation values ​​is set based on the fundamental principle that "monitoring parameters such as pollutant concentration, soil moisture, and runoff are all non-negative physical quantities," and its mathematical expression is: This constraint directly eliminates the incomprehensible possibility of negative restoration values. For example, the ammonia nitrogen concentration of water quality sensors and the humidity value of soil sensors cannot be negative. Therefore, this constraint is the basis for ensuring the physical validity of restoration values.

[0155] The upper limit constraint on remediation value concentration is set based on the actual physical characteristics of the monitoring parameters or environmental standards, and the mathematical expression is: In the formula This is the upper limit of the concentration of the parameter monitored by the abnormal sensor. This value is determined according to the parameter type. For water quality parameters, the limit in the national environmental quality standards or the saturation concentration corresponding to the water body's environmental capacity can be referenced. For soil parameters, the maximum carrying capacity of the soil's physical and chemical properties (such as the maximum water holding capacity of the soil) can be referenced. This constraint prevents the remediation value from exceeding the actual physical upper limit of the parameter. For example, if the solubility of a certain pollutant in water is limited, the remediation value cannot exceed its saturation concentration.

[0156] The constraint based on the conservation of pollutant mass is constructed based on the physical law of "conservation of total pollutant quantity within the watershed." Its core is to ensure that the remediation values ​​of anomaly sensors conform to the balance relationship of pollutant fluxes upstream and downstream. Anomaly sensors are used as water quality sensors (monitoring pollutant concentrations). Taking this as an example, its mass conservation constraint needs to be combined with the upstream runoff sensor (monitoring runoff volume) in the associated observation vector. ), upstream water quality sensors (monitoring upstream pollutant concentrations) ) and downstream runoff sensors (monitoring runoff volume) The mathematical expression for the derivation of the observed values ​​can be simplified to:

[0157]

[0158] In the formula, The output flux of pollutants at the abnormal sensor; For upstream pollutant input flux; This represents the amount of pollutant attenuation during transport (calculated using the attenuation coefficient in the watershed mechanism proxy model). This constraint ensures that the remediated pollutant concentration does not lead to an imbalance in the total amount of pollutants within the watershed. For example, when the upstream input flux is limited, the downstream remediated output flux cannot far exceed the input flux (when attenuation is ignored), which conforms to the law of conservation of mass.

[0159] By adding the above three types of constraints, the optimization problem model is transformed from "unconstrained mathematical optimization" to "mechanism optimization with physical constraints". The repair value must simultaneously satisfy the minimization of the mathematical objective and compliance with physical laws.

[0160] S44. Solve the constrained optimization repair model to obtain the optimal repair value that minimizes the objective function and satisfies all constraints; replace the corresponding abnormal values ​​in the sensor abnormal data with the optimal repair value.

[0161] Specifically, the objective function of the constrained optimization repair model is a quadratic function (both terms are squared), and the constraints are linear inequality constraints (non-negativity, upper concentration limit, and mass conservation constraints are all linear relationships). Therefore, optimization algorithms suitable for "quadratic objective function + linear constraints" can be used to solve the problem, such as the Lagrange multiplier method, interior point method, or quadratic programming algorithm. The core logic of the solution process is: within the feasible region that satisfies all constraints, find the value that minimizes the objective function. That is, the optimal repair value.

[0162] The specific solution process includes: first, determining the feasible region formed by the constraints (all conditions satisfying the constraints). , mass conservation constraint (Set); then calculate the gradient of the objective function within the feasible region and search for the minimum point along the gradient descent direction; if the minimum point is located inside the feasible region, then that point is the optimal repair value; if the minimum point is located at the boundary of the feasible region (such as touching the non-negativity constraint or concentration upper limit constraint), then the boundary point is the optimal repair value (the objective function must be minimized at that boundary point).

[0163] After obtaining the optimal repair value, it is replaced with the abnormal values ​​in the sensor's abnormal data. In the time series database of the abnormal sensor, the abnormal value record corresponding to the time of the fault occurrence is located, the record is overwritten with the optimal repair value, and the data is marked as "repaired reliable data". At the same time, the original abnormal value and the repair process log (including associated observation vectors, mechanism function calculation results, constraints, etc.) are retained to facilitate subsequent traceability and verification.

[0164] After the replacement is completed, briefly verify the rationality of the repaired values. Substitute the repaired values ​​into the dynamic cross-domain event graph and check their logical consistency with the upstream and downstream target sensor data (whether they conform to causal relationships and time delay parameters). If logical conflicts still exist, the weighting coefficients need to be readjusted. , or constraint parameters (such as) The model is then resolved until the repaired values ​​meet all physical laws and logical requirements, ensuring that the repaired data can be used for subsequent pollution monitoring analysis and decision support.

[0165] The aforementioned method for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors first standardizes the heterogeneous monitoring data and constructs a watershed proxy model that integrates physical mechanisms. Based on this, a dynamic event graph characterizing cross-domain causal relationships is established. When data anomalies are detected, the system uses this graph to trace the source and drives the proxy model to generate theoretical expected values. Logical consistency verification is performed through comparison to distinguish between real pollution events and sensor malfunctions. For data determined to be faulty, an optimization model integrating physical constraints is further established and solved to generate self-consistent repair values ​​that conform to the pollution transmission law. This achieves a paradigm shift in quality control from isolated threshold judgment to cross-domain logical verification, significantly improving the accuracy of identifying complex pollution events, the reliability of sensor fault diagnosis, and the physical rationality of abnormal data repair results.

[0166] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0167] Based on the same inventive concept, this application also provides an apparatus for implementing the above-described method for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors. The solution provided by this apparatus is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the apparatus for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors provided below can be found in the limitations of the above-described method for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors, and will not be repeated here.

[0168] In one exemplary embodiment, such as Figure 3 As shown, an agricultural non-point source pollution monitoring data fusion and quality control device 30 for heterogeneous sensors is provided to implement the methods in the above-described method embodiments. The device includes:

[0169] The multi-source data preprocessing module 31 is used to perform spatiotemporal alignment and feature extraction processing on the heterogeneous raw monitoring data in the agricultural non-point source pollution monitoring network to generate a standardized time-series feature dataset for each watershed sub-unit.

[0170] The pollution transmission modeling module 32 is used to construct a simplified mechanism module based on the physical mechanism of agricultural non-point source pollution transmission. Based on the simplified mechanism module, a machine learning model is trained using historical monitoring data to obtain a watershed mechanism proxy model that describes the watershed mechanism from upstream rainfall and soil characteristics to downstream water quality parameter response.

[0171] The dynamic relationship graph construction module 33 is used to construct a dynamic cross-domain event graph based on the driving relationship between input and output variables described by the watershed mechanism proxy model and the sensor network spatial topology of the agricultural non-point source pollution monitoring network. In the dynamic cross-domain event graph, the nodes represent sensor or hydrological events, and the edges represent the causal influence relationship and expected time delay between nodes.

[0172] The anomaly tracing and localization module 34 is used to trace the source in reverse along the causal influence edge in the dynamic cross-domain event graph when any sensor data is detected as abnormal, and obtain the upstream node data; and use the upstream node data to drive the watershed mechanism proxy model to calculate the theoretical expected value range of the current abnormal data.

[0173] The data consistency verification module 35 is used to perform cross-domain logical consistency verification and fault conflict location by comparing the actual observed value with the theoretical expected value range, and to obtain abnormal sensor data that is judged to be logically conflicted and the conflict is located as a sensor fault.

[0174] The self-consistent data repair module 36 is used to establish an optimized repair model based on abnormal sensor data, which integrates the mechanistic relationship described by the physical law constraint and the watershed mechanism proxy model; and generates a self-consistent repair value that conforms to the pollution transmission law to replace the abnormal sensor data by solving the optimized repair model.

[0175] Embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the aforementioned method embodiments.

[0176] Embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0177] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0178] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A method for data fusion and quality control of agricultural non-point source pollution monitoring based on heterogeneous sensors, characterized in that, The method includes: S1. Perform spatiotemporal alignment and feature extraction processing on the heterogeneous raw monitoring data in the agricultural non-point source pollution monitoring network to generate a standardized time-series feature dataset for each watershed sub-unit; S2. Construct a simplified mechanism module based on the physical mechanism of agricultural non-point source pollution transmission. Based on the simplified mechanism module, train a machine learning model using historical monitoring data to obtain a watershed mechanism proxy model that describes the watershed mechanism from upstream rainfall and soil characteristics to downstream water quality parameter response. S3. Based on the driving relationship between input and output variables described by the watershed mechanism proxy model, and combined with the sensor network spatial topology of the agricultural non-point source pollution monitoring network, a dynamic cross-domain event graph is constructed; wherein, the nodes of the dynamic cross-domain event graph represent sensor or hydrological events, and the edges of the dynamic cross-domain event graph represent the causal relationship between nodes and the expected time delay. S4. When any sensor data is detected as abnormal, reverse tracing is performed along the causal influence edge in the dynamic cross-domain event graph to obtain upstream node data; the upstream node data is used to drive the watershed mechanism proxy model to calculate the theoretical expected value range of the current abnormal data. S5. By comparing the actual observed values ​​with the theoretical expected value range, cross-domain logical consistency verification and fault conflict location are performed to obtain abnormal sensor data that are judged to be logically conflicted and the conflict is located as sensor fault. S6. Based on the abnormal sensor data, establish an optimized repair model that integrates physical law constraints and the mechanism relationship described by the watershed mechanism proxy model; generate self-consistent repair values ​​that conform to the pollution transport law by solving the optimized repair model to replace the abnormal sensor data.

2. The method according to claim 1, characterized in that, The simplified mechanism module is constructed based on the physical mechanism of agricultural non-point source pollution transmission. Based on the simplified mechanism module, a machine learning model is trained using historical monitoring data to obtain a watershed mechanism proxy model describing the response of water quality parameters from upstream rainfall and soil characteristics to downstream rainfall. This model includes: S11. Based on the principle of full storage and runoff generation, a simplified hydrological response module is constructed to describe the transformation process from rainfall to soil water and then to surface runoff; wherein, the simplified hydrological response module uses a water tank model to simulate the change of soil water storage and integrates a simplified form of the SCS curve number method to calculate surface runoff. S12. Based on empirical relationships of pollutant migration, a simplified pollution migration module is constructed to describe the migration and transformation process of pollutant concentration from surface runoff to river channels; wherein, the simplified pollution migration module uses a nonlinear function containing time delay parameters to simulate the flushing and decay process of pollutant concentration after the runoff peak. S13. Extract multi-source features from the standardized time-series feature dataset, including rainfall event pulse features, soil moisture steady-state features, and water quality parameter steady-state features from the historical sequence, and construct a training sample set based on the multi-source features; wherein, the input features of the training sample set are the multi-source features, and the output label of the training sample set is the water quality feature at a specific time lag in the future; S14. Using the gradient boosting decision tree algorithm, the undetermined parameters in the simplified hydrological response module and the simplified pollution migration module are jointly trained and fitted with the training sample set to obtain a machine learning model for predicting the water quality characteristics of the downstream based on the multi-source characteristics of the upstream. The machine learning model is used as the watershed mechanism proxy model.

3. The method according to claim 1, characterized in that, The construction of a dynamic cross-domain event graph, based on the driving relationship between input and output variables described in the watershed mechanism proxy model and combined with the sensor network spatial topology of the agricultural non-point source pollution monitoring network, includes: S21. Define each type of sensor at each monitoring point of the agricultural non-point source pollution monitoring network as a sensor node, and define each rainfall event exceeding the threshold as an independent event node. The node set of the dynamic cross-domain event map is composed of all the sensor nodes and the event nodes. S22. Based on the sub-basin confluence relationship determined by the watershed digital elevation model, establish a spatial topological edge representing the water flow transmission path between the sensor nodes in the upstream and downstream sub-basins. S23. Based on the physical driving relationship described by the watershed mechanism proxy model, establish causal influence edges between the sensor nodes and the event nodes that have causal relationships; wherein each causal influence edge has an initial time delay parameter representing the time required for the cause to lead to the result; S24. During the online operation of the agricultural non-point source pollution monitoring network, the data of the nodes at both ends of the causal influence edge are continuously monitored. By analyzing the sequential relationship of the node state changes, the initial time delay parameters of each causal influence edge are dynamically adjusted and updated to obtain the actual time delay parameters. S25. The sensor nodes, event nodes, spatial topology edges, causal influence edges, and the actual time delay parameters corresponding to each causal influence edge are combined to construct the dynamic cross-domain event graph.

4. The method according to claim 3, characterized in that, During the online operation of the agricultural non-point source pollution monitoring network, the data of the nodes at both ends of the causal influence edge are continuously monitored. By analyzing the temporal relationship of the node state changes, the initial time delay parameters of each causal influence edge are dynamically adjusted and updated to obtain the actual time delay parameters, including: S31. For a causal edge pointing from cause node A to result node B, continuously monitor the data of cause node A. When a change in the state of cause node A exceeding a first preset range is detected, record the start time of the corresponding change event. ; S32, at the starting time The next maximum possible time delay window Internally, continuously monitor the data of result node B, and when result node B is detected... When the state changes beyond the second preset range, record the start time of the corresponding response event. This forms a pair of causal event observation pairs. ;in, The maximum possible time delay; S33. Collect a series of historical causal event observation pairs to form a causal event observation pair set; calculate different time delays based on the causal event observation pair set. The statistical association strength is obtained by calculating the association strength sequence; wherein, mutual information is used as the statistical association strength, and the formula for calculating the mutual information is: in, Indicates time delay The mutual information mentioned below, Let A represent the set of all possible discrete states of cause node A. Let B represent the set of all possible discrete states of the result node B. Indicates that cause node A is at time [time]. The discrete state values, This indicates that the result node B is at time [time]. The discrete state values; express The marginal probability distribution, express The marginal probability distribution, express and The joint probability distribution of ; S34. Find the mutual information. Reaching maximum time delay Time lag This serves as the actual time delay parameter for the corresponding causal influence edge.

5. The method according to claim 3 or 4, characterized in that, The process involves establishing an optimized remediation model based on the abnormal sensor data, integrating physical constraints with the mechanistic relationship described by the watershed mechanism proxy model; and generating self-consistent remediation values ​​that conform to the pollution transport law to replace the abnormal sensor data by solving the optimized remediation model, including: S41. Determine the sensor corresponding to the abnormal data of the sensor to be repaired, and designate it as the abnormal sensor; determine the fault occurrence time and the abnormal value to be repaired corresponding to the abnormal data of the sensor to be repaired; based on the dynamic cross-domain event graph, query all other sensors that are directly connected to the abnormal sensor through the causal influence edge at the corresponding fault occurrence time and whose data status is reliable, and designate them as target sensors; construct an associated observation vector based on the observation value of the target sensor. S42. Based on the associated observation vector and the mechanism function derived from the watershed mechanism proxy model, an optimization problem model for data repair is established; wherein, the objective function of the optimization problem model aims to minimize the deviation between the repaired value and the initial estimated value, and the deviation between the repaired value and the theoretical value calculated by the mechanism function; the objective function is: in, The repair value to be solved. This is the initial estimate obtained by interpolation of the abnormal values. For the associated observation vector, The mechanism function is... For the corresponding theoretical value, and These are preset weighting coefficients; S43. Add constraints derived from physical laws to the optimization problem model to obtain a constrained optimization repair model; wherein, the constraints include: non-negativity constraint of repair value, upper limit constraint of repair value concentration, and constraint based on the conservation of pollutant mass; S44. Solve the constrained optimization repair model to obtain the optimal repair value that minimizes the objective function and satisfies all constraints; replace the corresponding abnormal value in the sensor abnormal data with the optimal repair value.

6. A data fusion and quality control device for agricultural non-point source pollution monitoring based on heterogeneous sensors, used to implement the method according to any one of claims 1 to 5, characterized in that, The device includes: The multi-source data preprocessing module is used to perform spatiotemporal alignment and feature extraction on heterogeneous raw monitoring data in the agricultural non-point source pollution monitoring network, generating a standardized time-series feature dataset for each watershed sub-unit; The pollution transmission modeling module is used to construct a simplified mechanism module based on the physical mechanism of agricultural non-point source pollution transmission. Based on the simplified mechanism module, a machine learning model is trained using historical monitoring data to obtain a watershed mechanism proxy model that describes the watershed mechanism from upstream rainfall and soil characteristics to downstream water quality parameter response. The dynamic relationship graph construction module is used to construct a dynamic cross-domain event graph based on the driving relationship between input and output variables described by the watershed mechanism proxy model and the sensor network spatial topology of the agricultural non-point source pollution monitoring network; wherein, the nodes of the dynamic cross-domain event graph represent sensor or hydrological events, and the edges of the dynamic cross-domain event graph represent the causal influence relationship and expected time delay between nodes; The anomaly tracing and localization module is used to trace the source in reverse along the causal influence edge in the dynamic cross-domain event graph when any sensor data is detected as abnormal, to obtain the upstream node data; and to use the upstream node data to drive the watershed mechanism proxy model to calculate the theoretical expected value range of the current abnormal data. The data consistency verification module is used to perform cross-domain logical consistency verification and fault conflict location by comparing the actual observed value with the theoretical expected value range, and to obtain abnormal sensor data that is judged to be logically conflicted and the conflict is located as a sensor fault. The self-consistent data repair module is used to establish an optimized repair model based on the abnormal sensor data, which integrates the constraints of physical laws and the mechanistic relationship described by the watershed mechanism proxy model; and to generate self-consistent repair values ​​that conform to the pollution transmission law by solving the optimized repair model to replace the abnormal sensor data.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.