Agricultural non-point source pollution risk assessment method fusing multi-source heterogeneous data analysis

By integrating multi-source heterogeneous data analysis, a cross-modal feature interaction model and a directed acyclic causal graph are constructed, which solves the inaccuracy problem of agricultural non-point source pollution risk assessment in existing technologies and achieves more accurate real-time response and pollution control.

CN121352467APending Publication Date: 2026-01-16ZHONGSHUI SANLI (BEIJING) TECH SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511404270.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing technologies for agricultural non-point source pollution risk assessment fail to reflect the complex changes within the watershed, resulting in inaccurate assessment results. They cannot support real-time forecasting and dynamic regulation, and the data is too singular and lacks representativeness, ignoring key meteorological factors and leading to overly empirical or parameterized process descriptions.

Method used

This study employs a multi-source heterogeneous data analysis method, constructs an indicator system using the analytic hierarchy process (AHP), introduces a spatiotemporal attention interaction submodule, dynamically adjusts weights, builds a cross-modal feature interaction model, extracts temporal and spatial feature vectors, forms a directed acyclic causal graph, identifies spurious correlation factors, corrects the causal association strength, and calculates the pollution risk index.

Benefits of technology

It significantly improves the real-time response capability of agricultural non-point source pollution risk assessment, enhances the positioning accuracy of key source areas and the effectiveness of pollution control, reduces prediction errors, adapts to sudden rainstorms and topographical differences, and strengthens the consistency between the weights after causal correction and the actual pollution driving mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121352467A_ABST
    Figure CN121352467A_ABST
Patent Text Reader

Abstract

The invention provides an agricultural non-point source pollution risk assessment method fusing multi-source heterogeneous data analysis, and the method comprises the steps: constructing an index system which comprises a target layer, a criterion layer and a scheme layer based on an analytic hierarchy process, introducing a space-time attention interaction sub-module into the system, and dynamically adjusting the initial weight of the criterion layer index; constructing a cross-modal feature interaction model, and performing deep feature fusion on time series data and spatial data of the standard data set; taking the pollution risk index as a target variable, taking a criterion layer index and an agricultural activity factor as reason variables, and taking seasons and terrains as hybrid variables to form a directed acyclic causal graph; fixing hybrid variables for the causal diagram, and evaluating direct causal influence of each reason variable on a target variable; and calculating the causal association strength of each criterion layer index by using the cross-modal features, correcting the space-time adaptive weight vector in combination with the causal association strength, and calculating the pollution risk index based on the correction result. According to the invention, the accuracy and timeliness of agricultural non-point source pollution prevention are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application belong to the field of artificial intelligence technology, and more specifically, they relate to a method for risk assessment of agricultural non-point source pollution by integrating multi-source heterogeneous data analysis. Background Technology

[0002] In the field of agricultural non-point source pollution monitoring, various sensor technologies involve multiple data sources, including meteorological data (rainfall, sunshine, air pressure, wind speed, temperature and humidity), hydrological data (runoff), water quality data (nitrogen, phosphorus, organic matter concentration), and geographic information data (digital elevation model DEM, land use, soil type).

[0003] Agricultural non-point source pollution is one of the major sources of water pollution, and its formation mechanism is complex and exhibits significant spatiotemporal heterogeneity. The generation, migration, and transformation of pollutants are influenced by a variety of natural factors and human activities, exhibiting characteristics of randomness, intermittency, and widespread occurrence.

[0004] Traditional monitoring methods have the following limitations: existing technologies rely on data from a limited number of meteorological stations, resulting in low spatiotemporal resolution and an inability to support real-time forecasting and dynamic control. The models cannot be integrated with real-time rainfall radar images or real-time wind speed and direction data to dynamically adjust forecasts, and the forecast results are difficult to reflect the complex spatial changes within the watershed. Furthermore, existing technologies have poor data representativeness, relatively simplified processes, and some processes are parameterized. They mainly rely on rainfall data while ignoring other key meteorological factors. The description of pollutant migration and transformation processes is too empirical or parameterized, ignoring key processes and leading to inaccurate assessment results. Summary of the Invention

[0005] The purpose of this invention is to provide a method for assessing agricultural non-point source pollution risk by integrating multi-source heterogeneous data analysis, so as to at least solve the technical problem that the prediction results of existing agricultural non-point source pollution risk assessment methods are difficult to reflect the complex changes within the watershed, resulting in inaccurate assessment results.

[0006] To achieve the above objectives, the embodiments of this application provide the following technical solutions.

[0007] According to one embodiment of this application, a method for risk assessment of agricultural non-point source pollution by integrating multi-source heterogeneous data analysis is provided, comprising the following steps:

[0008] Acquire multi-source heterogeneous data from agricultural non-point source pollution monitoring, and preprocess the multi-source heterogeneous data to obtain a standard dataset;

[0009] An indicator system comprising a target layer, a criterion layer, and a scheme layer is constructed based on the analytic hierarchy process. A spatiotemporal attention interaction submodule is introduced into the indicator system to dynamically adjust the initial weights of the indicators in the criterion layer.

[0010] A cross-modal feature interaction model is constructed to perform deep feature fusion on temporal and spatial data of a standard dataset. Specifically, a temporal processing network is used to process the temporal data and extract temporal feature vectors; a spatial feature extraction network is used to process the spatial data and extract spatial feature vectors; a modal correlation matrix is ​​calculated, and the temporal and spatial feature vectors are weighted and calculated through the modal correlation matrix to generate cross-modal features.

[0011] Using the pollution risk index as the target variable, the criterion-level indicators and agricultural activity factors as causal variables, and season and topography as confounding variables, a directed acyclic causal graph is formed. The confounding variables are fixed in the causal graph, and the direct causal impact of each causal variable on the target variable is evaluated separately to identify and eliminate spurious correlation factors. The causal association strength of each criterion-level indicator is calculated using cross-modal features, and the spatiotemporal adaptive weight vector is corrected in combination with the causal association strength to obtain the causal-corrected weight vector.

[0012] The pollution risk index is calculated and output based on the weight vector after causal correction.

[0013] Furthermore, the introduction of a spatiotemporal attention interaction submodule into the indicator system, and the steps for dynamically adjusting the initial weights of the criterion-level indicators, include:

[0014] In the time dimension of the spatiotemporal attention interaction submodule, the temporal correlation of meteorological time series data is captured through the attention mechanism, and the time attention weight coefficient is calculated.

[0015] In the spatial dimension of the spatiotemporal attention interaction submodule, terrain features are extracted by combining digital elevation model data, and different spatial attention weight coefficients are assigned to different terrain units.

[0016] In the dynamic adjustment of weights, based on the latest collected multi-source data, the weights of the criteria layer indicators are corrected by combining the temporal attention weight coefficient and the spatial attention weight coefficient, resulting in a spatiotemporal adaptive weight vector.

[0017] Furthermore, the step of using a time-series processing network to process time-series data and extract time-series feature vectors includes:

[0018] Variational mode decomposition (VMD) algorithm is used to adaptively divide multi-source time-series data into frequency bands, decomposing the original time-series data into multi-scale subsequences; wherein, the original time-series data is decomposed into K modal components. The objective function is expressed as:

[0019]

[0020] The constraints are expressed as follows:

[0021]

[0022] In the formula, Represents the time derivative. Represents the Dirac function, The center frequencies of each mode are represented by , and * indicates the convolution operation. The long-term trend component, seasonal cycle component, and sudden disturbance component are obtained by solving the alternating direction multiplier method.

[0023] A nested gated loop unit network is constructed, with the main branch using bidirectional gated loop units to capture full-time dependencies. The full-time data is modeled as follows:

[0024]

[0025] in, This represents the hidden state at time t in the main branch. The input data represents time t; This represents the hidden state of the main branch at time t-1. This represents the hidden state of the main branch at time t+1; Indicates a bidirectional gated loop unit;

[0026] The first sub-branch extracts trend features for the long-term trend component using a large-time-step gated cyclic unit, represented as:

[0027]

[0028] In the formula, This represents the long-term trend feature vector extracted at time t, and GRU stands for Gated Recurrent Unit. This represents the long-term trend component data input at time t; This represents the long-term trend feature vector input at time t-1;

[0029] The second sub-branch uses an adaptive periodic gating cyclic unit to extract periodic features for the seasonal periodic component, represented as:

[0030]

[0031]

[0032] In the formula, T represents the length of the seasonal cycle. This represents the Pearson coefficient, used to measure the linear correlation between two time series data. This represents the seasonal periodic component data input at time t; This represents the periodic feature vector extracted at time t; Represents the periodic eigenvector at time tT;

[0033] The third sub-branch employs an attention-enhanced gated recurrent unit for sudden disturbance components, represented as:

[0034]

[0035]

[0036] In the formula, This represents the attention weight coefficient. Indicates the attenuation coefficient. Indicates the moment when the sudden disturbance occurred; This represents the feature vector of the sudden disturbance extracted at time t. This represents the sudden disturbance component data input at time t. This represents the characteristic vector of the sudden disturbance at time tT. The characteristic variable representing the sudden disturbance at time t-1;

[0037] The output features of the main branch and sub-branch are weighted and integrated through an attention fusion mechanism. The weight coefficients of the attention fusion mechanism are dynamically adjusted based on the contribution of each component to historical pollution events, and finally a time-series feature vector containing multi-scale temporal correlation information is generated.

[0038] Among them, the weight coefficients of the attention fusion mechanism Represented as:

[0039]

[0040] In the formula, This represents the contribution coefficient of the branch. This represents the hidden state of the k-th sub-branch at time t; This represents the normalization operation on the hidden state; Let represent the temporal feature vector extracted by the i-th sub-branch at time t.

[0041] Furthermore, the step of using a spatial feature extraction network to process spatial data and extract spatial feature vectors includes:

[0042] The spatial data in the standard dataset are scaled to obtain micro-scale data of field plots and macro-scale data of sub-watersheds; the pollution correlation coefficients of spatial data at different scales are calculated and expressed as follows:

[0043]

[0044] In the formula, This represents spatial data at the s-th scale, where s=1 represents the field scale and s=2 represents the sub-basin scale. This represents historical pollution load data at the corresponding scale. This represents the Pearson correlation coefficient, used to quantify the strength of the association between spatial data at different scales and pollution risk;

[0045] A multi-branch spatial feature extraction network is constructed, including a field-scale branch and a sub-watershed-scale branch, and each branch embeds an agricultural non-point source pollution spatial feature encoding module;

[0046] For the field-scale branch, a convolutional neural network (CNN) combined with a soil texture coding module is used to extract features from the grid data.

[0047] For the sub-basin scale branch, a graph convolutional neural network (GCN) combined with a hydrological connectivity encoding module is used to convert the sub-basin spatial grid into a graph structure. The nodes of the graph structure are the sub-basin grid cells, and the edge weights are the hydrological connectivity between cells. The hydrological connectivity is determined based on the flow direction matrix calculated from the DEM and the cumulative discharge data, and is expressed as:

[0048]

[0049] In the formula, Let i be the hydrological connectivity between grid cells i and j in the sub-basin. The cumulative flow from unit i to j The maximum cumulative flow within the sub-basin; GCN captures the correlation characteristics between grid cells and pollutant confluence within the sub-basin through graph convolution operations;

[0050] Based on pollution correlation coefficient A dynamic attention fusion layer is constructed to weight and fuse the local feature vectors output by the field-scale branch and the global feature vectors output by the sub-watershed-scale branch; simultaneously, a soil-topography collaborative attention coefficient is introduced. , Based on the ternary correlation between "soil texture-slope-pollution loss" in historical data, it is expressed as:

[0051]

[0052] In the formula, The mutual information entropy is represented by T, slope data, soil texture data, land use data, hydrological connectivity data, and pollution loss data.

[0053] The weighted fusion spatial feature vector is represented as: ;

[0054] In the formula, Represents the fused spatial feature vector; Represents the collaborative attention coefficient. This represents the local feature vector output by the field-scale branch. This represents the global feature vector output by the sub-basin scale branch; , This represents the pollution correlation coefficient.

[0055] Furthermore, in the step of forming a directed acyclic causal graph with the pollution risk index as the target variable, criterion-level indicators and agricultural activity factors as causal variables, and season and topography as confounding variables:

[0056] In the target variable layer, the pollution risk index RI is calculated by coupling the weight vector after causal correction with the standardized values ​​of the criteria layer indicators, and is expressed as follows:

[0057]

[0058] In the formula, This represents the weight of the i-th indicator in the criterion layer after causal correction. This represents the standardized value of the criteria-level indicator;

[0059] In the causal variable layer, pollution processes are divided into generation stage nodes and migration stage nodes. Generation stage nodes include rainfall and soil moisture from the criterion layer, as well as agricultural activity factors such as fertilization intensity and irrigation frequency. Migration stage nodes include auxiliary indicators from the criterion layer such as wind speed, atmospheric pressure, and topographically derived factors. Topographically derived factors include runoff potential, which is calculated based on slope, aspect, and cumulative flow from Digital Elevation Model (DEM) data, and is expressed as...

[0060]

[0061] in, Accumulate flow for grid cells, This represents the maximum cumulative traffic in the region.

[0062] In the confounding variable layer, season and topography are decomposed into dynamic and static sub-variables that are strongly correlated with the pollution process. Seasonal variables are decomposed into crop growth period, and topographic variables are decomposed into soil texture and slope grade.

[0063] Furthermore, the step of forming the directed acyclic causal graph also includes:

[0064] Construct causal edges for the generation stage; including: directed edges between fertilization intensity and soil moisture, directed edges between fertilization intensity and pollution risk index, and directed edges between soil moisture and pollution risk index;

[0065] Construct causal edges for the migration phase; specifically, establish directed edges between wind speed C3 and atmospheric ammonia deposition risk, and between runoff potential T1 and surface runoff pollution risk; where the edge weights for wind speed and atmospheric ammonia deposition risk are defined. The coefficients were calibrated based on the linear regression coefficients of wind speed and atmospheric ammonia concentration; the side weights for runoff potential T1 and surface runoff pollution risk were... The linear regression coefficients based on slope and runoff nitrogen content were calibrated.

[0066] Construct causal edges for confounding variables, and establish directed edges between confounding sub-variables and causal variables, including edge weights for crop growth cycle and fertilization intensity. The marginal weights of soil texture and runoff potential .

[0067] Furthermore, the steps of fixing confounding variables in the causal graph, individually evaluating the direct causal impact of each causal variable on the target variable, and identifying and eliminating spurious correlation factors include:

[0068] Hierarchical fixation is implemented based on the attribute differences of confounding variables in the causal graph, including static confounding variable fixation and dynamic confounding variable fixation;

[0069] For the causal graph with confounding variables fixed, backdoor adjustment and causal effect decomposition methods are used to calculate the direct causal effect of each causal variable on the pollution risk index RI separately. Specifically, in the backdoor path blocking, for the causal variable X, all backdoor paths of causal variables, confounding variables, and the pollution risk index in the causal graph are identified, and these paths are blocked by fixing the confounding variables, retaining only the direct path from X to RI. The formula for calculating the average direct causal effect is as follows:

[0070]

[0071] In the formula, This indicates that the intervention causal variable X takes the actual observed value. Z = z represents the baseline value where the intervention X has a mean of 0, and Z = z represents the value of the confounding variable after fixing. Express conditional expectation; for continuous causal variables, use derivatives to calculate marginal effects;

[0072] A dual verification mechanism of causal strength and data consistency is introduced to identify and eliminate causal variables with spurious correlations.

[0073] Furthermore, the step of calculating the causal correlation strength of each criterion layer index using cross-modal features, and then correcting the spatiotemporal adaptive weight vector based on the causal correlation strength to obtain the causally corrected weight vector includes:

[0074] Cross-modal features of dimension d With criteria layer indicator set Perform association modeling;

[0075] Combining direct causal effect values Calculate the criteria layer indicators causal relationship strength , represented as:

[0076]

[0077] In the formula, The dynamic equilibrium coefficient is adaptively adjusted based on cross-modal features and the Pearson correlation coefficient of the causal graph. Indicates the dynamic correlation coefficient;

[0078] The spatiotemporal adaptive weight vector is modified.

[0079] Furthermore, cross-modal features of dimension d... With criteria layer indicator set The steps for performing association modeling include:

[0080] Cross-modal features are integrated using a fully connected neural network. Projecting onto the criterion layer index space yields the projected features. ; where the output of the i-th neuron Represented as:

[0081]

[0082] In the formula, This represents the learnable weight coefficients. Indicates the bias term. Indicates cross-modal features;

[0083] Calculate projection features Corresponding criteria level indicators Dynamic correlation coefficient , represented as:

[0084]

[0085] In the formula, T represents the length of the time series. Represents the characteristic mean. Indicates the mean of the indicators and the dynamic correlation coefficient. The value range is [-1, 1]. Positive values ​​indicate positive correlation, negative values ​​indicate negative correlation, and the larger the absolute value, the more significant the spatiotemporal correlation between the cross-modal projection features and the criterion layer indicators.

[0086] Furthermore, the step of correcting the spatiotemporal adaptive weight vector includes:

[0087] Causal-corrected weight vector In the diagram, the i-th element is represented as:

[0088]

[0089] In the formula, This represents the weight of the m-th indicator after causal correction. This represents the i-th element in the weight vector after causal correction; , These represent the spatiotemporal adaptive weights of the i-th and j-th indicators, respectively; , These represent the strength of the causal association between the i-th and j-th indicators and the risk of agricultural non-point source pollution, respectively. , Both represent exponential enhancement terms;

[0090] Rules for agricultural non-point source pollution are introduced to constrain the correction results. If the correction results violate the constraints, they are re-optimized using the Lagrange multiplier method.

[0091] Compared with existing technologies, the beneficial effects of the agricultural non-point source pollution risk assessment method integrating multi-source heterogeneous data analysis in the embodiments of this application are:

[0092] This application introduces a spatiotemporal attention interaction submodule, which captures the temporal correlation between the rate of change of continuous rainfall intensity and the time window after fertilization in the time dimension, and combines DEM to extract terrain features and assign differentiated weights in the spatial dimension, thus solving the problem that static weights cannot adapt to sudden rainstorms and terrain differences.

[0093] This application uses a time-series processing network to extract the time-series features of cumulative rainfall and changes in soil nitrogen content after fertilization, and a spatial feature extraction network to extract the spatial features of soil water retention capacity and crop root interception potential. Then, it uses a modal correlation matrix to quantify the synergistic effects of rainfall on sandy soil and irrigation on clay soil, thereby reducing the prediction error of cross-modal features on pollution load.

[0094] This application uses season and topography as confounding variables, and evaluates the direct causal impact of causal variables separately by fixing the confounding variables. It identifies rainfall, soil moisture, and fertilization intensity as driving factors, which greatly improves the consistency between the weights after causal correction and the actual pollution driving mechanism. Furthermore, it uses cross-modal features to calculate the causal association strength, which integrates the direct causal effect and the feature correlation degree. This can correct the spatiotemporal adaptive weights and avoid the problem of reduced accuracy of pollution risk assessment in extreme scenarios.

[0095] In summary, this application significantly improves the real-time response capability of agricultural non-point source pollution risk assessment, and enhances the accuracy of key source area location and pollution control effectiveness. Attached Figure Description

[0096] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0097] In the attached diagram:

[0098] Figure 1 This is a flowchart illustrating the implementation of the agricultural non-point source pollution risk assessment method integrating multi-source heterogeneous data analysis according to the present invention.

[0099] Figure 2 This is a structural block diagram of the agricultural non-point source pollution risk assessment device that integrates multi-source heterogeneous data analysis according to the present invention. Detailed Implementation

[0100] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0101] like Figure 1 The diagram shown illustrates a flowchart of an agricultural non-point source pollution risk assessment method based on multi-source heterogeneous data analysis, provided as an embodiment of the present invention. The method includes the following steps:

[0102] Step S101: Obtain multi-source heterogeneous data from agricultural non-point source pollution monitoring, and preprocess the multi-source heterogeneous data to obtain a standard dataset;

[0103] First, acquire multi-source heterogeneous data for monitoring agricultural non-point source pollution, specifically including meteorological data, soil data, agricultural activity data, geographic environment data, and water quality response data;

[0104] Meteorological data includes rainfall, sunshine duration, air pressure, wind speed, air temperature and humidity, precipitation intensity, and precipitation frequency; among them, rainfall and air temperature and humidity are collected by ground meteorological stations on an hourly basis; precipitation intensity and precipitation frequency are supplemented by microwave remote sensing satellites around the clock.

[0105] Soil data includes soil moisture, total nitrogen content, total phosphorus content, organic matter, mechanical composition, and field capacity. Soil moisture and total nitrogen / total phosphorus content were collected every 30 minutes by a soil sensor array buried at a depth of 20 cm; soil organic matter, mechanical composition, and field capacity were obtained through laboratory analysis of soil samples.

[0106] Agricultural activity data includes fertilizer application (including nitrogen and phosphorus ratio), irrigation duration, irrigation water volume, crop type, planting density, and livestock pollution discharge coefficient;

[0107] The geographic environment data includes digital elevation model (DEM), land use types (such as irrigated land, orchards, etc.) and vegetation cover (NDVI); the digital elevation model and land use types are acquired by satellite remote sensing with a spatial resolution of ≥2m (DEM terrain accuracy ±0.5m); the vegetation cover is acquired by satellite remote sensing combined with lidar, and terrain parameters such as slope and aspect are further extracted based on the DEM.

[0108] Water quality response data includes total nitrogen content, total phosphorus content, chemical oxygen demand (COD), ammonia nitrogen content, and suspended sediment content in surface runoff.

[0109] Furthermore, the acquired multi-source heterogeneous data is preprocessed to obtain a standard dataset. The preprocessing includes data cleaning, data imputation, data standardization, and spatiotemporal alignment. The preprocessing steps lay the data foundation for accurate assessment of agricultural non-point source pollution.

[0110] Please continue to refer to Figure 1 The agricultural non-point source pollution risk assessment method of this embodiment further includes the following steps:

[0111] Step S102: Construct an indicator system based on the analytic hierarchy process, which includes a target layer, a criterion layer, and a scheme layer. Introduce a spatiotemporal attention interaction submodule into the indicator system to dynamically adjust the initial weights of the criterion layer indicators.

[0112] The target layer includes the Agricultural Non-point Source Pollution Risk Index (RI), used to comprehensively quantify the risk level of agricultural non-point source pollution in a specific area (or field). The criterion layer includes rainfall factors, soil factors, agricultural activity factors, topographic factors, and water quality response factors. Among them, rainfall factors include indicators such as rainfall amount, rainfall intensity, and rainfall frequency; soil factors include indicators such as soil moisture, total soil nitrogen content, total soil phosphorus content, and soil organic matter; agricultural activity factors include indicators such as fertilizer application amount, nitrogen-phosphorus ratio of fertilizer, irrigation duration, irrigation water volume, and crop type; topographic factors include indicators such as slope, aspect, and runoff potential derived from the digital elevation model (DEM); and water quality response factors include indicators such as total nitrogen concentration, total phosphorus concentration, and chemical oxygen demand (COD) in surface runoff. The scenario layer consists of specific evaluation indicators corresponding to each criterion layer. For example, the scenario layer indicators for rainfall factors include daily rainfall, hourly rainfall intensity, and continuous rainfall duration; and the scenario layer indicators for soil factors include soil moisture and total soil nitrogen content.

[0113] Furthermore, the initial weights of the criteria layer indicators are determined using the analytic hierarchy process (AHP), and a judgment matrix is ​​constructed for the criteria layer indicators. , among which, element The importance of criterion Ci relative to criterion Cj is indicated by assigning values ​​using a 1-9 scale, where 1 indicates that both are equally important and 9 indicates that Ci is extremely more important than Cj.

[0114] Based on the judgment matrix, the initial weight vector of each criterion layer index is calculated using the eigenvalue method, and a consistency check is performed.

[0115] Furthermore, in this embodiment of the application, a spatiotemporal attention interaction submodule is introduced into the indicator system. The step of dynamically adjusting the initial weights of the criterion layer indicators includes:

[0116] In the time dimension of the spatiotemporal attention interaction submodule, the temporal correlation of meteorological time-series data is captured through an attention mechanism, and the temporal attention weight coefficient is calculated. Specifically, based on the temporal dependence of agricultural non-point source pollution pollutant generation and migration, the dynamic impact of criterion-level indicators on pollution risk under different time windows is quantified through an attention mechanism, as follows: Multi-scale time windows are set to cover different temporal correlation scenarios: In the short-term response window, preprocessed high-frequency meteorological time-series data within the window are selected; in the medium-term cumulative window, cumulative meteorological data and agricultural activity data within the window are selected, and further temporal correlation feature extraction is performed. For rainfall, soil moisture, wind speed, solar radiation intensity, and atmospheric pressure in the criterion layer, temporal correlation features are extracted for each time window to ensure that the features are directly related to the pollution mechanism.

[0117] Furthermore, in calculating the time attention weight coefficients of each criterion-level indicator, the correlation strength between the time-series correlation feature vector of the criterion-level indicator under the time window and the time query vector trained based on the time-series feature library of historical pollution events is first quantified. Then, the exp transformation value of this correlation strength is divided by the sum of the exp transformation values ​​of the corresponding correlation strengths of all criterion-level indicators (softmax normalization) to finally obtain the weight coefficients and ensure that the sum of all coefficients is 1.

[0118] In the spatial dimension of the spatiotemporal attention interaction submodule, topographic features are extracted using digital elevation model (DEM) data, and differentiated spatial attention weight coefficients are assigned to different topographic units. Specifically, based on the strong spatial heterogeneity of agricultural non-point source pollution, topographic features are extracted using DEM data, and differentiated weights are assigned to the criterion-level indicators of different topographic units. Among these, topographic parameters (slope, aspect, runoff potential) extracted from DEM data are used to divide the assessment area into gentle topographic units, mild-slope topographic units, and steep-slope topographic units to ensure matching with pollution migration mechanisms. At the same time, preprocessed spatial data is matched to each topographic unit to form a spatial feature combination of topography-soil-vegetation. For the criterion-level indicators, spatial correlation features are extracted based on the characteristics of the topographic units. Similarly, the spatial attention weight coefficients are calculated using spatial contribution and softmax normalization methods to calculate the spatial attention weight coefficients of each criterion-level indicator in different topographic units.

[0119] In the dynamic adjustment of weights, based on the latest collected multi-source data, the weights of the criteria layer indicators are corrected by combining the temporal attention weight coefficient and the spatial attention weight coefficient to obtain a spatiotemporally adaptive weight vector. Specifically, the weights of each criterion layer indicator are corrected by adopting a coupling mode of initial weights, temporal weights and spatial weights.

[0120] This application introduces a spatiotemporal attention interaction submodule, which captures the temporal correlation between the rate of change of continuous rainfall intensity and the time window after fertilization in the time dimension, and combines DEM to extract terrain features and assign differentiated weights in the spatial dimension, thus solving the problem that static weights cannot adapt to sudden rainstorms and terrain differences.

[0121] Please continue to refer to Figure 1 The agricultural non-point source pollution risk assessment method of this embodiment further includes the following steps:

[0122] Step S103: Construct a cross-modal feature interaction model to perform deep feature fusion on the temporal and spatial data of the standard dataset; wherein, a temporal processing network is used to process the temporal data and extract temporal feature vectors; a spatial feature extraction network is used to process the spatial data and extract spatial feature vectors; calculate the modal correlation matrix, and perform weighted operations on the temporal feature vectors and spatial feature vectors through the modal correlation matrix to generate cross-modal features;

[0123] Multi-source time-series data on agricultural non-point source pollution generally suffer from problems such as mixed multi-scale features and significant noise interference.

[0124] In one implementation of this application, the step of extracting the temporal feature vector includes:

[0125] A variational mode decomposition algorithm is used to adaptively divide multi-source time-series data into frequency bands, decomposing the original time-series data into multi-scale subsequences. In this embodiment, the variational mode decomposition is achieved by constructing and solving a constrained variational problem to realize the adaptive decomposition of the original time-series data. Let the original time-series data be x(t), and the original time-series data is decomposed into K modal components. Each modal component corresponds to a temporal feature at a certain scale;

[0126] The objective function is expressed as:

[0127] ;

[0128] The constraints are expressed as follows: ;

[0129] In the formula, Represents the set of modal components and center frequency set Optimize; Represents the time derivative operator. Represents the Dirac function, The center frequencies of each modality are represented by , and * indicates the convolution operation. Represents a complex exponential function; Let represent the square of the L2 norm, and j represent the imaginary unit. The long-term trend component, seasonal cycle component, and sudden disturbance component are obtained by solving the alternating direction multiplier method.

[0130] This application embodiment uses the Alternating Directional Multiplier Method (ADMM) to adaptively decompose multi-source time-series data of agricultural non-point source pollution into multi-scale components. Among them, the long-term trend component corresponds to the interannual cumulative effect of agricultural non-point source pollution, the seasonal cycle component corresponds to the periodic correlation between crop growth cycle and pollution emissions, and the sudden disturbance component corresponds to short-term strong impact events such as rainstorms and fertilization.

[0131] As a preferred embodiment, to address the shortcomings of traditional Variational Mode Decomposition (VMD) parameters which rely on empirical selection, this embodiment also employs the Crowned Porcupine Optimization (CPO) algorithm to achieve adaptive parameter optimization. This includes: constructing a fitness function with the dual objectives of minimizing mode reconstruction error and maximizing mode independence; during the optimization process, setting the population size to 30, the number of iterations to 50, and the parameter search range to include: the number of modes K and the quadratic penalty factor. The optimization operation includes: simulating the foraging and defense behaviors of hogs, obtaining the optimal parameter combination through population updates and local search; and outputting the parameters corresponding to the minimum value of the fitness function.

[0132] Furthermore, the step of extracting temporal feature vectors also includes: constructing a nested gated recurrent unit network, with the main branch using bidirectional gated recurrent units to capture full temporal dependencies, and modeling the full temporal data as follows:

[0133]

[0134] in, This represents the hidden state at time t in the main branch. The input data represents time t; This represents the hidden state of the main branch at time t-1. This represents the hidden state of the main branch at time t+1; Indicates a bidirectional gated loop unit;

[0135] The first sub-branch extracts trend features for the long-term trend component using a large-time-step gated cyclic unit, represented as:

[0136]

[0137] In the formula, This represents the long-term trend feature vector extracted at time t, and GRU stands for Gated Recurrent Unit. This represents the long-term trend component data input at time t; This represents the long-term trend feature vector input at time t-1;

[0138] The second sub-branch uses an adaptive periodic gating cyclic unit to extract periodic features for the seasonal periodic component, represented as:

[0139]

[0140]

[0141] In the formula, T represents the length of the seasonal cycle. This represents the Pearson coefficient, used to measure the linear correlation between two time series data. This represents the seasonal periodic component data input at time t; This represents the periodic feature vector extracted at time t; Represents the periodic eigenvector at time tT;

[0142] The third sub-branch employs an attention-enhanced gated recurrent unit for sudden disturbance components, represented as:

[0143]

[0144]

[0145] In the formula, This represents the attention weight coefficient. Indicates the attenuation coefficient. Indicates the moment when the sudden disturbance occurred; This represents the feature vector of the sudden disturbance extracted at time t. This represents the sudden disturbance component data input at time t. This represents the characteristic vector of the sudden disturbance at time tT. The characteristic variable representing the sudden disturbance at time t-1;

[0146] Furthermore, in order to perform attention fusion on multi-scale temporal features, this embodiment introduces an attention fusion mechanism, which dynamically adjusts the weights based on the contribution of each component to historical pollution events, thereby achieving the fusion of multi-scale temporal features.

[0147] Specifically, in this embodiment, the output features of the main branch and sub-branch are weighted and integrated through an attention fusion mechanism. The weight coefficients of the attention fusion mechanism are dynamically adjusted based on the contribution of each component to the historical pollution event, and finally a time-series feature vector containing multi-scale temporal correlation information is generated.

[0148] In this embodiment of the application, the weight coefficients of the attention fusion mechanism are... Represented as:

[0149]

[0150] In the formula, This represents the contribution coefficient of the branch. This represents the hidden state of the k-th sub-branch at time t; This represents the normalization operation on the hidden state; This represents the temporal feature vector extracted by the i-th sub-branch at time t;

[0151] In this embodiment, the contribution coefficient of the branch is expressed as: In the formula, N represents the number of historical pollution events. Let be the average expert score of the k-th sub-branch in the i-th historical event;

[0152] The final output of this embodiment after fusion is a multi-scale fused temporal feature vector. This temporal feature vector not only retains the temporal data of each sub-branch, but also highlights the component features with high contribution in the current scene through attention weights.

[0153] The spatial feature extraction network constructed in this embodiment uses a convolutional neural network (CNN) to extract local spatial features and a graph convolutional network (GCN) to capture global spatial relationships. This network architecture can learn high-dimensional and abstract spatial feature vectors from multi-source spatial data, providing spatial dimension input for subsequent cross-modal feature fusion.

[0154] Specifically, in one implementation of this application, the step of using a spatial feature extraction network to process spatial data and extract spatial feature vectors includes:

[0155] The spatial data in the standard dataset are scaled to obtain micro-scale data of field plots and macro-scale data of sub-watersheds; the pollution correlation coefficients of spatial data at different scales are calculated and expressed as follows:

[0156]

[0157] In the formula, This represents spatial data at the s-th scale, where s=1 represents the field scale and s=2 represents the sub-basin scale. This represents historical pollution load data at the corresponding scale. This represents the Pearson correlation coefficient, used to quantify the strength of the association between spatial data at different scales and pollution risk;

[0158] A multi-branch spatial feature extraction network is constructed, including a field-scale branch and a sub-watershed-scale branch, and each branch embeds an agricultural non-point source pollution spatial feature encoding module;

[0159] For the field-scale branch, a convolutional neural network (CNN) combined with a soil texture encoding module is used to extract features from the grid data. The CNN's convolutional layers use 3×3 kernels, ReLU activation function, stride 1, and padding 1 (to ensure the feature map size after convolution is consistent with the input). The pooling layers use 2×2 max pooling with a stride of 2 to downsample and retain key features. In the output layer, after CNN processing, a local spatial feature map is obtained.

[0160] For the sub-basin scale branch, a graph convolutional neural network (GCN) combined with a hydrological connectivity coding module is used to convert the sub-basin spatial grid into a graph structure. The nodes of the graph structure are sub-basin grid cells, and the edge weights are the hydrological connectivity between cells.

[0161] In this embodiment, the hydrological connectivity is determined based on the flow direction matrix calculated from the DEM and the cumulative discharge data, and is expressed as follows:

[0162]

[0163] In the formula, Let i be the hydrological connectivity between grid cells i and j in the sub-basin. The cumulative flow from unit i to j The maximum cumulative flow within the sub-basin; GCN captures the correlation characteristics between grid cells and pollutant confluence within the sub-basin through graph convolution operations;

[0164] Based on pollution correlation coefficient A dynamic attention fusion layer is constructed to perform weighted fusion of the local feature vectors output by the field-scale branch and the global feature vectors output by the sub-basin-scale branch;

[0165] Furthermore, in a preferred embodiment of this application, a soil-topography synergistic attention coefficient is introduced. , Based on the ternary correlation between soil texture, slope, and pollution loss in historical data, it is calculated and expressed as follows:

[0166]

[0167] In the formula, The mutual information entropy is represented by T, slope data, soil texture data, land use data, hydrological connectivity data, and pollution loss data.

[0168] The weighted fusion spatial feature vector is represented as: In the formula, Represents the fused spatial feature vector; Represents the collaborative attention coefficient. This represents the local feature vector output by the field-scale branch. This represents the global feature vector output by the sub-basin scale branch; , This represents the pollution correlation coefficient.

[0169] Therefore, this application uses a time-series processing network to extract the time-series characteristics of cumulative rainfall and changes in soil nitrogen content after fertilization, and a spatial feature extraction network to extract the spatial characteristics of soil water retention capacity and crop root interception potential. Then, it uses a modal correlation matrix to quantify the synergistic effects of rainfall on sandy soil and irrigation on clay soil, thereby reducing the prediction error of cross-modal features on pollution load.

[0170] Please continue to refer to Figure 1 The agricultural non-point source pollution risk assessment method of this embodiment further includes the following steps:

[0171] Step S104: Using the pollution risk index as the target variable, the criterion-level indicators and agricultural activity factors as causal variables, and season and topography as confounding variables, a directed acyclic causal graph is formed; the confounding variables are fixed in the causal graph, and the direct causal impact of each causal variable on the target variable is evaluated separately, identifying and eliminating spurious correlation factors; the causal association strength of each criterion-level indicator is calculated using cross-modal features, and the spatiotemporal adaptive weight vector is corrected in combination with the causal association strength to obtain the causal corrected weight vector;

[0172] Step S105: Calculate and output the pollution risk index based on the causal-corrected weight vector.

[0173] Furthermore, in step S104 of this application embodiment, in the step of forming a directed acyclic causal graph with pollution risk index as the target variable, criterion layer indicators and agricultural activity factors as causal variables, and season and topography as confounding variables:

[0174] In the target variable layer, the pollution risk index RI is calculated by coupling the weight vector after causal correction with the standardized values ​​of the criteria layer indicators, and is expressed as follows:

[0175]

[0176] In the formula, This represents the weight of the i-th indicator in the criterion layer after causal correction. This represents the standardized value of the criteria-level indicator;

[0177] In the causal variable layer, the pollution process is divided into generation stage nodes and migration stage nodes;

[0178] In this embodiment of the application, the generation stage node includes rainfall and soil moisture in the criterion layer, as well as fertilization intensity and irrigation frequency of agricultural activity factors;

[0179] In this embodiment of the application, the migration stage nodes include auxiliary indicators of the criterion layer such as wind speed, atmospheric pressure, and topographically derived factors. The topographically derived factors include runoff potential, which is obtained based on digital elevation model (DEM) data by calculating slope, aspect, and cumulative flow, and is expressed as:

[0180]

[0181] in, Accumulate flow for grid cells, This represents the maximum cumulative traffic in the region.

[0182] In the confounding variable layer, season and topography are decomposed into dynamic and static sub-variables that are strongly correlated with the pollution process. Seasonal variables are decomposed into crop growth period, and topographic variables are decomposed into soil texture and slope grade.

[0183] Furthermore, in this embodiment of the application, the step of forming a directed acyclic causal graph further includes:

[0184] Construct causal edges for the generation stage, including: directed edges between fertilization intensity and soil moisture, directed edges between fertilization intensity and pollution risk index, and directed edges between soil moisture and pollution risk index.

[0185] Among them, the marginal weights of fertilization intensity and soil moisture The mutual information entropy is calculated based on historical data of fertilizer application, soil moisture, and pollution load, and is expressed as:

[0186]

[0187] In the formula, Represents mutual information entropy. , Represents any two causal variables, This represents the sum of the mutual information entropy of all causal variable pairs; The pollution risk index is represented by F1, which represents the fertilization intensity, and C2 represents the soil moisture.

[0188] Marginal weights of fertilization intensity and pollution risk index Based on the conditional mutual information between fertilization intensity and pollution risk index from historical data, used to eliminate the influence of confounding factors such as soil moisture, it is expressed as:

[0189]

[0190] In the formula, For conditional mutual information, V represents the direct causal contribution of fertilization intensity to pollution risk after excluding the interference of soil moisture, with soil moisture C2 as the condition; V represents the causal variable. This represents confounding variables related to the causal variable V. It represents the sum of conditional mutual information between all causal variables and the RI under the corresponding confounding conditions;

[0191] The side weights of soil moisture C2 and pollution risk index RI are calculated based on the conditional mutual information between soil moisture and pollution risk index in historical data, and are used to eliminate the influence of confounding factors such as fertilization intensity, and are expressed as follows:

[0192]

[0193] In the formula, This represents the edge weights of soil moisture C2 and the pollution risk index RI.

[0194] Construct causal edges for the migration phase; specifically, establish directed edges between wind speed C3 and atmospheric ammonia deposition risk, and between runoff potential T1 and surface runoff pollution risk; where the edge weights for wind speed and atmospheric ammonia deposition risk are defined. The coefficients were calibrated based on the linear regression coefficients of wind speed and atmospheric ammonia concentration; the side weights for runoff potential T1 and surface runoff pollution risk were... The linear regression coefficients based on slope and runoff nitrogen content were calibrated.

[0195] Construct causal edges for confounding variables, and establish directed edges between confounding sub-variables and causal variables, including edge weights for crop growth cycle and fertilization intensity. The marginal weights of soil texture and runoff potential .

[0196] Specifically, the steps of fixing confounding variables in the causal diagram, individually evaluating the direct causal impact of each causal variable on the target variable, and identifying and eliminating spurious correlation factors include:

[0197] Hierarchical fixation is implemented based on the attribute differences of confounding variables in the causal graph, including fixation of static confounding variables and fixation of dynamic confounding variables. Static confounding variables refer to confounding factors whose attributes are stable in the long term and do not change dynamically with time / scene, such as soil texture, terrain slope, and long-term crop type. Fixation methods include stratification / grouping based on inherent attributes. Dynamic confounding variables refer to confounding factors whose attributes change dynamically with time and scene, such as soil moisture, wind speed, and crop growth period. Fixation methods include combining dynamic characteristics with time-series control / conditional constraints.

[0198] Furthermore, for the causal graph after fixing the confounding variables, the backdoor adjustment and causal effect decomposition methods are used to calculate the direct causal effect value of each causal variable on the pollution risk index RI separately.

[0199] In the backdoor path blocking, for the causal variable X, the backdoor paths of all causal variables, confounding variables and pollution risk indices in the causal graph are identified, and the path is blocked by fixing the confounding variables, leaving only the direct path from X to RI.

[0200] In average direct causal effect In the calculation, the formula is expressed as:

[0201]

[0202] In the formula, This indicates that the intervention causal variable X takes the actual observed value. Z = z represents the baseline value where the intervention X has a mean of 0, and Z = z represents the value of the confounding variable after fixing. This represents the conditional expectation; for continuous causal variables, the marginal effect is calculated using derivatives, expressed as:

[0203]

[0204] In the formula, This represents taking the partial derivative with respect to X; X represents the causal variable.

[0205] A dual verification mechanism of causal strength and data consistency is introduced to identify and eliminate causal variables with spurious correlations.

[0206] Furthermore, the step of calculating the causal correlation strength of each criterion layer index using cross-modal features, and then correcting the spatiotemporal adaptive weight vector based on the causal correlation strength to obtain the causally corrected weight vector includes:

[0207] Cross-modal features of dimension d With criteria layer indicator set Perform association modeling;

[0208] Combining direct causal effect values Calculate the criteria layer indicators causal relationship strength , represented as:

[0209]

[0210] In the formula, The dynamic equilibrium coefficient is adaptively adjusted based on cross-modal features and the Pearson correlation coefficient of the causal graph. Indicates the dynamic correlation coefficient;

[0211] Finally, the spatiotemporal adaptive weight vector is corrected.

[0212] Furthermore, in the embodiments of this application, the cross-modal features of dimension d are... With criteria layer indicator set The steps for performing association modeling include:

[0213] Cross-modal features are integrated using a fully connected neural network. Projecting onto the criterion layer index space yields the projected features. ;

[0214] Wherein, the output of the i-th neuron Represented as:

[0215]

[0216] In the formula, This represents the learnable weight coefficients. Indicates the bias term. Indicates cross-modal features;

[0217] Calculate projection features Corresponding criteria level indicators Dynamic correlation coefficient , represented as:

[0218]

[0219] In the formula, T represents the length of the time series. Represents the characteristic mean. Indicates the mean of the indicators and the dynamic correlation coefficient. The value range is [-1, 1]. Positive values ​​indicate positive correlation, negative values ​​indicate negative correlation, and the larger the absolute value, the more significant the spatiotemporal correlation between the cross-modal projection features and the criterion layer indicators.

[0220] Furthermore, in this embodiment of the application, the step of correcting the spatiotemporal adaptive weight vector includes:

[0221] Causal-corrected weight vector In the diagram, the i-th element is represented as:

[0222]

[0223] In the formula, in the formula, This represents the weight of the m-th indicator after causal correction. This represents the i-th element in the weight vector after causal correction; , These represent the spatiotemporal adaptive weights of the i-th and j-th indicators, respectively; , These represent the strength of the causal association between the i-th and j-th indicators and the risk of agricultural non-point source pollution, respectively. , Both represent exponential enhancement terms, used to strengthen the weighting of indicators with high causal correlation strength and suppress indicators with low correlation.

[0224] Furthermore, rules for agricultural non-point source pollution are introduced to constrain the correction results. If the correction results violate the constraints, they are re-optimized using the Lagrange multiplier method.

[0225] This application uses season and topography as confounding variables, and evaluates the direct causal impact of causal variables separately by fixing the confounding variables. It identifies rainfall, soil moisture, and fertilization intensity as driving factors, which greatly improves the consistency between the weights after causal correction and the actual pollution driving mechanism. Furthermore, it uses cross-modal features to calculate the causal association strength, which integrates the direct causal effect and the feature correlation degree. This can correct the spatiotemporal adaptive weights and avoid the problem of reduced accuracy of pollution risk assessment in extreme scenarios.

[0226] In summary, this application significantly improves the real-time response capability of agricultural non-point source pollution risk assessment, and enhances the accuracy of key source area location and pollution control effectiveness.

[0227] Figure 2 This is a structural block diagram of an agricultural non-point source pollution risk assessment device that integrates multi-source heterogeneous data analysis according to an embodiment of this application.

[0228] like Figure 2 As shown, the agricultural non-point source pollution risk assessment device includes:

[0229] The data processing module 201 is used to acquire multi-source heterogeneous data from agricultural non-point source pollution monitoring and to preprocess the multi-source heterogeneous data to obtain a standard dataset.

[0230] The indicator system construction module 202 is used to construct an indicator system containing a target layer, a criterion layer and a scheme layer based on the analytic hierarchy process. A spatiotemporal attention interaction submodule is introduced into the indicator system to dynamically adjust the initial weights of the indicators in the criterion layer.

[0231] The feature extraction and fusion module 203 is used to construct a cross-modal feature interaction model and perform deep feature fusion on temporal and spatial data of the standard dataset. Specifically, a temporal processing network is used to process the temporal data and extract temporal feature vectors; a spatial feature extraction network is used to process the spatial data and extract spatial feature vectors; a modal correlation matrix is ​​calculated, and the temporal feature vectors and spatial feature vectors are weighted through the modal correlation matrix to generate cross-modal features.

[0232] The causal association module 204 is used to form a directed acyclic causal graph with the pollution risk index as the target variable, the criterion-level indicators and agricultural activity factors as causal variables, and season and topography as confounding variables. The confounding variables are fixed in the causal graph, and the direct causal impact of each causal variable on the target variable is evaluated separately to identify and eliminate spurious correlation factors. The causal association strength of each criterion-level indicator is calculated using cross-modal features, and the spatiotemporal adaptive weight vector is corrected in combination with the causal association strength to obtain the causal corrected weight vector.

[0233] Risk assessment module 205 is used to calculate and output the pollution risk index based on the causal-corrected weight vector.

[0234] In one embodiment, a computer device is provided, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the agricultural non-point source pollution risk assessment method integrating multi-source heterogeneous data analysis provided in the above embodiments.

[0235] In another embodiment, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the agricultural non-point source pollution risk assessment method that integrates multi-source heterogeneous data analysis provided in the above embodiments.

[0236] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the various embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.

[0237] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the agricultural non-point source pollution risk assessment method that integrates multi-source heterogeneous data analysis described above.

[0238] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0239] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

[0240] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for agricultural non-point source pollution risk assessment by fusing multi-source heterogeneous data analysis, characterized in that, The method comprises the following steps: acquiring multi-source heterogeneous data of agricultural non-point source pollution monitoring, and pre-processing the multi-source heterogeneous data to obtain a standard data set; an index system including a target layer, a criterion layer and a scheme layer is constructed based on an analytic hierarchy process, a time-space attention interaction submodule is introduced in the index system, and initial weights of the criterion layer indexes are dynamically adjusted; a cross-modal feature interaction model is constructed to perform deep feature fusion on time-series data and spatial data of the standard data set; wherein, a time-series processing network is used to process the time-series data and extract a time-series feature vector; a spatial feature extraction network is used to process the spatial data and extract a spatial feature vector; a modal correlation matrix is calculated, and the time-series feature vector and the spatial feature vector are weighted and operated through the modal correlation matrix to generate cross-modal features; a pollution risk index is taken as a target variable, criterion layer indexes and agricultural activity factors are taken as cause variables, and seasons and terrain are taken as confounding variables to form a directed acyclic causal graph; the confounding variables of the causal graph are fixed, the direct causal influence of each cause variable on the target variable is evaluated separately, pseudo-correlation factors are identified and excluded; the causal correlation strength of each criterion layer index is calculated using the cross-modal features, and the time-space adaptive weight vector is corrected combined with the causal correlation strength to obtain a causal corrected weight vector; the pollution risk index is calculated based on the causal corrected weight vector and output. 2.The agricultural non-point source pollution risk assessment method of fusing multi-source heterogeneous data analysis according to claim 1, characterized in that, The step of introducing the time-space attention interaction submodule in the index system and dynamically adjusting the initial weights of the criterion layer indexes comprises: in the time dimension of the time-space attention interaction submodule, the time-series correlation of meteorological time-series data is captured through an attention mechanism, and a time attention weight coefficient is calculated; in the spatial dimension of the time-space attention interaction submodule, terrain features are extracted combined with digital elevation model data, and different spatial attention weight coefficients are allocated to different terrains; in the weight dynamic adjustment, the criterion layer index weights are corrected through the combination of the time attention weight coefficient and the spatial attention weight coefficient to obtain a time-space adaptive weight vector. 3.The agricultural non-point source pollution risk assessment method of fusing multi-source heterogeneous data analysis according to claim 2, characterized in that, The step of using the time-series processing network to process the time-series data and extract the time-series feature vector comprises: Adopting the variational mode decomposition algorithm to divide the multi-source time series data into adaptive frequency bands, the original time series data is decomposed into multi-scale sub-sequences; wherein, the original time series data is decomposed into K mode components , the objective function is represented as: The constraint condition is expressed as: wherein denotes the time derivative, denotes the Dirac function, denotes the center frequency of each mode, and denotes the convolution operation; the long-term trend component, the seasonal periodic component and the sudden disturbance component are obtained by solving the alternating direction multiplier method; a nested gated recurrent unit network is constructed, a bidirectional gated recurrent unit is used in a main branch to model the full time-series data; a first sub-branch uses a large time step gated recurrent unit to extract trend features for long-term trend components; a second sub-branch uses an adaptive period gated recurrent unit to extract periodic features for seasonal period components; a third sub-branch uses an attention enhanced gated recurrent unit; The output features of the main branch and the sub-branch are weighted and integrated through an attention fusion mechanism, a weight coefficient of the attention fusion mechanism is dynamically adjusted based on the contribution of each component to the historical pollution event, and a time sequence feature vector containing multi-scale time sequence correlation information is generated; wherein the weight coefficient of the attention fusion mechanism is represented as: In the formula, denotes the contribution degree coefficient of the branch; denotes the hidden state of the kth sub-branch at time t; denotes the normalization operation on the hidden state; denotes the time sequence feature vector extracted by the ith sub-branch at time t. 4.The agricultural non-point source pollution risk assessment method of fusing multi-source heterogeneous data analysis according to claim 3, characterized in that, The step of using the spatial feature extraction network to process the spatial data and extract the spatial feature vector comprises: the spatial data in the standard data set are divided into scales to obtain field micro-scale data and sub-basin macro-scale data; a pollution correlation degree coefficient of different scale spatial data is calculated and expressed as: wherein represents spatial data of the s-th scale; represents historical pollution load data of the corresponding scale, represents a Pearson correlation coefficient; a multi-branch spatial feature extraction network is constructed, including a field scale branch and a sub-basin scale branch, and each branch is embedded with an agricultural non-point source pollution spatial feature coding module; for the field scale branch, a convolutional neural network (CNN) combined with a soil texture coding module is used to extract features from the grid data; For the sub-basin scale branch, a graph convolutional neural network (GCN) combined with a hydrological connectivity coding module is used to convert the sub-basin spatial grid into a graph structure, with the nodes of the graph structure being the sub-basin grid cells and the edge weights being the hydrological connectivity between the cells; the hydrological connectivity is determined based on the flow direction matrix calculated from the DEM and the cumulative flow data, and is expressed as: wherein, is the hydrological connectivity between sub-basin grid cells i and j, is the cumulative flow from cell i to j, is the maximum cumulative flow within the sub-basin; GCN captures the associated features of the grid cell-pollutant confluence within the sub-basin through graph convolution operations; Based on pollution correlation coefficient A dynamic attention fusion layer is constructed to weight and fuse the local feature vectors output by the field scale branch and the global feature vectors output by the sub-basin scale branch. Introducing soil-terrain synergistic attention coefficient The weighted fused spatial feature vector is represented as: In the formula, denotes the spatial feature vector after fusion; denotes the cooperative attention coefficient, denotes the local feature vector of the field scale branch output, denotes the global feature vector of the sub-basin scale branch output; , denotes the pollution correlation coefficient.

5. The method of agricultural non-point source pollution risk assessment fusing multi-source heterogeneous data analysis according to claim 4, characterized in that, In the step of forming a directed acyclic causal graph with the pollution risk index as the target variable, the criterion layer indicators and the agricultural activity factors as the cause variables, and the seasons and the terrain as the confounding variables, the pollution risk index RI is calculated based on the coupling of the causal correction weight vector and the criterion layer indicator standardized value, and is expressed as: In the target variable layer, the pollution risk index RI is calculated based on the coupling of the causal correction weight vector and the criterion layer indicator standardized value, and is expressed as: In the formula, represents the weight of the i-th index of the post-cause-effect correction criterion layer, represents the criterion layer index standardization value; In the cause variable layer, it is divided into generation stage nodes and migration stage nodes according to the pollution process; In the confounding variable layer, the seasons and the terrain are decomposed into dynamic sub-variables and static sub-variables that are strongly associated with the pollution process, and the season variable is decomposed into crop growth periods, and the terrain variable is decomposed into soil texture and slope grade.

6. The method of agricultural non-point source pollution risk assessment fusing multi-source heterogeneous data analysis according to claim 5, characterized in that, The step of forming a directed acyclic causal graph also includes: The generation stage causal edges include: the directed edges of fertilizer intensity and soil moisture, the directed edges of fertilizer intensity and pollution risk index, and the directed edges of soil moisture and pollution risk index; The migration stage causal edge is constructed; wherein, the directed edge of wind speed C3 and atmospheric ammonia deposition risk, the directed edge of convergence potential T1 and surface runoff pollution risk are established; wherein, the edge weight of wind speed and atmospheric ammonia deposition risk The edge weight of convergence potential T1 and surface runoff pollution risk is calibrated based on the linear regression coefficient of slope and runoff nitrogen content The edge weight of convergence potential T1 and surface runoff pollution risk is calibrated based on the linear regression coefficient of slope and runoff nitrogen content Constructing the causal edges of the hybrid variable, establishing the directed edges of the hybrid sub-variables and the cause variables, including the edge weights of the crop growth cycle and the fertilization intensity and the edge weights of the soil texture and the convergence potential .

7. The method of agricultural non-point source pollution risk assessment fusing multi-source heterogeneous data analysis according to claim 6, characterized in that, The step of fixing the confounding variables of the causal graph, separately evaluating the direct causal effects of each cause variable on the target variable, and identifying and excluding pseudo-correlated factors includes: Based on the attribute differences of the confounding variables in the causal graph, hierarchical fixing is implemented, including static confounding variable fixing and dynamic confounding variable fixing; for the causal graph after fixing the confounding variables, the backdoor adjustment and causal effect decomposition methods are used to separately calculate the direct causal effect values of each cause variable on the pollution risk index RI; A double-checking mechanism of causal strength and data consistency is introduced to identify and exclude pseudo-correlated cause variables.

8. The method of agricultural non-point source pollution risk assessment fusing multi-source heterogeneous data analysis according to claim 7, characterized in that, The step of calculating the causal correlation strength of each criterion layer indicator using cross-modal feature and correcting the spatio-temporal adaptive weight vector to obtain the causal correction weight vector includes: cross-modal features of dimension d are obtained with the criterion layer indicator set correlation modeling is performed; combining the direct causal effect values , the criterion layer index of the causal correlation strength , is expressed as: In the formula, represents a dynamic balance coefficient, which is adaptively adjusted based on the Pearson correlation coefficient of the cross-modal features and the causal diagram; represents a dynamic correlation coefficient; The step of correcting the spatio-temporal adaptive weight vector includes:

9. The method of agricultural non-point source pollution risk assessment fusing multi-source heterogeneous data analysis according to claim 8, characterized in that, The cross-modal feature of dimension d is obtained by The criterion layer indicator set The step of performing correlation modeling comprises: The cross-modal features are projected to a criterion layer index space by a fully connected neural network to obtain projected features The cross-modal features are projected to a criterion layer index space by a fully connected neural network to obtain projected features ; wherein the output of the i-th neuron is represented as: ; wherein the output of the i-th neuron is represented as: In the formula, represents a learnable weight coefficient, represents a bias term, represents a cross-modal feature; Computing projection features Dynamic correlation with corresponding criteria layer indicators is expressed as:​ where T represents the length of time series, denotes the feature mean, denotes the indicator mean, dynamic correlation coefficient The value range of is [-1, 1], and the positive value represents the positive correlation, the negative value represents the negative correlation, and the greater the absolute value represents the more significant the spatiotemporal correlation between the cross-modal projection features and the criterion layer indicators.

10. The method of agricultural non-point source pollution risk assessment fusing multi-source heterogeneous data analysis according to claim 9, characterized in that, The step of correcting the spatio-temporal adaptive weight vector includes: causally corrected weight vector In particular, the i-th element is expressed as: In the formula, represents the weight of the mth index after causal correction, represents the ith element in the weight vector after causal correction; , respectively represent the ith and jth index spatiotemporal adaptive weights; , respectively represent the causal correlation strength between the ith and jth index and the agricultural non-point source pollution risk, , both represent exponential enhancement terms; The rule in the field of agricultural non-point source pollution is introduced to constrain the correction result, and if the correction result violates the constraint, the Lagrange multiplier method is used for re-optimization.