Machine learning-based method and system for predicting water quality of rural non-point sources
By constructing a dynamic migration network and fusing multimodal features, the pollution contribution is quantified, solving the problems of dynamic decay and lag in the migration process of rural non-point source pollutants in existing technologies, and achieving more accurate water quality prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN UNIV JINCHENG INST
- Filing Date
- 2026-02-05
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies struggle to accurately characterize the dynamic decay and lag-accumulation synergistic effects of non-point source pollutants in rural areas during their migration process, resulting in insufficient prediction accuracy across different regions and rainfall scenarios.
By acquiring basic data, a dynamic migration network is constructed, spatiotemporal state evolution and multimodal feature fusion are performed, pollution contribution characteristics are quantified, a water quality prediction function is established, and predictions are made in conjunction with real-time environmental data.
It significantly improves the accuracy and reliability of water quality prediction in complex scenarios of rural non-point source pollution, and overcomes the simplification dependence of existing mechanistic models on complex underlying surface conditions and migration processes.
Smart Images

Figure CN121641279B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and more specifically, to a method and system for predicting rural non-point source water quality based on machine learning. Background Technology
[0002] Rural non-point source pollution control is a core challenge for watershed water environment management and green agricultural development. Pollution sources are characterized by dispersion, wide reach, and randomness. Pollutant output intensity is driven by rainfall runoff and closely coupled with underlying surface conditions such as land use, soil properties, and farming activities, forming a complex "source-runoff-sink" relationship. This makes accurately predicting its impact on water bodies exceptionally difficult. Currently, mainstream prediction technologies in this field mainly rely on mechanistic models based on fixed parameters and simplified physical processes. While these models conceptually describe the pollution migration pathways, their construction heavily depends on homogenizing prior knowledge of highly heterogeneous farmland underlying surfaces and hydrological pathways. This makes it difficult to accurately characterize the synergistic effects of pollutant dynamic attenuation along spatial pathways, and the lag and accumulation over time during actual migration. Consequently, the adaptability and prediction accuracy of these models are significantly limited when dealing with different regions and rainfall characteristics.
[0003] Given the shortcomings of the existing technologies, there is an urgent need for a machine learning-based method and system for predicting non-point source water quality in rural areas. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for predicting rural non-point source water quality based on machine learning, in order to improve the aforementioned problems. To achieve the above objective, the technical solution adopted by this invention is as follows:
[0005] Firstly, this application provides a machine learning-based method for predicting non-point source water quality in rural areas, including:
[0006] Acquire basic data, including time series of nitrogen and phosphorus concentrations, time series of rainfall, data on the distribution of farmland land use types, and data on soil organic matter content at rural non-point source water quality monitoring stations;
[0007] Based on the aforementioned basic data, pollutant migration paths are constructed. By analyzing the hydrological connectivity between geographical units and simulating the attenuation process of pollutants with runoff, a dynamic migration network is obtained.
[0008] Spatiotemporal state evolution is performed based on the dynamic migration network. By coupling the pollutant flux of each node in the migration network with the time series, a dynamic graph structure is obtained.
[0009] Based on the dynamic graph structure, multimodal feature fusion is performed. By analyzing the spatiotemporal dependence of pollutant flux in the graph structure, the combined impact of factors such as farmland plots, soil properties, and crop types on water quality changes at monitoring points is quantified, and pollution contribution characteristics are obtained.
[0010] A prediction function is constructed based on the pollution contribution characteristics to obtain a water quality prediction function.
[0011] Water quality values are generated based on the water quality prediction function. The water quality prediction result is obtained by inputting real-time environmental data into the water quality prediction function and calculating.
[0012] Secondly, this application also provides a machine learning-based rural non-point source water quality prediction system, including:
[0013] The acquisition module is used to acquire basic data, which includes time series of nitrogen and phosphorus concentrations, time series of rainfall, distribution data of farmland land use types, and soil organic matter content data of rural non-point source water quality monitoring points.
[0014] The analysis module is used to construct pollutant migration paths based on the basic data. By analyzing the hydrological connectivity between geographical units and simulating the attenuation process of pollutants with runoff, a dynamic migration network is obtained.
[0015] The evolution module is used to perform spatiotemporal state evolution according to the dynamic migration network. By coupling the pollutant flux of each node in the migration network with the time series, a dynamic graph structure is obtained.
[0016] The fusion module is used to perform multimodal feature fusion based on the dynamic graph structure. By analyzing the spatiotemporal dependence of pollutant flux in the graph structure, it quantifies the joint impact of factors such as farmland plots, soil properties and crop types on water quality changes at monitoring points, and obtains pollution contribution characteristics.
[0017] A construction module is used to construct a prediction function based on the pollution contribution characteristics to obtain a water quality prediction function.
[0018] The output module generates water quality values based on the water quality prediction function. By inputting real-time environmental data into the water quality prediction function and calculating, the water quality prediction result is obtained.
[0019] The beneficial effects of this invention are as follows:
[0020] This invention dynamically constructs a pollutant migration network from basic data, and quantifies the pollution contribution based on the spatiotemporal state of the network's evolution and the fusion of multimodal features. Finally, it establishes a prediction function that can reflect the actual "source-pathway-sink" process, effectively overcoming the dependence of existing mechanistic models on complex underlying surface conditions and simplified migration processes, and significantly improving the accuracy and reliability of water quality prediction in complex scenarios of rural non-point source pollution. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating a machine learning-based method for predicting non-point source water quality in rural areas, as described in an embodiment of the present invention.
[0023] Figure 2 This is a schematic diagram of the structure of a rural non-point source water quality prediction system based on machine learning, as described in an embodiment of the present invention.
[0024] Figure 3 This is a schematic diagram of a rural non-point source water quality prediction device based on machine learning, as described in an embodiment of the present invention.
[0025] The diagram is labeled as follows: 800, a machine learning-based rural non-point source water quality prediction device; 801, processor; 802, memory; 803, multimedia component; 804, I / O interface; 805, communication component; 901, acquisition module; 902, analysis module; 903, evolution module; 904, fusion module; 905, construction module; 906, output module. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0027] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0028] Example 1:
[0029] This embodiment provides a machine learning-based method for predicting non-point source water quality in rural areas.
[0030] See Figure 1 The figure shows that the method includes steps S100 to S600.
[0031] Step S100: Obtain basic data, which includes time series of nitrogen and phosphorus concentrations, time series of rainfall, distribution data of farmland land use types, and data of soil organic matter content at rural non-point source water quality monitoring points;
[0032] Understandably, nitrogen and phosphorus concentration time series data come from regular manual or automatic monitoring at fixed monitoring points within the watershed, reflecting the final state of water quality; rainfall time series data originate from rain gauges or weather radar data, representing the primary driver of pollution migration; farmland land use type distribution data are obtained through remote sensing image interpretation to identify the spatial location of pollutant sources, such as paddy fields, dry land, and orchards; and soil organic matter content data are obtained through on-site sampling and testing or by querying soil type maps, serving as an intrinsic attribute for assessing the soil's own potential for pollutant retention and release. These data collectively constitute a complete dataset from driving factors and pollution source characteristics to environmental responses.
[0033] Step S200: Construct pollutant migration paths based on basic data. By analyzing the hydrological connectivity between geographical units and simulating the attenuation process of pollutants with runoff, a dynamic migration network is obtained.
[0034] It is important to note that the core of this step lies in transforming static geospatial data into a system that reflects the dynamic transport process of pollutants. The approach does not rely on pre-defined, simplified hydrological models, but rather employs a machine learning-driven approach. Starting with the actual spatial distribution of farmland plots and rainfall characteristics, a preliminary migration framework is constructed by analyzing the natural confluence paths and hydrological connectivity between plots. Then, factors such as soil properties are introduced, and machine learning methods are used to simulate the physicochemical attenuation of pollutants on this path network. This allows the network to not only characterize the direction of water flow but also embed the dynamic characteristics of pollutant intensity changes during migration, resulting in a more realistic dynamic migration network with attenuation properties.
[0035] Step S300: Perform spatiotemporal state evolution based on the dynamic migration network. By coupling the pollutant flux of each node in the migration network with the time series, a dynamic graph structure is obtained.
[0036] Understandably, this step unfolds and couples the pollutant flux carried by each node in the dynamic migration network (such as farmland exits and ditch nodes) in each rainfall event according to the time series, thereby elevating the static network topology into a four-dimensional spatiotemporal entity. This dynamic graph structure allows for a complete description of the state (pollutant flux) of each node at any given time and its connections with neighboring nodes, providing a foundation for subsequent analysis of pollutant propagation patterns in the spatiotemporal continuum.
[0037] Step S400: Perform multimodal feature fusion based on the dynamic graph structure. By analyzing the spatiotemporal dependence of pollutant flux in the graph structure, quantify the joint impact of factors such as farmland plots, soil properties, and crop types on water quality changes at monitoring points, and obtain pollution contribution characteristics.
[0038] It should be noted that the goal of this step is to extract the core features that play a decisive role in downstream water quality from the complex spatiotemporal dynamic graph. The method involves in-depth analysis of the complex relationships contained within the graph, including not only the water flow path associations determined by spatial proximity, but also the temporal pollution release patterns caused by crop growth cycles, fertilization activities, and the interception or promotion of pollutants by soil properties. By quantifying the combined impact of these multimodal factors on pollutants throughout their entire path from source to monitoring point, the complex graph structure information is ultimately condensed into a quantitative contribution of each pollution source to the water quality target, achieving data-driven dimensionality reduction and feature extraction of complex physical processes.
[0039] Step S500: Construct a prediction function based on the pollution contribution characteristics to obtain a water quality prediction function;
[0040] Understandably, this step involves constructing a mapping function that accurately captures the complex nonlinear relationship between "pollution contribution" and "water quality response." This process does not employ a fixed function form but rather adaptively learns, based on historical data, how the pollution contribution characteristics of the upstream region synergize under different rainfall, soil, and crop scenarios, ultimately manifesting as concentration changes at downstream monitoring points. The resulting water quality prediction function is essentially a computable mathematical relationship embedding the formation and migration mechanisms of non-point source pollution, possessing the ability to directly extrapolate water quality results from real-time environmental inputs.
[0041] Step S600: Generate water quality values based on the water quality prediction function. This is done by inputting real-time environmental data into the water quality prediction function and calculating the water quality prediction results.
[0042] It should be noted that this step represents the practical application phase of the model, the core of which lies in performing online predictions using the established prediction function. When new real-time environmental data (such as rainfall) is available, the system triggers the prediction process, substituting the real-time data into the established pollution migration network and contribution quantification relationship to generate a pollution contribution feature vector for the current moment. This vector is then input into the water quality prediction function for calculation. This process combines static historical data models with dynamic real-time observation data, ultimately outputting a quantitative prediction of water quality for a future period, providing a direct basis for risk management.
[0043] Further, step S200 includes steps S210 to S230.
[0044] Step S210: Based on the distribution data of farmland land use types and the time series of rainfall, perform hydrological response unit division processing, identify continuous farmland plots with similar runoff generation and confluence characteristics and determine the priority flow paths between them to obtain a preliminary hydrological connection network.
[0045] Step S220: Based on the preliminary hydrological connection network and soil organic matter content data, conduct pollutant migration and attenuation simulation processing. By simulating the concentration attenuation of dissolved pollutants along the migration path due to soil adsorption and biodegradation, a pollutant migration path with attenuation characteristics is obtained.
[0046] Step S230: Based on the pollutant migration path and nitrogen and phosphorus concentration time series, perform dynamic network weight correction processing, and obtain the actual flux efficiency of the path by using the time delay relationship of concentration changes at upstream and downstream monitoring points.
[0047] Specifically, in step S210, firstly, based on the distribution data of farmland land use types, continuous plots with similar cultivation patterns and topography are identified as basic hydrological response units. Then, their runoff characteristics are analyzed in conjunction with historical rainfall data. Based on surface runoff patterns, priority flow paths for water and pollutant migration between plots are delineated, forming a preliminary hydrological connection network. On this basis, in step S220, soil organic matter content data is introduced. For each migration path in the network, the concentration decay of dissolved pollutants during runoff migration is simulated due to adsorption by soil particles and biodegradation by microorganisms along the path. This transforms the abstract connection path into a pollutant migration path with physicochemical decay characteristics. Furthermore, in step S230, the theoretical path is empirically corrected using actual monitored nitrogen and phosphorus concentration time series. By analyzing the temporal delay and amplitude decay relationship of peak concentrations at upstream and downstream monitoring points, the pollutant transport efficiency of each path in the actual environment is calculated. Based on this, the path weights are dynamically corrected, ultimately resulting in a high-precision dynamic migration network that not only has a physical basis but is also calibrated with actual data.
[0048] Further, step S310 includes steps S310 to S330.
[0049] Step S310: Based on the dynamic migration network and rainfall time series, perform migration event segmentation processing, and obtain discrete pollution migration event sequences by identifying effective rainfall events and their driven independent pollution migration processes.
[0050] Step S320: Based on the pollution migration event sequence and dynamic migration network, perform node state temporal processing. By aligning and resampling the pollutant flux on each node in each event according to the time step, a node state evolution sequence is obtained.
[0051] Step S330: Based on the node state evolution sequence and the topological connection relationship of the dynamic migration network, perform spatiotemporal graph construction processing. By coupling the node state at each time step with the connection relationship between nodes, a dynamic graph structure is obtained.
[0052] Specifically, step S310 first performs structured processing on continuous rainfall and water quality data, identifying valid rainfall events exceeding the runoff generation threshold based on the rainfall time series. Each event and its triggered complete pollution migration process from pollution initiation to runoff completion are treated as an independent analysis unit, thus segmenting the continuous time series into discrete pollution migration event sequences. This process decomposes the complex continuous pollution load into multiple discrete, independently analyzable migration events, laying the foundation for subsequent refined analysis. Based on this, step S320, for each independent pollution migration event, utilizes the constructed dynamic migration network to align the pollutant flux data changing over time at each node in the network (such as the outlets of different farmland plots and the confluence of ditches) during the event's duration according to a unified time step. Interpolation and resampling are used to generate the state evolution sequence of each node on the standard time axis. This process solves the problem that data from different locations may be out of sync in time, allowing the states of all nodes to be compared and correlated on a unified time dimension. Finally, step S330 couples the state evolution sequences of all nodes at multiple time steps obtained above with the inherent topology of the dynamic migration network that represents spatial connectivity. That is, at each specific time step, not only is the pollutant flux state of each node recorded, but the spatial connectivity between nodes is also preserved, thereby constructing a graph structure that changes dynamically over time. This dynamic graph structure fully captures the temporal dependence of pollutant transmission on the spatial network and realizes the quantitative characterization of the spatiotemporal coordinated process of non-source pollution migration.
[0053] Further, step S410 includes steps S410 to S430.
[0054] Step S410: Based on the dynamic graph structure, perform spatiotemporal correlation pattern extraction processing. By separating the water flow path correlation caused by the spatial proximity of farmland plots and the pollution release correlation caused by the temporal sequence of crop growth cycle in the graph, the spatiotemporal correlation pattern is obtained.
[0055] Step S420: Based on the spatiotemporal correlation pattern and soil organic matter content data, perform source-sink relationship quantification. By calculating the balance between the interception effect of high organic matter soil on pollutants and the absorption efficiency of different crop types, obtain the intensity weight of each plot as a pollution source or sink.
[0056] Step S430: Based on the intensity weight and spatiotemporal correlation pattern of each plot, perform contribution transmission aggregation processing. By tracing back along the spatiotemporal correlation path and aggregating the weighted influence of all upstream plots on downstream monitoring points, the pollution contribution characteristics are obtained.
[0057] Step S410 first performs a deep analysis of the dynamic graph structure, decomposing the complex relationships implicit in the graph into two dominant modes: one is the water flow path relationship determined by the spatial proximity of farmland plots, and the other is the pollution release relationship determined by the crop growth cycle and the temporality of agricultural activities. By separating these two mechanisms that respectively dominate spatial transport and temporal release, a more essential spatiotemporal relationship mode can be extracted, thus clearly distinguishing the "channel" characteristics and "source" characteristics of pollutant migration. Preferably, the goal of this step is to analyze the dynamic graph structure... Decomposed into spatial association patterns and temporal association patterns, respectively using spatial adjacency matrices. Time-series correlation matrix express.
[0058] Spatial adjacency matrix : Indicates water flow path associations determined solely by geographic spatial connectivity. In the formula, For the plot of land and The distance of the water flow path between them The characteristic decay length represents the rate at which spatial influence decays with distance.
[0059] Temporal correlation matrix : Indicates the correlation of pollution release determined by the consistency of crop growth cycle. In the formula, and The parameters are used to characterize the crop growth stage or fertilization period on plots i and j (the range is 0 to 1, where 0 is the fallow period and 1 is the vigorous growth period after fertilization).
[0060] Ultimately, the spatiotemporal correlation pattern is represented by the Hadamard product (i.e., dot product) of these two matrices to capture the synergistic effect of space and time:
[0061] ;
[0062] In the formula, It is a spatiotemporal correlation pattern; This is the dot product operator.
[0063] Step S420 utilizes this spatiotemporal correlation model and combines it with soil organic matter content data to precisely quantify the "source-sink" role of each farmland plot. Specifically, this process calculates the dynamic balance between the adsorption and retention capacity of high-organic-matter soil for pollutants and the nutrient absorption efficiency of different crop types throughout their growth cycle. This transforms the qualitative judgment of whether a plot is a net pollution source or a net pollution sink into a quantifiable intensity weight. Preferably, the goal of this step is to calculate the source-sink role of each plot. Intensity weight Quantify its net effect as a pollution source or sink. The formula is:
[0064] ;
[0065] In the formula, For the plot of land Soil retention factors, , The retention factor is the unit organic matter content. For the plot of land Soil organic matter content; Assigned fixed empirical values to crop absorption factors based on different crop types.
[0066] Finally, step S430 uses the intensity weights of each plot as basic input. Based on the transmission path defined by the spatiotemporal correlation pattern extracted in the first step, it traces all upstream pollution sources backward along the network, starting from the downstream monitoring points. Then, using a weighted aggregation algorithm, it accumulates the contributions of each upstream plot along the migration path, thus obtaining an integrated pollution contribution feature that comprehensively reflects the pollution source, migration path, and spatiotemporal delay. The pollution contribution feature formula is:
[0067] ;
[0068] In the formula, All upstream land parcels to downstream monitoring points The characteristics of the total pollution contribution; upstream plot The intensity weight; For upstream plots to monitoring point The spatiotemporal correlation strength; For monitoring points The set of all upstream plots; From arrive The set of all intermediate geographical units (nodes) traversed along the migration path; This refers to the sequence number of the intermediate geographical unit (node); intermediate nodes on the path The rate of decay of pollutants.
[0069] Further, step S500 includes steps S510 to S530.
[0070] Step S510: Based on the pollution contribution characteristics and nitrogen and phosphorus concentration time series, perform contribution-response relationship modeling. By establishing a nonlinear mapping between the pollution contribution characteristics and the concentration change values of water quality monitoring points under different rainfall scenarios, the initial prediction function is obtained.
[0071] Further, step S510 includes steps S511 to S513.
[0072] Step S511: Based on the rainfall time series, perform rainfall scenario pattern classification processing. Through cluster analysis, historical rainfall events are divided into multiple typical scenario patterns according to their intensity, duration and spatial distribution characteristics to obtain rainfall scenario classification results.
[0073] Step S512: Based on the rainfall scenario classification results and pollution contribution characteristics, a scenario-based response surface is constructed. By identifying the dynamic balance between the ability of rainfall kinetic energy to initiate surface pollutants and the runoff dilution effect in each type of rainfall scenario, a pollution response model is obtained.
[0074] Step S513: Based on the pollution response mode, perform initial prediction function integration processing. By establishing a selector mechanism that dynamically maps real-time rainfall characteristics to the corresponding response mode and calculates its output, the initial prediction function is obtained.
[0075] In some preferred embodiments, step S511 first performs pattern mining on the historical rainfall time series, and uses cluster analysis to classify the rainfall events into several typical rainfall scenario patterns based on their intensity, duration, and spatial distribution characteristics, such as high-intensity short-duration and low-intensity long-duration categories, thereby transforming continuous rainfall data into discrete and representative rainfall scenario classification results. Based on this, step S512 analyzes the historical pollution contribution characteristics and actual concentration change data at water quality monitoring points for each identified rainfall scenario, focusing on identifying the initiation and scouring capacity of rainfall kinetic energy on surface pollutants and the resulting runoff under this type of rainfall characteristic. The dynamic balance between the dilution effect on pollutants and the pollution response model is used to construct a unique pollution response model for each type of rainfall scenario. This model quantitatively describes the intrinsic law of the conversion of pollution contribution to water quality concentration under specific rainfall conditions. Step S513 is responsible for integrating the above series of pollution response models for different scenarios into an operable prediction whole. The core of this whole is to construct an intelligent selector mechanism. This mechanism can dynamically map the real-time input rainfall characteristics to the most matching rainfall scenario category and activate the corresponding pollution response model for calculation. Finally, through this process of "scenario recognition - pattern matching - output calculation", it is integrated into an initial prediction function. The initial prediction function is:
[0076] ;
[0077] In the formula, Indicates the time step Predicted water quality concentration; An index for rainfall scenario types, These represent three typical scenarios: light rain with long duration, moderate rain with regular duration, and heavy rain with short duration. This is the scenario selector function; , , This is a scenario-specific coefficient, representing the combined effect of different rainfall types on pollutant transport, including rainfall kinetic energy and dilution capacity. For the first Pollution contribution characteristics of individual farmland plots; For the first Area (hectares) of each farmland plot. For the first Soil organic matter content of each farmland plot; For time The intensity of rainfall; For time The duration of rainfall; This represents the total number of farmland plots within the upstream catchment area.
[0078] Step S520: Based on the initial prediction function, optimize the function's generalization ability. By analyzing the distribution pattern of prediction errors under different combinations of soil properties and crop types, adaptively adjust the confidence weights of the contribution features to obtain the optimized prediction function.
[0079] Step S530: Based on the optimized prediction function, perform time-dependent embedding processing. By incorporating historical prediction results as feedback signals into the function, the cumulative and lag effects of pollution contribution are captured, and a water quality prediction function is obtained.
[0080] Specifically, step S510 first utilizes historical data to establish the core relationships of the prediction model. Instead of constructing a single global model, it categorizes historical scenarios based on characteristics such as rainfall intensity and duration. For each type of rainfall scenario, it establishes a unique nonlinear mapping relationship between pollution contribution characteristics and downstream water quality concentration changes, resulting in an initial prediction function composed of multiple sub-models. This function can more precisely respond to pollution output characteristics driven by different rainfall events. In step S520, to improve the generalization ability of this initial function in complex farmland environments, the system deeply analyzes the distribution patterns of its prediction errors under different combinations of soil properties and crop types, identifying those prediction errors that occur under specific soil-crop combinations (such as planting vegetables in sandy soil). The system identifies areas with systematic biases and adaptively adjusts the confidence weights of pollution contribution characteristics in the final prediction accordingly. This reduces interference from the specific underlying surface conditions, resulting in an optimized prediction function corrected for error patterns. Step S530 further considers the inherent temporal characteristics of non-point source pollution, treating the optimized prediction function as a system with memory capabilities. By incorporating historical prediction results as feedback signals into the current calculation, the function learns and simulates the cumulative and hysteresis effects of pollutants in the environment due to infiltration, slow flow, and other processes. The resulting water quality prediction function not only includes a static mapping between contribution and concentration but also embeds the dynamic temporal patterns of pollution migration, significantly improving the simulation capability for actual water quality changes. The water quality prediction function is:
[0081] ;
[0082] In the formula, For time step The final water quality prediction results; This is the error correction strength coefficient; , The prediction bias is specific to soil type and crop type. This value is obtained by analyzing the distribution of historical prediction errors under different soil-crop combinations. It is the hyperbolic tangent function; This represents the historical residual effect coefficient. The decay time constant of the residual effect; This forms an exponentially decaying kernel function to simulate the decay of pollutant residues over time. For time The predicted value is used to provide feedback on historical information; This is an index variable representing the number of time delay steps.
[0083] Further, step S600 includes steps S610 to S630.
[0084] Step S610: Based on real-time environmental data, perform prediction trigger condition judgment processing. By monitoring whether the real-time rainfall intensity and duration reach the preset runoff threshold, determine the timing of starting the prediction and obtain the prediction trigger signal.
[0085] Step S620: Based on the predicted trigger signal and real-time environmental data, perform real-time contribution feature generation processing, and substitute real-time rainfall and soil moisture data into the constructed pollution migration network and contribution quantification relationship to obtain the real-time pollution contribution feature vector.
[0086] Step S630: Based on the real-time pollution contribution feature vector and water quality prediction function, calculate the water quality results and perform lag effect compensation processing. Input the feature vector into the prediction function to calculate the initial value, and superimpose the residual impact of the previous pollution contribution to obtain the water quality prediction result.
[0087] Specifically, step S610 first defines an intelligent triggering mechanism for initiating the prediction task. This involves continuously monitoring rainfall intensity and duration in real-time environmental data and comparing it with a pre-set runoff threshold based on the characteristics of the watershed's underlying surface. Only when real-time rainfall reaches or exceeds this threshold, possessing the potential to form surface runoff and carry pollutants, will the system generate a prediction trigger signal. This ensures the economy and timeliness of the prediction process and avoids invalid calculations for processes without actual pollution migration. Once a prediction trigger signal is obtained, step S620 immediately begins, rapidly extrapolating by substituting the real-time collected rainfall, soil moisture, and other data into the pollution migration network and contribution quantification relationship constructed in the preceding steps. The system dynamically calculates the real-time pollution contribution feature vector of each pollution source in the entire watershed to the downstream monitoring points under the current meteorological and soil moisture conditions. This vector comprehensively reflects the pollution output potential at this moment. In step S630, the generated real-time pollution contribution feature vector is input into the finally trained water quality prediction function to calculate the initial predicted value of the water quality concentration. Subsequently, the system further introduces compensation for the lag effect. Through a feedback loop containing a decay function, the residual impact of the previous pollution contribution represented by the recent historical prediction value is quantified and superimposed on the initial value, thereby obtaining a more accurate and reliable final water quality prediction result that reflects both the current instantaneous input and the historical cumulative effect.
[0088] Example 2:
[0089] like Figure 2 As shown, this embodiment provides a rural non-point source water quality prediction system based on machine learning. The system includes:
[0090] The acquisition module 901 is used to acquire basic data, which includes time series of nitrogen and phosphorus concentrations, time series of rainfall, distribution data of farmland land use types, and data of soil organic matter content at rural non-point source water quality monitoring points.
[0091] Analysis module 902 is used to construct pollutant migration paths based on basic data. By analyzing the hydrological connectivity between geographical units and simulating the attenuation process of pollutants with runoff, a dynamic migration network is obtained.
[0092] Evolution module 903 is used to perform spatiotemporal state evolution based on the dynamic migration network. By coupling the pollutant flux of each node in the migration network with the time series, a dynamic graph structure is obtained.
[0093] The fusion module 904 is used to perform multimodal feature fusion based on the dynamic graph structure. By analyzing the spatiotemporal dependence of pollutant flux in the graph structure, it quantifies the joint impact of factors such as farmland plots, soil properties and crop types on water quality changes at monitoring points, and obtains pollution contribution characteristics.
[0094] Module 905 is used to construct a prediction function based on the pollution contribution characteristics to obtain a water quality prediction function.
[0095] The output module 906 generates water quality values based on the water quality prediction function. It obtains water quality prediction results by inputting real-time environmental data into the water quality prediction function and performing calculations.
[0096] In one specific embodiment of this application, the analysis module includes:
[0097] The first analysis unit is used to divide the hydrological response units based on the distribution data of farmland land use types and the time series of rainfall. By identifying continuous farmland plots with similar runoff generation and confluence characteristics and determining the priority flow paths between them, a preliminary hydrological connection network is obtained.
[0098] The second analysis unit is used to simulate pollutant migration and attenuation based on preliminary hydrological connection network and soil organic matter content data. By simulating the concentration attenuation of dissolved pollutants along the migration path due to soil adsorption and biodegradation, the pollutant migration path with attenuation characteristics is obtained.
[0099] The third analysis unit is used to perform dynamic network weight correction processing based on the pollutant migration path and nitrogen and phosphorus concentration time series. By using the time delay relationship of concentration changes at upstream and downstream monitoring points to invert the actual flux efficiency of the path, the dynamic migration network is obtained.
[0100] In one specific embodiment of this application, the evolution module includes:
[0101] The first evolutionary unit is used to segment migration events based on the dynamic migration network and rainfall time series. By identifying effective rainfall events and the independent pollution migration processes they drive, discrete pollution migration event sequences are obtained.
[0102] The second evolution unit is used to perform temporal processing of node states based on the pollution migration event sequence and dynamic migration network. By aligning and resampling the pollutant flux on each node in each event according to the time step, the node state evolution sequence is obtained.
[0103] The third evolutionary unit is used to construct a spatiotemporal graph based on the node state evolution sequence and the topological connection relationship of the dynamic migration network. By coupling the node state at each time step with the connection relationship between nodes, a dynamic graph structure is obtained.
[0104] Example 3:
[0105] Corresponding to the above method embodiments, this embodiment also provides a rural non-point source water quality prediction device based on machine learning. The rural non-point source water quality prediction device based on machine learning described below and the rural non-point source water quality prediction method based on machine learning described above can be referred to and corresponded to each other.
[0106] Figure 3 This is a block diagram illustrating a machine learning-based rural non-point source water quality prediction device 800 according to an exemplary embodiment. Figure 3 As shown, the rural non-point source water quality prediction device 800 based on machine learning may include: a processor 801 and a memory 802. The rural non-point source water quality prediction device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0107] The processor 801 controls the overall operation of the machine learning-based rural non-point source water quality prediction device 800 to complete all or part of the steps in the aforementioned machine learning-based rural non-point source water quality prediction method. The memory 802 stores various types of data to support the operation of the machine learning-based rural non-point source water quality prediction device 800. This data may include, for example, instructions for any application or method operating on the machine learning-based rural non-point source water quality prediction device 800, as well as application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 805 is used for wired or wireless communication between the machine learning-based rural non-point source water quality prediction device 800 and other devices. Wireless communication methods include Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, and an NFC module.
[0108] In an exemplary embodiment, a machine learning-based rural non-point source water quality prediction device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the aforementioned machine learning-based rural non-point source water quality prediction method.
[0109] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the above-described machine learning-based rural non-point source water quality prediction method. For example, the computer-readable storage medium may be the memory 802 including the program instructions, which may be executed by a processor 801 of a machine learning-based rural non-point source water quality prediction device 800 to complete the above-described machine learning-based rural non-point source water quality prediction method.
[0110] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A machine learning-based method for predicting non-point source water quality in rural areas, characterized in that, include: Acquire basic data, including time series of nitrogen and phosphorus concentrations, time series of rainfall, data on the distribution of farmland land use types, and data on soil organic matter content at rural non-point source water quality monitoring points; Based on the aforementioned basic data, pollutant migration paths are constructed. By analyzing the hydrological connectivity between geographical units and simulating the attenuation process of pollutants with runoff, a dynamic migration network is obtained. Spatiotemporal state evolution is performed based on the dynamic migration network. By coupling the pollutant flux of each node in the migration network with the time series, a dynamic graph structure is obtained. Based on the dynamic graph structure, multimodal feature fusion is performed. By analyzing the spatiotemporal dependence of pollutant flux in the graph structure, the combined impact of farmland plots, soil properties, and crop type factors on water quality changes at monitoring points is quantified, and pollution contribution characteristics are obtained. A prediction function is constructed based on the pollution contribution characteristics to obtain a water quality prediction function. Water quality values are generated based on the water quality prediction function. The water quality prediction result is obtained by inputting real-time environmental data into the water quality prediction function and calculating. The process of constructing pollutant migration pathways based on the aforementioned basic data includes: Based on the farmland land use type distribution data and the rainfall time series, hydrological response units are divided. By identifying continuous farmland plots with similar runoff generation and confluence characteristics and determining the priority flow paths between them, a preliminary hydrological connection network is obtained. Based on the preliminary hydrological connection network and the soil organic matter content data, pollutant migration and attenuation simulation processing is carried out. By simulating the concentration attenuation of dissolved pollutants along the migration path due to soil adsorption and biodegradation, a pollutant migration path with attenuation characteristics is obtained. Based on the pollutant migration path and nitrogen and phosphorus concentration time series, dynamic network weight correction processing is performed. The actual flux efficiency of the path is obtained by inverting the time-series relationship of concentration changes at upstream and downstream monitoring points. The spatiotemporal state evolution based on the dynamic migration network includes: Based on the dynamic migration network and the rainfall time series, migration event segmentation is performed. By identifying effective rainfall events and the independent pollution migration processes they drive, discrete pollution migration event sequences are obtained. Based on the pollution migration event sequence and the dynamic migration network, node state temporal processing is performed. By aligning and resampling the pollutant flux on each node in each event according to the time step, a node state evolution sequence is obtained. Based on the node state evolution sequence and the topological connection relationship of the dynamic migration network, a spatiotemporal graph construction process is performed. By coupling the node state at each time step with the connection relationship between nodes, a dynamic graph structure is obtained. The multimodal feature fusion based on the dynamic graph structure includes: Based on the dynamic graph structure, spatiotemporal correlation pattern extraction processing is performed. By separating the water flow path correlation caused by the spatial proximity of farmland plots and the pollution release correlation caused by the temporal sequence of crop growth cycle in the dynamic graph structure, the spatiotemporal correlation pattern is obtained. Based on the spatiotemporal correlation pattern and the soil organic matter content data, the source-sink relationship is quantified. By calculating the balance between the interception effect of high organic matter soil on pollutants and the absorption efficiency of different crop types, the intensity weight of each plot as a pollution source or sink is obtained. Based on the intensity weights and spatiotemporal correlation patterns of each land parcel, contribution transmission and aggregation processing is performed. By tracing back along the spatiotemporal correlation path and aggregating the weighted influence of all upstream land parcels on downstream monitoring points, pollution contribution characteristics are obtained.
2. The method for predicting rural non-point source water quality based on machine learning according to claim 1, characterized in that, The prediction function is constructed based on the pollution contribution characteristics, including: Based on the pollution contribution characteristics and the nitrogen and phosphorus concentration time series, a contribution-response relationship modeling process is performed. By establishing a nonlinear mapping between the pollution contribution characteristics and the concentration change values at water quality monitoring points under different rainfall scenarios, an initial prediction function is obtained. Based on the initial prediction function, the function generalization ability is optimized. By analyzing the distribution pattern of prediction error under different combinations of soil properties and crop types, the confidence weight of the contribution feature is adaptively adjusted to obtain the optimized prediction function. Based on the optimized prediction function, time-dependent embedding processing is performed. By incorporating historical prediction results as feedback signals into the function, the cumulative and lag effects of pollution contributions are captured, resulting in a water quality prediction function.
3. The method for predicting rural non-point source water quality based on machine learning according to claim 1, characterized in that, The generation of water quality values based on the water quality prediction function includes: Based on the real-time environmental data, a prediction trigger condition judgment process is performed. By monitoring whether the real-time rainfall intensity and duration reach the preset runoff threshold, the timing for initiating the prediction is determined, and a prediction trigger signal is obtained. Based on the predicted trigger signal and real-time environmental data, real-time contribution feature generation processing is performed. Real-time rainfall and soil moisture data are substituted into the constructed pollution migration network and contribution quantification relationship to obtain the real-time pollution contribution feature vector. Based on the real-time pollution contribution feature vector and the water quality prediction function, water quality results are calculated and lag effect compensation is performed. The feature vector is input into the prediction function to calculate the initial value, and the residual influence of previous pollution contributions is superimposed to obtain the water quality prediction result.
4. The method for predicting rural non-point source water quality based on machine learning according to claim 2, characterized in that, Based on the pollution contribution characteristics and the nitrogen and phosphorus concentration time series, contribution-response relationship modeling is performed, including: Based on the rainfall time series, rainfall scenario pattern classification is performed. Through cluster analysis, historical rainfall events are divided into multiple typical scenario patterns according to their intensity, duration and spatial distribution characteristics, and the rainfall scenario classification results are obtained. Based on the rainfall scenario classification results and the pollution contribution characteristics, a scenario-based response surface is constructed. By identifying the dynamic balance between the ability of rainfall kinetic energy to initiate surface pollutants and the runoff dilution effect in each type of rainfall scenario, a pollution response model is obtained. Based on the pollution response mode, an initial prediction function integration process is performed. The initial prediction function is obtained by establishing a selector mechanism that dynamically maps real-time rainfall characteristics to the corresponding response mode and calculates its output.
5. A rural non-point source water quality prediction system based on machine learning, characterized in that, include: The acquisition module is used to acquire basic data, which includes time series of nitrogen and phosphorus concentrations, time series of rainfall, distribution data of farmland land use types, and soil organic matter content data of rural non-point source water quality monitoring points. The analysis module is used to construct pollutant migration paths based on the basic data. By analyzing the hydrological connectivity between geographical units and simulating the attenuation process of pollutants with runoff, a dynamic migration network is obtained. The evolution module is used to perform spatiotemporal state evolution according to the dynamic migration network. By coupling the pollutant flux of each node in the migration network with the time series, a dynamic graph structure is obtained. The fusion module is used to perform multimodal feature fusion based on the dynamic graph structure. By analyzing the spatiotemporal dependence of pollutant flux in the graph structure, it quantifies the joint impact of farmland plots, soil properties and crop types on water quality changes at monitoring points, and obtains pollution contribution characteristics. A construction module is used to construct a prediction function based on the pollution contribution characteristics to obtain a water quality prediction function. The output module generates water quality values based on the water quality prediction function. It obtains the water quality prediction results by inputting real-time environmental data into the water quality prediction function and performing calculations. The analysis module includes: The first analysis unit is used to perform hydrological response unit division processing based on the farmland land use type distribution data and the rainfall time series. By identifying continuous farmland plots with similar runoff generation and confluence characteristics and determining the priority flow paths between them, a preliminary hydrological connection network is obtained. The second analysis unit is used to perform pollutant migration and attenuation simulation processing based on the preliminary hydrological connection network and the soil organic matter content data. By simulating the concentration attenuation of dissolved pollutants along the migration path due to soil adsorption and biodegradation, a pollutant migration path with attenuation characteristics is obtained. The third analysis unit is used to perform dynamic network weight correction processing based on the pollutant migration path and nitrogen and phosphorus concentration time series. By using the time delay relationship of concentration changes at upstream and downstream monitoring points to invert the actual flux efficiency of the path, the dynamic migration network is obtained. The evolution module includes: The first evolutionary unit is used to perform migration event segmentation processing based on the dynamic migration network and the rainfall time series, and to obtain a discrete pollution migration event sequence by identifying effective rainfall events and the independent pollution migration processes they drive. The second evolution unit is used to perform node state temporal processing based on the pollution migration event sequence and the dynamic migration network. By aligning and resampling the pollutant flux on each node in each event according to the time step, a node state evolution sequence is obtained. The third evolution unit is used to perform spatiotemporal graph construction processing based on the node state evolution sequence and the topological connection relationship of the dynamic migration network. By coupling the node state at each time step with the connection relationship between nodes, a dynamic graph structure is obtained. The multimodal feature fusion based on the dynamic graph structure includes: Based on the dynamic graph structure, spatiotemporal correlation pattern extraction processing is performed. By separating the water flow path correlation caused by the spatial proximity of farmland plots and the pollution release correlation caused by the temporal sequence of crop growth cycle in the dynamic graph structure, the spatiotemporal correlation pattern is obtained. Based on the spatiotemporal correlation pattern and the soil organic matter content data, the source-sink relationship is quantified. By calculating the balance between the interception effect of high organic matter soil on pollutants and the absorption efficiency of different crop types, the intensity weight of each plot as a pollution source or sink is obtained. Based on the intensity weights and spatiotemporal correlation patterns of each land parcel, contribution transmission and aggregation processing is performed. By tracing back along the spatiotemporal correlation path and aggregating the weighted influence of all upstream land parcels on downstream monitoring points, pollution contribution characteristics are obtained.
Citation Information
Patent Citations
Space-time big data fused drainage basin water quality pollution traceability analysis method and system
CN119646471A
Method for determining contribution rate of pollution load in water quality assessment section of annular river network system based on water quantity constitute
US20230054713A1