Agricultural non-point source pollution monitoring data quality evaluation method
By using grid division and dual-track digital archive construction, combined with aerial mobile monitoring, we have achieved accurate assessment of the data quality of agricultural non-point source pollution monitoring. This solves the problems of insufficient targeting of assessment objects and data reliability in existing technologies, and provides an efficient data quality optimization solution.
Patent Information
- Application Number
- CN202511640481.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-02-06
AI Technical Summary
Existing methods for assessing the quality of agricultural non-point source pollution monitoring data lack consideration of homogeneity factors such as land use, planting structure, and topographic slope. This results in insufficient targeting of the assessment objects, inadequate comprehensive consideration of multi-dimensional indicators, and a lack of on-site verification for judging data anomalies, thus affecting the reliability and application value of the monitoring data.
The system employs grid division and dual-track digital archive construction. By calculating scores based on three dimensions—data coverage and representativeness, data intrinsic correlation, and spatial logic between units—a comprehensive quality score threshold is set. The system utilizes aerial mobile monitoring units for close-range monitoring and combines data from fixed monitoring units to calculate differences. A comprehensive data deviation threshold is then set to determine data compliance.
It improves the targeting and accuracy of monitoring data assessment, reduces misjudgments, ensures the reliability of data quality, provides credible support for the control of agricultural non-point source pollution, and outputs targeted rectification suggestions.
Smart Images

Figure CN121481333A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural environmental monitoring technology, and in particular to a method for assessing the quality of agricultural non-point source pollution monitoring data. Background Technology
[0002] Agricultural non-point source pollution monitoring is a crucial foundation for agricultural environmental governance and management decisions. The quality of monitoring data directly impacts the accuracy of pollution assessment and remediation plan formulation. Current methods for assessing the quality of agricultural non-point source pollution monitoring data often lack homogeneous consideration of factors such as land use, planting structure, and topographic slope when dividing monitoring areas, resulting in insufficient targeting of assessment objects. The assessment process relies heavily on single data sources from fixed monitoring units, lacking comprehensive consideration of multi-dimensional indicators, making it difficult to fully reflect the data quality status. Furthermore, existing methods lack effective on-site verification tools for judging anomalies in monitoring data; analysis based solely on a single data level is prone to misjudgment, and the accuracy of data compliance determination is insufficient, ultimately affecting the reliability and subsequent application value of the monitoring data. Summary of the Invention
[0003] To address the aforementioned technical problems, this invention provides a method for assessing the quality of agricultural non-point source pollution monitoring data. The technical solution adopted is as follows: A method for assessing the quality of agricultural non-point source pollution monitoring data includes the following steps: Step 1: Divide the agricultural non-point source pollution area to be monitored into several grid areas to be monitored, and arrange a fixed monitoring unit at the physical center of each grid area; Step 2: The server creates a dual-track digital profile for each grid area. The dual-track digital profile includes static and dynamic attributes. The dual-track digital profile is stored in the database and updated according to a preset period, and the update version record is retained. Step 3: Calculate the index scores for the three dimensions of data coverage and representativeness, data internal correlation, and spatial logic between units, respectively, and normalize the index scores. Provide multiple sets of preset weights and calculate the comprehensive quality score of each monitoring unit based on the selected weights. Step 4: Set a comprehensive quality score threshold. When the comprehensive quality score calculated in Step 4 is lower than the comprehensive quality score threshold, it is determined that the agricultural non-point source pollution monitoring data of the corresponding grid area is abnormal. Step 5: Control the aerial mobile monitoring unit to fly over the grid area where anomalies exist for close-range monitoring; Step 6: The server simultaneously receives monitoring data from both the fixed monitoring unit and the aerial mobile monitoring unit, calculates the difference for each monitoring data item, and calculates the comprehensive data deviation based on the differences of all data items. Step 7: Set the comprehensive data deviation threshold. If the comprehensive data deviation calculated in Step 6 is greater than the comprehensive data deviation threshold, the evaluation result of the monitoring data of the fixed monitoring unit in the corresponding grid area being unqualified will be output.
[0004] Optionally, based on the administrative boundaries, watershed range, and land use contiguousness of the area to be monitored for agricultural non-point source pollution, GIS spatial analysis tools can be used to divide the area into regular grids to ensure the homogeneity of land use type, planting structure, and terrain slope within a single grid.
[0005] Optionally, the fixed monitoring unit includes a multi-parameter water quality sensor and a flow sensor, and the collected data is uploaded to the server in real time.
[0006] Optional, the static attributes of the dual-track digital archive include grid area, average elevation, average slope, main soil type, dominant land use type, and sub-basin number. Dynamic attributes include crop type, fertilization and pesticide application cycle, monitoring data from fixed monitoring units, data collection frequency, and distribution of surrounding pollution sources; The database is a relational database. Static attributes are verified once a year, while dynamic attributes are updated in a hierarchical manner on an annual, quarterly, or real-time basis.
[0007] Optionally, the data coverage and characterization dimensions mentioned in step 3 include three indicators: spatiotemporal coverage, source data integrity, and data continuity. The data intrinsic correlation dimension includes three indicators: covariance of physicochemical parameters, rationality of temporal changes, and effectiveness of extreme values. The spatial logic dimension between units includes three indicators: neighborhood grid consistency, land use-concentration matching degree, and rationality of watershed topology. All indicators are normalized to a score of 0-100, with higher scores indicating better data quality.
[0008] Optionally, in step 3, the data coverage and representativeness dimension scores are compared. The calculation formula is: ; in, , and These are scores for spatiotemporal coverage, source data integrity, and data continuity indicators, respectively. Data intrinsic correlation dimension score The calculation formula is: ; in, , and The scores are for the covariance of physicochemical parameters, the rationality of time-series changes, and the effectiveness of extreme values, respectively. Inter-unit spatial logical dimension score The calculation formula is: ; in, , and These are the scores for neighborhood grid consistency, land use-concentration matching degree, and watershed topological rationality. The overall quality score S is a weighted sum of the scores from three dimensions: data coverage and representativeness, data intrinsic correlation, and spatial logic between units. The calculation formula is as follows: ; , , These are the weights of the three dimensions.
[0009] Optionally, the spatiotemporal coverage score is calculated by dividing the number of valid data points after removing outliers by the expected number of data points obtained by multiplying the number of statistical days by the daily collection frequency, and then multiplying by 100. The source data integrity score is calculated by dividing the number of effectively collected core indicators by the total number of core indicators that include at least pH, dissolved oxygen, turbidity, total phosphorus, and ammonia nitrogen, and then multiplying by 100. The data continuity score is calculated by dividing the longest period of continuous data without anomalies within the statistical period by the total monitoring period, and then multiplying by 100. The covariance score of chemical parameters is determined based on the correlation strength of typical parameter pairs. A correlation strength of not less than 0.6 and the correlation direction conforms to common sense is 100 points, a correlation strength between 0.3 and 0.6 is 60 points, and all other cases are 0 points. The score for the rationality of temporal changes is based on the degree of matching between the data trend and agricultural production rhythms such as fertilization cycle, rainfall events, and crop growth stages, and is awarded as 100, 50, or 0 points respectively. The extreme value validity score is determined by whether the extreme value identified according to the 3σ rule falls within the range of 10%-90% of the instrument's range and corresponds to a driving event, and is scored as 100 or 0 points. The neighborhood grid consistency score is calculated by subtracting the relative deviation of the mean of the core index of this grid from the mean of the same index of the three adjacent grids by 100, with a minimum score of 0. The land use-concentration matching score is determined by whether the concentration in this grid falls within the statistical quartile range of the corresponding land use type, and is awarded 100 or 0 points. The watershed topology rationality score is determined by whether the concentration of the downstream grid is not less than 70% of the area-weighted average concentration of the upstream adjacent grid, and is either 100 points or 0 points.
[0010] Optionally, the aerial mobile monitoring unit mentioned in step 5 is a multi-rotor UAV equipped with a water quality monitoring module and a GPS positioning module. The monitoring indicators of the water quality monitoring module are consistent with those of the fixed monitoring unit. The approach monitoring path includes five monitoring points, including the grid center point and four vertices, with a flight altitude of 5-10 meters above the ground.
[0011] Optionally, the monitoring data items in step 6 include pH, dissolved oxygen, turbidity, total phosphorus, ammonia nitrogen, and flow rate; Differences in individual data items The formula for calculation using the relative deviation method is: ; This represents the average value of data from the aerial motion monitoring unit. The average value of data from the same period in a fixed monitoring unit is used; the comprehensive data deviation is the weighted average of the differences among all data items, with total phosphorus and ammonia nitrogen having a weight of 0.25 and the remaining data items having a weight of 0.1.
[0012] Optionally, the comprehensive data deviation threshold mentioned in step 7 is 15%, and the comprehensive data deviation threshold is adjusted according to the accuracy of the dual-source monitoring unit equipment and the monitoring environment error. The output evaluation results include a single grid evaluation report, a summary table of all grids, and rectification suggestions for non-compliant data. Rectification suggestions include calibrating fixed monitoring units, replacing sensors, or adjusting deployment locations.
[0013] In summary, the present invention has at least one of the following beneficial technical effects: This invention provides a method for assessing the quality of agricultural non-point source pollution monitoring data. Through homogeneous grid division and dual-track digital archive construction, it provides a precise foundation for data quality assessment, enhancing its relevance. Employing a three-dimensional, multi-indicator comprehensive scoring system and a fixed-mobile monitoring collaborative verification model, it comprehensively covers key dimensions of data quality, reducing misjudgments and improving assessment accuracy. The method clearly defines indicator calculation methods, weight configurations, and threshold standards, resulting in a clear and operable process that facilitates practical application. It outputs targeted rectification suggestions, directly guiding the optimization of monitoring data quality, ensuring the reliability of monitoring data, and providing strong support for decision-making in agricultural non-point source pollution control. Attached Figure Description
[0014] Figure 1 This is a flowchart illustrating a method for assessing the quality of agricultural non-point source pollution monitoring data according to the present invention. Detailed Implementation
[0015] The present invention will be further described in detail below with reference to the accompanying drawings.
[0016] This invention discloses a method for assessing the quality of agricultural non-point source pollution monitoring data.
[0017] Reference Figure 1 Example 1: A method for assessing the quality of agricultural non-point source pollution monitoring data, comprising the following steps: Step 1: Divide the agricultural non-point source pollution area to be monitored into several grid areas to be monitored, and arrange a fixed monitoring unit at the physical center of each grid area; Step 2: The server creates a dual-track digital profile for each grid area. The dual-track digital profile includes static and dynamic attributes. The dual-track digital profile is stored in the database and updated according to a preset period, and the update version record is retained. Step 3: Calculate the index scores for the three dimensions of data coverage and representativeness, data internal correlation, and spatial logic between units, respectively, and normalize the index scores. Provide multiple sets of preset weights and calculate the comprehensive quality score of each monitoring unit based on the selected weights. Step 4: Set a comprehensive quality score threshold. When the comprehensive quality score calculated in Step 4 is lower than the comprehensive quality score threshold, it is determined that the agricultural non-point source pollution monitoring data of the corresponding grid area is abnormal. Step 5: Control the aerial mobile monitoring unit to fly over the grid area where anomalies exist for close-range monitoring; Step 6: The server simultaneously receives monitoring data from both the fixed monitoring unit and the aerial mobile monitoring unit, calculates the difference for each monitoring data item, and calculates the comprehensive data deviation based on the differences of all data items. Step 7: Set the comprehensive data deviation threshold. If the comprehensive data deviation calculated in Step 6 is greater than the comprehensive data deviation threshold, the evaluation result of the monitoring data of the fixed monitoring unit in the corresponding grid area being unqualified will be output.
[0018] Optionally, based on the administrative boundaries, watershed range, and land use contiguousness of the area to be monitored for agricultural non-point source pollution, GIS spatial analysis tools can be used to divide the area into regular grids to ensure the homogeneity of land use type, planting structure, and terrain slope within a single grid.
[0019] Optionally, the fixed monitoring unit includes a multi-parameter water quality sensor and a flow sensor, and the collected data is uploaded to the server in real time.
[0020] Optional, the static attributes of the dual-track digital archive include grid area, average elevation, average slope, main soil type, dominant land use type, and sub-basin number. Dynamic attributes include crop type, fertilization and pesticide application cycle, monitoring data from fixed monitoring units, data collection frequency, and distribution of surrounding pollution sources; The database is a relational database. Static attributes are verified once a year, while dynamic attributes are updated in a hierarchical manner on an annual, quarterly, or real-time basis.
[0021] By adopting the above technical solution, firstly, the area to be monitored is divided into regular grids. Grid units are determined based on administrative boundaries, watershed range, and land use contiguousness. This ensures that factors affecting pollution monitoring, such as land use type, planting structure, and topographic slope, are homogeneous within each grid. When fixed monitoring units are arranged at the physical center of the grid, the data they collect can effectively represent the pollution status of that grid, laying an accurate spatial foundation for subsequent assessments.
[0022] Secondly, a dual-track digital archive is constructed, with static attributes recording stable basic information of the grid and dynamic attributes tracking agricultural production and monitoring-related data that change over time. Through regular updates and version retention, a comprehensive background reference is provided for data quality assessment, giving specific attribute basis for indicator calculation and logical judgment.
[0023] Furthermore, a multi-indicator evaluation system is designed from three dimensions: data coverage and representativeness, data intrinsic correlation, and spatial logic between units. The data coverage and representativeness dimension measures the completeness and continuity of the data, ensuring that the data fully reflects the monitored objects in time and space. The data intrinsic correlation dimension examines the logical rationality of the data itself and the parameters, conforming to the physicochemical laws of pollution changes. The spatial logic between units verifies the coordination between the data and the surrounding grid, aligning with the geographical patterns of pollution transmission. A comprehensive quality score is calculated through weighted averages to achieve a complete quantification of data quality.
[0024] Then, an abnormal data is identified by setting a comprehensive quality score threshold, which triggers the aerial mobile monitoring unit to conduct close-range monitoring. By leveraging the synergy between mobile and fixed monitoring, the risk of misjudgment based on a single data level is reduced through on-site verification.
[0025] Finally, by calculating the comprehensive deviation between fixed and mobile monitoring data and combining it with preset thresholds to determine the qualification of fixed monitoring data, targeted rectification suggestions are output, forming a closed loop of assessment-verification-judgment-optimization to ensure the reliability of monitoring data and provide credible data support for the control of agricultural non-point source pollution.
[0026] Optionally, the data coverage and characterization dimensions mentioned in step 3 include three indicators: spatiotemporal coverage, source data integrity, and data continuity. The data intrinsic correlation dimension includes three indicators: covariance of physicochemical parameters, rationality of temporal changes, and effectiveness of extreme values. The spatial logic dimension between units includes three indicators: neighborhood grid consistency, land use-concentration matching degree, and rationality of watershed topology. All indicators are normalized to a score of 0-100, with higher scores indicating better data quality.
[0027] Optionally, in step 3, the data coverage and representativeness dimension scores are compared. The calculation formula is: ; in, , and These are scores for spatiotemporal coverage, source data integrity, and data continuity indicators, respectively. Data intrinsic correlation dimension score The calculation formula is: ; in, , and The scores are for the covariance of physicochemical parameters, the rationality of time-series changes, and the effectiveness of extreme values, respectively. Inter-unit spatial logical dimension score The calculation formula is: ; in, , and These are the scores for neighborhood grid consistency, land use-concentration matching degree, and watershed topological rationality. The overall quality score S is a weighted sum of the scores from three dimensions: data coverage and representativeness, data intrinsic correlation, and spatial logic between units. The calculation formula is as follows: ; , , These are the weights of the three dimensions.
[0028] Optionally, the spatiotemporal coverage score is calculated by dividing the number of valid data points after removing outliers by the expected number of data points obtained by multiplying the number of statistical days by the daily collection frequency, and then multiplying by 100. The source data integrity score is calculated by dividing the number of effectively collected core indicators by the total number of core indicators that include at least pH, dissolved oxygen, turbidity, total phosphorus, and ammonia nitrogen, and then multiplying by 100. The data continuity score is calculated by dividing the longest period of continuous data without anomalies within the statistical period by the total monitoring period, and then multiplying by 100. The covariance score of chemical parameters is determined based on the correlation strength of typical parameter pairs. A correlation strength of not less than 0.6 and the correlation direction conforms to common sense is 100 points, a correlation strength between 0.3 and 0.6 is 60 points, and all other cases are 0 points. The score for the rationality of temporal changes is based on the degree of matching between the data trend and agricultural production rhythms such as fertilization cycle, rainfall events, and crop growth stages, and is awarded as 100, 50, or 0 points respectively. The extreme value validity score is determined by whether the extreme value identified according to the 3σ rule falls within the range of 10%-90% of the instrument's range and corresponds to a driving event, and is scored as 100 or 0 points. The neighborhood grid consistency score is calculated by subtracting the relative deviation of the mean of the core index of this grid from the mean of the same index of the three adjacent grids by 100, with a minimum score of 0. The land use-concentration matching score is determined by whether the concentration in this grid falls within the statistical quartile range of the corresponding land use type, and is awarded 100 or 0 points. The watershed topology rationality score is determined by whether the concentration of the downstream grid is not less than 70% of the area-weighted average concentration of the upstream adjacent grid, and is either 100 points or 0 points.
[0029] By adopting the above technical solutions, a three-dimensional evaluation system is constructed based on the core impact dimensions of data quality. The data coverage and representativeness dimension focuses on the basic usability of the data. Through three indicators—spatiotemporal coverage, source data integrity, and data continuity—it verifies the sufficiency of data coverage in the spatiotemporal range, the comprehensiveness of core monitoring indicator collection, and the uninterrupted stability of the data sequence, ensuring that the basic data for evaluation is sufficient and usable. The data intrinsic correlation dimension focuses on the logical rationality of the data. Through three indicators—physicochemical parameter covariance, temporal variation rationality, and extreme value effectiveness—it verifies the scientific correlation between monitoring parameters, the degree of fit between data changes and the actual rhythm of agricultural production, and the rationality of the generation of abnormal data, ensuring that the data itself is free of logical contradictions and conforms to objective laws. The inter-unit spatial logic dimension focuses on the scenario adaptability of the data. Through three indicators—neighborhood grid consistency, land use-concentration matching degree, and watershed topological rationality—it verifies the spatial coordination between the data and adjacent grids, the adaptability of the data to the land use attributes of the region, and the consistency of the data with the pollution transmission patterns of the upstream and downstream of the watershed, ensuring that the data conforms to the actual geographical space and pollution diffusion scenarios.
[0030] Standardized indicator scores are achieved through unified quantitative rules. All indicators are normalized to a score of 0-100, ensuring comparability and summarization of scores for indicators of different types and verification logics. Scores for each dimension are calculated using an arithmetic mean, as the three indicators within the same dimension are equally important in representing the quality of that dimension, collectively forming a complete evaluation of the dimension's quality. The overall quality score is calculated using a weighted sum, with weight allocation reflecting the emphasis on different dimensions of data quality in various monitoring scenarios, making the evaluation results more aligned with practical application scenarios.
[0031] The calculation logic for each indicator is closely aligned with its quality characterization objective. Spatiotemporal coverage, source data integrity, and data continuity directly reflect the basic completeness of the data through proportional calculations; the covariance of physicochemical parameters, the rationality of temporal variations, and the validity of extreme values verify the data logic through scientific principles and real-world scenarios; neighborhood grid consistency, land use-concentration matching, and watershed topological rationality verify the data scenario adaptability through spatial correlation and attribute adaptation. This entire set of calculation logic forms a progressive verification of basic completeness, internal logic, and scenario adaptability, ultimately achieving a precise quantitative assessment of data quality.
[0032] Optionally, the aerial mobile monitoring unit mentioned in step 5 is a multi-rotor UAV equipped with a water quality monitoring module and a GPS positioning module. The monitoring indicators of the water quality monitoring module are consistent with those of the fixed monitoring unit. The approach monitoring path includes five monitoring points, including the grid center point and four vertices, with a flight altitude of 5-10 meters above the ground.
[0033] By adopting the above technical solution, multi-rotor UAVs were selected as the aerial mobile monitoring unit. Their flexibility and maneuverability allow for rapid response to monitoring needs in abnormal grids, precise arrival at designated areas, and suitability for agricultural non-point source pollution monitoring areas with dispersed terrain and potentially complex topography. Equipping the fixed monitoring unit with the same water quality monitoring module ensures complete consistency of the indicator dimensions for dual-source monitoring, avoiding data comparison failures due to differences in monitoring indicators and ensuring the accuracy of difference calculations.
[0034] The design includes five monitoring points, including the grid center point and four vertices, which can fully cover the entire grid area, avoid the randomness of single-point sampling, and enable mobile monitoring data to comprehensively reflect the actual pollution situation within the grid, forming a comprehensive comparison basis with the single-point data of fixed monitoring units.
[0035] Setting the flight altitude to 5-10 meters above the ground ensures that the monitoring module is in full contact with the surface water to obtain accurate and effective water quality data. It also avoids the impact of terrain obstacles at too low an altitude, or the decrease in monitoring accuracy at too high an altitude, thus ensuring the reliability of mobile monitoring data and providing strong support for subsequent comprehensive data deviation calculation and qualification determination of fixed monitoring data.
[0036] Optionally, the monitoring data items in step 6 include pH, dissolved oxygen, turbidity, total phosphorus, ammonia nitrogen, and flow rate; Differences in individual data items The formula for calculation using the relative deviation method is: ; This represents the average value of data from the aerial motion monitoring unit. The average value of data from the same period in a fixed monitoring unit is used; the comprehensive data deviation is the weighted average of the differences among all data items, with total phosphorus and ammonia nitrogen having a weight of 0.25 and the remaining data items having a weight of 0.1.
[0037] By adopting the above technical solutions, pH, dissolved oxygen, turbidity, total phosphorus, ammonia nitrogen, and flow rate are selected as monitoring data items. These indicators are core parameters reflecting the status of agricultural non-point source pollution, covering basic water quality properties, pollutant concentrations, and hydrological conditions. They can comprehensively cover the key dimensions of pollution monitoring and ensure the comprehensiveness of difference comparisons.
[0038] The relative deviation method is used to calculate the difference for a single data item. This method can eliminate the influence of the difference in magnitude of different indicators, intuitively reflect the relative deviation of dual-source data, make the differences of different types of monitoring indicators comparable, and avoid the distortion of difference judgment due to the difference in magnitude of indicators.
[0039] The overall data deviation is calculated using a weighted average, with total phosphorus and ammonia nitrogen given higher weights. These two are core indicators of agricultural non-point source pollution, and their data accuracy is crucial for assessing the pollution situation, highlighting the impact of differences in core indicators on overall data quality. Other indicators are allocated equal weights to ensure the comprehensiveness of data differences. This entire calculation logic ensures the scientific rigor of difference measurement while aligning with the core needs of agricultural non-point source pollution monitoring, providing precise quantitative support for subsequent compliance determination.
[0040] Optionally, the comprehensive data deviation threshold mentioned in step 7 is 15%, and the comprehensive data deviation threshold is adjusted according to the accuracy of the dual-source monitoring unit equipment and the monitoring environment error. The output evaluation results include a single grid evaluation report, a summary table of all grids, and rectification suggestions for non-compliant data. Rectification suggestions include calibrating fixed monitoring units, replacing sensors, or adjusting deployment locations.
[0041] By adopting the above technical solution, the comprehensive data deviation threshold is set at 15%. This value is determined based on the accuracy level of conventional equipment in the dual-source monitoring unit and the environmental error range of agricultural non-point source pollution monitoring. It can effectively distinguish between normal data deviations and abnormal deviations caused by equipment failure, improper deployment, etc., providing a clear and scientific quantitative standard for qualification judgment. At the same time, it allows for adjustment of the threshold according to the differences in the accuracy of the monitoring equipment actually used and the environmental complexity of the monitoring area, ensuring that the judgment standard is adapted to the specific application scenario and improving the objectivity and rationality of the judgment results.
[0042] It outputs a single grid assessment report and a summary table of all grids, which not only meets the needs of refined query of single grid data quality, but also enables global control of the overall data quality of the monitoring area, allowing users to clearly understand the quality status in different dimensions.
[0043] For non-compliant data, we provide rectification suggestions such as calibrating fixed monitoring units, replacing sensors, or adjusting deployment locations. These suggestions directly address common causes of non-compliant fixed monitoring data, providing users with readily implementable optimization measures. This transforms data quality assessment results into practical improvement actions, forming a complete quality optimization loop and continuously improving the reliability of agricultural non-point source pollution monitoring data.
[0044] The following describes the implementation principle of the present invention using specific embodiments: The total area of the agricultural non-point source pollution monitoring area in a certain watershed is approximately 50 square kilometers, encompassing various land use types such as paddy fields, cornfields, and aquaculture areas. This method was used to conduct a monitoring data quality assessment, and the specific implementation process is as follows: Step 1: Based on the administrative boundaries, watershed area, and land use contiguousness of the region, a regular grid was created using GIS spatial analysis tools. The grid size was set to 1000m × 1000m, resulting in 50 monitoring grid areas. After division, the homogeneity of land use type, planting structure, and terrain slope within each grid was verified to ensure they all met the assessment requirements. A fixed monitoring unit was deployed at the physical center of each grid area. The fixed monitoring unit included a multi-parameter water quality sensor and a flow sensor. The multi-parameter water quality sensor monitored indicators such as pH, dissolved oxygen, turbidity, total phosphorus, and ammonia nitrogen, while the flow sensor monitored the regional runoff. All collected data was uploaded to the server in real time, with a data collection frequency set to once every 2 hours.
[0045] Step 2: The server creates a dual-track digital profile for each grid area and stores it in a relational database. Static attributes include the area, average elevation, average slope, main soil type, dominant land use type, and sub-basin number of each grid area. Static attributes are verified annually. Dynamic attributes include crop type, fertilization and pesticide application cycle, monitoring data from fixed monitoring units, data collection frequency, and distribution of surrounding pollution sources. Crop type and fertilization and pesticide application cycle are updated annually, data collection frequency is updated quarterly, monitoring data from fixed monitoring units is updated in real time, and the distribution of surrounding pollution sources is updated quarterly. Version records are retained after each update.
[0046] Step 3: Calculate the three-dimensional index scores for each grid fixed monitoring unit's data. The data coverage and representativeness dimension includes three indicators: spatiotemporal coverage, source data integrity, and data continuity. The data intrinsic correlation dimension includes three indicators: physicochemical parameter covariance, temporal variation rationality, and extreme value effectiveness. The inter-unit spatial logic dimension includes three indicators: neighborhood grid consistency, land use-concentration matching degree, and watershed topological rationality. All index scores are normalized to 0-100. When calculating the scores for each dimension, the data coverage and representativeness dimension score is the arithmetic mean of the three index scores; the data intrinsic correlation dimension score is the arithmetic mean of the three index scores; and the inter-unit spatial logic dimension score is the arithmetic mean of the three index scores. This evaluation selects a preset weight set for conventional monitoring scenarios: the data coverage and representativeness dimension weight is 0.35, the data intrinsic correlation dimension weight is 0.35, and the inter-unit spatial logic dimension weight is 0.3. Based on the selected weights, the comprehensive quality score for each fixed monitoring unit is calculated.
[0047] Step 4: The box plot method is used to determine the comprehensive quality score threshold as 60 points. The server compares the comprehensive quality score of each grid with this threshold and determines that the comprehensive quality score of 3 grid areas is lower than 60 points, indicating that the agricultural non-point source pollution monitoring data of the corresponding grid areas are abnormal.
[0048] Step 5: Control the aerial mobile monitoring unit to fly over the three grid areas exhibiting anomalies for close-range monitoring. The aerial mobile monitoring unit is a multi-rotor UAV equipped with a water quality monitoring module and a GPS positioning module. The monitoring indicators of the water quality monitoring module are consistent with those of the fixed monitoring unit. The close-range monitoring path includes five monitoring points: the center point and four vertices of each grid. The flight altitude is controlled at 6-8 meters above the ground. The unit stays at each monitoring point for 4 minutes to collect data. The average value of the monitoring data from each monitoring point is taken as the final data for that point and transmitted back to the server in real time.
[0049] Step 6: The server simultaneously receives monitoring data from the fixed monitoring units and the aerial mobile monitoring units in these three grid areas. It selects pH, dissolved oxygen, turbidity, total phosphorus, ammonia nitrogen, and flow rate as monitoring data items. The relative deviation method is used to calculate the difference for each data item. Then, with a weight of 0.25 for total phosphorus and ammonia nitrogen and a weight of 0.1 for the other data items, the weighted average of the differences of all data items is calculated to obtain the comprehensive data deviation.
[0050] Step 7: Set the comprehensive data deviation threshold to 15%, and compare the calculated comprehensive data deviation with this threshold. The comprehensive data deviations for two grids are 18% and 22%, respectively, both greater than 15%. The comprehensive data deviation for one grid is 12%, less than 15%. The final output evaluation results include individual grid evaluation reports for these three grids and a summary table of all 50 grids. The monitoring data of the fixed monitoring units corresponding to the two grids with comprehensive data deviations greater than the threshold are deemed unqualified, and rectification suggestions are given. It is recommended that the fixed monitoring units of these two grids be calibrated; if the data still does not meet the standards after calibration, the sensors should be replaced.
[0051] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for assessing the quality of agricultural non-point source pollution monitoring data, characterized in that, Includes the following steps: Step 1: Divide the agricultural non-point source pollution area to be monitored into several grid areas to be monitored, and arrange a fixed monitoring unit at the physical center of each grid area; Step 2: The server creates a dual-track digital profile for each grid area. The dual-track digital profile includes static and dynamic attributes. The dual-track digital profile is stored in the database and updated according to a preset period, and the update version record is retained. Step 3: Calculate the index scores for the three dimensions of data coverage and representativeness, data internal correlation, and spatial logic between units, respectively, and normalize the index scores. Provide multiple sets of preset weights and calculate the comprehensive quality score of each monitoring unit based on the selected weights. Step 4: Set a comprehensive quality score threshold. When the comprehensive quality score calculated in Step 4 is lower than the comprehensive quality score threshold, it is determined that there is an anomaly in the agricultural non-point source pollution monitoring data of the corresponding grid area. Step 5: Control the aerial mobile monitoring unit to fly over the grid area where anomalies exist for close-range monitoring; Step 6: The server simultaneously receives monitoring data from both the fixed monitoring unit and the aerial mobile monitoring unit, calculates the difference for each monitoring data item, and calculates the comprehensive data deviation based on the differences of all data items. Step 7: Set the comprehensive data deviation threshold. If the comprehensive data deviation calculated in Step 6 is greater than the comprehensive data deviation threshold, the evaluation result of the monitoring data of the fixed monitoring unit in the corresponding grid area being unqualified will be output.
2. The method for assessing the quality of agricultural non-point source pollution monitoring data according to claim 1, characterized in that, Based on the administrative boundaries, watershed range, and land use contiguousness of the areas to be monitored for agricultural non-point source pollution, GIS spatial analysis tools are used to divide the area into regular grids to ensure the homogeneity of land use types, planting structures, and topographic slopes within each grid.
3. The method for assessing the quality of agricultural non-point source pollution monitoring data according to claim 2, characterized in that, The fixed monitoring unit includes a multi-parameter water quality sensor and a flow sensor, and the collected data is uploaded to the server in real time.
4. The method for assessing the quality of agricultural non-point source pollution monitoring data according to claim 3, characterized in that, The static attributes of the dual-track digital archive include grid area, average elevation, average slope, main soil type, dominant land use type, and sub-basin number. Dynamic attributes include crop type, fertilization and pesticide application cycle, monitoring data from fixed monitoring units, data collection frequency, and distribution of surrounding pollution sources; The database is a relational database. Static attributes are verified once a year, while dynamic attributes are updated in a hierarchical manner on an annual, quarterly, or real-time basis.
5. The method for assessing the quality of agricultural non-point source pollution monitoring data according to claim 4, characterized in that, The data coverage and representativeness dimensions mentioned in step 3 include three indicators: spatiotemporal coverage, source data integrity, and data continuity. The data intrinsic correlation dimension includes three indicators: covariance of physicochemical parameters, rationality of temporal changes, and effectiveness of extreme values. The spatial logic dimension between units includes three indicators: neighborhood grid consistency, land use-concentration matching degree, and rationality of watershed topology. All indicators are normalized to a score of 0-100, with higher scores indicating better data quality.
6. The method for assessing the quality of agricultural non-point source pollution monitoring data according to claim 5, characterized in that, In step 3, the scores for data coverage and representativeness dimensions are calculated. The calculation formula is: ; in, , and These are scores for spatiotemporal coverage, source data integrity, and data continuity indicators, respectively. Data intrinsic correlation dimension score The calculation formula is: ; in, , and The scores are for the covariance of physicochemical parameters, the rationality of time-series changes, and the effectiveness of extreme values, respectively. Inter-unit spatial logical dimension score The calculation formula is: ; in, , and These are the scores for neighborhood grid consistency, land use-concentration matching degree, and watershed topological rationality indicators, respectively. The overall quality score S is a weighted sum of the scores from three dimensions: data coverage and representativeness, data intrinsic correlation, and spatial logic between units. The calculation formula is as follows: ; , , These are the weights of the three dimensions.
7. The method for assessing the quality of agricultural non-point source pollution monitoring data according to claim 6, characterized in that, The spatiotemporal coverage score is calculated by dividing the number of valid data points after removing outliers by the expected number of data points obtained by multiplying the number of statistical days by the daily collection frequency, and then multiplying by 100. The source data integrity score is calculated by dividing the number of effectively collected core indicators by the total number of core indicators that include at least pH, dissolved oxygen, turbidity, total phosphorus, and ammonia nitrogen, and then multiplying by 100. The data continuity score is calculated by dividing the longest period of continuous data without abnormalities within the statistical period by the total monitoring period, and then multiplying by 100. The covariance score of chemical parameters is determined based on the correlation strength of typical parameter pairs. A correlation strength of not less than 0.6 and the correlation direction conforms to common sense is 100 points, a correlation strength between 0.3 and 0.6 is 60 points, and all other cases are 0 points. The score for the rationality of temporal changes is based on the degree of matching between the data trend and agricultural production rhythms such as fertilization cycle, rainfall events, and crop growth stages, and is awarded as 100, 50, or 0 points respectively. The extreme value validity score is determined by whether the extreme value identified according to the 3σ rule falls within the range of 10%-90% of the instrument's range and corresponds to a driving event, and is scored as 100 or 0 points. The neighborhood grid consistency score is calculated by subtracting the relative deviation of the mean of the core index of this grid from the mean of the same index of the three adjacent grids by 100, with a minimum score of 0. The land use-concentration matching score is determined by whether the concentration in this grid falls within the statistical quartile range of the corresponding land use type, and is awarded 100 or 0 points. The watershed topology rationality score is determined by whether the concentration of the downstream grid is not less than 70% of the area-weighted average concentration of the upstream adjacent grid, and is either 100 points or 0 points.
8. The method for assessing the quality of agricultural non-point source pollution monitoring data according to claim 7, characterized in that, The aerial mobile monitoring unit mentioned in step 5 is a multi-rotor UAV equipped with a water quality monitoring module and a GPS positioning module. The monitoring indicators of the water quality monitoring module are the same as those of the fixed monitoring unit. The approach monitoring path includes five monitoring points, including the grid center point and four vertices, with a flight altitude of 5-10 meters above the ground.
9. The method for assessing the quality of agricultural non-point source pollution monitoring data according to claim 8, characterized in that, The monitoring data items mentioned in step 6 include pH, dissolved oxygen, turbidity, total phosphorus, ammonia nitrogen, and flow rate; Differences in individual data items The formula for calculation using the relative deviation method is: ; This represents the average value of data from the aerial motion monitoring unit. The average value of data from the same period in a fixed monitoring unit is used; the comprehensive data deviation is the weighted average of the differences among all data items, with total phosphorus and ammonia nitrogen having a weight of 0.25 and the remaining data items having a weight of 0.
1.
10. The method for assessing the quality of agricultural non-point source pollution monitoring data according to claim 9, characterized in that, The comprehensive data deviation threshold mentioned in step 7 is 15%, and the comprehensive data deviation threshold is adjusted according to the accuracy of the dual-source monitoring unit equipment and the monitoring environment error. The output evaluation results include a single grid evaluation report, a summary table of all grids, and rectification suggestions for non-compliant data. Rectification suggestions include calibrating fixed monitoring units, replacing sensors, or adjusting deployment locations.