Water quality data daily scale processing method based on hourly level automatic monitoring
By adopting a daily-scale processing method for water quality data based on hourly automatic monitoring, the problem of inaccurate results in high-frequency water quality monitoring using traditional methods is solved. This method enables time-weighted calculation at the daily scale, improving the comparability and stability of the results and making it suitable for batch processing of multiple stations and parameters.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HOHAI UNIV
- Filing Date
- 2025-11-19
- Publication Date
- 2026-07-24
AI Technical Summary
Traditional methods for calculating daily average concentrations are difficult to accurately reflect the true changes in water quality parameters in high-frequency continuous water quality monitoring, especially under conditions of non-uniform sampling, missing measurements, noise, and drift, which leads to systematic biases and inaccurate results.
A daily-scale water quality data processing method based on hourly automatic monitoring is adopted, including steps such as invalid data identification, data standardization, construction of adjacent valid segments, trapezoidal integration and boundary closure, and time-weighted calculation to obtain daily-scale monitoring data.
It eliminates omissions/miscalculations caused by the day boundary truncation, improves the comparability and stability of the results, ensures the consistency and verifiability of the calculations, adapts to different watersheds and parameter systems, facilitates rapid adaptation and maintenance, and supports batch processing of multiple sites and multiple parameters.
Smart Images

Figure CN121501880B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of environmental monitoring and intelligent data processing technology, specifically relating to a daily-scale time-weighted programmed processing method for hourly continuous automatic monitoring water quality data. Background Technology
[0002] With the accelerating digitalization and intelligentization of global water environment monitoring systems, continuous monitoring and high-frequency online observation have gradually become important technical means for watershed water quality supervision. In recent years, automated monitoring equipment has been widely deployed in surface water sections, lakes, and reservoirs, increasing the monitoring frequency from traditional periodic sampling to real-time observation at hourly or even minute intervals. This high-frequency continuous monitoring mode can reveal the dynamic changes of water quality parameters on diurnal and short-term scales, providing key data support for pollution load estimation, water quality evolution pattern analysis, and ecological risk assessment. However, with the increase in monitoring frequency, data characteristics and processing requirements have also changed significantly. On the one hand, high-frequency monitoring data exhibits non-uniformity in time and multi-point asynchrony in space, accompanied by issues such as missing data, noise, and drift, resulting in a huge data volume and complex structure. On the other hand, water quality parameters (such as dissolved oxygen, ammonia nitrogen, and conductivity) are affected by multiple factors such as rainfall, runoff, temperature, meteorological changes, and sudden emissions, exhibiting strong non-stationarity and abrupt changes on hourly or even minute-scales, making it difficult for traditional low-frequency mean calculation methods to accurately reflect their true changes.
[0003] Traditional methods for calculating daily average concentrations mainly include the arithmetic mean at equal intervals and the flow-weighted mean (FWM). The arithmetic mean method relies on the premise that the concentration at each sampling point represents the average level within the sampling interval. However, this premise is often difficult to meet under high-frequency continuous monitoring conditions. Furthermore, when water bodies are affected by disturbances such as rainfall, sudden increases in runoff, or short-term discharges, concentrations can exhibit significant peak fluctuations within a short period. If the arithmetic mean at equal intervals is still used to calculate the daily representative value, the statistical process will ignore the non-uniformity of the sampling time interval, leading to underestimation of high concentration periods and overestimation of low concentration periods, thus introducing systematic bias and causing the results to deviate from the true temporal variation characteristics of the water body. The flow-weighted mean (FWM) theoretically reflects the comprehensive characteristics of concentration changes with hydrodynamics, but its application depends on synchronous, high-precision flow data. In actual monitoring, flow and concentration records often suffer from differences in temporal resolution, missing measurements, or asynchrony, and some cross-sections are not equipped with flow monitoring equipment, making this method difficult to widely apply.
[0004] Given the limitations of traditional methods, researchers have begun exploring concentration estimation methods that better reflect the characteristics of monitoring time series. Altier et al. verified the rationality of Time-Weighted Average (TWA) in reflecting pollutant exposure levels on continuous time scales. As a weighted statistical form that considers differences in sampling intervals, TWA can overcome the representativeness bias caused by uneven distribution of sampling time to some extent. However, the engineering implementation of TWA is not intuitive: First, the cross-day boundary and the final closure are asynchronous and intermittent at different stations, which easily leads to systematic errors at the day boundary; second, heterogeneity in sampling step length and the prevalence of missing / abnormal measurements at different stations make it difficult for general tools to automatically identify the effective interval and dynamically allocate time weights; third, insufficient coverage and equipment maintenance downtime occur frequently, urgently requiring quality labeling and auditing mechanisms coupled with computational logic; fourth, batch processing requires consistent rules and traceable output under conditions of multiple stations, multiple parameters, and long time series.
[0005] Based on this, an integrated rule system of "coverage constraint + effective endpoint removal + cross-day closure + structured archiving" has been formed, becoming a key technical path for high-frequency continuous monitoring data to move from "collection" to "intelligent computing". This invention proposes a systematic solution to the above-mentioned pain points. Summary of the Invention
[0006] Purpose of the invention: In view of the fact that the traditional low-frequency mean calculation method of the existing technology is difficult to accurately reflect the real change process of water quality monitoring in the basin, the present invention proposes a daily scale processing method for water quality data based on hourly automatic monitoring.
[0007] Technical Solution: To achieve the above-mentioned objectives, the present invention adopts the following technical solution: a daily-scale processing method for water quality data based on hourly automatic monitoring, comprising the following steps:
[0008] S1, Invalid Data Determination: Define an invalid value set N, and define the date using the local time zone for all monitoring data. All timestamps are unified to the same time zone and deduplicated and sorted in ascending order;
[0009] S2, Data Standardization Processing: After standardizing the monitoring data according to the format {station, parameter, datetime, value}, invalid monitoring records are identified through the invalid value set N and removed or replaced. Here, station represents the station name, parameter represents the monitoring indicator, datetime represents the local time zone timestamp, and value represents the indicator monitoring value.
[0010] S3, Data Day Grouping and Adjacency Valid Construction: Grouping by station-parameter-datetime as the key to obtain the daily sequence. and in the daily sequence Construct adjacent valid pairs;
[0011] S4, Integral of adjacent segments within the daily sequence: In the daily sequence In the middle, the integral of adjacent segments inside is calculated by the trapezoidal integral method;
[0012] S5, Symmetrically Closed Start and End Boundaries of Daily Sequence: When calculating the start and end segments of the daily sequence, for the two cases where the effective values of the start and end points of the daily sequence exist or are missing, the integrals of the start and end segments of the daily sequence are obtained by the trapezoidal integral of the first segment respectively.
[0013] S6. The internal adjacent segment integrals and start-end segment integrals of the monitoring data obtained in steps S4 and S5 are fused to obtain the daily-scale monitoring data.
[0014] Furthermore, the set of invalid values mentioned in step S1 , , -1 indicates missing data, -1 indicates instrument malfunction, and 0 indicates that the data could not be detected. This indicates that information is missing.
[0015] Furthermore, in step S2, invalid monitoring records are removed, and for multiple monitoring records at the same time, only one is retained based on the source credibility or the "most recent" principle.
[0016] Further, in step S4, for any adjacent valid pair and The integral of adjacent segments is calculated using the trapezoidal integral method with non-uniform sampling. And merge to obtain the integral of the internal adjacent segments of the sequence for that day. ,
[0017] ,
[0018] In the formula, Indicates two consecutive valid sampling times within a day and The time integral contribution of water quality indicators between different periods. Indicates two valid sampling times and The length of the time interval between them Indicates at time Valid monitoring values of water quality indicators obtained from the site Indicates at time The effective monitoring value of this water quality indicator obtained from the site, This represents the integral contribution of all adjacent valid sampling time pairs within the same day. The total amount of internal time points obtained by summing them up for the day.
[0019] Furthermore, in step S5, the integral of the starting segment of the daily sequence at 0:00 is obtained as follows:
[0020] When a valid value exists, take... With the first subsequent valid point of the day The first segment of the trapezoid is formed and included in the daily score:
[0021] ,
[0022] It indicates the time from the start of the day 0:00 to the first subsequent valid time of the day. The contribution of time integrals between them This indicates the valid monitoring value of the water quality indicator at 0:00 on the day of the event.
[0023] If no valid value exists at the starting point 0:00 of the day, a valid value will be collected by holding at zero order. Push forward to And obtained by the following calculation,
[0024] ,
[0025] The integral of the initial segment of the sequence at 0:00 on the current day is obtained as follows:
[0026] The next day If it exists and is valid, take the last valid point of the day. and Forming a trans-sun trapezoid, its contribution Counted towards daily points:
[0027] ,
[0028] Indicates the last valid time of the day. The amount of time points contributed up to 24:00 on the same day. Indicates the last valid time of the day. The effective monitoring value of this water quality indicator is specified. This indicates the valid monitoring value of the water quality indicator at 0:00 the following day. This indicates the last valid sampling time in the sequence for that day;
[0029] The next day If missing or invalid, it is preserved at order zero. Closed to :
[0030] .
[0031] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0032] (1) This invention eliminates omissions / miscalculations caused by the truncation of the day boundary, adapts to short-term equipment stops, ensures consistent and verifiable standards, and improves the consistency of cross-day alignment.
[0033] (2) This invention avoids the systematic bias of arithmetic mean under uneven sampling, and makes the representative value consistent with the sampling rhythm; the result has stronger comparability and stability at different stations and during different days.
[0034] (3) This invention separates invalidation determination, adjacent segment construction, internal integration, boundary closure, normalization, and audit write-back into independent modules, arranged with standard interfaces; it performs batch calculations according to "date × parameter × site", with parallel calculations and serial write-back, supports windowing and breakpoint resumption, and structured disk storage. This invention has high reusability and low coupling, which facilitates rapid adaptation to different watersheds and parameter systems; it reduces maintenance and upgrade costs, improves engineering stability, and facilitates integration with ETL / data lake.
[0035] (4) This invention configures invalid sets and out-of-bounds thresholds according to parameters; anomalies only affect the selection of adjacent segments and do not spread to the entire daily caliber; it generates first / last segment closed mode labels and audit records of key variables, and generates fingerprint binding versions and rule configurations for the results. This invention has controllable input quality and does not spread anomalies; the results are recalculated, auditable, and comparable, meeting engineering audit and compliance requirements, and supporting long-term operation and maintenance and quality assessment. Attached Figure Description
[0036] Figure 1 This is a flowchart illustrating the present invention;
[0037] Figure 2 This is a comparison chart of the calculation results of the embodiments of the present invention and the results of manual calculation (Shiyan City);
[0038] Figure 3 This is a comparison chart of the calculation results of the embodiments of the present invention and the results of manual calculation (Nanyang City);
[0039] Figure 4 This is a comparison chart of the calculation results of the embodiments of the present invention and the results of manual calculation (Shangluo City);
[0040] Figure 5 This is a graph showing the consistency evaluation results of daily manual and code calculations over a year according to an embodiment of the present invention. Detailed Implementation
[0041] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0042] like Figure 1 As shown, the present invention provides a daily-scale time-weighted programmed processing method for water quality based on hourly automatic monitoring, comprising the following steps:
[0043] S1 Data Objects and Rule Settings
[0044] S1.1 Abstract the monitoring records into a sequence of quadruples arranged in ascending time order. .
[0045] S1.2 Define the set of invalid values N (including but not limited to) -1, 0 and the out-of-range value determined according to the range / standard; if If the content exceeds the boundary, it will be considered invalid.
[0046] S1.3 The local time zone is used to define the date. All timestamps are unified to the same time zone and deduplicated and ordered in ascending order.
[0047] S2 Standardization and Cleaning
[0048] S2.1 Field standardization: Unify to the {station, parameter, datetime, value} pattern; fill in and mark abnormal fields.
[0049] S2.2 Invalidity Determination: Remove or mark invalid and out-of-bounds records; if multiple records exist at the same time, retain one based on the source credibility or the "most recent" principle.
[0050] S3 Day Grouping and Adjacent Valid Segment Construction
[0051] S3.1 Grouping by "site-parameter-date" to obtain the daily sequence. .
[0052] S3.2 Constructing valid adjacent pairs: For any adjacent pair A pair is considered a valid segment and its duration is recorded only when both endpoints are valid values. .
[0053] S4 Integral of adjacent segments within the same day (time weight adaptive)
[0054] S4.1 For each valid segment, accumulate the daily integral using the trapezoidal integral of non-uniform sampling:
[0055] ,
[0056] S4.2 The above weights The adaptive variation with the sampling interval naturally characterizes the non-uniformity of the monitoring time distribution.
[0057] S5 is symmetrically closed at its starting and ending boundaries (example demonstration of caliber calculation)
[0058] Example A: Valid from 0:00 on the current day to 0:00 the following day (cross-day trapezoidal)
[0059] Valid points for the day: ; 0:00 the next day: (All units are mg / L).
[0060] Internal section:
[0061] .
[0062] Final closure (spanning day trapezoid, included in the current day):
[0063] ,
[0064] Example B: Missing at 0:00 on the same day, missing at 0:00 on the next day (first / last paragraph zero-order preserved)
[0065] Valid points for the day: Missing at 0:00 the next day.
[0066] The starting point is closed (zero-order preservation):
[0067] .
[0068] Internal section:
[0069] .
[0070] Final closure (zero-order hold):
[0071] .
[0072] S6 Daily Value Normalization and Output
[0073] S6.1 The first paragraph The sum of integrals of adjacent segments within the same segment and the last segment Accumulated .
[0074] Example A, Daily Integral:
[0075] ,
[0076] Note: 0:00 already exists and is valid, so the first segment naturally participates in the integration as an internal adjacent segment; the last segment is closed across days according to the "priority of 0:00 the next day".
[0077] Example B, daily points:
[0078] ,
[0079] Explanation: If 0:00 is missing, the first value is shifted forward; if 0:00 is missing the next day, the last value is retained until 24:00, which fully conforms to the symmetrical closure rule.
[0080] S6.2 The time-weighted average of the day is calculated using a fixed day length and normalized:
[0081] ,
[0082] Example A
[0083] ,
[0084] Example B
[0085] ,
[0086] S6.3 If (i.e., no valid adjacent segments are available on the current day and a head-tail closure cannot be achieved), output the placeholder value and record the reason; otherwise, output Structured results.
[0087] S7 Batch Processing and Parallel Processing
[0088] S7.1 performs batch traversal and calculations on multiple files, multiple sites, and multiple parameter sets along the "date × parameter" dimension, and unifies naming and archiving.
[0089] S7.2 supports site-level / parameter-level parallelism; serialization of write-back and disk write-to-disk phases to avoid resource contention; and adoption of date window streaming processing to reduce memory usage.
[0090] S8 Auditing and Fault Tolerance
[0091] S8.1 records key events such as invalid value removal, boundary closure method, and time alignment throughout the process; the output is bound to the session identifier to ensure that any result can be recalculated and traced.
[0092] S8.2 When a large number of consecutive test failures or maintenance shutdowns occur at the same site, the start, stop, and closure logic will still be executed; if ultimately... It only outputs placeholder and audit information, and does not generate representative values.
[0093] Example 2:
[0094] To verify the feasibility and effectiveness of the method of this invention, continuous monitoring data of river sections published by the "National Surface Water Quality Automatic Monitoring Real-time Data Release System" were selected for verification. Using data from one river monitoring section each in Shiyan, Nanyang, and Shangluo cities in a specific year as samples, six indicators—water temperature, pH, dissolved oxygen (mg / L), conductivity (μS / cm), total phosphorus (mg / L), and total nitrogen (mg / L)—were selected as processing targets. Based on the technical solution of this invention, data standardization, invalid value identification, construction of adjacent valid segments, symmetrical closure of start and end boundaries, and time-weighted mean calculation were completed, and structured results were output.
[0095] Application effect verification:
[0096] Accuracy verification: Taking one month of data as an example ( Figure 2 – Figure 4 The dataset covers daily data from three cities, three rivers, and three water quality indicators. For each river section and each indicator, the time-weighted average output by this method is compared daily with the results of manual segmental integration using the same rules. The data points completely overlap, indicating consistency between the calculation method and the implementation. Further consistency evaluation is conducted based on one year's (365 days) of hourly automatic monitoring data. Figure 5 The mean difference between the code and the manually calculated results is 0, the consistency limit is 0, and all sample points fall on the "difference = 0" line, indicating that the two types of results are completely consistent, thus verifying the calculation accuracy and reliability of this method.
[0097] Efficiency verification: Under the same data scale and processing scope, the method of this invention achieves fully automated operation, with a single batch processing time of about 30 seconds. Compared with the time required for manual calculation of parameters day by day, which takes more than ten days, this method can complete the same workload in more than ten minutes, significantly reducing labor and time costs.
[0098] Error avoidance verification: The data processing process requires no manual intervention. Key steps such as invalid value screening, cross-day connection, and result writing back are all completed automatically by the program, avoiding omissions, miscalculations, and inconsistencies in the calculation of boundary completion, cross-day closure, and result summarization by humans, ensuring the stability of the processing process and the traceability of the results.
[0099] This embodiment demonstrates that, under conditions of multi-site, multi-parameter, and non-uniform sampling, the method of this invention can rapidly generate daily time-weighted averages with uniform rules and fixed caliber, possessing reliable accuracy, significant efficiency, and auditable and verifiable engineering application value. Implementation details not explicitly stated can be implemented using existing technologies.
[0100] The present invention has the following advantages:
[0101] (1) Adaptive time weighting for stronger representativeness. Under non-uniform sampling conditions, this invention uses the length of adjacent sampling periods as the time weight and employs trapezoidal integrals to obtain the daily representative value, which is then output as a fixed 24-hour output. This mechanism does not require manual weighting or data interpolation and can adaptively respond to situations such as sparse-dense uneven sampling and extended / shortened sampling intervals caused by sudden events. It avoids the systematic bias of "spreading out peaks" and "raising troughs" by equal-interval averaging, and more realistically depicts the continuous change process throughout the day, significantly improving the statistical validity and physical rationality of the concentration representative value.
[0102] (2) Comprehensive handling of day boundaries and rigorous cross-day connections. This invention proposes a boundary handling mechanism of "symmetrical closure from 0:00 to 24:00" and "priority cross-day at 0:00 of the next day": when a valid value exists at 0:00 (or 0:00 of the next day), it forms a trapezoid with adjacent valid points and is included in the daily integral; when it is missing, it remains closed to the boundary at the zero order. This rule systematically eliminates omissions / miscalculations caused by day boundary truncation, takes into account the engineering realities such as clock synchronization differences between different stations and intermittent shutdowns of monitoring equipment, and ensures that the calculation caliber is consistent across stations and across dates, the results are verifiable, and the errors are predictable, making it suitable for inclusion in standardized assessments and comparative evaluations.
[0103] (3) Validity constraints and anomaly protection ensure traceable and reliable results. This invention only integrates adjacent pairs where both endpoints are valid, and adopts a conservative strategy of "marking but not including" for missing measurements and out-of-bounds errors. Closure patterns and key variables are recorded at the beginning and end boundaries and at cross-day closures to form an auditable computational chain. This design avoids the uncertainty introduced by subjective interpolation and achieves deterministic output without relying on machine learning / deep learning. It also supports batch processing and linear expansion of multiple sites, multiple parameters, and long-term series, facilitating long-term, stable, and low-maintenance deployment in watershed-level monitoring networks.
[0104] (4) Strong batch deployment and engineering availability. It automatically traverses and calculates according to "date × parameter × station" and writes back to the table, outputting a unified structure (such as date, station, parameter, weighted average, etc.), with standardized naming and disk storage, supporting large-scale processing of multiple stations, multiple parameters, and long-term series; the algorithm complexity is linear for a single day, and the computing resources are controllable.
[0105] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for daily-scale processing of water quality data based on hourly automatic monitoring, characterized in that... Includes the following steps: S1, Invalid Data Determination: Define the set of invalid values. N The local time zone is used to define the date for all monitoring data. All timestamps are unified to the same time zone and deduplicated and sorted in ascending order; S2, Data Standardization Processing: After standardizing the monitoring data according to the {station, parameter, datetime, value} format, the invalid value set is used... N Identify invalid monitoring records and remove or replace them. Here, station represents the station name, parameter represents the monitoring indicator, datetime represents the local time zone timestamp, and value represents the indicator monitoring value. S3, Data Day Grouping and Adjacency Valid Construction: Grouping by station-parameter-datetime as the key to obtain the daily sequence. and in the daily sequence Construct adjacent valid pairs; S4, Integral of adjacent segments within the daily sequence: In the daily sequence In the middle, the integral of adjacent segments inside is calculated by the trapezoidal integral method; S5, Symmetrically Closed Start and End Boundaries of Daily Sequence: When calculating the start and end segments of the daily sequence, for the two cases where the effective values of the start and end points of the daily sequence exist or are missing, the integrals of the start and end segments of the daily sequence are obtained by the trapezoidal integral of the first segment respectively. S6, fuse the internal adjacent segment integrals and start-end segment integrals of the monitoring data obtained in steps S4 and S5 for the day to obtain the daily-scale monitoring data; Step S4, for any adjacent valid pair and The integral of adjacent segments is calculated using the trapezoidal integral method with non-uniform sampling. And merge to obtain the integral of the internal adjacent segments of the sequence for that day. , , In the formula, Indicates two consecutive valid sampling times within a day and The time integral contribution of water quality indicators between different periods. Indicates two valid sampling times and The length of the time interval between them Indicates at time Valid monitoring values of water quality indicators obtained from the site Indicates at time The effective monitoring value of this water quality indicator obtained from the site, This represents the integral contribution of all adjacent valid sampling time pairs within the same day. The total amount of internal time points obtained by summing them up for the day.
2. The method for daily-scale processing of water quality data based on hourly automatic monitoring according to claim 1, characterized in that: The set of invalid values in step S1 , , This indicates that data is missing. This indicates an instrument malfunction. This indicates that it cannot be detected. This indicates that information is missing.
3. The method for daily-scale processing of water quality data based on hourly automatic monitoring according to claim 1, characterized in that: In step S2, invalid monitoring records are removed, and for multiple monitoring records at the same time, only one is retained based on the source credibility or the "most recent" principle.
4. The method for daily-scale processing of water quality data based on hourly automatic monitoring according to claim 1, characterized in that: In step S5, the integral of the starting segment of the sequence at 0:00 on the current day is obtained as follows: When a valid value exists, take... With the first subsequent valid point of the day The first segment of the trapezoid is formed and included in the daily score: , It indicates the time from the start of the day 0:00 to the first subsequent valid time of the day. The contribution of time integrals between them This indicates the valid monitoring value of the water quality indicator at 0:00 on the day of the event. If no valid value exists at the starting point 0:00 of the day, a valid value will be collected by holding at zero order. Push forward to And obtained by the following calculation, , The integral of the initial segment of the sequence at 0:00 on the current day is obtained as follows: The next day If it exists and is valid, take the last valid point of the day. and Forming a trans-sun trapezoid, its contribution Counted towards daily points: , Indicates the last valid time of the day. The amount of time points contributed up to 24:00 on the same day. Indicates the last valid time of the day. The effective monitoring value of this water quality indicator is specified. This indicates the valid monitoring value of the water quality indicator at 0:00 the following day. This indicates the last valid sampling time in the sequence for that day; The next day If missing or invalid, it is preserved at order zero. Closed to : 。
Citation Information
Patent Citations
CN114564883A
US20200027096A1