Water body pollution monitoring and early warning method based on big data analysis
By constructing a historical fingerprint database and combining it with the matching and confidence quantification of real-time hydrological scenario data, the problem of water pollution early warning under sparse data was solved, achieving low-cost, high-timeliness, and reliable predictive early warning, and avoiding misleading predictions.
Patent Information
- Application Number
- CN202511663354.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-02-17
AI Technical Summary
Existing technologies struggle to provide effective predictive warnings of water pollution under sparse monitoring data conditions, particularly in their inability to effectively organize and reuse propagation characteristics from sparse historical data. Furthermore, existing methods are prone to overfitting or prediction failure when faced with sparse data.
By constructing a historical fingerprint database, we can mine the correlation characteristics between pollution fluctuation events at upstream monitoring points and downstream response signals. By combining this with real-time hydrological scenario data for matching and confidence quantification, we can achieve low-cost and high-timeliness predictive early warning.
Under sparse data conditions, an engineerable prediction method is provided, which can efficiently and reliably predict water pollution, avoid misleading predictions caused by blindly applying historical experience, and has comprehensive monitoring capabilities for different types of pollution.
Smart Images

Figure CN121543798A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a water pollution monitoring and early warning method based on big data analysis, belonging to the field of information processing and data mining technology. Background Technology
[0002] Currently, for state prediction of complex physical systems such as hydrology, meteorology, and industrial processes, two widely used mainstream technical approaches have emerged: one is the physical mechanism-based modeling approach, which attempts to fully reproduce the dynamic evolution process inside the system through complex partial differential equations and other means; the other is the data-driven approach represented by deep learning, which attempts to learn and fit the complex input-output relationship of the system by analyzing massive amounts of historical data and training models. However, when these mainstream data analysis methods are applied to real-world engineering scenarios such as watershed monitoring, a long-standing constraint, namely the spatiotemporal sparsity of monitoring data, limits the engineering application of the above methods. Due to economic and physical limitations, the practice of deploying high-density sensors in a vast space is difficult to implement on a large scale, which means that data analysis systems must operate under the real constraint of data sparsity. Under these constraints, both of the aforementioned mainstream approaches have exposed their inherent limitations: the physical mechanism-based modeling approach, due to its extremely stringent requirements for boundary conditions and medium parameters, is difficult to effectively calibrate in sparse data environments and has high computational costs, making it unsuitable for rapid and low-cost emergency deployment; while the data-driven approach, whose prediction accuracy is highly dependent on massive, high-density training data, presents a fundamental contradiction between this data-driven tool and the reality of sparse data when faced with sparse data, leading to the model being prone to overfitting or prediction failure.
[0003] Faced with this situation, the field is also exploring the use of lighter data mining algorithms. For example, Chinese invention patent CN120142596A discloses a water monitoring and early warning system and method based on big data. This method analyzes historical early warning records to statistically determine the frequency of specific pollutants appearing at different times of the day to identify their characteristic time periods. Based on this, subsequent real-time monitoring time points are dynamically adjusted. However, the essence of this approach is still a simple statistical analysis based on the temporal clustering of historical occurrences, and its data analysis process completely ignores the pollution at upstream and downstream monitoring points. The physical processes of pollution propagation also fail to take into account the decisive influence of dynamic hydrological conditions such as flow rate and water level on pollution propagation. Therefore, this method can only be used to adjust the sampling frequency of local monitoring points, and cannot solve the core problem of how to predict and warn of the downstream propagation of upstream pollution events based on dynamic hydrological conditions among sparse monitoring points. However, these attempts are often limited to simple statistics or isolated matching of historical data, and fail to effectively solve a deeper problem, namely, how to structure and reuse the fragmented historical experience fragments observed at sparse monitoring points that are bound to specific conditions.
[0004] Therefore, the technical problem to be solved by this invention is how to construct a brand-new data analysis and early warning method that can fundamentally accept the reality of sparse monitoring data, actively avoid the dependence on fitting complex physical processes or data-driven models, and instead extract, organize and construct a reusable experience knowledge base by deeply mining sparse historical data and extracting the propagation characteristics bound to dynamic environmental conditions. On this basis, through simple matching of real-time conditions and historical experience, a low-cost, high-timeliness and high-reliability predictive early warning can be achieved. Summary of the Invention
[0005] This invention provides a water pollution monitoring and early warning method based on big data analysis. Its main purpose is to solve the problem that existing complex modeling or data-driven methods are difficult to apply under the inherent sparsity constraints of monitoring data. There is an urgent need for a simple and effective data analysis path, which can achieve low-cost and high-timeliness prediction by mining and reusing historical propagation experience bound to dynamic operating conditions.
[0006] To achieve the above objectives, the present invention provides a water pollution monitoring and early warning method based on big data analysis, the method comprising: Step 101, Historical Fingerprint Database Construction Steps: Obtain historical water quality data from at least one upstream monitoring point and at least one downstream monitoring point within the watershed, along with historical hydrological scenario data for the corresponding time periods; mine and identify historical pollution fluctuation events from the historical water quality data of the upstream monitoring points; for each historical pollution fluctuation event, search and associate its historical response signal with the historical water quality data of the downstream monitoring points; extract pollution propagation feature fingerprints based on historical pollution fluctuation events and historical response signals; bind the pollution propagation feature fingerprints with the historical hydrological scenario data corresponding to the event and store them in the fingerprint database. Step 102, Real-time matching and confidence quantification steps: Real-time monitoring of upstream monitoring points to detect current pollution fluctuation events; When a current pollution fluctuation event is detected, current hydrological scenario data is acquired; Based on the current hydrological scenario data, in the fingerprint database, one or more pollution propagation feature fingerprints corresponding to the historical hydrological scenario data most similar to the current hydrological scenario data are matched and retrieved, and the matching similarity score representing the distance between the current hydrological scenario data and the most similar historical hydrological scenario data is calculated and output simultaneously. Step 103, the safety gating and early warning release step, compares the matching similarity score with a preset confidence threshold; when the matching similarity score meets the confidence threshold, the pollution propagation feature fingerprint retrieved in step 102 is applied to quantitatively predict the state of the current pollution fluctuation event propagating to downstream monitoring points and generate quantitative early warning information; when the matching similarity score does not meet the confidence threshold, the quantitative prediction is stopped and non-quantitative robust early warning information is generated instead.
[0007] Preferably, step 101, which involves mining and identifying historical pollution fluctuation events, specifically includes: step 201, using peak detection logic to identify historical peak events representing sudden pollution from historical water quality data at upstream monitoring points; step 202, using statistical trend analysis logic to mine and identify historical baseline drift trends representing chronic leakage from historical water quality data at upstream monitoring points; and step 101, which involves binding pollution propagation characteristic fingerprints with historical hydrological scenario data corresponding to the event and storing them in a fingerprint database, further includes: when a historical pollution fluctuation event is a historical peak event, adding an event pattern identifier to the pollution propagation characteristic fingerprint before storing it in the fingerprint database; and when a historical pollution fluctuation event is a historical baseline drift trend, adding a trend pattern identifier to the pollution propagation characteristic fingerprint before storing it in the fingerprint database.
[0008] Preferably, both historical hydrological scenario data and current hydrological scenario data are scenario feature vectors containing at least two hydrological parameters; the matching and retrieval in step 102 specifically involves: using the current scenario feature vector corresponding to the current hydrological scenario data... Based on the retrieval criteria, calculations are performed in an N-dimensional feature space within the fingerprint database. The historical scenario feature vector corresponding to each historical hydrological scenario data stored in the fingerprint database. Weighted Euclidean distance between and retrieve distance The pollution propagation feature fingerprint corresponding to one or more of the smallest historical scenario feature vectors; where, distance Calculated according to the following rules: ,in, The number of dimensions of the scenario feature vector. and They are respectively and The Dimensional parameter values, For the first The preset weight coefficients corresponding to the dimension parameters.
[0009] Preferably, the confidence threshold is a preset maximum tolerance distance threshold; in step 103, the matching similarity score is compared with the confidence threshold, specifically: the matching similarity score is compared with the maximum tolerance distance threshold, the matching similarity score represents the distance between the current hydrological scenario data and the closest historical hydrological scenario data; when the matching similarity score is less than the maximum tolerance distance threshold, it is determined that the matching similarity score meets the confidence threshold; when the matching similarity score is greater than or equal to the maximum tolerance distance threshold, it is determined that the matching similarity score does not meet the confidence threshold.
[0010] Preferably, the pollution propagation characteristic fingerprint is a multivariate feature vector, which includes at least one of the following: the propagation delay time of the historical response signal relative to the historical pollution fluctuation event; the peak attenuation rate of the peak concentration of the historical response signal relative to the peak concentration of the historical pollution fluctuation event; and the diffusion parameter of the signal shape of the historical response signal relative to the signal shape of the historical pollution fluctuation event.
[0011] Preferably, the scenario feature vector contains at least two hydrological parameters selected from at least one of the following: river flow data from upstream monitoring points, river water level data from downstream monitoring points, and rainfall data from the watershed.
[0012] Preferably, step 102, detecting the current pollution fluctuation event, specifically includes: step 701, running peak detection logic and statistical trend analysis logic in parallel to monitor water quality data at upstream monitoring points in real time; step 102, matching and retrieving, further includes: step 702, when the peak detection logic is triggered, using the current hydrological scenario data as the retrieval basis, specifically retrieving pollution propagation characteristic fingerprints with event pattern identifiers from the fingerprint database; step 703, when the statistical trend analysis logic is triggered, using the current hydrological scenario data as the retrieval basis, specifically retrieving pollution propagation characteristic fingerprints with trend pattern identifiers from the fingerprint database.
[0013] Preferably, when multiple pollution transmission feature fingerprints are retrieved in step 102, step 103 applies the pollution transmission feature fingerprints for quantitative prediction, specifically including: step 801, calculating a weighted average of the multiple pollution transmission feature fingerprints based on their respective matching similarity scores to generate a final prediction fingerprint; step 802, applying the final prediction fingerprint for quantitative prediction.
[0014] Preferably, the non-quantitative robust early warning information generated in step 103 is as follows: Step 901, when the matching similarity score does not meet the confidence threshold, a preset, globally unified fastest emergency propagation time is applied; Step 902, a boundary early warning information containing the fastest emergency propagation time and clearly indicating that the current situation exceeds the range of historical experience is generated and published.
[0015] Preferably, in step 101, historical pollution fluctuation events are mined and identified by using a sliding window-based Z-score detection algorithm to identify concentration peaks in historical water quality data of upstream monitoring points; historical response signals are searched and associated with these signals; and in the historical water quality data of downstream monitoring points, within a preset delay time window in which the concentration peak occurs, a cross-correlation analysis algorithm is used to search for and match a response signal peak with the highest morphological correlation.
[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. By mining historical water quality data, pollution fluctuation events at upstream monitoring points are correlated with response signals at downstream monitoring points, and pollution propagation feature fingerprints characterizing the relationship between the two are extracted. The key is to bind each extracted propagation feature fingerprint with historical hydrological scenario data at the time of the event, storing them together in a fingerprint database. This construction method allows the propagation patterns implicit in historical data to be transformed and solidified into a data structure that can be retrieved based on hydrological scenarios. In real-time monitoring, when an upstream event occurs, the system uses the currently acquired real-time hydrological scenario data as the retrieval basis to match the most relevant propagation feature fingerprint from historical experience in the database, and generates predictive warnings based on this. This method establishes an alternative path for data analysis, which does not rely on the construction of complex mechanistic models or large-scale data training. Instead, it provides an engineering-implementable prediction method under the constraint of data sparsity by extracting and reusing the correspondence between input, response, and data in sparse historical data.
[0017] 2. In the retrieval and matching process, a quantitative evaluation mechanism for the reliability of the matching results is further introduced. When the system retrieves the most similar historical hydrological scenario data based on the current hydrological scenario data, it simultaneously calculates the matching similarity score between the two. By comparing this score with a preset confidence threshold, the system is able to identify out-of-scenario conditions. When the matching result is reliable, i.e., the similarity score meets the threshold requirement, the system applies the retrieved fingerprint to perform quantitative prediction steps. However, when the matching result is unreliable, such as when the current hydrological scenario exceeds the coverage of historical experience, the system stops quantitative prediction and instead generates a non-quantitative robust early warning information based on boundary conditions. This design avoids the inherent risk that data analysis models may output misleading prediction information due to blindly applying historical experience when facing unknown scenarios, and constructs a dual-mode working method that combines high-confidence quantitative prediction with low-confidence safety alarms.
[0018] 3. During the data mining phase, two feature recognition logics were employed in parallel. Not only were sudden pollution events identified through fluctuation detection, but statistical trend analysis was also used to mine and identify historical chronic leakage trends characterized by a continuous rise in the water quality baseline. The propagation feature fingerprints corresponding to these two different modes were extracted and stored in a fingerprint database, and distinguished by pattern identifiers. During the real-time monitoring phase, the system also operated peak detectors and baseline trend detectors in parallel. Once a specific pollution pattern was detected, the system would specifically retrieve the fingerprint with the corresponding pattern identifier from the fingerprint database based on the current hydrological scenario. This design allowed a single data analysis framework and fingerprint database structure to be reused by two orthogonal pollution recognition logics, expanding the system's monitoring and cognitive capabilities from a single event dimension to a dual-modal dimension of event-trend coexistence, achieving comprehensive monitoring of different pollution types. Attached Figure Description
[0019] Figure 1 This is a flowchart of the early warning method based on fingerprint matching and security gating of the present invention; Figure 2 This is a schematic diagram showing the frequency distribution of the pollution propagation characteristic fingerprint of the present invention; Figure 3 This is a schematic diagram illustrating the limitations of existing technologies under sparse data constraints. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the scope of protection of the invention.
[0021] The water pollution monitoring and early warning method based on big data analysis provided by this invention has a core technical framework that constructs an offline mining-online matching closed-loop data processing and decision-making system. This system executes within a computational intelligence framework and mainly includes two interconnected stages: the historical fingerprint database construction step defined in step 101, and the real-time matching, confidence quantification, and security gating early warning steps defined in steps 102 and 103. The former stage, as a data mining and knowledge base construction process, is responsible for extracting and structurally storing scenario-response experience from sparse historical data. The latter stage, as an online pattern matching and decision support process, is responsible for retrieving and applying historical experience in real-time monitoring and ensuring the safety of decision-making through the confidence quantification mechanism. In step 101, the historical fingerprint database construction step, the system acquires and integrates two types of historical data: one is watershed data... The system collects all historical water quality data from at least one upstream monitoring point (e.g., station A) and at least one downstream monitoring point (e.g., station C). This data is a high-frequency time series, including pollution indicators such as COD and ammonia nitrogen. Another type of data is historical hydrological scenario data for the corresponding time period, such as river flow data at station A and river water level data at station C. The system performs data mining and feature extraction to identify and correlate historical pollution propagation events. Two feature identification logics are used in parallel: the first is a peak detection logic for sudden pollution events, used to identify historical peak events; the second is a statistical trend analysis logic for chronic leaks, used to mine historical baseline drift trends. When executing the peak detection logic to mine and identify historical pollution fluctuation events at upstream monitoring point A, the system uses a sliding window-based Z-score detection algorithm. The input to this algorithm is the historical water quality time series of station A. The processing logic is as follows: set a sliding window for statistics, for example... Hours, and calculate the mean within that window. with standard deviation For each data point within the window Calculate its Z-score The system presets a Z-score threshold. This threshold can be calibrated based on the distribution of historical data, for example... ,when When the event occurs, the system determines that the point is a historical pollution fluctuation event and outputs the timestamp of the event. With peak concentration .
[0022] For each identified The system initiates the associated steps for downstream response, based on historical water quality data from downstream station C. The system searches for its historical response signal. This process presents an objective challenge: downstream signals can be distorted due to dispersion and attenuation. To achieve reliable correlation, the system employs a cross-correlation analysis algorithm. The input to this algorithm is... as well as The nearby upstream signal segment; its processing logic defines a preset delay time window based on physical common sense, for example... From the following 1 hour to 72 hours, the system calculates the upstream signal segment and the downstream signal segment only within this window. The morphological cross-correlation coefficients between sliding segments were analyzed, and the time points with the highest correlations were searched. and its peak to regard it as a counterpart to The system retrieves associated historical response signals. Once the association is successful, it immediately extracts a pollution propagation feature fingerprint based on the historical pollution fluctuation event and the historical response signal. This fingerprint is a multivariate feature vector, and its input is... , , , And the morphological parameters of the two signals, such as their respective half-width and full height. and Its processing logic involves calculating and outputting a set of quantitative indicators characterizing the propagation process. This multivariate feature vector includes at least: propagation delay time. Peak attenuation rate ; and the dispersion parameter When performing statistical trend analysis logic to uncover historical baseline drift trends, the system employs time series decomposition or moving average analysis methods, for example, for... The system calculates a 7-day moving average and identifies persistent upward trends in this average sequence that last for more than a threshold (e.g., more than 48 hours), classifying them as historical chronic leakage trends. The system employs similar correlation logic to search for the responsiveness trends corresponding to downstream station C and extracts the corresponding trend propagation feature fingerprints, such as trend propagation delay. and trend decay rate The final step in building the knowledge base is to bind the extracted fingerprints to the scenarios and store them in the database. To address the impact of complex hydrological conditions, such as upstream flow and downstream backwater effects, the historical hydrological scenario data acquired by the system is constructed into a scenario feature vector containing at least two hydrological parameters. Taking a two-dimensional vector as an example, ,in This refers to the upstream river flow data for station A. The system will extract pollution propagation characteristic fingerprints from the downstream river channel of station C. Corresponding to this event The system binds the structured knowledge together; simultaneously, it attaches a pattern identifier to the entry, such as an event pattern identifier or a trend pattern identifier, ultimately binding the structured knowledge entry to this pattern. The fingerprint is stored in the fingerprint database, and the knowledge base is constructed by exhaustively searching through historical data.
[0023] In step 102, the real-time matching and confidence quantification step, the system operates in online monitoring mode. The system runs the aforementioned peak detection logic and statistical trend analysis logic in parallel to monitor the water quality data of upstream station A in real time. When any detector detects a current pollution fluctuation event, such as when the peak detection logic is triggered, the system immediately acquires the current hydrological scenario data and constructs the current scenario feature vector. ,For example The system is based on Based on the retrieval criteria, pattern matching is performed in the fingerprint database; this process is adaptive: based on the triggered detector, i.e., the peak detection logic in this example, the system specifically searches the fingerprint database for historical fingerprints with event pattern identifiers; this matching retrieval process is performed within a... In the 3D feature space (in this example) ) Execution -Nearest neighbor query; system calculation Compare with all historical scenario feature vectors in the fingerprint database that meet the criteria. Weighted Euclidean distance between The distance Calculated according to the following rules: in, The number of dimensions of the scenario feature vector ( ), and They are respectively and The Dimensional parameter values (such as flow rate or water level). For the first The preset weighting coefficients corresponding to the dimension parameters; these weighting coefficients The determination of this parameter is a parameter calibration procedure: it can be quantified by analyzing historical fingerprint database data and using multiple regression or sensitivity analysis. or right The degree of sensitivity to change, and setting accordingly. Values, for example, if the flow rate The impact is water level 1.5 times, which can be set after data normalization. and System retrieval distance The smallest one or more, A historical fingerprint During the search, the system simultaneously calculates and outputs a matching similarity score. The score The minimum distance retrieved Characterization; the score It quantifies the distance between the current situation and historical experience, a smaller A value indicates a high degree of reliability in the match.
[0024] In step 103, the security gating and early warning release step, the system activates the decision gating mechanism; the system uses the matching similarity score output in step 102. (Right now ), and a preset confidence threshold Compare; the confidence threshold It is a maximum tolerance distance threshold, which represents the minimum threshold at which the system acknowledges the reference value of historical experience; The calibration is also a deterministic procedure: the system can analyze all fingerprints in the database offline. The distance distribution between each pair, and A statistical value, such as the 95th percentile, is set for this distance distribution; any match greater than this distance is considered an out-of-context event. Based on this comparison, the system performs a bimodal decision: First path: when When the matching similarity score meets the confidence threshold, it proves that the historical experience is reliable; the system uses the retrieved pollution transmission feature fingerprints for quantitative prediction; if multiple fingerprints are retrieved in step 102, There are , and their distances are respectively . The system then performs a weighted average calculation: calculating the weights of each component based on the matching similarity scores. ,For example This assigns higher weight to fingerprints that are closer in distance; then the final fingerprint for prediction is calculated. Finally, system application. (For example and Quantitative prediction of current pollution fluctuation events, for example, predicted arrival time = current time + Predicted peak concentration = current peak concentration And generate quantitative early warning information; the second path: when If the matching similarity score fails to meet the confidence threshold, indicating that the current scenario is an out-of-context event and historical experience is unreliable, the system immediately suspends quantitative prediction to avoid releasing misleading information. Instead, the system generates non-quantitative robust early warning information, which uses a preset, globally unified fastest emergency propagation time. ;Should The value is a conservative lower bound based on the historical maximum flow velocity under physical boundary conditions, obtained through offline calibration; the system issues a boundary warning message, which includes... It also explicitly states that the current scenario exceeds the scope of historical experience.
[0025] Example 1: The method described in this example is a specific application of a water pollution monitoring and early warning method based on big data analysis under a specific working condition. This method operates in a tidal river section with complex hydrological conditions. There is an industrial discharge outlet near station A upstream of this river section, and station C downstream is an important drinking water intake. The water quality safety of this section faces the dual pressure of sudden pollution upstream and tidal backwater downstream. The system has constructed a fingerprint database based on historical data of the past 5 years for this watershed, according to step 101 in the specific implementation method. The historical hydrological scenario data in this database are all constructed to include the river flow at station A upstream. and the river water level at downstream station C Two-dimensional scene feature vector In order to address the impact of tidal backflow on the duration of pollution transmission The impact; meanwhile, the system has been calibrated offline and stored a confidence threshold for step 103. ,Should The value is set to all values in the historical database. Weighted Euclidean distance between The 95th percentile represents the coverage boundary of historical experience. On a specific workday, during real-time monitoring in step 102, the online monitoring instrument at upstream station A detected a sudden COD concentration peak, indicating a current pollution fluctuation event. The system immediately triggered the early warning process. Simultaneously, the current hydrological scenario data acquired by the system showed that the current operating conditions were an extreme flood period occurring once every few decades, leading to an increase in the river flow at station A. At its highest historical level, and with the downstream C station experiencing both a spring tide and a strong storm surge, the river level at station C has risen significantly. This data is also at a level far exceeding historical records; the system immediately constructs a current scenario feature vector from this set of data. Based on this, the system specifically searches the fingerprint database for fingerprints with event pattern identifiers; the system follows the definition in the specific implementation method. 3D space matching logic, calculation With all fingerprints in the database Weighted Euclidean distance between And retrieve the historical fingerprint with the smallest distance. ,Should This corresponds to a historical high-flow-high-water-level condition; during this process, the system simultaneously calculates and outputs a matching similarity score. ,Right now and minimum distance between The system then enters the security gating logic in step 103, which will... Compared with the preset confidence threshold Comparison; due to The scenario represented, namely extreme flooding superimposed on extreme storm surge, has never occurred in the historical database of the past 5 years, and its vector position is... The out-of-context region of the feature space leads to its similarity to the most similar historical fingerprint among the least similar. Distance between (Right now It is still greater than the confidence threshold. ,Right now .
[0026] This comparison result activates the decision gate for matches with similarity scores that do not meet the confidence threshold; the system therefore suspends the application of the retrieved data. The system avoids issuing a highly misleading quantitative warning by blindly applying irrelevant historical experience, instead employing a safety strategy for low-confidence paths, generating robust, non-quantitative warning information. The system also retrieves a globally preset, physically boundary-calibrated fastest emergency propagation time. The system immediately issued a boundary warning to downstream station C, clearly indicating that the current hydrological situation exceeded the historical experience database and could not be accurately predicted quantitatively. The fastest emergency communication time of 2.5 hours had been activated. This data analysis method, through the matching of multi-dimensional scenario feature vectors in step 102 and the coordinated operation of matching confidence quantification and safety gating in step 103, enables the system to identify out-of-scenario conditions without relying on complex physical models. A decision-making balance mechanism was established between the two data analysis needs of reusing historical experience for prediction and avoiding misjudgments caused by unknown scenarios. This ensures that when facing real-world challenges beyond data distribution, the decision support information output by the early warning system always prioritizes safety.
[0027] Example 2: This example objectively verifies the effectiveness and security of data processing, particularly the synergistic effect of multi-dimensional scenario feature vectors and security gating mechanisms in prediction. The experimental data comes from a hydrodynamic simulation model, which has been calibrated based on the physical parameters and historical hydrological data of the target watershed, from upstream station A to downstream station C. This model can generate predictions with real propagation time, including the interaction of tides and flow. The system generates a hydrological and water quality dataset; using 5 years of historical data generated by the model, and following the data mining and knowledge base construction procedure in step 101, it offline constructs a fingerprint database for testing; the system also generates a 1-year test dataset, which includes 100 known occurrence times and upstream peak concentrations. and the actual downstream transmission time The current pollution fluctuation events; these 100 events are divided into two groups: 80 in-situ events, whose current situation feature vectors All of these can be found in the historical fingerprint database with high confidence; and there are 20 out-of-context events, whose All of these are extreme operating condition combinations not included in the historical database.
[0028] To compare performance, three data analysis test groups were set up: Control Group 1: A simplified matching method was used, in which fingerprint database construction and real-time matching only used upstream traffic. As a one-dimensional scenario feature vector, it lacks the security gating mechanism of step 103; Control group 2: uses an incomplete matching method, which adopts a specific implementation method. 3D context feature vector However, it lacks the security gating mechanism of step 103, meaning the system always trusts the nearest neighbor matching result and forces a quantitative prediction output; the sample group of this invention adopts the complete technical solution disclosed, including based on 3D Context Feature Vector Matching, and matching similarity scores With confidence threshold Step 103 of the comparison involves a security gating mechanism; the evaluation indicators for the experiment include: the average value of quantitative predictions for 80 scenarios. Prediction error, in hours, is calculated as: Error = |Prediction - For 20 out-of-context events, the number of severely misleading predictions was counted. A method that outputs a quantitative prediction for an out-of-context event with a prediction error greater than [a certain value] is considered a successful prediction. When the accuracy rate is 50%, it is recorded as one serious misleading prediction; the sample group of this invention is in the case of security gating activation ( The non-quantitative robust early warning information output is not considered a misleading prediction.
[0029] Table 1: Results of the Comparative Experiment on Data Analysis Methods
[0030] Referring to Table 1, the data analysis results show that: comparing the performance of control group 1 and control group 2 on in-scenario events, control group 1 (using only flow rate) performed better. The prediction error for the control group 2 was 4.52 hours, while the prediction error for the control group 2 was 4.52 hours. and water level The prediction error of the 2D vector was reduced to 0.68 hours; this objectively confirms that in watersheds with complex hydrological conditions, the specific implementation method, Data mining and matching of 1D scenario feature vectors is a prerequisite for achieving high prediction accuracy. 1D matching suffers from poor prediction results due to index confounding issues. Secondly, comparing the performance of control group 2 and the sample group of this invention on out-of-scenario events, their prediction accuracy within the scenario is similar (0.68 hours vs. 0.71 hours), but their performance differs under out-of-scenario conditions. Control group 2, lacking security gating, applied fingerprints from distant locations in the feature space, resulting in severely misleading predictions in all 20 out-of-scenario events. In contrast, the security gating mechanism of the sample group of this invention was activated in all 20 out-of-scenario events. As a result, the system stopped quantitative prediction and instead issued non-quantitative robust early warning information, with zero instances of seriously misleading predictions. Experimental data confirms that the data analysis method disclosed in this invention, by combining the matching of multi-dimensional scenario feature vectors with the security gating of matching confidence quantification, ensures the accuracy of predictions within the scenario while avoiding the risk of the model outputting misleading information when facing data outside the scenario.
[0031] Example 3: To further verify the necessity of step 103 of the above-mentioned security gating mechanism, the following comparative example is set up. This example uses the same simulation dataset and 100 test events as Example 2 to set up a control group; the data analysis method of this control group is the same as that of the present invention sample group in Example 2 in steps 101 and 102, both using flow-based... and water level The method uses 2D contextual feature vectors for data mining and matching; its only key difference is that the control group's method does not include the security gating and early warning release mechanism defined in step 103; that is, after retrieving the most similar historical fingerprint, this method does not consider its matching similarity score. (Right now Does it meet the confidence threshold? The system performed quantitative predictions in all cases. The method of the control group was applied to the same group of 100 test events, using the same evaluation metrics as in Example 2. The experimental results are recorded as follows: For 80 in-scenario events, the average prediction rate of the control group method was... The prediction error was 0.68 hours; for 20 out-of-concept events, the number of seriously misleading predictions was 20.
[0032] Data analysis revealed that while the control group method ensured in-context prediction accuracy with an error of only 0.68 hours through multi-dimensional scenario feature vectors, its data processing logic exhibited a flaw under specific conditions: when faced with 20 out-of-context events, its matching algorithm still returned a mathematical nearest neighbor fingerprint, but the fingerprint corresponding to... The values do not meet the requirements ( The confidence threshold calibrated in Example 2 Due to the lack of safety gating, this method uses historical fingerprints that are no longer relevant to the current operating conditions for quantitative prediction, resulting in seriously misleading predictions for all 20 out-of-context events. The results of this comparison confirm that multidimensional scenario matching steps 101 and 102 alone are insufficient to fully address the data analysis challenges of hydrological early warning. It is necessary to introduce the matching confidence quantification and safety gating mechanism defined in step 103 in a coordinated manner to avoid the risk of the system outputting misleading predictions when facing unknown operating conditions.
[0033] Example 4: This example combines Figures 1 to 3 This section describes a water pollution monitoring and early warning method based on big data analysis, such as... Figure 1 As shown, the historical data acquisition step acquires upstream / downstream historical water quality data and historical hydrological scenario data. This process uses peak detection logic in parallel to identify historical peak events and statistical trend analysis logic to mine historical baseline drift trends. Both are used to extract feature fingerprints, which are then bound to historical hydrological scenarios and stored in a historical fingerprint database. In the real-time monitoring stage, the real-time monitoring and event detection steps are used to detect current pollution fluctuation events and acquire current hydrological scenario data, thereby triggering real-time matching and confidence quantification, i.e., step 102. This step retrieves the historical fingerprint database and calculates the matching similarity score, entering the security gating, i.e., step 103. Here, the matching similarity score is compared with the confidence threshold. When the threshold is met, the system applies the retrieved fingerprint to perform quantitative prediction and generate quantitative early warning information. When the threshold is not met, the system stops quantitative prediction and generates non-quantitative robust early warning information. Finally, both paths execute the release of early warning information.
[0034] like Figure 2 As shown, the horizontal axis represents parameter values, ranging from 0 to 50, and the vertical axis represents frequency, ranging from 0 to 40. The figure exemplarily illustrates the distribution of three core characteristic parameters, namely propagation delay time. (hours) (represented by solid line), peak decay rate (represented by dashed lines) and dispersion parameters (Represented by dotted lines). The frequency distribution of these three parameters all exhibits a single-peak shape, with the frequency lowest at both ends of the parameter value, close to 0 and 50. Specifically, the propagation delay time... The frequency of the peak value reaches a peak of 40 when the parameter value is 25, while the frequencies of the peak attenuation rate and the dispersion parameter both reach their respective peak values of 35 and 30 when the parameter value is 20.
[0035] like Figure 3 As shown in the figure, this illustrates the technical background that the present invention aims to address: the limitations of traditional monitoring and early warning methods under sparse data. These limitations stem from the real-world constraint of the spatiotemporal sparsity of monitoring data. This makes physical mechanism modeling methods extremely demanding in terms of boundary conditions and medium parameter inputs, and difficult to effectively calibrate in sparse data environments. Furthermore, data-driven methods, due to their high dependence on massive, high-density training data, are prone to overfitting or prediction failures under sparse data conditions. At the same time, existing data mining methods also limit their application effectiveness due to the lack of simplified data processing methods and the inability to effectively organize and reuse historical experience fragments.
[0036] Example 5: To ensure the reproducibility of water pollution monitoring and early warning methods based on big data analysis when deployed in specific watersheds, a standardized calibration procedure for data mining parameters is implemented. This procedure eliminates the dependence of the data mining algorithm on empirical values in step 101. This calibration procedure is executed during the offline analysis phase when the system is first deployed in a new watershed, taking upstream station A and downstream station C as an example. Its input is at least three years of historical water quality and hydrological data for the watershed. First, the key parameters of the peak detection algorithm for historical pollution fluctuation events in step 101, i.e., Z-score detection, are calibrated to determine the sliding window. The length of the system is used to analyze the autocorrelation of water quality data. The window is set to the average time length during which the data correlation drops to 0.5, thus ensuring that the window represents a statistically stable baseline, for example, when calculated in this watershed. Hours; to determine the Z-score threshold The system calculates in The Z-score distribution of all historical data within the hour was analyzed, and the top 0.1% of extreme data (i.e., known pollution events) were removed. Then, The 99.9th percentile of the remaining data distribution is used as the statistical boundary for the algorithm to distinguish random noise from events. For example, in this watershed, it is calibrated as... Second, the system calibrates the preset delay time window for the correlation search algorithm of historical response signals in step 101, i.e., the cross-correlation analysis. The system retrieves the operating conditions with the highest flow velocity, such as during flood season, and the lowest flow velocity, such as during dry season, from historical hydrological data, and combines this with the river channel distance between station A and station C. Estimate the fastest physical propagation time and slowest propagation time To accommodate data errors, the algorithm's preset delay time window is set to [ ]. , This range is, for example, [0.8 hours, 65 hours].
[0037] Third, the statistical trend analysis logic for the historical baseline drift trend in steps 202 and 701 is made transparent and its parameters are calibrated. In this embodiment, the logic is implemented using a double moving average algorithm, and its input is the water quality time series of upstream station A. The processing logic includes: calculating a long-term moving average. its window Set to 168 hours, or 7 days, and a short-term moving average. its window The time frame is set to 24 hours; the algorithm's trend recognition and judgment rules are quantified as follows: when continuous Data points, The calibration is set for 12 hours, and the duration is greater than 12 hours. At that time, the algorithm determines the start of a historical chronic leakage trend, where The baseline sensitivity is calibrated to 0.15; when Falling back to When the following occurs, the trend ends; the algorithm extracts the trend propagation feature fingerprint of this trend, which... Calculated as downstream station C Sequence and upstream station A The time delay corresponding to the peak cross-correlation of the sequence during the trend event; by executing the above parameter calibration procedure, all key parameters in step 101 of the data mining algorithm, including , Delay window , , , All of these are converted from empirical values to calculated values based on the statistical characteristics of the data of the deployed watershed itself, such as autocorrelation, data distribution, physical boundaries, and trend characteristics, ensuring the reproducibility of step 101 in the fingerprint database construction process.
[0038] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A water pollution monitoring and early warning method based on big data analysis, characterized in that, The method comprises: Step 101, a historical fingerprint library construction step, obtaining historical water quality data of at least one upstream monitoring point and at least one downstream monitoring point in a river basin, and historical hydrological scenario data corresponding to a time period; from the historical water quality data of the upstream monitoring point, a historical pollution fluctuation event is mined and identified; for each historical pollution fluctuation event, a historical response signal is searched and associated in the historical water quality data of the downstream monitoring point; based on the historical pollution fluctuation event and the historical response signal, a pollution propagation characteristic fingerprint is extracted; the pollution propagation characteristic fingerprint is bound with the historical hydrological scenario data corresponding to the event, and stored in the fingerprint library; Step 102, a real-time matching and confidence quantification step, real-time monitoring of the upstream monitoring point to detect a current pollution fluctuation event; when the current pollution fluctuation event is detected, current hydrological scenario data is obtained; with the current hydrological scenario data as the retrieval basis, one or more pollution propagation characteristic fingerprints corresponding to the historical hydrological scenario data most similar to the current hydrological scenario data are matched and retrieved in the fingerprint library, and a matching similarity score representing the distance between the current hydrological scenario data and the most similar historical hydrological scenario data is calculated and output simultaneously; Step 103, a safety gating and early warning release step, comparing the matching similarity score with a preset confidence threshold; when the matching similarity score meets the confidence threshold, the pollution propagation characteristic fingerprint retrieved in step 102 is applied to quantitatively predict the state of the current pollution fluctuation event propagating to the downstream monitoring point, and quantitative early warning information is generated; when the matching similarity score does not meet the confidence threshold, the quantitative prediction is aborted, and non-quantitative robust early warning information is generated instead. 2.The water pollution monitoring and early warning method based on big data analysis according to claim 1, characterized in that, In step 101, the historical pollution fluctuation event is mined and identified, specifically including: step 201, using peak detection logic to identify a historical peak event representing sudden pollution from the historical water quality data of the upstream monitoring point; step 202, using statistical trend analysis logic to mine and identify a historical baseline drift trend representing chronic leakage from the historical water quality data of the upstream monitoring point; in step 101, the pollution propagation characteristic fingerprint is bound with the historical hydrological scenario data corresponding to the event, and stored in the fingerprint library, further including: when the historical pollution fluctuation event is a historical peak event, the pollution propagation characteristic fingerprint is stored in the fingerprint library after being attached with an event mode identifier; when the historical pollution fluctuation event is a historical baseline drift trend, the pollution propagation characteristic fingerprint is stored in the fingerprint library after being attached with a trend mode identifier. 3.The water pollution monitoring and early warning method based on big data analysis according to claim 1, characterized in that, Both historical and current hydrological scenario data are scenario feature vectors containing at least two hydrological parameters; the matching and retrieval in step 102 specifically involves: using the current scenario feature vector corresponding to the current hydrological scenario data... Based on the retrieval criteria, calculations are performed in an N-dimensional feature space within the fingerprint database. The historical scenario feature vector corresponding to each historical hydrological scenario data stored in the fingerprint database. Weighted Euclidean distance between and retrieve distance The pollution propagation feature fingerprint corresponding to one or more of the smallest historical scenario feature vectors; where, distance Calculated according to the following rules: ,in, The number of dimensions of the scenario feature vector. and They are respectively and The Dimensional parameter values, For the first The preset weight coefficients corresponding to the dimension parameters.
4. The water pollution monitoring and early warning method based on big data analysis according to claim 1, characterized in that, The confidence threshold is a preset maximum tolerance distance threshold; in step 103, the matching similarity score is compared with the confidence threshold, specifically: the matching similarity score is compared with the maximum tolerance distance threshold, and the matching similarity score represents the distance between the current hydrological scenario data and the most similar historical hydrological scenario data; when the matching similarity score is less than the maximum tolerance distance threshold, it is determined that the matching similarity score meets the confidence threshold; When the matching similarity score is greater than or equal to the maximum tolerance distance threshold, it is determined that the matching similarity score does not meet the confidence threshold. 5.The water pollution monitoring and early warning method based on big data analysis according to claim 1, characterized in that, The pollution propagation characteristic fingerprint is a multi-feature vector, and the multi-feature vector includes at least one selected from the following: a propagation delay time of a historical response signal relative to a historical pollution fluctuation event; a peak value decay rate of a peak value concentration of the historical response signal relative to a peak value concentration of the historical pollution fluctuation event; and a dispersion degree parameter of a signal form of the historical response signal relative to a signal form of the historical pollution fluctuation event.
6. The water pollution monitoring and early warning method based on big data analysis according to claim 3, characterized in that the scenario The at least two hydrological parameters included in the feature vector are selected from at least one of the following: river flow data of an upstream monitoring point, river water level data of a downstream monitoring point, and rainfall data of a river basin.
7. The water pollution monitoring and early warning method based on big data analysis according to claim 2, characterized in that, The detection of the current pollution fluctuation event in step 102 specifically includes: step 701, running peak value detection logic and statistical trend analysis logic in parallel to monitor water quality data of the upstream monitoring point in real time; and step 102, matching and searching further includes: step 702, when the peak value detection logic is triggered, searching for a pollution propagation characteristic fingerprint with an event mode identifier in the fingerprint library based on the current hydrological scenario data; and step 703, when the statistical trend analysis logic is triggered, searching for a pollution propagation characteristic fingerprint with a trend mode identifier in the fingerprint library based on the current hydrological scenario data. 8.The water pollution monitoring and early warning method based on big data analysis according to claim 1, characterized in that, When multiple pollution propagation characteristic fingerprints are searched in step 102, the pollution propagation characteristic fingerprint is applied for quantitative prediction in step 103, specifically including: step 801, based on the respective matching similarity scores of the multiple pollution propagation characteristic fingerprints, performing weighted average calculation on the multiple pollution propagation characteristic fingerprints to generate a final prediction fingerprint; and step 802, applying the final prediction fingerprint for quantitative prediction. 9.The water pollution monitoring and early warning method based on big data analysis of claim 1, wherein, The non-quantitative robust warning information generated in step 103 is specifically: step 901, when the matching similarity score does not satisfy the confidence threshold, applying a preset, globally unified fastest emergency propagation time; and step 902, generating and publishing a boundary warning information containing the fastest emergency propagation time and explicitly prompting that the current scenario is out of the range of historical experience. 10.The water pollution monitoring and early warning method based on big data analysis according to claim 1, characterized in that, In step 101, the historical pollution fluctuation event is mined and identified, a Z-score detection algorithm based on a sliding window is used to identify a concentration peak value in historical water quality data of the upstream monitoring point; and a historical response signal is searched and associated, a cross-correlation analysis algorithm is used to search and match a response signal peak with the highest form correlation in a preset delay time window of the concentration peak value in the historical water quality data of the downstream monitoring point.
Citation Information
Patent Citations
Water body monitoring and early warning system and method based on big data
CN120142596A