Hydrological data management method and system based on machine learning

Through machine learning-based hydrological data governance methods, combined with fuzzy entropy and Granger causal analysis, the shortcomings of existing hydrological data anomaly detection and traceability methods are solved, and efficient and accurate hydrological data anomaly identification and traceability are achieved.

CN120104967AActive Publication Date: 2025-06-06NANJING HYDRAULIC RES INST

Patent Information

Application Number
CN202510174505.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing hydrological data anomaly detection and traceability methods have shortcomings in accuracy, real-time and intelligence, and it is difficult to adapt to dynamic hydrological anomaly recognition in complex environments, with a high misjudgment rate, and lack of a causal inference mechanism, so it is impossible to accurately locate the root cause of the abnormality.

Method used

Using a hydrological data governance method based on machine learning, we use the hydrological monitoring data to obtain and preprocess hydrological monitoring data, calculate the fuzzy entropy value and build a causal relationship model, and combine Granger's causal analysis and fuzzy entropy optimization strategy to realize dynamic anomaly detection and reverse causal analysis.

Benefits of technology

It improves the accuracy and adaptability of abnormal identification of hydrological data, can more accurately identify the true causal relationship between hydrological variables, automatically trace the abnormal trigger factors and quantify their impact, and improves the reliability and efficiency of abnormal traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104967A_ABST
    Figure CN120104967A_ABST
Patent Text Reader

Abstract

The invention discloses a hydrological data management method and system based on machine learning. The method comprises the following steps: S1, generating a hydrological monitoring data set with a unified temporal-spatial resolution; s2, constructing a fuzzy entropy matrix of the hydrological monitoring data; s3, generating a hydrological monitoring data causal model expressed in a directed acyclic graph form; s4, identifying the abnormal hydrological monitoring data in real time by dynamically adjusting the abnormal detection threshold value; s5, generating an abnormal traceability candidate path, and forming an abnormal hydrological monitoring data causal chain; and S6, performing fuzzy entropy optimization processing on the abnormal traceability candidate path, screening out a key abnormal causal path according to the fuzzy entropy value of the hydrological variable in each causal chain and the influence weight of the hydrological variable in the causal chain, and further positioning the root cause of the abnormal hydrological monitoring data. The invention provides a more efficient and intelligent technical scheme for water resource management, disaster prevention early warning and environment monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of hydrological technology, and in particular to a hydrological data management method and system based on machine learning. Background Art

[0002] With the continuous development of hydrological monitoring technology, various types of hydrological sensors have been widely used in water resources management, disaster prevention and mitigation, and ecological environment monitoring to realize real-time collection of key hydrological variables. However, due to the complexity of the natural environment, limitations on sensor accuracy, and interference in the data transmission process, there are often noise, missing values, or abnormal data in the hydrological monitoring data. Abnormal data may be caused by equipment failure, data collection errors, and extreme climate events, which poses a great challenge to the accuracy and reliability of hydrological monitoring data.

[0003] At present, the anomaly detection methods of hydrological data mainly rely on time series analysis, statistical methods and machine learning technology. For example, the method based on threshold setting can quickly detect abnormal data beyond the normal range, but this method relies on fixed thresholds set by experience and is difficult to adapt to the dynamically changing hydrological environment. The method based on time series prediction uses historical data to establish trend models, such as autoregressive moving average models or long short-term memory networks, but when the data anomalies are more complex or are greatly affected by the external environment, the prediction accuracy is likely to decrease; the anomaly detection method based on clustering and classification can detect abnormal data points, but due to the lack of in-depth understanding of the causal relationship between hydrological variables, it is easy to misjudge emergencies as anomalies, or cannot accurately distinguish the true source of abnormal data.

[0004] In addition, in terms of tracing the source of hydrological data anomalies, existing technologies mainly rely on statistical correlation analysis or simple rule inference, which makes it difficult to effectively identify the true causal relationship between variables. At the same time, traditional anomaly tracing methods lack intelligent processing methods for large-scale hydrological data, and it is difficult to automatically track and trace anomalies when faced with complex hydrological systems.

[0005] In summary, the existing methods for detecting and tracing anomalies in hydrological data still have many deficiencies in terms of accuracy, real-time and intelligence: on the one hand, existing methods often rely on fixed thresholds or statistical correlations, which are difficult to adapt to dynamic hydrological anomaly identification in complex environments, resulting in a high misjudgment rate; on the other hand, the lack of causal inference mechanism makes it impossible to accurately locate the root cause of the anomaly, resulting in a lack of scientific basis for the anomaly tracing process. In addition, in the context of massive hydrological data, traditional methods are difficult to achieve efficient automated anomaly detection and tracing, resulting in a lag in response speed, affecting the accuracy of water resources management and disaster prevention decisions. Summary of the invention

[0006] One purpose of the present invention is to propose a hydrological data management method and system based on machine learning. The present invention provides a more efficient and intelligent technical solution for water resources management, disaster prevention and early warning, and environmental monitoring.

[0007] A hydrological data management method based on machine learning according to an embodiment of the present invention includes the following steps:

[0008] S1. Obtain and preprocess hydrological monitoring data to generate a hydrological monitoring data set with uniform temporal and spatial resolution;

[0009] S2. Calculate the fuzzy entropy value of each hydrological variable for the hydrological monitoring data set, form a fuzzy entropy evaluation result for measuring the uncertainty of each hydrological variable, and construct a fuzzy entropy matrix of the hydrological monitoring data;

[0010] S3. Based on the hydrological monitoring data set and its fuzzy entropy matrix, the Granger causality analysis method is used to construct a causal relationship model between various hydrological variables in the hydrological monitoring data. The direct and indirect causal relationships between various hydrological variables are determined by causal tests under different time lag conditions, and a causal model of hydrological monitoring data represented by a directed acyclic graph is generated;

[0011] S4. Continuously monitor the hydrological monitoring data set, and use the hydrological monitoring data causal model combined with the fuzzy entropy optimization strategy to detect anomalies in the real-time hydrological monitoring data set, and dynamically adjust the anomaly detection threshold to identify abnormal hydrological monitoring data in real time;

[0012] S5. When abnormal hydrological monitoring data is detected, reverse causal analysis is performed on the abnormal hydrological monitoring data according to the hydrological monitoring data causal model, and candidate abnormal source tracing paths are generated based on the Granger causal analysis results and fuzzy entropy evaluation results to form a causal chain of abnormal hydrological monitoring data;

[0013] S6. Perform fuzzy entropy optimization on the candidate paths for anomaly tracing, screen out the key anomaly causal paths based on the fuzzy entropy values ​​of the hydrological variables in each causal chain and their influence weights in the causal chain, and then locate the root causes of the abnormal hydrological monitoring data.

[0014] Optionally, the step S1 includes:

[0015] S11. Acquire hydrological monitoring data in real time through the hydrological monitoring sensor network, the hydrological monitoring data including rainfall, flow rate, water level, evaporation and temperature, and preliminarily store the hydrological monitoring data according to the timestamp t to form an original hydrological data set:

[0016] D raw ={X i (t)|X i∈{P,V,H,E,T},t∈T 1};

[0017] Among them, P is rainfall, V is flow velocity, H is water level, E is evaporation, T is temperature, and X is the average temperature. i (t) represents the value of the hydrological variable at time t, T 1 is the time collection;

[0018] S12. Perform data cleaning on the original hydrological data set, including detecting and removing outliers, filling missing data, and removing sensor noise to form a preliminary cleaned hydrological monitoring data set;

[0019] S15. Perform data standardization on the hydrological monitoring data set after preliminary cleaning so that each hydrological variable has the same numerical scale;

[0020] S16. Reconstruct the standardized hydrological monitoring dataset according to the timestamp so that all hydrological variables are aligned within the same time step to form a hydrological monitoring dataset D with a unified spatiotemporal resolution. final :

[0021]

[0022] Among them, T 1 ′ is the reconstructed time set.

[0023] Optionally, step S2 includes the following steps:

[0024] S21. For each hydrological variable in the hydrological monitoring data set At each sampling time t j Calculate its membership value μ according to the fuzzy membership function ij :

[0025]

[0026] in, Represents the hydrological variable X i At sampling time t j The standardized value of is the hydrological variable X i In the dataset D final The mean value in N i is the hydrological variable X i The total number of samples, α i is an adjustment parameter used to reflect the hydrological variable X i The degree of sensitivity to its deviation from the mean strengthens or suppresses the influence of abnormal fluctuations in the calculation of fuzzy membership;

[0027] S22. Based on the membership value μij For each hydrological variable The fuzzy entropy value E is calculated by combining the membership of each sampling point and its deviation from the mean. f,i :

[0028]

[0029] Where ln(·) represents the natural logarithm, β i is the weighting factor used to adjust the hydrological variable X i The influence of extreme deviation value on fuzzy entropy contribution, σ i is the hydrological variable X i The standard deviation reflects the degree of data dispersion;

[0030] S23. The fuzzy entropy value E of each hydrological variable f,i Arranged in the order of the index of hydrological variables, the fuzzy entropy matrix F of hydrological monitoring data is constructed:

[0031]

[0032] Among them, E f,1 、E f,2 、E f,3 、E f,4 and E f,5 They correspond to the fuzzy entropy values ​​of rainfall P, flow velocity V, water level H, evaporation E and temperature T respectively.

[0033] Optionally, step S3 includes the following steps:

[0034] S31. Based on the hydrological monitoring data set and its fuzzy entropy matrix, for each pair of hydrological variables and i≠j, under different time lags L, unweighted vector autoregression models and weighted vector autoregression models are constructed and Granger causality tests are performed:

[0035]

[0036] Among them, a i,k is the hydrological variable X i The autoregressive coefficient at lag k, b ij,k is the hydrological variable X j At lag k, i The influence coefficient of i (t) and ε ij (t)) are the error terms of the model, ω ij is the weighting factor based on the fuzzy entropy matrix;

[0037] S32. Set the null hypothesis H 0 0 is the hydrological variable right There is no Granger causality, that is, only relying on the unweighted vector autoregression model, setting the alternative hypothesis H 1 Hydrological variables right It has Granger causality, that is, the weighted vector autoregression model is used, and the F statistic test is performed at the preset significance level α. If the null hypothesis H is rejected, 0 :b ij,k =0, then the hydrological variable is determined Granger causality affects hydrological variables And record the direct causal relationship and the corresponding weighting factor ω ij ;

[0038] S33. Based on the test results of step S32, a causal model of hydrological monitoring data represented by a directed acyclic graph G = (V, E) is constructed, wherein the vertex set corresponds to rainfall P, flow velocity V, water level H, evaporation E and temperature T, and the edge set E includes the causal model from the hydrological variable X to the edge set E. j To hydrological variable X i The directed edge of Granger causality affects And the weight of the edge is determined by the corresponding weighting factor ω ij Determine and generate the final causal model of hydrological monitoring data.

[0039] Optionally, step S4 includes the following steps:

[0040] S41. Continuously monitor the real-time hydrological monitoring data set during the monitoring period and calculate the j Calculating hydrological variables Deviation measure d from the historical distribution of hydrological variables i (t j );

[0041] S42. Combine the hydrological monitoring data causal model G for each hydrological variable X i The set S of causal influencing variables within the lag time window L i Perform causal weighted deviation calculation:

[0042]

[0043] Among them, S i is the influence X in the causal model i The variable set, ω ij is the causal weighting factor, τ ij is the hydrological variable X j X i The lag time step that has an impact, as a measure of the deviation of hydrological variables after integrating causal influences;

[0044] S43. Combine fuzzy entropy optimization strategy to optimize each hydrological variable X i Calculating dynamic anomaly detection thresholds

[0045]

[0046] in, is the hydrological variable X i The average deviation measure in historical data, E f,i is the fuzzy entropy value of hydrological variables, γ i is the benchmark anomaly detection coefficient, ξ i is the fuzzy entropy influencing factor;

[0047] S44. At every moment t j , if satisfied Then determine the hydrological variable X i At time t j When an abnormality occurs, the abnormal hydrological variable, the time of occurrence and the causal impact path are recorded.

[0048] Optionally, step S5 includes the following steps:

[0049] S51. After detecting abnormal hydrological monitoring data, a causal tracing search space of abnormal hydrological monitoring data is constructed based on the causal model of hydrological monitoring data, a set of abnormal variables is defined, and the tracing variable set is expanded according to the level of reverse causal relationship in the causal model:

[0050] A m ={X k ∣(X k ,X l )∈E,X l ∈A m-1},m=1,2,...,M;

[0051] Among them, A 0 =A represents the abnormal initial variable set, A m It represents the set of abnormal influencing variables traced back to the mth layer, where M is the maximum tracing depth;

[0052] S52. Calculate the causal influence strength C of each variable in the anomaly tracing path i,j , and weighted sorting of abnormal causal relationships:

[0053]

[0054] in, is the dynamic anomaly detection threshold, τ ij is the hydrological variable Xj X i The lag time step that has an impact, C i,j Reflects the variable X j For abnormal hydrological variables X i the extent of the impact;

[0055] S53. Sort the abnormality tracing paths according to the causal influence strength, and generate the abnormality tracing candidate path set P anom}:

[0056] P anom ={(X k ,X l )∣C k,l >θ,X k ∈A m ,X l ∈A m-1};

[0057] Among them, θ is the threshold of causal traceability strength, and only the traceability paths with causal influence strength greater than θ are retained;

[0058] S54. Combine fuzzy entropy optimization strategy to identify abnormal source candidate path set P anom Optimize and calculate the comprehensive entropy weight score S of each traceability path p ;

[0059] S55. Based on the comprehensive entropy weight score, select the highest score as the final anomaly tracing path and construct the causal chain of abnormal hydrological monitoring data C anom :

[0060]

[0061] Abnormal causal chain C anom Record the layer-by-layer impact path of abnormal variables to form the final causal analysis results of abnormal hydrological monitoring data.

[0062] A hydrological data governance system based on machine learning, used to execute a hydrological data governance method based on machine learning, including the following modules:

[0063] A data acquisition module is used to acquire hydrological monitoring data in real time through a hydrological monitoring sensor network. The hydrological monitoring data includes rainfall, flow rate, water level, evaporation and temperature, and to store the acquired data according to timestamps to form a hydrological monitoring data set;

[0064] The data preprocessing module is used to clean, denoise, fill missing values, remove outliers and standardize the hydrological monitoring data set to form a hydrological monitoring data set with uniform temporal and spatial resolution;

[0065] The fuzzy entropy calculation module is used to calculate the fuzzy entropy value of each hydrological variable in the hydrological monitoring data set and construct the fuzzy entropy matrix of the hydrological monitoring data to measure the uncertainty of the hydrological variables;

[0066] The causal relationship modeling module is used to construct the causal relationship model between the hydrological variables in the hydrological monitoring data based on the hydrological monitoring data set and its fuzzy entropy matrix using the Granger causal analysis method, and to perform causal tests using the unweighted vector autoregression model and the weighted vector autoregression model to determine the direct and indirect causal relationships between the hydrological variables, and to generate a causal model of the hydrological monitoring data represented in the form of a directed acyclic graph;

[0067] The real-time anomaly detection module is used to continuously monitor the hydrological monitoring data set, and perform anomaly detection on the real-time hydrological monitoring data based on the hydrological monitoring data causal model combined with the fuzzy entropy optimization strategy, calculate the causal weighted deviation, and identify abnormal hydrological monitoring data based on the dynamically adjusted anomaly detection threshold, and record abnormal variables and occurrence time;

[0068] The anomaly tracing analysis module is used to perform reverse causal analysis based on the causal model of hydrological monitoring data after abnormal hydrological monitoring data is detected. It generates candidate anomaly tracing paths based on the Granger causal analysis results and fuzzy entropy evaluation results, and calculates the comprehensive entropy weight score in combination with the fuzzy entropy optimization strategy to optimize the anomaly tracing path. Finally, it constructs the causal chain of abnormal hydrological monitoring data and forms a complete anomaly analysis result.

[0069] The beneficial effects of the present invention are:

[0070] The present invention constructs a causal model of hydrological monitoring data based on Granger causality analysis, and adjusts the weight of causal relationship in combination with fuzzy entropy optimization. It can more accurately identify the true causal relationship between hydrological variables and avoid misjudgment caused by pseudo-correlation between dependent variables. By introducing causal inference technology, the system can automatically trace back the most critical abnormal triggering factors when an abnormality occurs, and quantify their impact, thereby improving the reliability of abnormal tracing.

[0071] In the process of anomaly detection in real-time hydrological monitoring data, the present invention combines causal weighted deviation calculation and fuzzy entropy optimization strategy, and dynamically adjusts the anomaly detection threshold to make anomaly identification more accurate and adaptive. The present invention calculates the causal weighted deviation of hydrological variables and combines it with fuzzy entropy to dynamically optimize the threshold, so that the system can adaptively adjust the anomaly judgment criteria according to the real-time data characteristics, ensuring stable operation under different hydrological environments.

[0072] The present invention constructs a set of abnormality tracing candidate paths in the abnormality tracing process, and optimizes and screens the abnormality tracing paths by calculating the comprehensive entropy weight score through the fuzzy entropy optimization strategy, so that the final abnormality causal chain can more accurately reflect the real abnormality conduction process. By calculating the causal influence intensity and combining fuzzy entropy optimization to remove low-correlation variables, the conciseness and effectiveness of the tracing path are ensured, thereby improving the efficiency and explainability of abnormality tracing. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0074] Figure 1 This is a flow chart of a hydrological data management method and system based on machine learning proposed in the present invention. DETAILED DESCRIPTION

[0075] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0076] refer to Figure 1 , a hydrological data governance method based on machine learning, comprising the following steps:

[0077] S1. Obtain and preprocess hydrological monitoring data to generate a hydrological monitoring data set with uniform temporal and spatial resolution;

[0078] S2. Calculate the fuzzy entropy value of each hydrological variable for the hydrological monitoring data set, form a fuzzy entropy evaluation result for measuring the uncertainty of each hydrological variable, and construct a fuzzy entropy matrix of the hydrological monitoring data;

[0079] S3. Based on the hydrological monitoring data set and its fuzzy entropy matrix, the Granger causality analysis method is used to construct a causal relationship model between various hydrological variables in the hydrological monitoring data. The direct and indirect causal relationships between various hydrological variables are determined by causal tests under different time lag conditions, and a causal model of hydrological monitoring data represented by a directed acyclic graph is generated;

[0080] S4. Continuously monitor the hydrological monitoring data set, and use the hydrological monitoring data causal model combined with the fuzzy entropy optimization strategy to detect anomalies in the real-time hydrological monitoring data set, and dynamically adjust the anomaly detection threshold to identify abnormal hydrological monitoring data in real time;

[0081] S5. When abnormal hydrological monitoring data is detected, reverse causal analysis is performed on the abnormal hydrological monitoring data according to the hydrological monitoring data causal model, and candidate abnormal source tracing paths are generated based on the Granger causal analysis results and fuzzy entropy evaluation results to form a causal chain of abnormal hydrological monitoring data;

[0082] S6. Perform fuzzy entropy optimization on the candidate paths for anomaly tracing, screen out the key anomaly causal paths based on the fuzzy entropy values ​​of the hydrological variables in each causal chain and their influence weights in the causal chain, and then locate the root causes of the abnormal hydrological monitoring data.

[0083] In this implementation, step S1 includes:

[0084] S11. Obtain hydrological monitoring data in real time through the hydrological monitoring sensor network. The hydrological monitoring data includes rainfall, flow rate, water level, evaporation and temperature. The hydrological monitoring data is initially stored according to the timestamp t to form the original hydrological data set:

[0085] D raw ={X i (t)|X i ∈{P,V,H,E,T},t∈T 1};

[0086] Among them, P is rainfall, V is flow velocity, H is water level, E is evaporation, T is temperature, and X is the average temperature. i (t) represents the value of the hydrological variable at time t, T 1 is the time collection;

[0087] S12. Perform data cleaning on the original hydrological data set, including detecting and removing outliers, filling missing data, and removing sensor noise to form a preliminary cleaned hydrological monitoring data set;

[0088] S15. Perform data standardization on the hydrological monitoring data set after preliminary cleaning so that each hydrological variable has the same numerical scale;

[0089] S16. Reconstruct the standardized hydrological monitoring dataset according to the timestamp so that all hydrological variables are aligned within the same time step to form a hydrological monitoring dataset D with a unified spatiotemporal resolution. final :

[0090]

[0091] Among them, T 1 ′ is the reconstructed time set.

[0092] In this implementation, step S2 includes the following steps:

[0093] S21. For each hydrological variable in the hydrological monitoring data set At each sampling time t j Calculate its membership value μ according to the fuzzy membership function ij :

[0094]

[0095] in, Represents the hydrological variable X i At sampling time t j The standardized value of is the hydrological variable X i In the dataset D final The mean value in N i is the hydrological variable X i The total number of samples, α i is an adjustment parameter used to reflect the hydrological variable X i The degree of sensitivity to its deviation from the mean strengthens or suppresses the influence of abnormal fluctuations in the calculation of fuzzy membership;

[0096] S22. Based on the membership value μ ij For each hydrological variable The fuzzy entropy value E is calculated by combining the membership of each sampling point and its deviation from the mean. f,i :

[0097]

[0098] Where ln(·) represents the natural logarithm, β i is the weighting factor used to adjust the hydrological variable X i The influence of extreme deviation value on fuzzy entropy contribution, σ i is the hydrological variable X i The standard deviation reflects the degree of data dispersion;

[0099] S23. The fuzzy entropy value E of each hydrological variable f,i Arranged in the order of the index of hydrological variables, the fuzzy entropy matrix F of hydrological monitoring data is constructed:

[0100]

[0101] Among them, E f,1 、E f,2 、E f,3 、E f,4 and E f,5 They correspond to the fuzzy entropy values ​​of rainfall P, flow velocity V, water level H, evaporation E and temperature T respectively.

[0102] In this implementation, step S3 includes the following steps:

[0103] S31. Based on the hydrological monitoring data set and its fuzzy entropy matrix, for each pair of hydrological variables and i≠j, under different time lags L, unweighted vector autoregression models and weighted vector autoregression models are constructed and Granger causality tests are performed:

[0104]

[0105]

[0106] Among them, a i,k is the hydrological variable X i The autoregressive coefficient at lag k, b ij,k is the hydrological variable X j At lag k, i The influence coefficient of i (t) and ε ij (t)) are the error terms of the model, ω ij is the weighting factor based on the fuzzy entropy matrix;

[0107] S32. Set the null hypothesis H 0 0 is the hydrological variable right There is no Granger causality, that is, only relying on the unweighted vector autoregression model, setting the alternative hypothesis H 1 Hydrological variables right It has Granger causality, that is, the weighted vector autoregression model is used, and the F statistic test is performed at the preset significance level α. If the null hypothesis H is rejected, 0 :b ij,k =0, then the hydrological variable is determined Granger causality affects hydrological variables And record the direct causal relationship and the corresponding weighting factor ω ij ;

[0108] S33. Based on the test results of step S32, a causal model of hydrological monitoring data represented by a directed acyclic graph G = (V, E) is constructed, wherein the vertex set corresponds to rainfall P, flow velocity V, water level H, evaporation E and temperature T, and the edge set E includes the causal model from the hydrological variable X to the edge set E. j To hydrological variable X i The directed edge of Granger causality affects And the weight of the edge is determined by the corresponding weighting factor ω ij Determine and generate the final causal model of hydrological monitoring data.

[0109] In this implementation, step S4 includes the following steps:

[0110] S41. Continuously monitor the real-time hydrological monitoring data set during the monitoring period and calculate the j Calculating hydrological variables Deviation measure d from the historical distribution of hydrological variables i (t j );

[0111] S42. Combine the hydrological monitoring data causal model G for each hydrological variable X i The set S of causal influencing variables within the lag time window L i Perform causal weighted deviation calculation:

[0112]

[0113] Among them, S i is the influence X in the causal model i The variable set, ω ij is the causal weighting factor, τ ij is the hydrological variable X j X i The lag time step that has an impact, as a measure of the deviation of hydrological variables after integrating causal influences;

[0114] S43. Combine fuzzy entropy optimization strategy to optimize each hydrological variable X i Calculating dynamic anomaly detection thresholds

[0115]

[0116] in, is the hydrological variable X i The average deviation measure in historical data, E f,i is the fuzzy entropy value of hydrological variables, γ i is the benchmark anomaly detection coefficient, ξ i is the fuzzy entropy influencing factor;

[0117] S44. At every moment t j , if satisfied Then determine the hydrological variable X i At time t j When an abnormality occurs, the abnormal hydrological variable, the time of occurrence and the causal impact path are recorded.

[0118] In this implementation, step S5 includes the following steps:

[0119] S51. After detecting abnormal hydrological monitoring data, a causal tracing search space of abnormal hydrological monitoring data is constructed based on the causal model of hydrological monitoring data, a set of abnormal variables is defined, and the tracing variable set is expanded according to the level of reverse causal relationship in the causal model:

[0120] A m ={X k ∣(X k ,X l )∈E,X l ∈A m-1},m=1,2,...,M;

[0121] Among them, A 0 =A represents the abnormal initial variable set, A m It represents the set of abnormal influencing variables traced back to the mth layer, where M is the maximum tracing depth;

[0122] S52. Calculate the causal influence strength C of each variable in the anomaly tracing path i,j , and weighted sorting of abnormal causal relationships:

[0123]

[0124] in, is the dynamic anomaly detection threshold, τ ij is the hydrological variable X j X i The lag time step that has an impact, C i,j Reflects the variable X j For abnormal hydrological variables X i the extent of the impact;

[0125] S53. Sort the abnormality tracing paths according to the causal influence strength, and generate the abnormality tracing candidate path set P anom}:

[0126] P anom ={(X k ,X l )∣C k,l >θ,X k ∈A m ,X l ∈A m-1};

[0127] Among them, θ is the threshold of causal traceability strength, and only the traceability paths with causal influence strength greater than θ are retained;

[0128] S54. Combine fuzzy entropy optimization strategy to identify abnormal source candidate path set P anom Optimize and calculate the comprehensive entropy weight score S of each traceability path p ;

[0129] S55. Based on the comprehensive entropy weight score, select the highest score as the final anomaly tracing path and construct the causal chain of abnormal hydrological monitoring data C anom :

[0130]

[0131] Abnormal causal chain C anom Record the layer-by-layer impact path of abnormal variables to form the final causal analysis results of abnormal hydrological monitoring data.

[0132] A hydrological data governance system based on machine learning, used to execute a hydrological data governance method based on machine learning, including the following modules:

[0133] The data acquisition module is used to obtain hydrological monitoring data in real time through the hydrological monitoring sensor network. The hydrological monitoring data includes rainfall, flow rate, water level, evaporation and temperature, and the acquired data is stored according to the timestamp to form a hydrological monitoring data set;

[0134] The data preprocessing module is used to clean, denoise, fill missing values, remove outliers and standardize the hydrological monitoring data set to form a hydrological monitoring data set with uniform temporal and spatial resolution;

[0135] The fuzzy entropy calculation module is used to calculate the fuzzy entropy value of each hydrological variable in the hydrological monitoring data set and construct the fuzzy entropy matrix of the hydrological monitoring data to measure the uncertainty of the hydrological variables;

[0136] The causal relationship modeling module is used to construct the causal relationship model between the hydrological variables in the hydrological monitoring data based on the hydrological monitoring data set and its fuzzy entropy matrix using the Granger causal analysis method, and to perform causal tests using the unweighted vector autoregression model and the weighted vector autoregression model to determine the direct and indirect causal relationships between the hydrological variables, and to generate a causal model of the hydrological monitoring data represented in the form of a directed acyclic graph;

[0137] The real-time anomaly detection module is used to continuously monitor the hydrological monitoring data set, and perform anomaly detection on the real-time hydrological monitoring data based on the hydrological monitoring data causal model combined with the fuzzy entropy optimization strategy, calculate the causal weighted deviation, and identify abnormal hydrological monitoring data based on the dynamically adjusted anomaly detection threshold, and record abnormal variables and occurrence time;

[0138] The anomaly tracing analysis module is used to perform reverse causal analysis based on the causal model of hydrological monitoring data after abnormal hydrological monitoring data is detected. It generates candidate anomaly tracing paths based on the Granger causal analysis results and fuzzy entropy evaluation results, and calculates the comprehensive entropy weight score in combination with the fuzzy entropy optimization strategy to optimize the anomaly tracing path. Finally, it constructs the causal chain of abnormal hydrological monitoring data and forms a complete anomaly analysis result.

[0139] Embodiment 1:

[0140] On July 15, 2024, the S12 hydrological monitoring station in a tributary basin in the middle reaches of the A River detected a sudden increase in water levels at 15:35:00. The water level jumped from 3.21 meters to 3.96 meters in a short period of time, and rose by 0.75 meters in 10 minutes. According to the historical hydrological data statistics of the basin, the water level in this area under similar rainfall conditions has never risen by more than 0.45 meters / 10 minutes in the past five years. At the same time, the flow velocity of the S09 hydrological monitoring station upstream of the station increased significantly in a short period of time, jumping from 1.2m / s to 2.8m / s, and then recovered to 1.3m / s within 20 minutes. Such drastic fluctuations are usually related to sudden flood discharge, extreme rainfall or abnormal monitoring equipment. Therefore, it is necessary to automatically detect abnormal data and further trace the source analysis to determine its root cause.

[0141] The system first continuously monitors the hydrological data of hydrological monitoring station S12 and its surrounding stations through the data acquisition module. The data sampling interval is 5 minutes. At 15:35:00, the system finds that the water level data of station S12 deviates by 3.4 times the historical standard deviation during the same period, exceeding the dynamically set anomaly detection threshold, triggering the real-time evaluation mechanism of the anomaly detection module.

[0142] Real-time calculation results:

[0143] Site number: S12;

[0144] Monitoring time: 15:35:00 on July 15, 2024;

[0145] Recorded water level: 3.96 meters (water level in the first 5 minutes: 3.21 meters);

[0146] The largest increase in the same period in history: 0.45 meters / 10 minutes;

[0147] Current increase: 0.75 meters / 10 minutes (exceeding the historical maximum increase by 66.7%);

[0148] Computational deviation metric: 3.4 (exceeds the dynamic anomaly threshold of 3.1);

[0149] The system combined causal weighted deviation calculation to confirm the water level abnormality and automatically generated an abnormal data report at 15:36:00, including abnormal variables, water level surge time, abnormal amplitude and influencing factors. The report was immediately uploaded to the hydrological management system and triggered the abnormal tracing module.

[0150] The system starts the anomaly tracing analysis module, and first traces the possible sources of the anomaly based on the causal model of hydrological monitoring data. The causal network analysis shows that the water level at station S12 and the flow velocity at the upstream station S09 have a causal weight of 0.92 and a lag time of about 15 minutes.

[0151] Tracing path construction:

[0152] At 15:20:00, the flow velocity at the S09 station suddenly increased from 1.2m / s to 2.8m / s.

[0153] At 15:35:00, the water level at the S12 station suddenly rose from 3.21 meters to 3.96 meters.

[0154] Since the flow velocity change at the S09 site has a delayed impact on the water level at the S12 site, the system calculates its causal weighted deviation and traces the source to determine that the S09 site may be the key influencing factor of the abnormal water level event.

[0155] Traceability data analysis:

[0156] The abnormal flow rate at S09 site will occur at 15:20:00 on July 15, 2024;

[0157] The abnormal water level at S12 site will occur at 15:35:00 on July 15, 2024;

[0158] Calculated lag time: 15 minutes (consistent with historical causality);

[0159] Causal influence strength: 0.92 (high causal weight);

[0160] In order to further determine the cause of the abnormal flow velocity, the rainfall data of the S09 station was systematically analyzed, and it was found that the rainfall at the upstream S03 station had an abnormal peak at 14:50:00. The rainfall increased from 1.5 mm to 48.7 mm within 10 minutes. The short-term precipitation intensity reached above the historical 95% quantile, which was an extreme rainfall event.

[0161] The system ultimately determined that the extreme rainfall event at station S03 was the root cause of the abnormal water level rise and constructed a complete abnormal source tracing path:

[0162] C anom ={S03 (peak rainfall at 14:50:00) → S09 (sudden increase in flow rate at 15:20:00) →

[0163] S12 (water level suddenly increased at 15:35:00)};

[0164] The system generated a complete abnormality tracing report at 15:37:00 and pushed it to the Hydrological Management Center. As the rainfall anomaly at the S03 site may affect multiple downstream basins, the Management Center dispatched the downstream reservoir flood discharge strategy in advance based on the results of this analysis, thus avoiding the possible risk of exceeding the warning water level.

[0165] This example proves that the method of the present invention can achieve efficient and accurate anomaly detection and source tracing analysis in actual hydrological monitoring scenarios, with higher detection accuracy, shorter source tracing path, and faster anomaly location capability. In the basin hydrological anomaly warning, this method can detect anomalies 15 minutes in advance and trace the cause of the anomaly within 10 minutes, providing strong data support for water resources management and disaster prevention warning.

[0166] The present invention constructs a causal model of hydrological monitoring data based on Granger causality analysis, and adjusts the weight of causal relationship in combination with fuzzy entropy optimization. It can more accurately identify the true causal relationship between hydrological variables and avoid misjudgment caused by pseudo-correlation between dependent variables. By introducing causal inference technology, the system can automatically trace back the most critical abnormal triggering factors when an abnormality occurs, and quantify their impact, thereby improving the reliability of abnormal tracing.

[0167] In the process of anomaly detection in real-time hydrological monitoring data, the present invention combines causal weighted deviation calculation and fuzzy entropy optimization strategy, and dynamically adjusts the anomaly detection threshold to make anomaly identification more accurate and adaptive. The present invention calculates the causal weighted deviation of hydrological variables and combines it with fuzzy entropy to dynamically optimize the threshold, so that the system can adaptively adjust the anomaly judgment criteria according to the real-time data characteristics, ensuring stable operation under different hydrological environments.

[0168] The present invention constructs a set of abnormality tracing candidate paths in the abnormality tracing process, and optimizes and screens the abnormality tracing paths by calculating the comprehensive entropy weight score through the fuzzy entropy optimization strategy, so that the final abnormality causal chain can more accurately reflect the real abnormality conduction process. By calculating the causal influence intensity and combining fuzzy entropy optimization to remove low-correlation variables, the conciseness and effectiveness of the tracing path are ensured, thereby improving the efficiency and explainability of abnormality tracing.

[0169] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A hydrological data management method based on machine learning, characterized in that: The steps include: S1. Obtain and preprocess hydrological monitoring data to generate a hydrological monitoring data set with uniform temporal and spatial resolution; S2. Calculate the fuzzy entropy value of each hydrological variable for the hydrological monitoring data set, form a fuzzy entropy evaluation result for measuring the uncertainty of each hydrological variable, and construct a fuzzy entropy matrix of the hydrological monitoring data; S3. Based on the hydrological monitoring data set and its fuzzy entropy matrix, the Granger causality analysis method is used to construct a causal relationship model between various hydrological variables in the hydrological monitoring data. The direct and indirect causal relationships between various hydrological variables are determined by causal tests under different time lag conditions, and a causal model of hydrological monitoring data represented by a directed acyclic graph is generated; S4. Continuously monitor the hydrological monitoring data set, and use the hydrological monitoring data causal model combined with the fuzzy entropy optimization strategy to detect anomalies in the real-time hydrological monitoring data set, and dynamically adjust the anomaly detection threshold to identify abnormal hydrological monitoring data in real time; S5. When abnormal hydrological monitoring data is detected, reverse causal analysis is performed on the abnormal hydrological monitoring data according to the hydrological monitoring data causal model, and candidate abnormal source tracing paths are generated based on the Granger causal analysis results and fuzzy entropy evaluation results to form a causal chain of abnormal hydrological monitoring data; S6. Perform fuzzy entropy optimization on the candidate paths for anomaly tracing, screen out the key anomaly causal paths based on the fuzzy entropy values ​​of the hydrological variables in each causal chain and their influence weights in the causal chain, and then locate the root causes of the abnormal hydrological monitoring data.

2. According to the method of hydrological data management based on machine learning in claim 1, it is characterized in that: The step S1 comprises: S11. Acquire hydrological monitoring data in real time through the hydrological monitoring sensor network, the hydrological monitoring data including rainfall, flow rate, water level, evaporation and temperature, and preliminarily store the hydrological monitoring data according to the timestamp t to form an original hydrological data set: D raw ={X i (t)∣X i ∈{P,V,H,E,T},t∈T1}; Among them, P is rainfall, V is flow velocity, H is water level, E is evaporation, T is temperature, and X is the average temperature. i (t) represents the value of the hydrological variable at time t, and T1 is the time set; S12. Perform data cleaning on the original hydrological data set, including detecting and removing outliers, filling missing data, and removing sensor noise to form a preliminary cleaned hydrological monitoring data set; S15. Perform data standardization on the hydrological monitoring data set after preliminary cleaning so that each hydrological variable has the same numerical scale; S16. Reconstruct the standardized hydrological monitoring dataset according to the timestamp so that all hydrological variables are aligned within the same time step to form a hydrological monitoring dataset D with a unified spatiotemporal resolution. final : Among them, T1 ′ is the reconstructed time set.

3. According to the method of hydrological data management based on machine learning in claim 1, it is characterized in that: The step S2 comprises the following steps: S21. For each hydrological variable in the hydrological monitoring data set At each sampling time t j Calculate its membership value μ according to the fuzzy membership function ij : in, Represents the hydrological variable X i At sampling time t j The standardized value of is the hydrological variable X i In the dataset D final The mean value in N i is the hydrological variable X i The total number of samples, α i is an adjustment parameter used to reflect the hydrological variable X i The degree of sensitivity to its deviation from the mean strengthens or suppresses the influence of abnormal fluctuations in the calculation of fuzzy membership; S22. Based on the membership value μ ij For each hydrological variable The fuzzy entropy value E is calculated by combining the membership of each sampling point and its deviation from the mean. f,i : Where ln(·) represents the natural logarithm, β i is the weighting factor used to adjust the hydrological variable X i The influence of extreme deviation value on fuzzy entropy contribution, σ i is the hydrological variable X i The standard deviation reflects the degree of data dispersion; S23. The fuzzy entropy value E of each hydrological variable f,i Arranged in the order of the index of hydrological variables, the fuzzy entropy matrix F of hydrological monitoring data is constructed: Among them, E f,1 、E f,2 、E f,3 、E f,4 and E f,5 They correspond to the fuzzy entropy values ​​of rainfall P, flow velocity V, water level H, evaporation E and temperature T respectively.

4. The hydrological data management method based on machine learning according to claim 1 is characterized in that: The step S3 comprises the following steps: S31. Based on the hydrological monitoring data set and its fuzzy entropy matrix, for each pair of hydrological variables and i≠j, under different time lags L, unweighted vector autoregression models and weighted vector autoregression models are constructed and Granger causality tests are performed: Among them, a i,k is the hydrological variable X i The autoregressive coefficient at lag k, b ij,k is the hydrological variable X j At lag k, i The influence coefficient of i (t) and ε ij (t)) are the error terms of the model, ω ij is the weighting factor based on the fuzzy entropy matrix; S32. Set the null hypothesis H00 as the hydrological variable right There is no Granger causality, that is, it only relies on the unweighted vector autoregression model, and the alternative hypothesis H1 is set as the hydrological variable right It has Granger causality, that is, the weighted vector autoregression model is used, and the F statistic test is performed at the preset significance level α. If the null hypothesis H0:b is rejected ij,k =0, then the hydrological variable is determined Granger causality affects hydrological variables And record the direct causal relationship and the corresponding weighting factor ω ij ; S33. Based on the test results of step S32, a causal model of hydrological monitoring data represented by a directed acyclic graph G = (V, E) is constructed, wherein the vertex set corresponds to rainfall P, flow velocity V, water level H, evaporation E and temperature T, and the edge set E includes the causal model from the hydrological variable X to the edge set E. j To hydrological variable X i The directed edge of Granger causality affects And the weight of the edge is determined by the corresponding weighting factor ω ij Determine and generate the final causal model of hydrological monitoring data.

5. The hydrological data management method based on machine learning according to claim 4 is characterized in that: The step S4 comprises the following steps: S41. Continuously monitor the real-time hydrological monitoring data set during the monitoring period and calculate the j Calculating hydrological variables Deviation measure d from the historical distribution of hydrological variables i (t j ); S42. Combine the hydrological monitoring data causal model G for each hydrological variable X i The set S of causal influencing variables within the lag time window L i Perform causal weighted deviation calculation: Among them, S i is the influence X in the causal model i The variable set, ω ij is the causal weighting factor, τ ij is the hydrological variable X j X i The lag time step that has an impact, as a measure of the deviation of hydrological variables after integrating causal influences; S43. Combine fuzzy entropy optimization strategy to optimize each hydrological variable X i Calculating dynamic anomaly detection thresholds in, is the hydrological variable X i The average deviation measure in historical data, E f,i is the fuzzy entropy value of hydrological variables, γ i is the benchmark anomaly detection coefficient, ξ i is the fuzzy entropy influencing factor; S44. At every moment t j , if satisfied Then determine the hydrological variable X i At time t j When an abnormality occurs, the abnormal hydrological variable, the time of occurrence and the causal impact path are recorded.

6. The hydrological data management method based on machine learning according to claim 5 is characterized in that: The step S5 comprises the following steps: S51. After detecting abnormal hydrological monitoring data, a causal tracing search space of abnormal hydrological monitoring data is constructed based on the causal model of hydrological monitoring data, a set of abnormal variables is defined, and the tracing variable set is expanded according to the level of reverse causal relationship in the causal model: A m ={X k ∣(X k ,X l )∈E,X l ∈A m-1 },m=1,2,...,M; Among them, A0=A represents the abnormal initial variable set, A m It represents the set of abnormal influencing variables traced back to the mth layer, where M is the maximum tracing depth; S52. Calculate the causal influence strength C of each variable in the anomaly tracing path i,j , and weighted sorting of abnormal causal relationships: in, is the dynamic anomaly detection threshold, τ ij is the hydrological variable X j X i The lag time step that has an impact, C i,j Reflects the variable X j For abnormal hydrological variables X i the extent of the impact; S53. Sort the abnormality tracing paths according to the causal influence strength, and generate the abnormality tracing candidate path set P anom }: P anom ={(X k ,X l )∣C k,l >θ,X k ∈A m ,X l ∈A m-1 }; Among them, θ is the threshold of causal traceability strength, and only the traceability paths with causal influence strength greater than θ are retained; S54. Combine fuzzy entropy optimization strategy to identify abnormal source candidate path set P anom Optimize and calculate the comprehensive entropy weight score S of each traceability path p ; S55. Based on the comprehensive entropy weight score, select the highest score as the final anomaly tracing path and construct the causal chain of abnormal hydrological monitoring data C anom : Abnormal causal chain C anom Record the layer-by-layer impact path of abnormal variables to form the final causal analysis results of abnormal hydrological monitoring data.

7. A hydrological data management system based on machine learning, used to execute the hydrological data management method based on machine learning according to claims 1-6, characterized in that: Includes the following modules: A data acquisition module is used to acquire hydrological monitoring data in real time through a hydrological monitoring sensor network. The hydrological monitoring data includes rainfall, flow rate, water level, evaporation and temperature, and to store the acquired data according to timestamps to form a hydrological monitoring data set; The data preprocessing module is used to clean, denoise, fill missing values, remove outliers and standardize the hydrological monitoring data set to form a hydrological monitoring data set with uniform temporal and spatial resolution; The fuzzy entropy calculation module is used to calculate the fuzzy entropy value of each hydrological variable in the hydrological monitoring data set and construct the fuzzy entropy matrix of the hydrological monitoring data to measure the uncertainty of the hydrological variables; The causal relationship modeling module is used to construct the causal relationship model between the hydrological variables in the hydrological monitoring data based on the hydrological monitoring data set and its fuzzy entropy matrix using the Granger causal analysis method, and to perform causal tests using the unweighted vector autoregression model and the weighted vector autoregression model to determine the direct and indirect causal relationships between the hydrological variables, and to generate a causal model of the hydrological monitoring data represented in the form of a directed acyclic graph; The real-time anomaly detection module is used to continuously monitor the hydrological monitoring data set, and perform anomaly detection on the real-time hydrological monitoring data based on the hydrological monitoring data causal model combined with the fuzzy entropy optimization strategy, calculate the causal weighted deviation, and identify abnormal hydrological monitoring data based on the dynamically adjusted anomaly detection threshold, and record abnormal variables and occurrence time; The anomaly tracing analysis module is used to perform reverse causal analysis based on the causal model of hydrological monitoring data after abnormal hydrological monitoring data is detected. It generates candidate anomaly tracing paths based on the Granger causal analysis results and fuzzy entropy evaluation results, and calculates the comprehensive entropy weight score in combination with the fuzzy entropy optimization strategy to optimize the anomaly tracing path. Finally, it constructs the causal chain of abnormal hydrological monitoring data and forms a complete anomaly analysis result.

Citation Information

Patent Citations

  • Sewage treatment fault diagnosis method and system based on data analysis

    CN118230069A

  • Unmanned aerial vehicle-based river hydrological sampling inspection method and system

    CN119151387A

  • Urban flood control toughness critical state identification method and system considering multiple pressure coupling

    CN119226782A

  • Time-considered learning system of fuzzy cognitive maps for decision making support

    KR1020120108459A

Cited By

  • Document compliance audit data management system and method based on artificial intelligence

    CN120744798A

  • Storage data analysis system and method for intelligent monitoring of agricultural product warehouse

    CN121212974A

  • Warehouse data analysis system and methods for intelligent monitoring of agricultural product warehouses

    CN121212974B

  • Underground equipment fault real-time diagnosis method and system based on edge calculation

    CN121580337A

  • Edge computing-based real-time diagnosis method and system for downhole equipment failure

    CN121580337B