Hydrological sequence data anomaly detection method and system based on cloud computing

Through the cloud-based hydrological sequence data anomaly detection method, the convolutional neural network calculates the abnormal score and generates dynamic thresholds, the problems of insufficient real-time and lack of dynamic thresholds in the previous technology are solved, and the accuracy and real-timeness of detection are improved.

CN119961842AInactive Publication Date: 2025-05-09SU QIAN SHI NIAN KE JI YOU XIAN GONG SI

Patent Information

Application Number
CN202510211161.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has strong data distribution hypothesis dependence in the detection of hydrological data anomalies, which is difficult to deal with complex and changeable hydrological data. It requires a large amount of labeled data for model optimization, and the application cost is high. In the process of abnormal detection, fixed thresholds are usually used, which cannot dynamically adapt to the real-time fluctuation of data, and is prone to missed or missed detection, resulting in insufficient reliability and accuracy of abnormal detection results.

Method used

The abnormal detection method of hydrological sequence data based on cloud computing is adopted. By collecting and preprocessing hydrological sequence data, time series data is generated, and abnormal scores are calculated using convolutional neural networks, dynamic thresholds are generated based on abnormal scores to identify abnormal data, a visual interface is constructed to display abnormal data, and the generated hydrological sequence data is stored.

Benefits of technology

The problems of insufficient real-time and lack of dynamic thresholds in the detection of hydrological data anomalies are solved, the interdependence between the characteristics of hydrological sequence data is enhanced, the model's ability to detect minor hydrological sequence anomalies is improved, and the accuracy and real-time detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961842A_ABST
    Figure CN119961842A_ABST
Patent Text Reader

Abstract

The invention discloses a hydrological sequence data anomaly detection method and system based on cloud computing, and relates to the technical field of hydrological data analysis and cloud computing, and the method comprises the steps: collecting and transmitting hydrological sequence data, preprocessing the received transmission data, and generating time sequence data; calculating an abnormal score based on the time sequence data, and generating a dynamic threshold value according to the abnormal score to recognize abnormal data; constructing a visual interface to display the abnormal data; and storing hydrological sequence data generated by collection and analysis. Through a two-step normalization method, hydrological sequence data are normalized, interdependence between hydrological sequence data features is enhanced, two decoders are used for adversarial training, features of hydrological sequence abnormal data are amplified, the detection capacity of a model for slight hydrological sequence anomalies is improved, and the abnormal score is compared with a dynamic threshold value, so that the detection accuracy of the hydrological sequence anomalies is improved. And the accuracy and the real-time performance of detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hydrological data analysis and cloud computing, and in particular to a method and system for detecting anomalies in hydrological sequence data based on cloud computing. Background Art

[0002] With the rapid development of modern information technology, research in the field of hydrological monitoring plays an indispensable role in water resources management, disaster prevention and mitigation, and ecological protection. Traditional hydrological monitoring technology mainly relies on fixed manual observation and simple sensor data collection methods. Due to the long data collection cycle, poor real-time performance and high manual participation, it is difficult to meet the needs of modern refined hydrological management. With the rise of sensor technology, the Internet of Things and cloud computing technology, the collection and analysis of hydrological sequence data are gradually developing towards high precision, real-time and intelligent directions. The emergence of hydrological data processing technology based on cloud computing makes real-time transmission, storage and analysis of massive data possible. However, in actual hydrological monitoring systems, there are still limitations. Hydrological sequence data usually contain a lot of noise, missing values ​​and abnormal information. These factors will have an adverse impact on subsequent hydrological model prediction and decision support.

[0003] The existing technology still has many shortcomings in the actual application of hydrological data. The existing technology relies heavily on the assumptions of data distribution, and it is difficult to cope with complex and changeable hydrological data. It requires a large amount of labeled data for model optimization, and the application cost is high. In the anomaly detection process, fixed thresholds are usually used for judgment, which cannot dynamically adapt to the real-time volatility of data, and are prone to missed detections or false detections, resulting in insufficient reliability and accuracy of anomaly detection results. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a method and system for anomaly detection of hydrological sequence data based on cloud computing, which solves the problems of strong dependence on assumptions about data distribution, difficulty in coping with complex and changeable hydrological data, need for a large amount of labeled data for model optimization, and high application cost. In the anomaly detection process, fixed thresholds are usually used for judgment, which cannot dynamically adapt to the real-time volatility of data, and are prone to missed detections or false detections, resulting in insufficient reliability and accuracy of anomaly detection results.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a method for detecting anomalies in hydrological series data based on cloud computing, which includes collecting and transmitting hydrological series data, preprocessing the received transmission data and generating time series data; calculating anomaly scores based on time series data, generating dynamic thresholds based on the anomaly scores to identify abnormal data; constructing a visual interface to display abnormal data; and storing the collected and analyzed hydrological series data.

[0008] As a preferred solution of the hydrological sequence data anomaly detection method based on cloud computing of the present invention, the collection and transmission of the hydrological sequence data includes:

[0009] Select river sections, rain gauges and reservoirs as monitoring points, deploy flow sensors to detect water flow per unit time, deploy rain sensors to measure rainfall, and deploy water level sensors to record water level height;

[0010] The sensor model collects data once per second;

[0011] Deploy edge computing nodes such as Raspberry Pi 4B in the monitoring equipment room of the hydrological station to receive data collected by sensors in real time and aggregate the collected data to generate raw data of multivariate time series in a unified format;

[0012] Use discrete Fourier transform DFT to perform frequency domain compression on the time series of the original data, use TLS to perform hash signature on the time series of the original data, and use MQTT protocol to transmit the hash-signed data to the cloud computing platform;

[0013] The cloud computing platform ensures that the data has not been tampered with by comparing the received hash signature with the hash signature before transmission, and generates time series data from the time series when the data has not been tampered with, which is recorded as transmission data.

[0014] As a preferred solution of the hydrological sequence data anomaly detection method based on cloud computing of the present invention, the preprocessing of the received transmission data and generating time series data includes:

[0015] Based on the distributed computing capability provided by the cloud computing platform, the missing values ​​of the transmission data are detected using statistical analysis methods, the missing values ​​are filled by the sliding mean interpolation method, the noise of the transmission data is removed by the wavelet transform method, and the transmission data is normalized;

[0016] The normalization processing of the transmission data includes defining a multidimensional time series M, wherein each time point t contains three characteristic dimension values ​​of the hydrological series data, namely, a water level value, a rainfall value, and a flow rate value;

[0017] Divide the time series data into sliding windows with a window length of w and a step length of s to obtain the time window vector M t ;

[0018] For the time window vector M t The feature columns in the MaxAbs Scaler operation are performed to obtain the column normalization value x′ i,j , the formula is as follows:

[0019]

[0020] Among them, x i,j For window M t The value of the i-th row and j-th column in Indicates window M t The maximum absolute value of the jth column, k is the number of characteristic dimensions of the hydrological series data;

[0021] Normalize the value x′ based on the column i,j , use the VectorScaler method to calculate the window M t For each row of i ||, the formula is as follows:

[0022]

[0023] Calculate the normalized value x″ of the feature row based on the L2 norm i,j , the formula is as follows:

[0024]

[0025] Define new feature dimension x″ i,n+1 =||X′ i ||;

[0026] The new feature dimension x″ i,4 Add to x″ i,j In the vector, the formula is as follows:

[0027] Y″ i,j =[h 1i ,h 2i ,h 3i ,||X′ i ||],

[0028] Among them, h 1i is the water level value at time i, h 2i is the rainfall value at time i, h 3i is the flow velocity value at time i;

[0029] For the added vector Y″ i,j Use normalization to get x″′ i,j , the formula is as follows:

[0030]

[0031] Based on the normalized value, the normalized time window vector W is obtained t .

[0032] As a preferred solution of the hydrological sequence data anomaly detection method based on cloud computing of the present invention, the calculation of the anomaly score based on the preprocessed time series data includes:

[0033] Collect historical hydrological sequence data with labels and preprocess them to generate a training set. Define the time window vector in the training set as W i ;

[0034] In the forward propagation stage, the Encoder of the convolutional neural network CNN is used as the feature extractor, including the input layer, convolution layer, pooling layer and fully connected output layer;

[0035] Input layer input time window vector W i ;

[0036] The convolution layer applies a one-dimensional convolution kernel to transform the input time window vector W i Generate feature map;

[0037] The pooling layer downsamples the feature map of the convolutional layer to obtain the spatial and temporal features of the time series;

[0038] The fully connected output layer flattens the pooled features and maps them to the latent variable space to obtain the latent variable Z;

[0039] In the loss calculation stage, the latent variable Z is input into the decoder Decoder1, and the time series is gradually restored through the fully connected layer and the deconvolution layer. The gradient descent method is used to obtain the restored time series to minimize the reconstruction error L1.

[0040] Input the restored time series into the decoder Decoder2 to generate an adversarial time series, and use the gradient descent method to maximize the reconstruction error L2 for the adversarial time series;

[0041] Use the loss mean ratio method to set weight parameters Balance the importance of minimizing the reconstruction error L1 and maximizing the reconstruction error L2;

[0042] In the back propagation phase, decoder Decoder1 is used to decode the input data W i Perform the first dimensionality reduction and decoding to obtain ED1(W i ), ED1(W i ) is further decoded as the input of Decoder2 to obtain ED2(W i), where ED1 is a combination of Encoder and Decoder1, and ED2 is a combination of Encoder and Decoder2;

[0043] Using the L2 norm method to calculate the difference between the outputs of Decoder1 and Decoder2, we get Q = || ED1(W i )-ED2(W i )||2, Q is a nested constraint item;

[0044] The joint objective function is defined based on the nested constraints, and the formula is as follows:

[0045]

[0046] Where L is the joint objective function;

[0047] Use the optimization algorithm Adam to update the parameters of Encoder, Decoder1 and Decoder2;

[0048] The threshold of the change amplitude of the joint objective function L is set by the statistical division method, which is marked as σ;

[0049] Iterate the model training, repeat the forward propagation phase, loss calculation phase, and back propagation phase until the change amplitude of the joint objective function L is less than σ, and then stop the iteration;

[0050] The weight parameters α' and β' are introduced as learnable parameters, and the initial values ​​are set using empirical rules. Based on the constraints of the trained minimum reconstruction error L1 and maximum reconstruction error L2, the weight parameters α and β are gradient updated and iteratively optimized to obtain the optimized weight parameters α and β.

[0051] According to the optimized weight parameters α and β, the real-time time window vector W is input. t Calculate the anomaly score S t , the formula is as follows:

[0052] S t =α||W t -ED1(W t )||2+β||W t -ED2(ED1(W t ))||2.

[0053] As a preferred solution of the hydrological sequence data anomaly detection method based on cloud computing of the present invention, the method of generating a dynamic threshold according to the anomaly score to identify the anomaly data includes:

[0054] For the anomaly score S t Use the distribution characteristic method to obtain the initial static threshold T init, filter out the abnormal scores higher than the initial static threshold as the tail score set, denoted as S filter ;

[0055] Define the tail fraction set S filter The number of middle tail fractions m and the total number of abnormal fractions h are used to calculate the probability of exceeding the limit. The formula is as follows:

[0056]

[0057] Among them, q is the probability of exceeding the limit;

[0058] Based on the anomaly score S t The probability density function of the EGEP distribution is defined as follows:

[0059]

[0060] Among them, f(S t ) is the probability density function of EGEP distribution, a is the shape parameter controlling the central entropy, a>0, b is the shape parameter controlling the tail behavior, b>0, ξ is the thickness parameter of the tail distribution, ξ>0, λ is the scale parameter, λ>0;

[0061] Use the maximum likelihood estimation MLE to estimate the probability density function f(S t ) is fitted with the shape parameter a that controls the central entropy, the shape parameter b that controls the tail behavior, the thickness parameter ξ of the tail distribution, and the scale parameter λ, and the formula is as follows:

[0062]

[0063] Among them, l(a,b,ξ,λ) is the maximum likelihood function, ln is the natural logarithm, and g is the index variable;

[0064] Calculate the dynamic threshold z based on the fitting parameters and the limit-exceeding probability q q , the formula is as follows:

[0065]

[0066] According to the anomaly score S t and dynamic threshold z q Determine whether the current time point is abnormal:

[0067] When S t >z q , judged as abnormal, output abnormal time point, defined as z;

[0068] When S t ≤z q , judged as normal, continue testing;

[0069] Real-time updates of anomaly scores and dynamic thresholds through a sliding window mechanism;

[0070] Based on the abnormal time point z, the two parts of the abnormal score are calculated as follows:

[0071]

[0072] in, is the first part anomaly score at time point z, w z,j is the value of feature j at time point z, and d is W z The characteristic dimension of

[0073]

[0074] in, is the anomaly score of the second part at time point z;

[0075] based on and Get the total anomaly score, the formula is as follows:

[0076]

[0077] Among them, S z is the total abnormal score at abnormal time point z;

[0078] The total anomaly score S is calculated using the squared distance attribution method z The feature contribution of the first part is obtained And the second part of the feature contribution

[0079] Contribution to the first part of the feature And the second part of the feature contribution Use the summation method to get the contribution value of each feature C z,j ;

[0080] According to the contribution value C of each feature z,h Sort all features in descending order, select the feature with the largest contribution value as the main abnormal feature, and output the abnormal time point and the main abnormal feature corresponding to the abnormal time point.

[0081] As a preferred solution of the hydrological sequence data anomaly detection method based on cloud computing of the present invention, the construction of a visual interface to display abnormal data includes:

[0082] Anomaly trend graphs are formed based on comprehensive anomaly scores and corresponding time points, and anomaly distribution heat maps are formed based on anomaly time points and corresponding anomaly features. All graphs are updated in real time based on real-time data.

[0083] Use Streamlit to build a visual interface and divide it into two areas;

[0084] The two areas include an area for an abnormal trend graph displayed by Plotly and another area for an abnormal distribution heat map displayed by Plotly.

[0085] As a preferred solution of the hydrological sequence data anomaly detection method based on cloud computing described in the present invention, wherein: the storage of the hydrological sequence data generated by collection and analysis refers to sorting the collected hydrological sequence data and the data results generated during the analysis process according to timestamps and storing them in a cloud database, marking the stored data, synchronously backing up the stored data regularly, regularly performing security and integrity checks on the stored data and backup data, and generating a test report.

[0086] In a second aspect, the present invention provides a hydrological sequence data anomaly detection system based on cloud computing, comprising:

[0087] The acquisition and preprocessing module is used for the acquisition and transmission of hydrological sequence data, preprocesses the received transmission data and generates time series data;

[0088] The model threshold module is used to calculate the anomaly score based on the preprocessed time series data and generate a dynamic threshold to identify abnormal data according to the anomaly score;

[0089] The visualization storage module is used to build a visualization interface to display abnormal data and store the hydrological sequence data generated by collection and analysis.

[0090] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the cloud computing-based hydrological sequence data anomaly detection method as described in the first aspect of the present invention is implemented.

[0091] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the cloud computing-based hydrological sequence data anomaly detection system as described in the first aspect of the present invention is implemented.

[0092] The beneficial effects of the present invention are as follows: the present invention collects hydrological sequence data and transmits it, pre-processes the received transmission data and generates time series data; calculates anomaly scores based on time series data, and generates dynamic thresholds according to the anomaly scores to identify abnormal data; solves the problems of insufficient real-time performance and lack of dynamic thresholds in hydrological data anomaly detection, enhances the interdependence between hydrological sequence data features, improves the model's ability to detect slight hydrological sequence anomalies, and improves the accuracy and real-time performance of detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0094] Figure 1 This is a flow chart of a hydrological sequence data anomaly detection method based on cloud computing in Example 1.

[0095] Figure 2 This is a schematic diagram of a hydrological sequence data anomaly detection system based on cloud computing in Example 1. DETAILED DESCRIPTION

[0096] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0097] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0098] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0099] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, and provides a hydrological sequence data anomaly detection method based on cloud computing, comprising the following steps:

[0100] S1, collect hydrological sequence data and transmit it, pre-process the received transmission data and generate time series data;

[0101] Specifically, the collection and transmission of hydrological sequence data includes:

[0102] Select river sections, rain gauges and reservoirs as monitoring points, deploy flow sensors to detect water flow per unit time, deploy rain sensors to measure rainfall, and deploy water level sensors to record water level height;

[0103] The sensor model collects data once per second;

[0104] Deploy edge computing nodes such as Raspberry Pi 4B in the monitoring equipment room of the hydrological station to receive data collected by sensors in real time and aggregate the collected data to generate raw data of multivariate time series in a unified format;

[0105] Use discrete Fourier transform DFT to perform frequency domain compression on the time series of the original data, use TLS to perform hash signature on the time series of the original data, and use MQTT protocol to transmit the hash-signed data to the cloud computing platform;

[0106] The cloud computing platform ensures that the data has not been tampered with by comparing the received hash signature with the hash signature before transmission, and generates time series data from the time series when the data has not been tampered with, which is recorded as transmission data.

[0107] By carefully selecting key locations as monitoring points, the dynamic changes of water resources can be accurately captured. By arranging sensors to collect multi-dimensional hydrological parameters at a high frequency (once per second), not only the temporal resolution of the data is enhanced, but also the speed and accuracy of responding to emergencies (such as floods and droughts) are improved. By introducing edge computing nodes to process field data, the initial cleaning and structuring of the data is achieved, which reduces the pressure on the cloud computing platform, shortens the data processing path, reduces the network bandwidth requirement, and improves the response speed of the system. By using discrete Fourier transform DFT to perform frequency domain compression on the time series of the original data, the main components of the signal are effectively extracted, and noise interference is eliminated, thereby simplifying the data volume without losing relevant information. Key information, hash signature the time series of the original data through TLS, and use the MQTT protocol to transmit the hash-signed data to the cloud computing platform. The TLS encryption technology and hash signature mechanism are used to ensure the security and integrity of data transmission, and prevent the risk of eavesdropping or tampering in the middle. By selecting the lightweight and efficient MQTT protocol, it is ensured that the data upload task can be completed stably and reliably even under poor network conditions. Hash value comparison and verification are performed through the cloud to ensure the consistency and non-repudiation of data in the entire link from source to destination, further consolidating the security of the system and providing users with a trusted data platform to support subsequent in-depth data mining, analysis and auxiliary decision-making.

[0108] Further, preprocessing the received transmission data and generating time series data includes:

[0109] Based on the distributed computing capability provided by the cloud computing platform, the missing values ​​of the transmission data are detected using statistical analysis methods, the missing values ​​are filled by the sliding mean interpolation method, the noise of the transmission data is removed by the wavelet transform method, and the transmission data is normalized;

[0110] The normalization process of the transmission data includes defining a multidimensional time series M, where each time point t contains three characteristic dimension values ​​of the hydrological series data, and the formula is as follows:

[0111]

[0112] x t =[h 1t ,h 2t ,h 3t ],

[0113] Among them, h 1t is the water level at time t, h 2t is the rainfall value at time t, h 3t is the flow velocity value at time t;

[0114] The time series data is divided into sliding windows with a window length of w and a step length of s to obtain a time window vector. The formula is as follows:

[0115]

[0116] Among them, M t Represented as the time window vector at the current moment;

[0117] For the time window vector M t The feature columns in the Max Abs Scaler operation are performed to obtain the column normalization value x′ i,j , the formula is as follows:

[0118]

[0119] Among them, x i,j For window M t The value of the i-th row and j-th column in Indicates window M t The maximum absolute value of the jth column, k is the number of characteristic dimensions of the hydrological series data;

[0120] Normalize the value x′ based on the column i,j , use the VectorScaler method to calculate the window M t For each row of i ||, the formula is as follows:

[0121]

[0122] Calculate the normalized value x″ of the feature row based on the L2 norm i,j , the formula is as follows:

[0123]

[0124] This operation maps each eigenvector onto the unit circle, ensuring consistency between the eigenvalue scales;

[0125] Define new feature dimension x″ i,4 =||X′ i ||;

[0126] The new feature dimension x″ i,4 Add to x″ i,j In the vector, the formula is as follows:

[0127] Y″ i,j =[h 1i ,h 2i ,h 3i ,||X′ i ||],

[0128] For the added vector Y″ i,j Use normalization to get x″′ i,j , the formula is as follows:

[0129]

[0130] The newly added feature dimension reflects the absolute size of the collected data, thereby enhancing the classifier's ability to distinguish data distribution. The final normalization maps the data to a high-dimensional unit sphere, further reducing the problem of data extrapolation.

[0131] Based on the normalized value, the normalized time window vector W is obtained t .

[0132] The existing normalization processing operation only focuses on the relative value of the feature and ignores the absolute value of the feature, which may lead to information loss, especially when the feature has actual physical meaning. When processing data far away from the training samples, extrapolation problems may occur, resulting in inaccurate prediction results. Min-Max normalization simplifies the data range, but may lose important proportional relationships and feature absolute size information when processing multidimensional data. The proposed two-step normalization method not only considers the interdependence between features, but also the absolute value of each feature. Ultimately, all samples are mapped to the high-dimensional unit sphere, which not only enhances the classifier's ability to distinguish data distribution, but also effectively solves the problem of data extrapolation, improving prediction performance and reliability.

[0133] Each feature column is normalized separately through a two-step normalization method, which maintains the proportional relationship of the original data and avoids the influence of extreme values. By calculating the L2 norm, the feature vector of each sample is mapped to the unit circle to ensure the consistency between the eigenvalue proportions. This normalization method further enhances the consistency of feature representation and promotes the learning efficiency of model parameters. By defining the new feature dimension as a newly added column, a new feature dimension that can reflect the absolute size is added, and then normalization is performed again.

[0134] S2. Calculate the anomaly score based on the time series data, and generate a dynamic threshold to identify abnormal data based on the anomaly score;

[0135] Specifically, the calculation of anomaly scores based on preprocessed time series data includes:

[0136] Collect historical hydrological sequence data with labels and preprocess them to generate a training set. Define the time window vector in the training set as W i ;

[0137] In the forward propagation stage, the Encoder of the convolutional neural network CNN is used as the feature extractor, including the input layer, convolution layer, pooling layer and fully connected output layer;

[0138] Input layer input time window vector W i ;

[0139] The convolution layer applies a one-dimensional convolution kernel to transform the input time window vector W i Generate feature map;

[0140] The pooling layer downsamples the feature map of the convolutional layer to obtain the spatial and temporal features of the time series;

[0141] The fully connected output layer flattens the pooled features and maps them to the latent variable space to obtain the latent variable Z;

[0142] In the loss calculation stage, the latent variable Z is input into the decoder Decoder1, and the time series is gradually restored through the fully connected layer and the deconvolution layer. The gradient descent method is used to obtain the restored time series to minimize the reconstruction error L1.

[0143] L1=||W i -ED1(W i )||2,

[0144] Among them, ED1 is the combination of Encoder and Decoder1;

[0145] Input the restored time series into the decoder Decoder2 to generate an adversarial time series, and use the gradient descent method to maximize the reconstruction error L2 for the adversarial time series;

[0146] L2=||W i -ED2(ED1(W i ))||2,

[0147] Among them, ED2 is the combination of Encoder and Decoder2;

[0148] Use the loss mean ratio method to set weight parameters Balance the importance of minimizing the reconstruction error L1 and maximizing the reconstruction error L2;

[0149] Using two decoders (Decoder1 and Decoder2) for adversarial training can effectively amplify the features of abnormal data and improve the model's ability to detect minor anomalies;

[0150] In the back propagation stage, the objective function is defined based on minimizing the reconstruction error L1 and maximizing the reconstruction error L2. The formula is as follows:

[0151]

[0152] in, is the objective function of decoder Decoder1, is the objective function of decoder Decoder2;

[0153] The above formula mainly optimizes the error of a single decoder, but fails to consider the problems existing between multi-level decoders: lack of correlation between decoders, insufficient sensitivity of anomaly detection, and inability to constrain the consistency between decoders;

[0154] Use decoder Decoder1 to decode the input data W i Perform the first dimensionality reduction and decoding to obtain ED1(W i ), ED1(W i ) is further decoded as the input of Decoder2 to obtain ED2(W i );

[0155] Using the L2 norm method to calculate the difference between the outputs of Decoder1 and Decoder2, we get Q = || ED1(W i )-ED2(W i )||2, Q is a nested constraint item;

[0156] The joint objective function is defined based on the nested constraints, and the formula is as follows:

[0157]

[0158] Where L is the joint objective function;

[0159] The joint objective function reflects the collaborative relationship between Decoder1 and Decoder2. The outputs of Decoder1 and Decoder2 are directly compared, avoiding the redundant steps of optimizing L2 alone. The optimization goal of Decoder2 changes from reconstruction to amplification of Decoder1 output, further enhancing the sensitivity to abnormal patterns.

[0160] Use the optimization algorithm Adam to update the parameters of Encoder, Decoder1 and Decoder2;

[0161] The threshold of the change amplitude of the joint objective function L is set by the statistical division method, which is marked as σ;

[0162] Iterate the model training, repeat the forward propagation phase, loss calculation phase, and back propagation phase until the change amplitude of the joint objective function L is less than σ, and then stop the iteration;

[0163] The weight parameters α' and β' are introduced as learnable parameters, and the initial values ​​are set using empirical rules. Based on the constraints of the trained minimum reconstruction error L1 and maximum reconstruction error L2, the weight parameters α and β are gradient updated and iteratively optimized to obtain the optimized weight parameters α and β.

[0164] According to the optimized weight parameters α and β, the real-time time window vector W is input. t Calculate the anomaly score S t , the formula is as follows:

[0165] S t =α||W t -ED1(W t )||2+β||W t -ED2(ED1(W t ))||2.

[0166] By designing an adversarial training architecture consisting of three convolutional neural networks, including an encoder and two decoders, and training the model in two stages (reconstruction minimization stage and adversarial training stage), the model can not only accurately reconstruct normal time series, but also effectively amplify the characteristics of abnormal data, thereby improving the ability to detect minor anomalies.

[0167] Furthermore, generating a dynamic threshold to identify abnormal data according to the abnormal score includes:

[0168] For the anomaly score S t Use the distribution characteristic method to obtain the initial static threshold T init , filter out the abnormal scores higher than the initial static threshold as the tail score set, denoted as S filter ;

[0169] Define the tail fraction set S filter The number of middle tail fractions m and the total number of abnormal fractions h are used to calculate the probability of exceeding the limit. The formula is as follows:

[0170]

[0171] Among them, q is the probability of exceeding the limit;

[0172] Based on the anomaly score S t The probability density function of the EGEP distribution is defined as follows:

[0173]

[0174] Among them, f(S t ) is the probability density function of EGEP distribution, a is the shape parameter controlling the central entropy, a>0, b is the shape parameter controlling the tail behavior, b>0, ξ is the thickness parameter of the tail distribution, ξ>0, λ is the scale parameter, λ>0;

[0175] Use the maximum likelihood estimation MLE to estimate the probability density function f(S t ) is fitted with the shape parameter a that controls the central entropy, the shape parameter b that controls the tail behavior, the thickness parameter ξ of the tail distribution, and the scale parameter λ, and the formula is as follows:

[0176]

[0177] Among them, l(a,b,ξ,λ) is the maximum likelihood function, ln is the natural logarithm, and g is the index variable;

[0178] Calculate the dynamic threshold z based on the fitting parameters and the limit-exceeding probability q q , the formula is as follows:

[0179]

[0180] According to the anomaly score S t and dynamic threshold z q Determine whether the current time point is abnormal:

[0181] When S t >z q , judged as abnormal, output abnormal time point, defined as z;

[0182] When S t ≤z q , judged as normal, continue testing;

[0183] Real-time updates of anomaly scores and dynamic thresholds through a sliding window mechanism;

[0184] Based on the abnormal time point z, the two parts of the abnormal score are calculated as follows:

[0185]

[0186] in, is the first part anomaly score at time point z, w z,j is the value of feature j at time point z, and d is W z The characteristic dimension of

[0187]

[0188] in, is the anomaly score of the second part at time point z;

[0189] based on and Get the total anomaly score, the formula is as follows:

[0190]

[0191] Among them, S z is the total abnormal score at abnormal time point z;

[0192] The total anomaly score S is calculated using the squared distance attribution method z The feature contribution of the first part is obtained And the second part of the feature contribution The formula is as follows:

[0193]

[0194] Contribution to the first part of the feature And the second part of the feature contribution Use the summation method to get the contribution value of each feature C z,j , the formula is as follows:

[0195]

[0196] According to the contribution value C of each feature z,j Sort all features in descending order, select the feature with the largest contribution value as the main abnormal feature, and output the abnormal time point and the main abnormal feature corresponding to the abnormal time point.

[0197] By introducing the GPD model to fit the anomaly scores that exceed the static threshold, the probability distribution of extreme events can be accurately described. By using the MLE method to optimize the parameter estimation of GPD, the accuracy of parameter estimation is ensured. By maximizing the joint probability of samples to find the optimal parameters, the model can be more in line with the actual data distribution and improve the reliability of anomaly detection. By combining the GPD parameters and the set probability level, the dynamic threshold that changes over time is calculated. This dynamic adjustment mechanism can adapt to the changes in abnormal characteristics in different time periods, improve the flexibility and response speed of anomaly detection, and avoid the problem of premature alarm or missed alarm caused by fixed thresholds. Through the sliding window technology, the system can continuously receive new data and update the anomaly score and dynamic threshold accordingly. This real-time update mechanism ensures that the model is always in the latest state. For the time point marked as abnormal, the abnormal contribution of each feature dimension (water level, rainfall and flow rate) is further decomposed, and the contribution of each feature dimension to the anomaly score is sorted to determine the most significant abnormal feature.

[0198] S3, build a visual interface to display abnormal data and store the hydrological sequence data generated by collection and analysis;

[0199] Specifically, building a visual interface to display abnormal data includes:

[0200] Anomaly trend graphs are formed based on comprehensive anomaly scores and corresponding time points, and anomaly distribution heat maps are formed based on anomaly time points and corresponding anomaly features. All graphs are updated in real time based on real-time data.

[0201] Use Streamlit to build a visual interface and divide it into two areas;

[0202] The two areas include an area for an abnormal trend graph displayed by Plotly and another area for an abnormal distribution heat map displayed by Plotly.

[0203] By using Streamlit to build a visualization interface, we can quickly build an easy-to-use interactive visualization platform. Streamlit is an application framework designed specifically for Python developers. It allows users to create beautiful and feature-rich web applications with a minimum of code. By introducing the Plotly library to draw anomaly trend graphs, we can intuitively display the changing trend of anomalies over time, and mark significant anomaly points to provide users with a clear timeline perspective, helping them to quickly identify the time and frequency of abnormal events. Using Plotly to generate anomaly distribution heat maps shows the time points when anomalies occur and their main characteristic dimensions (water level, rainfall, and flow rate). The heat map represents the intensity of anomalies by color depth, compressing multidimensional data into two-dimensional graphics. This visualization method is particularly helpful in quickly locating the source of anomalies, reducing cognitive burden, and highlighting key information.

[0204] Furthermore, storing the hydrological sequence data generated by collection and analysis means sorting the collected hydrological sequence data and the data results generated during the analysis process according to timestamps and storing them in a cloud database, marking the stored data, synchronously backing up the stored data regularly, regularly checking the security and integrity of the stored data and the backup data, and generating a test report.

[0205] By using timestamps to sort all the data generated by collection and analysis and storing them in the database, the temporal consistency of the data is guaranteed, and the time-ordered storage of the data is achieved, which is convenient for subsequent data retrieval and analysis. By adding tags to each data, data classification management and rapid retrieval are achieved. Through a regular periodic backup strategy, the data in the cloud database is copied to an independent backup system, which realizes redundant storage of data and ensures data security and recovery capabilities. By setting a periodic security and integrity detection mechanism, it is ensured that the data has not been tampered with or leaked, thereby improving the security and credibility of the data.

[0206] This embodiment also provides a hydrological sequence data anomaly detection system based on cloud computing, including:

[0207] The acquisition and preprocessing module is used for the acquisition and transmission of hydrological sequence data, preprocesses the received transmission data and generates time series data;

[0208] The model threshold module is used to calculate the anomaly score based on the preprocessed time series data and generate a dynamic threshold to identify abnormal data according to the anomaly score;

[0209] The visualization storage module is used to build a visualization interface to display abnormal data and store the hydrological sequence data generated by collection and analysis.

[0210] This embodiment also provides a computer device, which is suitable for the case of a hydrological sequence data anomaly detection method based on cloud computing, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the hydrological sequence data anomaly detection method based on cloud computing proposed in the above embodiment.

[0211] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0212] The present embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the hydrological sequence data anomaly detection method based on cloud computing as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, disk or optical disk.

[0213] In summary, the present invention collects hydrological sequence data and transmits it, preprocesses the received transmission data and generates time series data; calculates anomaly scores based on time series data, and generates dynamic thresholds to identify abnormal data according to the anomaly scores; solves the problems of insufficient real-time performance and lack of dynamic thresholds in hydrological data anomaly detection, enhances the interdependence between hydrological sequence data features, improves the model's ability to detect slight hydrological sequence anomalies, and improves the accuracy and real-time performance of detection.

[0214] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A hydrological sequence data anomaly detection method based on cloud computing, characterized by: include: Collect and transmit hydrological series data, pre-process the received transmission data and generate time series data; Calculate anomaly scores based on time series data, and generate dynamic thresholds based on anomaly scores to identify abnormal data; Build a visual interface to display abnormal data and store the hydrological series data generated by collection and analysis.

2. The hydrological sequence data anomaly detection method based on cloud computing according to claim 1, characterized in that: The collection and transmission of the hydrological sequence data includes: Select river sections, rain gauges and reservoirs as monitoring points, deploy flow sensors to detect water flow per unit time, deploy rain sensors to measure rainfall, and deploy water level sensors to record water level height; The sensor model collects data once per second; Deploy edge computing nodes such as Raspberry Pi 4B in the monitoring equipment room of the hydrological station to receive data collected by sensors in real time and aggregate the collected data to generate raw data of multivariate time series in a unified format; Use discrete Fourier transform DFT to perform frequency domain compression on the time series of the original data, use TLS to perform hash signature on the time series of the original data, and use MQTT protocol to transmit the hash-signed data to the cloud computing platform; The cloud computing platform ensures that the data has not been tampered with by comparing the received hash signature with the hash signature before transmission, and generates time series data from the time series when the data has not been tampered with, which is recorded as transmission data.

3. The hydrological sequence data anomaly detection method based on cloud computing according to claim 2, characterized in that: The step of preprocessing the received transmission data and generating time series data comprises: Based on the distributed computing capability provided by the cloud computing platform, the missing values ​​of the transmission data are detected using statistical analysis methods, the missing values ​​are filled by the sliding mean interpolation method, the noise of the transmission data is removed by the wavelet transform method, and the transmission data is normalized; The normalization processing of the transmission data includes defining a multidimensional time series M, wherein each time point t contains three characteristic dimension values ​​of the hydrological series data, namely, a water level value, a rainfall value, and a flow rate value; Divide the time series data into sliding windows with a window length of w and a step length of s to obtain the time window vector M t ; For the time window vector M t The feature columns in the MaxAbs Scaler operation are performed to obtain the column normalization value x′ i,j , the formula is as follows: Among them, x i,j For window M t The value of the i-th row and j-th column in Indicates window M t The maximum absolute value of the jth column, k is the number of characteristic dimensions of the hydrological series data; Normalize the value x′ based on the column i,j , use the VectorScaler method to calculate the window M t For each row of i ||, the formula is as follows: Calculate the normalized value x″ of the feature row based on the L2 norm i,j , the formula is as follows: Define new feature dimension x″ i,n+1 =||X′ i ||; The new feature dimension x″ i,4 Add to x″ i,j In the vector, the formula is as follows: Among them, h 1i is the water level value at time i, h 2i is the rainfall value at time i, j 3i is the flow velocity value at time i; For the added vector Y″ i,j Use normalization to get x″ i,j , the formula is as follows: Based on the normalized value, the normalized time window vector W is obtained t。 4. The hydrological sequence data anomaly detection method based on cloud computing according to claim 3, characterized in that: The calculation of the anomaly score based on the preprocessed time series data includes: Collect historical hydrological sequence data with labels and preprocess them to generate a training set. Define the time window vector in the training set as W i ; In the forward propagation stage, the Encoder of the convolutional neural network CNN is used as the feature extractor, including the input layer, convolution layer, pooling layer and fully connected output layer; Input layer input time window vector W i ; The convolution layer applies a one-dimensional convolution kernel to transform the input time window vector W i Generate feature map; The pooling layer downsamples the feature map of the convolutional layer to obtain the spatial and temporal features of the time series; The fully connected output layer flattens the pooled features and maps them to the latent variable space to obtain the latent variable Z; In the loss calculation stage, the latent variable Z is input into the decoder Decoder1, and the time series is gradually restored through the fully connected layer and the deconvolution layer. The gradient descent method is used to obtain the restored time series to minimize the reconstruction error L1. Input the restored time series into the decoder Decoder2 to generate an adversarial time series, and use the gradient descent method to maximize the reconstruction error L2 for the adversarial time series; Use the loss mean ratio method to set weight parameters Balance the importance of minimizing the reconstruction error L1 and maximizing the reconstruction error L2; In the back propagation phase, decoder Decoder1 is used to decode the input data W i Perform the first dimensionality reduction and decoding to obtain ED1(W i ), ED1(W i ) is further decoded as the input of Decoder2 to obtain ED2(W i ), where ED1 is a combination of Encoder and Decoder1, and ED2 is a combination of Encoder and Decoder2; Using the L2 norm method to calculate the difference between the outputs of Decoder1 and Decoder2, we get Q = || ED1(W i )-ED2(W i )||2, Q is a nested constraint item; The joint objective function is defined based on the nested constraints, and the formula is as follows: Where L is the joint objective function; Use the optimization algorithm Adam to update the parameters of Encoder, Decoder1 and Decoder2; The threshold of the change amplitude of the joint objective function L is set by the statistical division method, marked as σ; Iterate the model training, repeat the forward propagation phase, loss calculation phase, and back propagation phase until the change amplitude of the joint objective function L is less than σ, and then stop the iteration; The weight parameters α' and β' are introduced as learnable parameters, and the initial values ​​are set using empirical rules. Based on the constraints of the trained minimum reconstruction error L1 and maximum reconstruction error L2, the weight parameters α and β are gradient updated and iteratively optimized to obtain the optimized weight parameters α and β. According to the optimized weight parameters α and β, the real-time time window vector W is input. t Calculate the anomaly score S t , the formula is as follows: S t =α||W t -ED1(W t )||2+β||In t -ED2(ED1(In t ))||2。 5. The hydrological sequence data anomaly detection method based on cloud computing according to claim 4, characterized in that: Generating a dynamic threshold according to the anomaly score to identify abnormal data includes: For the anomaly score S t Use the distribution characteristic method to obtain the initial static threshold T init , filter out the abnormal scores higher than the initial static threshold as the tail score set, denoted as S filter ; Define the tail fraction set S filter The number of middle tail fractions m and the total number of abnormal fractions h are used to calculate the exceedance probability formula, which is as follows: Among them, q is the probability of exceeding the limit; Based on the anomaly score S t The probability density function of the EGEP distribution is defined as follows: Among them, f(S t ) is the probability density function of EGEP distribution, a is the shape parameter controlling the central entropy, a>0, b is the shape parameter controlling the tail behavior, b>0, ξ is the thickness parameter of the tail distribution, ξ>0, λ is the scale parameter, λ>0; Use the maximum likelihood estimation MLE to estimate the probability density function f(S t ) is fitted with the shape parameter a that controls the central entropy, the shape parameter b that controls the tail behavior, the thickness parameter ξ of the tail distribution, and the scale parameter λ, and the formula is as follows: Among them, l(a,b,ξ,λ) is the maximum likelihood function, ln is the natural logarithm, and g is the index variable; Calculate the dynamic threshold z based on the fitting parameters and the limit-exceeding probability q q , the formula is as follows: According to the anomaly score S t and dynamic threshold z q Determine whether the current time point is abnormal: When S t >z q , judged as abnormal, output abnormal time point, defined as z; When S t ≤z q , judged as normal, continue testing; Real-time updates of anomaly scores and dynamic thresholds through a sliding window mechanism; Based on the abnormal time point z, the two parts of the abnormal score are calculated as follows: in, is the first part anomaly score at time point z, w z,j is the value of feature j at time point z, and d is W z The characteristic dimension of in, is the anomaly score of the second part at time point z; based on and Get the total anomaly score, the formula is as follows: Among them, S z is the total abnormal score at abnormal time point z; The total anomaly score S is calculated using the squared distance attribution method z The feature contribution of the first part is obtained And the second part of the feature contribution Contribution to the first part of the feature And the second part of the feature contribution Use the summation method to get the contribution value of each feature C z,j ; According to the contribution value C of each feature z,j Sort all features in descending order, select the feature with the largest contribution value as the main abnormal feature, and output the abnormal time point and the main abnormal feature corresponding to the abnormal time point.

6. The hydrological sequence data anomaly detection method based on cloud computing according to claim 5, characterized in that: The construction of a visual interface to display abnormal data includes: Anomaly trend graphs are formed based on comprehensive anomaly scores and corresponding time points, and anomaly distribution heat maps are formed based on anomaly time points and corresponding anomaly features. All graphs are updated in real time based on real-time data. Use Streamlit to build a visual interface and divide it into two areas; The two areas include an area for an abnormal trend graph displayed by Plotly and another area for an abnormal distribution heat map displayed by Plotly.

7. The hydrological sequence data anomaly detection method based on cloud computing according to claim 6, characterized in that: The storage of the hydrological sequence data generated by collection and analysis refers to sorting the collected hydrological sequence data and the data results generated during the analysis process according to timestamps and storing them in a cloud database, marking the stored data, synchronously backing up the stored data regularly, regularly checking the security and integrity of the stored data and the backup data, and generating a test report.

8. A hydrological sequence data anomaly detection system based on cloud computing, based on the hydrological sequence data anomaly detection method based on cloud computing according to any one of claims 1 to 7, characterized in that: include: The acquisition and preprocessing module is used for the acquisition and transmission of hydrological sequence data, preprocesses the received transmission data and generates time series data; The model threshold module is used to calculate the anomaly score based on the preprocessed time series data and generate a dynamic threshold to identify abnormal data according to the anomaly score; The visualization storage module is used to build a visualization interface to display abnormal data and store the hydrological sequence data generated by collection and analysis.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the hydrological sequence data anomaly detection method based on cloud computing according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the hydrological sequence data anomaly detection method based on cloud computing according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Industrial internet time series data anomaly detection method and device

    CN116522265A

  • Time sequence anomaly detection method and system based on time sequence and multiple variables

    CN117313015A

  • Abnormal data detection method for mechanical state monitoring

    CN118094421A

  • KPIs anomaly detection method based on VAE-GRU and adversarial training

    CN118227366A

  • Anomaly detection method for large-scale multivariate time series data in cloud environment

    WO2022160902A1

Cited By

  • Cross-border e-commerce information risk analysis method in combination with cloud computing

    CN120875888A

  • Dynamic visualization and intelligent alarm method and system based on multi-source water and rain condition monitoring data

    CN122281841A