Hydropower station fault prediction method for bias error data

The trusted data is filtered through PCA, the CNN-LSTM process timing characteristics and the GAN generates an expanded data set, combined with a multi-scale DNN network for fault identification, solving the problem of data imbalance and abnormality in hydropower station fault detection, achieving higher detection accuracy and robustness.

CN120257079APending Publication Date: 2025-07-04SDIC GANSU XIAOSANXIA POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510137647.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When traditional hydropower station fault detection methods face large-scale, complex data and nonlinear dynamic changes in systems, it is difficult to fully and accurately reflect the real-time operating status of the equipment, especially when dealing with error problems such as data imbalance and data abnormalities, they are poorly adaptable.

Method used

The controllable trustworthiness detection mechanism based on PCA is used to filter trustworthy data, combine the CNN-LSTM network to process timing characteristics, use GAN to generate new samples similar to the real fault data to expand the data set, and feature extraction and fault identification are performed through multi-scale DNN networks.

Benefits of technology

It improves the robustness and accuracy of hydropower station fault detection, can effectively deal with data imbalance and abnormal problems, and improves the accuracy and comprehensiveness of fault classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257079A_ABST
    Figure CN120257079A_ABST
Patent Text Reader

Abstract

The invention discloses a hydropower station fault prediction method for bias error data, and the method comprises the following steps: designing a controllable credibility detection mechanism based on a PCA (principal component analysis) algorithm, analyzing the change direction of the principal component of monitoring data, calculating the credibility of the data, screening out credible data, and eliminating abnormal values; a CNN-LSTM (Convolutional Neural Network-Long Short Term Memory) (Convolutional Neural Network-Long Short Term Memory) network is adopted to process and predict the time sequence data, and a gating mechanism is utilized to process the operation state data of the hydropower station; the method comprises the following steps: generating an abnormal sample in combination with a GAN (Generative Advanced Network) technology, and expanding a training data set; and designing a multi-scale feature extraction network model, carrying out data hierarchical processing by using a plurality of branch networks, extracting features of different scales, and carrying out fault detection. The method is suitable for fault prediction conditions of data exception, fault sample imbalance and the like with bias error data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hydropower station fault detection, and particularly to a hydropower station fault prediction method for biased data. Background Art

[0002] Fault detection in hydropower stations is crucial for ensuring the safe and reliable operation of equipment. Because equipment failures not only cause economic losses, but may also lead to serious safety accidents and even have negative impacts on society and the environment. Traditional fault detection methods mainly rely on signal processing and statistical techniques, such as wavelet transform, NMF (Non-negative Matrix Factorization), and PLS (Partial Least Squares), etc. These methods show certain limitations when facing large-scale, complex data and the non-linear dynamic changes of the system, and it is difficult to comprehensively and accurately reflect the real-time operating state of the equipment. In recent years, fault detection technologies based on machine learning and deep learning have made remarkable progress. Especially neural network models, with their excellent feature extraction capabilities, can extract deep-level fault information from a large amount of data. Although deep neural networks have shown significant advantages in processing complex data, traditional neural networks still have certain limitations in dealing with the dynamic changes of long time series data, especially in dealing with bias problems such as data imbalance and data anomalies, and their adaptability is poor. Summary of the Invention

[0003] The purpose of the present invention is to propose a hydropower station fault prediction method for biased data in view of the defects and deficiencies of the above-mentioned existing technologies. This method realizes that the model is still applicable to hydropower station fault prediction in the presence of data bias problems such as data anomalies and sample imbalance.

[0004] The technical solution adopted by the present invention to solve its technical problems is: a hydropower station fault prediction method for biased data, including the following steps:

[0005] Step S1: Analyze the mainstream data trend of the monitoring data based on a controllable credibility detection mechanism, calculate the credibility of the data under different monitoring times, and screen the data according to the credibility;

[0006] Step S2: Normalize the screened data set, perform time window sampling to capture the time series characteristics of the data, and the sampled data will be sent into a CNN-LSTM network to predict future time series data;

[0007] Step S3: Use a GAN network to generate new samples similar to the real fault data in statistical characteristics to expand the data.

[0008] Step S4: The multi-scale DNN network model extracts features and classifies the data predicted by the CNN-LSTM network, determines whether there is faulty data in the predicted data, and outputs the fault identification result. If no fault is detected, return to the sampling window to continue the prediction and check the next batch of data.

[0009] Further, in step S1 of the present invention, the data credibility detection includes the following:

[0010] The credibility detection process includes data analysis and feature selection. Principal component analysis is used to identify the main change directions in the repeated monitoring data, calculate the component matrix P and the proportion of explained variance λ, project the data into a new coordinate system to obtain the direction with the largest variance in the data, that is, the principal component, calculate the contribution degree of each group of monitoring data to the PCA result, so as to evaluate the credibility of the data:

[0011]

[0012] where C i is the contribution degree of the i-th group of monitoring data to the PCA result, n is the number of principal components, P ji is the load of the j-th principal component on the corresponding variable of the i-th group of data, and λ j is the variance of the j-th principal component. Sort each group of monitoring data according to these contribution degrees and assign selection probabilities. The higher the contribution degree of the data, the greater the probability of being selected.

[0013] Further, in step S2 of the present invention, capturing the temporal features of the data and predicting future temporal data includes the following:

[0014] After data normalization preprocessing and time window sampling are completed, it is input into the CNN-LSTM network to extract temporal features and make predictions. The CNN-LSTM network structure mainly includes a CNN layer, an LSTM layer, and a fully connected layer. The CNN layer is used to extract local features of the data, the LSTM layer is used to capture the temporal features of the data, and the final prediction result is output to the fully connected layer.

[0015] In the CNN layer, the output feature map V is as follows:

[0016]

[0017] where z is the activation function, w i and x i represent the convolutional kernel weight and the input matrix respectively, b is the bias parameter, ∑ is the summation symbol, and i is the element index.

[0018] After the calculation of the convolutional layer, the obtained feature maps are sequentially concatenated into vectors and output to the fully connected layer. The convolution of the CNN layer weights the data at each time point, that is, extracts features from the data at each time point, and the weights remain consistent across all time points. The output layer provides the data after linear transformation for the subsequent LSTM layer.

[0019] Next, the LSTM layer receives the output of the CNN layer as input, selects the information useful for prediction from it, and integrates it into the new candidate cell state. This process can be described as follows:

[0020] I t = σ(W i [x t , h t-1 +b i )

[0021] where I t is the value of the input gate at the current time t, x t is the input at the current time t, h t-1 is the hidden state vector at the previous time t-1, σ is the activation function, W i is the weight matrix, and b i is the bias vector.

[0022] The LSTM layer determines which information should be retained and which should be discarded through the forget gate. The forget gate calculates a value between 0 and 1 based on the current input and the cell state at the previous time, which represents the degree of forgetting of each information unit. The output of the forget gate is multiplied by the cell state at the previous time to control the degree of information forgetting. The update process of the forget gate can be described as follows:

[0023] F t = σ(W f [x t , h t-1 +b f )

[0024] where F t is the value of the forget gate at the current time t, σ is the activation function, W f is the weight matrix of the forget gate, x t is the input vector at the current time t, h t-1 is the hidden state vector at the previous time t-1, and b f is the bias vector of the forget gate.

[0025] The output gate determines which information in the cell state needs to be passed to the hidden state at the next time or used as the output of the current time. These updated information will be passed to the next layer of the network or directly used as the output of the model. This process can be described as follows:

[0026] O t = σ(W o [x t ,h t-1 +b o )

[0027] where O t is the value of the output gate at the current time t, σ is the activation function, W o is the weight matrix of the output gate, x t is the input vector at the current time t, h t-1 is the hidden state vector at the previous time t-1, and b o is the bias vector of the output gate.

[0028] The update of the hidden state h t and the cell state C t can be expressed as follows:

[0029] h t = O t ⊙ tanh(C t )

[0030] C t = F t ⊙ C t-1 + I t ⊙ C′ t-1

[0031] where h t is the hidden state at the current time t, O t is the output gate, ⊙ is element-wise multiplication, tanh is the hyperbolic tangent function, C t is the cell state at the current time t, F t is the forget gate, C t-1 is the cell state at the previous time t-1, I t is the input gate, and C′ t-1 is the processed cell state at the previous time t-1.

[0032] After the LSTM layer, the output is connected to a fully connected layer, which maps the output of the LSTM layer to the target prediction value or the final neuron output. In this way, the fully connected layer converts the time series features extracted by the LSTM into specific prediction results.

[0033] Furthermore, in step S3 of the present invention, the GAN generation and augmentation of fault data include the following:

[0034] The GAN network mainly consists of a generator and a discriminator. Its objective function is as follows:

[0035]

[0036] where D is the discriminator and G is the generator, denotes taking the expectation, x is the real data, p data (x) is the real data distribution, z is the random Gaussian noise, p z (z) is the noise distribution, G(z) is the data generated by the generator based on the noise, D(G(z)) is the discrimination result of the discriminator on the generated data, and D(x) is the discrimination result of the discriminator on the real data.

[0037] The input of the generator network model is a random noise vector, which is mapped to an intermediate representation vector through a fully connected layer. The BN layer normalizes the output. The LeakyReLU activation function provides a non-zero gradient for negative input values, alleviating the vanishing gradient problem and maintaining the non-linearity of the network. Then it is connected to the second fully connected layer. The output layer uses the tanh activation function to map the generated data values to the specified output range.

[0038] The discriminator is used to distinguish between real data and fake data generated by the generator. The input layer of the discriminator receives samples from the generator or the real data set and extracts features through a fully connected layer. Subsequently, the LeakyReLU activation layer and the Dropout layer introduce non-linear transformation and regularization techniques respectively. The output layer generates a scalar value representing the probability that the model judges whether the input data is real data. The Sigmoid function compresses the output to the range of [0, 1], where being close to 1 indicates that the input data is real data, and being close to 0 indicates that the input data is generated data.

[0039] Subsequently, the output of the discriminator is fed back into the generator again to help the generator generate more real data. Through the adversarial training process, the generator and the discriminator are continuously optimized in the mutual game. When the objective function of the GAN is stable, the generator can generate more and more real data and is finally difficult to be accurately distinguished by the discriminator, thus achieving the global optimum.

[0040] Furthermore, in step S4 of the present invention above, the feature extraction and fault identification include the following:

[0041] The DNN fault detection network for multi-scale feature extraction combines the prediction results of the CNN-LSTM network and the fault data expanded during the GAN training process to deeply analyze the fault features, accurately identify the fault types, and finally output the detection results.

[0042] The DNN network contains a multi-branch structure. The input data is separately sent into three independent branches, and the sub-networks of each branch independently process the received input and perform fault identification. Each sub-network consists of multiple hidden layers, and the hidden layers perform feature extraction and learning through fully connected layers and Dropout layers. The fully connected layers capture the non-linear relationships of the input data, and the Dropout layers randomly discard some neurons during each training process to prevent the model from overfitting during training. Sub-networks with different hierarchical structures learn to capture data features at different levels. After being processed by the hidden layers, the fault identification results output by each sub-network will be connected through a merging layer to form a comprehensive feature representation, and further processed to obtain the final fault detection result. If a fault is detected, the system will give an early warning and output the fault identification result; if there is no fault, the system will return to the sampling window and continue to predict and detect the next batch of data.

[0043] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0044] 1. The present invention can address the problems of unbalanced hydropower station fault data and data anomalies. The solution uses a GAN to generate fault samples to expand the training data set and solve the small sample learning problem. At the same time, combined with the controllable credibility detection mechanism of the PCA algorithm, it analyzes the main component change direction of the data, screens out the credible data and excludes the outliers, thereby improving the robustness and accuracy of fault detection and reducing the influence of environmental and equipment errors on the detection results.

[0045] 2. The present invention adopts a structure combining CNN and LSTM. The CNN extracts the local features of the data, and the LSTM captures the long-term dependence relationships to process the complex time series features in the operation data of hydropower station equipment. Compared with traditional methods, the solution is more superior in dynamic and non-linear data analysis.

[0046] 3. The present invention adopts a DNN model based on multi-scale feature extraction. Multiple branch networks extract features at different scales, integrating local features and global information. This design improves the accuracy of fault classification and can comprehensively understand the complex patterns in the data more than traditional single models or shallow networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is the overall system model diagram of the present invention.

[0048] Figure 2 is the structural diagram of the CNN-LSTM network model of the present invention.

[0049] Figure 3 is the structural diagram of the multi-scale feature extraction fault detection network of the present invention.

[0050] Figure 4It is the structural diagram of the GAN generator and discriminator model of the present invention.

[0051] Figure 5 It is the comparison diagram of the ROC (Receiver Operating Characteristic Curve) of the embodiment of the present invention.

[0052] Figure 6 It is the comparison diagram of the predicted value and the true value of the embodiment of the present invention.

[0053] Among them, Figure 6 (a), Figure 6 (b), Figure 6 (c), Figure 6 (d) are the comparison diagrams of the predicted value and the true value of the embodiment of the present invention.

[0054] Figure 7 It is the scatter comparison diagram of the normal data sample and the abnormal data sample of the embodiment of the present invention.

[0055] Among them, Figure 7 (a), Figure 7 (b), Figure 7 (c) are the scatter comparison diagrams of the normal data sample and the abnormal data sample of the embodiment of the present invention. Detailed implementation manners

[0056] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0057] As Figure 1 shown, the present invention proposes a hydropower station fault prediction method for biased data, which improves the accuracy and robustness of data prediction, and improves the accuracy of fault classification, and has stronger fault detection ability. The specific steps are as follows:

[0058] Step S1: Analyze the mainstream data trend of the monitoring data based on the controllable credibility detection mechanism, calculate the credibility of the data under different monitoring times, and screen the data according to the credibility:

[0059] The credibility detection process includes data analysis and feature selection. The principal component analysis is used to identify the main change direction in the repeated monitoring data, calculate the component matrix P and the proportion of explained variance λ, project the data into a new coordinate system to obtain the direction with the largest variance in the data, that is, the principal component, and calculate the contribution degree of each group of monitoring data to the PCA result, so as to evaluate the credibility of the data:

[0060]

[0061] Among them, Ci is the contribution degree of the i-th group of monitoring data to the PCA result, n is the number of principal components, and P ji is the load of the j-th principal component on the corresponding variable of the i-th group of data, and λ j is the variance of the j-th principal component. According to these contribution degrees, the monitoring data of each group are sorted, and selection probabilities are assigned. The higher the contribution degree of the data, the greater the probability of being selected. This part is as Figure 1 shown in the first two steps shown, mainly to select normal data from all data before data processing.

[0062] Step S2: Normalize the filtered data set, perform time window sampling to capture the time series characteristics of the data, and the sampled data will be sent into the CNN-LSTM network to predict future time series data:

[0063] The specific network model structure of CNN-LSTM is as Figure 2 shown. Before data prediction, normalization and window sampling are performed first. The sampling window n is 30, which means that the network predicts the data at the 31st time point by learning the data of the first 30 time points.

[0064] The model parameters are as follows: the hidden layer size is 128, and the network structure contains 3 hidden layers. A total of 1000 rounds of iteration are performed during the training process, and each round is trained with a batch size of 32. The learning rate is set to 0.001, and the momentum parameters (0.9, 0.999) are used for optimization.

[0065] After data normalization preprocessing and time window sampling are completed, they are input into the CNN-LSTM network to extract time series characteristics and make predictions. The CNN-LSTM network structure mainly includes a CNN layer, an LSTM layer, and a fully connected layer. The CNN layer is used to extract local features of the data, the LSTM layer is used to capture time series characteristics of the data, and the final prediction result is output to the fully connected layer. The input data shape of this network is (batch size, channels, sequence length), where batch size is the batch size of 32, channels is the number of channels of the input data, that is, the feature dimension, and sequence length is the sequence length of 30.

[0066] In the CNN layer, the output feature map V is as follows:

[0067]

[0068] where z is the activation function, w i and x i represent the convolutional kernel weight and the input matrix respectively, b is the bias parameter, ∑ is the summation symbol, and i is the element index.

[0069] After the calculation of the convolutional layer, the obtained feature maps are successively concatenated into a vector and output to the fully connected layer. The convolution of the CNN layer weights the data at each time point, that is, extracts features from the data at each time point, and the weights remain consistent across all time points. The output layer provides the data after linear transformation for the subsequent LSTM layer.

[0070] Next, the LSTM layer receives the output of the CNN layer as input, selects the information useful for prediction from it, and integrates it into a new candidate cell state. This process can be described as follows:

[0071] I t = σ(W i [x t , h t-1 + b i )

[0072] where I t is the value of the input gate at the current time t, x t is the input at the current time t, h t-1 is the hidden state vector at the previous time t - 1, σ is the activation function, W i is the weight matrix, and b i is the bias vector. The shape of the LSTM layer is (num_layers, batch_size, hidden_size). Among them, there are three hidden layers, and each hidden layer contains 128 neurons.

[0073] The LSTM layer determines which information should be retained and which should be discarded through the forget gate. The forget gate calculates a value between 0 and 1 based on the current input and the cell state at the previous time, which represents the degree of forgetting of each information unit. The output of the forget gate is multiplied by the cell state at the previous time to control the degree of information forgetting. The update process of the forget gate can be described as follows:

[0074] F t = σ(W f [x t , h t-1 + b f )

[0075] where F t is the value of the forget gate at the current time t, σ is the activation function, W f is the weight matrix of the forget gate, x t is the input vector at the current time t, h t-1 is the hidden state vector at the previous time t - 1, and b f is the bias vector of the forget gate.

[0076] The output gate determines which information in the cell state needs to be passed to the hidden state at the next time step or be used as the model output at the current time step. These updated information will be passed to the next layer of the network or directly used as the model output. This process can be described as follows:

[0077] O t =σ(W o [x t ,h t-1 +b o )

[0078] Where O t is the value of the output gate at the current time step t, σ is the activation function, W o is the weight matrix of the output gate, x t is the input vector at the current time step t, h t-1 is the hidden state vector at the previous time step t - 1, and b o is the bias vector of the output gate.

[0079] The update of the hidden state h t and the cell state C t can be described as follows:

[0080] h t =O t ⊙tanh(C t )

[0081] C t =F t ⊙C t-1 +I t ⊙C′ t-1

[0082] Where h t is the hidden state at the current time step t, O t is the output gate, ⊙ is element-wise multiplication, tanh is the hyperbolic tangent function, C t is the cell state at the current time step t, F t is the forget gate, C t-1 is the cell state at the previous time step t - 1, I t is the input gate, and C′ t-1 is the processed cell state at the previous time step t - 1.

[0083] After the LSTM layer, the output is connected to a fully connected layer. The shape of this layer is (batch size, hidden size). The hidden layer has 128 neurons. The fully connected layer maps the output of the LSTM layer to the target prediction value or the final neuron output. In this way, the fully connected layer converts the time series features extracted by the LSTM into specific prediction results.

[0084] Step S3: Use a GAN network to generate new samples similar to real fault data in statistical characteristics to augment the data:

[0085] The GAN network mainly consists of a generator and a discriminator. The specific structure is as Figure 3 shown. The model parameters are as follows: the noise dimension is 100, and the output layer of the generator uses the tanh activation function. The loss function uses binary cross-entropy, and the learning rate is set to 0.0001. During the training process, 256 samples are used for each batch training, and the total number of training epochs is 100 rounds.

[0086] The objective function optimized by the GAN network is as follows:

[0087]

[0088] where D is the discriminator and G is the generator, denotes the expectation, x is the real data, p data (x) is the real data distribution, z is the random Gaussian noise, p z (z) is the noise distribution, G(z) is the data generated by the generator according to the noise, D(G(z)) is the discrimination result of the discriminator on the generated data, and D(x) is the discrimination result of the discriminator on the real data.

[0089] The generator model is a feed-forward neural network that generates data points with a distribution similar to the real data from random noise. The input of the generator network model is a random noise vector, which is mapped to a (7,7,128) vector through a fully connected layer. The BN layer normalizes the output, and the LeakyReLU activation function provides a non-zero gradient for negative input values, alleviating the vanishing gradient problem and maintaining the non-linear characteristics of the network. Then it is connected to the second fully connected layer, and the number of units in this layer is 256. The output layer uses the tanh activation function to map the generated data values to limit the output value range between [-1,1].

[0090] The discriminator is used to distinguish between real data and fake data generated by the generator. The input layer of the discriminator receives samples from the generator or the real data set and extracts features through a fully connected layer. Subsequently, the LeakyReLU activation layer and the Dropout layer introduce non-linear transformation and regularization techniques respectively. The output layer generates a scalar value representing the probability that the model judges whether the input data is real data. The Sigmoid function compresses the output to the range of [0,1], where a value close to 1 indicates that the input data is real data, and a value close to 0 indicates that the input data is generated data.

[0091] Subsequently, the output of the discriminator is fed back into the generator as input again to help the generator generate more realistic data. Through the adversarial training process, the generator and the discriminator are continuously optimized in the mutual game. When the objective function of the GAN is stable, the generator can generate increasingly realistic data and eventually becomes difficult to be accurately distinguished by the discriminator, thus achieving the global optimum.

[0092] Step S4: The multi-scale DNN network model extracts features and classifies the data predicted by the CNN-LSTM network, determines whether there is faulty data in the predicted data, and outputs the fault identification result. If no fault is detected, return to the sampling window to continue the prediction and examine the next batch of data.

[0093] The DNN fault detection network with multi-scale feature extraction combines the prediction results of the CNN-LSTM network and the faulty data augmented during the GAN training process to deeply analyze the fault features, accurately identify the fault type, and finally output the detection result. The network model structure is as Figure 4 shown, and the specific simulation experiment data includes the following:

[0094] The model parameters are as follows: the hidden layer size is 128, the activation function uses ReLU, and a Dropout rate of 0.2 is used in the hidden layer to prevent overfitting. The activation function of the output layer is Sigmoid, the optimizer adopts Adam (Adaptive Moment Estimation), and the learning rate is set to 0.0001. The loss function selects Binary Cross-Entropy Loss, the total number of training rounds is 100 rounds, and 16 samples are used for batch training in each round.

[0095] The DNN network contains a multi-branch structure. The input data is respectively fed into three independent branches. The sub-networks of each branch independently process the received input and perform fault identification. In this network, from top to bottom, the sub-networks respectively contain 5, 3, and 8 hidden layers. The hidden layers perform feature extraction and learning through fully connected layers and Dropout layers. The fully connected layers capture the non-linear relationships of the input data, and the Dropout layer randomly discards some neurons during each training process to prevent the model from overfitting during training. The sub-networks with different hierarchical structures learn to capture data features at different levels. After being processed by the hidden layers, the fault identification results output by each sub-network will be connected through the merging layer to form a comprehensive feature representation, and further processed to obtain the final fault detection result. If a fault is detected, the system will give an early warning and output the fault identification result; if there is no fault, the system will return to the sampling window and continue to predict and detect the next batch of data.

[0096] Figure 5The figure shows the comparison results of the ROC curves of the fault prediction solutions of the present invention and those based on RNN (Recurrent Neural Network), GRU (Gated Recurrent Unit), and XGBoost (Extreme Gradient Boosting). It can be seen that the fault prediction result of the present invention is better than other solutions, with high prediction and detection accuracy. Figure 6 The figure shows the comparison chart of the predicted values and the true values of the CNN-LSTM model in this solution, where Figure 6 (a) is the comparison chart of data sample 1, Figure 6 (b) is the comparison chart of data sample 2, Figure 6 (c) is the comparison chart of data sample 3, Figure 6 (d) is the comparison chart of data sample 4. It can be seen from the figure that the prediction results are in good agreement with the true data, and this consistency reflects the good performance of the model in the prediction task. Figure 7 The figure shows the scatter plot comparison chart of normal data samples and abnormal data samples in this solution, where Figure 7 (a) and Figure 7 (b) present the scatter distribution diagrams of the normal monitoring data obtained after repeated monitoring of data sample 5 and data sample 6, Figure 7 (c) is the scatter distribution diagram of abnormal monitoring data. It can be clearly seen that the concentration trend of abnormal data samples is different from that of normal data samples. In the present invention, the data credibility detection part will screen and eliminate the abnormal data samples as shown in the figure to ensure the accuracy and reliability of the subsequent analysis results.

[0097] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A hydropower station fault prediction method for biased data, characterized in that, It includes the following steps: Step S1: Analyze the mainstream data trend of the monitoring data based on a controllable credibility detection mechanism, calculate the credibility of the data under different monitoring times, and screen the data according to the credibility; Step S2: Perform normalization processing on the screened data set, perform time window sampling to capture the time series characteristics of the data. The sampled data will be fed into a CNN-LSTM (Convolutional Neural Network-Long Short-Term Memory) network to predict future time series data. The capturing of the time series characteristics of the data and the prediction of future time series data include the following content: After the data normalization preprocessing and time window sampling are completed, it is input into the CNN-LSTM network to extract time series characteristics and make predictions. The CNN-LSTM network structure includes a CNN layer, an LSTM layer, and a fully connected layer. The CNN layer is used to extract local characteristics of the data, the LSTM layer is used to capture the time series characteristics of the data, and the final prediction result is output to the fully connected layer; In the CNN layer, the output feature map V is as follows: where z is the activation function, w i and x i represent the convolutional kernel weights and the input matrix respectively, b is the bias parameter, ∑ is the summation symbol, and i is the element index; After being calculated by the convolutional layer, the obtained feature maps are sequentially connected into a vector and output to the fully connected layer. The convolution of the CNN layer weights the data at each time point, that is, extracts features from the data at each time point. The weights remain consistent at all time points. The output layer provides the data after linear transformation for the subsequent LSTM layer; The LSTM layer receives the output of the CNN layer as input, selects the information useful for prediction from it, and integrates it into a new candidate cell state. This process can be described as follows: I t = σ(W i [x t , h t-1 + b i ) where I t is the value of the input gate at the current time t, x t is the input at the current time t, h t-1 is the hidden state vector at the previous time t-1, σ is the activation function, W i is the weight matrix, b i is the bias vector; The LSTM layer determines which information should be retained and which should be discarded through a forget gate. The forget gate calculates a value between 0 and 1 based on the current input and the cell state at the previous moment. This value represents the forgetting degree of each information unit. The output of the forget gate is multiplied by the cell state at the previous moment to control the forgetting degree of the information. The update process of the forget gate can be described as follows: F t = σ(W f [x t , h t-1 + b f ) where F t is the value of the forget gate at the current time t, σ is the activation function, W f is the weight matrix of the forget gate, x t is the input vector at the current time t, h t-1 is the hidden state vector at the previous time t - 1, b f is the bias vector of the forget gate; The output gate determines which information in the cell state needs to be passed to the hidden state at the next moment or used as the output of the current moment. These updated information will be passed to the next layer of the network or directly used as the output of the model. This process can be described as follows: O t = σ(W o [x t ,h t-1 + b o ) where O t is the value of the output gate at the current time t, σ is the activation function, and W o is the weight matrix of the output gate, x t is the input vector at the current time t, h t-1 is the hidden state vector at the previous time t - 1, and b o is the bias vector of the output gate; Hidden state h t and cell state C t can be updated as follows: h t = O t ⊙tanh(C t ) C t = F t ⊙C t-1 + I t ⊙C t ' -1 where h t is the hidden state at the current time t, O t is the output gate, ⊙ is element-wise multiplication, tanh is the hyperbolic tangent function, C t is the cell state at the current time t, F t is the forget gate, C t-1 is the cell state at the previous time t - 1, I t is the input gate, C t ' -1 is the processed cell state at the previous time t - 1; After the LSTM layer, the output is connected to the fully connected layer. The fully connected layer maps the output of the LSTM layer to the target prediction value or the final neuron output. The fully connected layer converts the time series characteristics extracted by the LSTM into specific prediction results; Step S3: Use a GAN (Generative Adversarial Networks) network to generate new samples similar to the real fault data in statistical characteristics to expand the data; Step S4: The multi-scale DNN (Deep Neural Networks) network model extracts features and classifies the data predicted by the CNN-LSTM network, determines whether there is faulty data in the predicted data, and outputs the fault identification result. If no fault is detected, it returns to the sampling window to continue the prediction and examines the next batch of data.

2. The method for predicting hydropower station faults facing biased data according to claim 1, wherein, In the said step S1, the data credibility detection includes the following: The credibility detection process includes data analysis and feature selection. Principal component analysis is used to identify the main change direction in the repeatedly monitored data, calculate the component matrix P and the proportion of explained variance λ, project the data into a new coordinate system to obtain the direction with the largest variance in the data, that is, the principal component, and calculate the contribution degree of each group of monitored data to the PCA (Principal Component Analysis) result, so as to evaluate the credibility of the data: Among them, C i is the contribution degree of the i-th group of monitoring data to the PCA result, n is the number of principal components, and P ji is the load of the j-th principal component on the corresponding variable of the i-th group of data. λ j is the variance of the j-th principal component. According to these contribution degrees, the monitoring data of each group are sorted and selection probabilities are assigned. The higher the contribution degree of the data, the greater the probability of being selected.

3. A hydropower station fault prediction method for biased data according to claim 1, characterized in that, In the said step S3, the GAN generation and augmentation of faulty data includes the following: The GAN network consists of a generator and a discriminator, and its objective function is as follows: Among them, D is the discriminator and G is the generator. denotes the expectation, x is the real data, p data (x) is the real data distribution, z is the random Gaussian noise, p z (z) is the noise distribution, G(z) is the data generated by the generator according to the noise, D(G(z)) is the discrimination result of the discriminator on the generated data, and D(x) is the discrimination result of the discriminator on the real data; The input of the generator network model is a random noise vector, which is mapped to an intermediate representation vector through the fully connected layer. The BN (Batch Normalization layer) layer normalizes the output. The LeakyReLU (Leaky Rectified Linear Unit) activation function provides a non-zero gradient for negative input values, alleviates the gradient vanishing problem and maintains the non-linear characteristics of the network. Then it is connected to the second fully connected layer, and the output layer uses the tanh (Hyperbolic Tangent Function) activation function to map the generated data value to the specified output range; The discriminator is used to distinguish between real data and fake data generated by the generator. The input layer of the discriminator receives samples from the generator or the real data set, and extracts features through the fully connected layer. The LeakyReLU activation layer and the Dropout (inactivation layer) layer introduce non-linear transformation and regularization techniques respectively. The output layer generates a scalar value, indicating the probability judgment of the model on whether the input data is real data. The Sigmoid (S-shaped function) function compresses the output to the range of [0,1], where approaching 1 indicates that the input data is real data, and approaching 0 indicates that the input data is generated data; The output of the discriminator is input back into the generator as feedback to help the generator generate more real data.

4. A method for predicting hydropower station faults facing biased data according to claim 1, characterized in that, In the said step S4, the feature extraction and fault identification includes the following: The DNN fault detection network for multi-scale feature extraction combines the prediction results of the CNN-LSTM network and the faulty data augmented during the GAN training process, deeply analyzes the fault features, accurately identifies the fault type, and finally outputs the detection result; The DNN network contains a multi-branch structure. The input data is respectively sent into three independent branches. The sub-networks of each branch independently process the received input and perform fault identification. Each sub-network consists of multiple hidden layers. The hidden layers perform feature extraction and learning through fully connected layers and Dropout layers. The fully connected layers capture the non-linear relationships of the input data. The Dropout layer randomly discards some neurons during each training process to prevent the model from overfitting during training. The sub-networks with different hierarchical structures learn to capture data features at different levels. After being processed by the hidden layers, the fault identification results output by each sub-network will be connected through a merging layer to form a comprehensive feature representation, and further processed to obtain the final fault detection result. If a fault is detected, the system will give an early warning and output the fault identification result; if there is no fault, the system will return to the sampling window and continue to predict and detect the next batch of data.