Power distribution network voltage quality prediction method based on improved CNN-LSTM neural network

By using an improved CNN-LSTM neural network, combined with Pearson correlation coefficient analysis, attention mechanism and Monte Carlo method, the problems of data stability and feature capture in distribution network voltage quality prediction are solved, and high-precision and robust voltage quality prediction is achieved.

CN120995193APending Publication Date: 2025-11-21STATE GRID HUBEI ELECTRIC POWER RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510957688.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing voltage quality prediction methods suffer from high data quality dependence, poor stability, and poor interpretability when dealing with distribution network voltage quality data that is highly random and strongly time-series-based. Furthermore, they struggle to effectively capture multi-scale features, resulting in poor prediction accuracy.

Method used

An improved CNN-LSTM neural network is adopted, and features are selected through Pearson correlation coefficient analysis. Combining convolutional neural network (CNN) and long short-term memory neural network (LSTM), attention mechanism and Monte Carlo method are introduced to construct confidence intervals. L2 regularization and voltage limit constraints are incorporated to optimize the loss function to improve prediction accuracy and robustness.

Benefits of technology

It effectively improves the accuracy and stability of distribution network voltage quality prediction, reduces the risk of model overfitting, improves prediction accuracy and computational efficiency, and provides more reliable prediction results and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995193A_ABST
    Figure CN120995193A_ABST
Patent Text Reader

Abstract

The invention provides a power distribution network voltage quality prediction method based on an improved CNN-LSTM neural network, and the method is used for power distribution network voltage quality prediction through a convolutional neural network-long and short term memory neural network hybrid model of an attention mechanism. Multi-dimensional spatial-temporal feature collaborative mining is carried out by fusing CNN and LSTM, an attention mechanism module is introduced to focus key period features, L2 regularization and voltage limit constraint are fused to improve a loss function, Monte Carlo Dropout is added to generate a 95% confidence interval, it is ensured that a predicted value conforms to a physical boundary while overfitting is suppressed, and the prediction accuracy is improved. And the prediction precision and the convergence speed are effectively improved. Through actual measurement data analysis and comparison, the advantages of the model provided by the invention in accuracy and calculation efficiency compared with similar methods are verified.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the field of power distribution network voltage quality prediction, and particularly relates to a power distribution network voltage quality prediction method based on an improved CNN-LSTM neural network. BACKGROUND

[0002] Voltage quality prediction is a quantitative estimation of the future trend of the effective value of the grid voltage to identify voltage quality problems such as voltage deviation, fluctuation and three-phase imbalance in advance. Accurate and efficient voltage quality prediction on the power distribution network side can provide more accurate scheduling basis for power system operators, thereby better meeting the power consumption behavior of users and improving the operation efficiency of the power grid. Research shows that every 1% improvement in the accuracy of power distribution network voltage quality prediction can save up to 10 million pounds of energy costs for the British power system every year. In recent years, the power generation and power consumption characteristics of the power distribution system in China have shown a trend of diversification and complication. A large number of distributed power sources and multiple types of loads are connected to the power distribution network, and the randomness, intermittency and nonlinearity of their operation lead to more complex voltage fluctuations on the power distribution network side, and the voltage quality problems such as high and low voltage at the end and harmonics caused thereby are increasingly serious, and the difficulty of voltage quality prediction is increasing. The traditional prediction method has been difficult to meet the prediction needs of the current power distribution network voltage quality in terms of prediction accuracy and speed.

[0003] The current voltage quality prediction methods are mainly divided into two categories, namely the method based on physical modeling and the method based on data-driven. The method based on physical modeling provides a solid theoretical basis for voltage quality prediction, and can accurately process the static analysis of the power network and the voltage distribution under specific working conditions, but when facing large-scale, high-dimensional and dynamically changing data, the prediction ability of the physical model is limited, especially in dealing with complex nonlinear relationships and time series dynamic changes.

[0004] The data-driven method is the development trend of voltage quality prediction in recent years, and its core lies in using neural network models to mine the voltage spatio-temporal variation law from massive data. The existing neural network-based methods have improved the accuracy of voltage quality prediction to some extent, but when dealing with the current high randomness and strong time series of power distribution network voltage quality data, there are still challenges such as high data quality dependence, poor stability and poor interpretability.

[0005] The hybrid prediction model composed of multiple single prediction models is an effective way to solve the stability, accuracy and interpretability problems of traditional models. The hybrid model can couple two or more neural networks together to complement each other's advantages. However, when selecting and designing a hybrid model to deal with time series prediction problems such as voltage quality prediction, the following three difficulties still need to be faced.

[0006] First, the problem of high-dimensional feature processing. Voltage quality is usually affected by historical voltage, weather and other multi-source data. How to reduce the dimension of high-dimensional data without losing key information, while avoiding overfitting of the model when dealing with sparse or noisy data, is the main strategy when selecting and designing feature processing models. If the key information cannot be balanced during the dimension reduction process, it is easy to cause model overfitting or feature distortion.

[0007] Second, the problem of feature correlation analysis. There may be a certain degree of linear or nonlinear relationship between different features, such as voltage and temperature. Accurate allocation of feature weights during model initialization or optimization parameters plays a positive role in improving prediction ability and speed. If the model fails to accurately allocate feature weights, it will not only reduce training efficiency, but also weaken the recognition of key influencing factors and increase the risk of prediction bias.

[0008] Third, the problem of capturing time dependence. There are trends, cycles and mutations in voltage data, and traditional models may not be able to capture multi-scale features, resulting in poor prediction accuracy. Especially in medium and long-term prediction tasks, if the model has weak ability to capture time dependence, it is likely to forget the influence of previous historical features, resulting in underfitting problem of the model.

[0009] Therefore, it is urgent to develop a voltage quality prediction method with high accuracy and high robustness. SUMMARY

[0010] To solve the above problems, the purpose of the present application is to provide an improved CNN-LSTM neural network-based power distribution network voltage quality prediction method with high accuracy and high robustness.

[0011] The technical scheme adopted by the present application is as follows:

[0012] An improved CNN-LSTM neural network-based power distribution network voltage quality prediction method, comprising the following steps:

[0013] Step one, select composite feature data, introduce Pearson correlation coefficient to analyze the influence of different features on voltage, and construct input matrix based on linear interpolation and sliding window method;

[0014] Step two, based on convolutional neural network CNN and long short-term memory neural network LSTM, CNN-LSTM power distribution network quality prediction model;

[0015] Step three, integrate attention mechanism into long short-term memory neural network LSTM, based on the hidden state sequence output by LSTM, generate attention score through dot product of query matrix and hidden state, and generate context vector as subsequent probability prediction input;

[0016] Step four, introduce Monte Carlo method to construct confidence interval, integrate L2 regularization and voltage limit constraint to improve the loss function of CNN-LSTM power grid quality prediction model, and use the improved CNN-LSTM power grid quality prediction model to predict voltage quality.

[0017] Further, step one specifically includes:

[0018] 1.1 Composite feature data selection: select voltage data V, weather data and calendar effect data, wherein the weather data adopts weather forecast data superimposed with normal distribution noise;

[0019] 1.2 Pearson correlation analysis: calculate the Pearson correlation coefficient of the composite feature data and the voltage fluctuation, and select the features significantly related to the voltage fluctuation as the model input;

[0020] 1.3 Data preprocessing:

[0021] Missing value processing: the voltage data is linearly interpolated based on the mean value of the previous and next 3 hours, and the missing value of solar irradiance is set to 0;

[0022] Abnormal value processing: delete the abnormal value which is more than 15% different from the neighborhood mean value and then re-insert it;

[0023] Normalization processing: adopt Min-Max method to map the data to [0, 1] interval, save the minimum and maximum parameters for subsequent prediction result inverse transformation;

[0024] 1.4 Sample generation: generate input matrix based on sliding window.

[0025] Further, the weather data includes temperature T, humidity H, wind speed W and solar irradiance I, and the calendar effect data includes weekend label R, summer label S and holiday label F; the voltage fluctuation is defined as the absolute difference between the last half sampling value of the historical voltage and its average value.

[0026] Further, the window size of the sliding window is 3 days, the dimension of the input matrix is 144x8x1, the input of the input matrix is the voltage, weather and calendar data of the previous three days, and the output is the fourth day voltage data.

[0027] Further, the calculation formula of the Pearson correlation coefficient is as follows:

[0028]

[0029] Wherein, x i , y i represent the feature value and the voltage fluctuation respectively, and respectively represent the mean value of the eigenvalue and the voltage fluctuation; the voltage fluctuation is equal to the absolute value of the historical voltage data of the last half of the sampling data minus the mean value thereof; when r = 1, it indicates that the feature and the voltage fluctuation are in complete positive correlation; when r = -1, it indicates that the feature and the voltage fluctuation are in complete negative correlation; when r = 0, it indicates that the feature and the voltage fluctuation have no correlation.

[0030] Further, the step two comprises:

[0031] A CNN network processing high-dimensional features is built to perform multiple convolution and pooling on the input matrix constructed in step one, extract high-dimensional features and output the feature map after dimension reduction; when building the CNN network, five batches are set, the samples of each batch are the same, and the dimension of the complete input matrix is 5*144*8*1; both the convolution and the pooling adopt the Tensorflow architecture, the convolution kernel is uniformly selected as a 3*3 dimensional matrix, the pooling kernel matrix is 2*2 in dimension, and the Adam optimizer is adopted for parameter optimization;

[0032] An LSTM network is built to capture time dependence, the LSTM network contains four layers, the first layer and the second layer are both 128 hidden units, the third layer and the fourth layer are both 64 hidden units; the output matrix dimension of each layer changes with the number of hidden units; after passing through the four-layer LSTM network, each hidden state in the output matrix contains the influence information of the corresponding time on the future voltage quality prediction.

[0033] Further, the transfer formula of the gating mechanism of the LSTM network is as follows:

[0034]

[0035]

[0036]

[0037]

[0038]

[0039] wherein, represents a convolution operation, represents a multiplication operation, represents a sigmoid activation function, tanh represents a hyperbolic tangent function, i t represents an input gate state, f t represents a forget gate state, C t represents a storage cell tensor, o t represents an output gate state, x t represents an input tensor at time t, h tdenotes the hidden state tensor; W f , W i , W C , W o are the weights of the forget gate, input gate, state gate and output gate respectively, b f , b i , b C , b o are the corresponding biases.

[0040] Further, the step three is specifically as follows:

[0041] 1) Calculate attention weight

[0042] The weight of each time step is determined by its attention score, and the calculation formula of the attention score is as follows:

[0043]

[0044] Where h t is the LSTM output of the t-th time step, W h is the weight matrix, and b h is the bias;

[0045] For the attention scores of all time steps, the softmax function is used to normalize them into the form of probability, so that the sum of the attention weights of all time steps is 1:

[0046]

[0047] Where α t is the attention weight of the t-th time step, e t is the attention score of the corresponding time step, and T is the total length of the sequence;

[0048] 2) Weighted summation

[0049] After calculating the attention weights of each time step, the output of the LSTM network is weighted and summed to obtain the final context vector; the context vector contains information of all time steps, and the information weights of different time steps are different; the calculation formula of the context vector c is as follows:

[0050]

[0051] Finally, the context vector c weighted by the attention mechanism will be transmitted to the subsequent probability prediction network through the fully connected layer to continue the prediction task, and the output of the probability prediction network is the weighted result based on the importance of each time step in the time series.

[0052] Further, the Monte Carlo method is introduced in step four to construct the confidence interval, and the L2 regularization and voltage limit constraint are integrated to improve the loss function of the CNN-LSTM power grid quality prediction model, which specifically includes:

[0053] Monte Carlo probability prediction: enable MC Dropout in the LSTM inference stage, output the prediction mean, variance and 95% confidence interval. When the confidence interval is narrow, the prediction result is stable and the decision confidence is high. Conversely, when the confidence interval is wide, related control measures should be handled with caution. The probability prediction network with MC Dropout is constructed as follows: for the same input x, the model performs N forward propagations in the inference stage, and each time a different prediction value is obtained , i = 1, 2, …, N, the mean and variance of these samples are calculated to describe the uncertainty of the prediction,

[0054]

[0055]

[0056] Using the properties of normal distribution, further construct the 95% confidence interval:

[0057]

[0058] Improved loss function: define the loss function to include the MSE error term, L2 regularization term and voltage limit penalty term, and the specific calculation formula of the loss function is as follows:

[0059]

[0060] wherein, represents the weight coefficient of the L2 regularization term, represents the weight coefficient of the voltage limit constraint, represents the voltage prediction value of the i-th sample, V i represents the actual voltage value of the i-th sample, N represents the total number of samples, V max and V min represent the upper and lower limit values of the voltage, respectively.

[0061] Further, the improved CNN-LSTM power grid quality prediction model is used to predict the voltage quality in step four, which specifically includes:

[0062] According to step one, process the actual grid data and weather data and use the sliding window to construct the input matrix;

[0063] The processed feature data is input into the built CNN network for processing high-dimensional features, and the high-dimensional features are subjected to multiple convolution and pooling to obtain a feature matrix after dimension reduction;

[0064] The feature matrix after dimension reduction is input into an LSTM network to capture the time dependence between features, and a preliminary prediction value is obtained;

[0065] The preliminary prediction value is input into an attention mechanism module coupled after the LSTM network to reflect the feature influence weight of the corresponding time step in real time;

[0066] The Monte Carlo method is introduced to generate the mean, variance and 95% confidence interval of the prediction value, and L2 regularization and voltage limit constraint are integrated to improve the loss function and optimize the voltage quality prediction result.

[0067] The hybrid model fusing the convolutional neural network and the long short-term memory neural network is constructed, and the attention mechanism with strong explanation is introduced for prediction, so that the shortcomings of single model in stability and generalization ability are effectively solved; the multi-dimensional space-time feature is co-mined by fusing CNN and LSTM, the attention mechanism module is introduced to focus on the key period feature, the L2 regularization and voltage limit constraint are integrated to improve the loss function, and the Monte Carlo Dropout is added to generate the 95% confidence interval, so that the overfitting is inhibited and the prediction value meets the physical boundary, the prediction accuracy and convergence speed are effectively improved. Through the analysis and comparison of the measured data, it is verified that the model compared with the similar method has the advantages of accuracy and calculation efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0068] Figure 1 A flow chart of a power distribution network voltage quality prediction method based on an improved CNN-LSTM neural network is provided for the embodiment of the present application.

[0069] Figure 2 A neural network architecture diagram of the power distribution network voltage quality prediction provided for the embodiment of the present application.

[0070] Figure 3 A comparison diagram of the power distribution network voltage quality prediction result provided for the embodiment of the present application.

[0071] Figure 4 A power distribution network voltage quality prediction attention weight thermal map provided for the embodiment of the present application. DETAILED DESCRIPTION

[0072] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present application.

[0073] As shown in Figure 1 and Figure 2 , the embodiments of the present application provide a power distribution network voltage quality prediction method based on an improved CNN-LSTM neural network, comprising the following steps:

[0074] Step 1.1: Composite feature data selection and description

[0075] The selected data includes three parts: voltage data, weather data and calendar effect data. The voltage data is the historical voltage data (V) of the power distribution network. The weather data includes temperature (T), humidity (H), wind speed (W) and solar irradiance (I). These four data are relatively important influencing factors in the voltage quality prediction of the power distribution network with photovoltaic output.

[0076] In addition, when studying the user's power consumption behavior, it is found that the voltage data on weekends fluctuates more than on weekdays, the overall voltage data in summer fluctuates more than in other seasons, and the voltage data on holidays fluctuates more than on ordinary days. Therefore, three labels are set to distinguish whether a sample belongs to the weekend (R), summer (S) and holiday (F) (0 means no, 1 means yes). Finally, the input data includes eight features: one voltage data, four weather data and three calendar label data.

[0077] In practice, the real weather data of the next day cannot be known, so only the weather forecast data can be used to replace the real weather data for prediction. It is worth noting that the weather forecast data is not included in the training data and test data, only the real historical weather data is included. Therefore, the following method can be used to simulate the weather forecast data:

[0078] (1)

[0079] where W d and W d ´ represent the real weather data and the weather forecast data of the dth day respectively, N represents the normal distribution noise, and P represents the proportion of adding noise.

[0080] Therefore, each sample is converted into a matrix in the following form:

[0081] (2)

[0082] Step 1.2: Pearson correlation analysis of feature data

[0083] In order to better understand the fluctuation law of power distribution network voltage and its influencing factors, correlation analysis is performed on 8 features in the data set. The Pearson correlation coefficient is used to represent the degree of correlation between each feature and voltage fluctuation. The calculation formula of Pearson correlation coefficient is as follows:

[0084] (3)

[0085] Where x i , y i represent the feature value and voltage fluctuation, and represent the average value of the feature value and voltage fluctuation. Voltage fluctuation is equal to the absolute value of the historical voltage data minus its average value of the latter half of the sampling data. When r = 1, it means that the feature and voltage fluctuation are completely positively correlated; when r = -1, it means that the feature and voltage fluctuation are completely negatively correlated; when r = 0, it means that the feature and voltage fluctuation have no correlation. According to the correlation analysis, there is a high positive correlation between historical voltage and voltage fluctuation, indicating that voltage fluctuation has strong time continuity. Solar irradiance and voltage have a certain positive correlation, especially when solar irradiance is high during the day, photovoltaic output is high, which has a greater impact on voltage stability. The correlation between temperature and voltage fluctuation is relatively complex, and generally shows a strong positive correlation. The correlation between humidity and wind speed and voltage fluctuation is relatively weak, but in extreme weather conditions (such as high humidity or strong wind), it may have a certain impact on the operation of power grid equipment and voltage quality.

[0086] Step 1.3: Feature data processing based on linear interpolation

[0087] 1) Missing value processing

[0088] In the original data, there are a small number of missing values and outliers in the voltage data and weather data. Through cross-validation, it is found that the linear interpolation method for processing missing values has relatively optimal effect and algorithm speed, so the linear interpolation method is selected,

[0089] (4)

[0090] Where, V1, V´, V2 represent the average value of the voltage data of the previous three hours, the voltage data of the interpolation point, and the average value of the voltage data of the next three hours, respectively, t1, t´, t2 represent the middle time point of the previous three hours, the interpolation point time, and the middle time point of the next three hours, respectively, so t2-t1=3. The weather data first needs to delete the repeated value, and the missing value of solar irradiance is directly set to 0. Because in practical application, the solar irradiance in rainy weather and at night is very small, and the photovoltaic output is almost zero.

[0091] 2) Abnormal value processing

[0092] Abnormal values are found by data traversal algorithm. When the voltage data or weather data at a certain place is 15% different from the average value of other data in the neighborhood (selected for 3 hours before and after), it is considered as an abnormal value. Finally, the abnormal value needs to be deleted and the data is supplemented again by linear interpolation.

[0093] 3) Normalization processing

[0094] Normalization processing of data can reduce the difference between data, thereby simplifying the calculation amount and improving the convergence speed of the model. According to the characteristics of the feature data, the Min-Max normalization method is used to process all data with dimensions, which can linearly map the same kind of data to the range of [0, 1], and the calculation formula is as follows:

[0095] (5)

[0096] Where, x, x min , x max represent the original data value, the minimum value in the data, and the maximum value in the data value, respectively, and x´ represents the normalized value. All data is replaced by its normalized value.

[0097] At this point, all steps of data processing are completed.

[0098] Step 1.4: Sample generation based on sliding window

[0099] The method of sliding window is selected to generate input samples and output samples. Through comparative experiments, it is found that the prediction effect of selecting 3-day sliding window is better than that of 4-day and 5-day, and it is better in balancing between capturing long-term dependence and maintaining data amount. Therefore, the voltage data, weather data and calendar data of the previous three days are selected as the initial input, and the initial output is the voltage data of the fourth day, which constitutes a complete data input-output sample. The input window slides forward one day, and the output window also slides forward one day, so a total of 361 samples can be formed.

[0100] By constantly changing the generation dimension of the sample, it is found that when combining three days of samples into one large matrix, the training time can be saved by 10% to 15% compared with inputting the three-day sample matrix separately. Therefore, the input matrix dimension is 144*8*1, and the dimension "1" here is the number of channels added to adapt to the subsequent CNN network. The output matrix dimension is 48*1. The specific form of the first two dimensions X of the input matrix is as follows:

[0101] (6)

[0102] The output matrix O is

[0103] (7)

[0104] Where, V d,n represents the nth data of the dth day, and the superscript pred represents the predicted value of the voltage.

[0105] Step 2.1: Local feature extraction and dimension reduction modeling based on CNN

[0106] When constructing the CNN network, the present application sets 5 batches, and the samples in each batch are exactly the same. Therefore, the dimension of the complete input matrix is 5*144*8*1. Both times of convolution and pooling adopt the Tensorflow architecture, the convolution kernel is uniformly selected as a 3*3 dimension matrix, the pooling kernel matrix dimension is 2*2, and the Adam optimizer is used for parameter optimization.

[0107] First convolution:

[0108] The number of convolution kernels is 32, the step is 1, the same padding method is selected to keep the original dimension, and the RELU activation function is used, and the convolution operation is performed by using the following formula:

[0109] (8)

[0110] (9)

[0111] Where, W represents the convolution kernel weight, b represents the bias, and * represents the convolution operation. The RELU activation function image used by the network is as shown in Figure 1 . The output matrix dimension obtained by the first layer is 5*144*8*32, that is, the time step and feature number of the original matrix remain unchanged, and the channel number increases to 32.

[0112] First pooling:

[0113] The maximum pooling is adopted, and the step is 2. The final output matrix dimension is 5*72*4*32. As can be seen, the data dimension is reduced by the maximum pooling operation, the calculation amount of the model is reduced, and the most significant features are extracted.

[0114] Second convolution:

[0115] The number of convolution kernels is increased to 64 compared to the first convolution, and the step, padding method and activation function remain unchanged. The convolution operation is still performed according to formulas (8)~(9). The final output matrix has a dimension of 5*72*4*64.

[0116] Second pooling:

[0117] The method and step of pooling are the same as the first pooling. Through the second pooling, the time step and the number of features are further reduced. The final output matrix has a dimension of 5*36*2*64.

[0118] In order to meet the requirements of LSTM for input data and simplify the calculation amount, the above output matrix is flattened to obtain a flattened matrix with a dimension of 5*4608, and this feature matrix is taken as the input of the subsequent LSTM network.

[0119] As can be seen, through the CNN network for dimension reduction and extraction of features, the problem of model overfitting is alleviated, which is more conducive to the subsequent LSTM network for prediction tasks. The specific steps and methods of constructing the CNN network of the present application are similar to those of the classical CNN network.

[0120] Step 2.2: Time series dependence modeling based on LSTM

[0121] Long short-term memory neural network (LSTM) has absolute advantages in processing time series voltage data. Compared with traditional recurrent neural networks, LSTM introduces a special gating mechanism, which can effectively preserve and transmit information in a long time range and alleviate the problem of gradient disappearance or gradient explosion. Its transmission formula is as follows:

[0122] (10)

[0123] (11)

[0124] (12)

[0125] (13)

[0126] (14)

[0127] wherein, represents convolution operation, represents multiplication operation, represents sigmoid activation function, tanh represents hyperbolic tangent function, i t represents input gate state, f tdenotes the forget gate state, C t denotes the storage cell tensor, o t denotes the output gate state, x t denotes the input tensor at time t, h t denotes the hidden state tensor. W f , W i , W C , W o are the weights of the forget gate, input gate, state gate and output gate respectively, b f , b i , b C , b o are all the corresponding biases. The working principle of the four gates is as follows:

[0128] Forget gate f: decides which information to discard from the cell state. It generates a value between 0 and 1 through a sigmoid function, indicating the degree of retention of each state value.

[0129] Input gate i: consists of two parts: a sigmoid layer decides which values will be updated, and a tanh layer generates a new candidate value vector. The output of the sigmoid layer and the tanh layer of the input gate are multiplied to obtain the updated candidate value.

[0130] State gate C: is the core of LSTM, which carries the information of previous time steps. The update of the cell state is obtained by adding the output of the forget gate and the output of the input gate.

[0131] Output gate o: decides the value of the next hidden state. It decides which cell state will be output through a sigmoid layer, and then generates a candidate value for the output state through a tanh layer, and finally combines the two parts to form the final output.

[0132] The LSTM network constructed by the present application contains four layers. The construction method of the four layers is the same, except that the number of hidden units is different. The first layer and the second layer are both 128 hidden units, and the third layer and the fourth layer are both 64 hidden units. The output matrix dimension of each layer changes with the number of hidden units. Finally, after passing through the four-layer LSTM network, the output matrix dimension is 5*64. Each hidden state in the output matrix contains the influence information of the corresponding time on the future voltage quality prediction.

[0133] Step 3: Attention mechanism-based explainability enhancement algorithm

[0134] Since the voltage data often has a relatively long historical dependence, the sensitivity of traditional LSTM to distant key time steps is limited, which may lead to information loss problems. In addition, the contribution of each feature to the voltage varies at different time steps. The introduction of attention mechanism can dynamically calculate the importance weight of each feature at different time steps, so that the model can focus more on the data that plays a key role in the prediction result, thereby improving the prediction accuracy. At the same time, the weight matrix generated by the attention mechanism provides an intuitive basis for analyzing the contribution of input data, enhancing the interpretability of the model.

[0135] Without attention mechanism, the LSTM network usually directly outputs the final hidden state as the final prediction result. After introducing the attention mechanism, the importance of each time step in the final prediction can be dynamically calculated according to its hidden state.

[0136] After the attention mechanism is integrated into the four-layer LSTM network, the attention score is generated by the dot product of the query matrix and the output matrix corresponding to all features after the fourth-layer LSTM network outputs the hidden state sequence. Finally, the context vector is obtained after normalization and weighted summation, which is used for subsequent prediction tasks. The specific steps are as follows:

[0137] 1) Calculate attention weight

[0138] The weight of each time step is determined by its attention score. The calculation formula of the attention score is as follows:

[0139] (15)

[0140] where h t is the LSTM output at the t-th time step, W h is the weight matrix, and b h is the bias.

[0141] For all time steps, the softmax function is used to normalize them into the form of probability, so that the sum of the attention weights of all time steps is 1. The purpose of this step is to convert the score of each time step into a weight representing its importance:

[0142] (16)

[0143] where α t is the attention weight of the t-th time step, e t is the attention score of the corresponding time step, and T is the total length of the sequence.

[0144] 2) Weighted summation

[0145] After calculating the attention weights of each time step, the output of the LSTM network is weighted and summed to obtain the final context vector. This vector contains information from all time steps, but the information weight of different time steps is different. The calculation formula of the context vector c is as follows:

[0146] (17)

[0147] Finally, the context vector c weighted by the attention mechanism is transmitted to the subsequent probability prediction network through the fully connected layer to continue the prediction task. The output is the weighted result based on the importance of each time step in the time series, providing more accurate and interpretable prediction guarantee for voltage quality prediction.

[0148] Step 4.1: Probability prediction modeling based on Monte Carlo method

[0149] In the voltage quality prediction task, similar to the traditional deterministic point estimation prediction result of single value voltage prediction, the reliability of the model affected by input data noise and system uncertainty cannot be effectively represented. Therefore, the Monte Carlo (MCDropout) method is introduced to construct a probability prediction framework, which outputs the prediction result in the form of probability distribution, so as to quantify the confidence interval of the model prediction. The present invention adds MC Dropout to the LSTM network and continues to perform inference and prediction, generates multiple forward propagation outputs, approximates the prediction distribution, calculates the mean and variance of the prediction result, and finally constructs a 95% confidence interval, enhancing the reliability of decision-making.

[0150] The probability prediction network with MC Dropout is constructed as follows: for the same input x, the model performs N forward propagations in the inference stage, and each time gets a different prediction value (i=1, 2, …, N). This set of prediction values can be regarded as a sample of the probability distribution of the model output. Due to the difference in each time of randomly discarded neurons, each prediction value will be slightly different. Then, the mean and variance of these samples can be calculated to describe the uncertainty of the prediction,

[0151] (18)

[0152] (19)

[0153] Using the properties of normal distribution, a 95% confidence interval can be further constructed

[0154] (20)

[0155] To obtain stable estimates, N = 100 was selected through multiple averaging comparative experiments. Finally, the output of the probability prediction network is a matrix with a dimension of 5*48.

[0156] This method not only gives the predicted value, but also quantitatively describes the uncertainty of the model in prediction, providing more abundant information for power grid operation. When the confidence interval is narrow, the prediction result is stable, and the decision confidence is high; on the contrary, when the confidence interval is wide, the related control measures should be handled with caution. This probability prediction can significantly improve the safety and stability of the system under the background of a large number of new energy access and complex load fluctuations.

[0157] Step 4.2: Improved loss function based on L2 regularization and voltage limit constraint

[0158] In order to enhance the robustness of the model and prevent overfitting, two new constraints are added to the loss function of the model: one is L2 regularization, which can prevent the model from entering local optimal solution and improve the robustness of the model; the other is voltage limit constraint, which is used for further testing of MC Dropout. The specific calculation formula of the loss function is as follows:

[0159] (21)

[0160] wherein, represents the weight coefficient of the L2 regularization term, represents the weight coefficient of the voltage limit constraint, represents the voltage prediction value of the i th sample, V i represents the actual voltage value of the i th sample, N represents the total number of samples, V max and V min respectively represent the upper and lower limit values of the voltage. The present application provides that in the loss function, V max = 1.05 p.u., V min = 0.95 p.u.

[0161] Figure 3 The comparative diagram of the power distribution network voltage quality prediction result provided by the embodiment of the present application. The curve fitting degree of the model of the present application is the best, followed by the CNN-LSTM model without integrating the attention mechanism, and the other three models are single models, which are slightly inferior to the hybrid model combining the advantages of the two models, and the last is the random forest model of the traditional machine learning. The data shows that the overall predicted voltage of the model of the present application compared with the true voltage, the error is not more than 0.4V. This is because the present application integrates the attention mechanism which can pay attention to the dominant features of different time steps, and also adds MC Dropout for probability prediction, and improves the loss function, so that the accuracy of the model of the present application is generally higher than that of other models.

[0162] Figure 4 The power distribution network voltage quality prediction attention weight heat map provided for the embodiments of the present application. Features 0~7 respectively represent historical voltage data, temperature, humidity, wind speed, solar irradiance, weekend label, summer label and holiday label. As can be seen from the figure, different features have different influence weights in different time steps. For example, on November 5, due to the large diurnal temperature difference, the night electricity consumption is more than the daytime electricity consumption, resulting in larger voltage fluctuation. Near the Spring Festival, the electricity consumption of every household is basically the highest in the year, and the influence weight of all features on voltage fluctuation will not have too much difference (except for the solar irradiance at night). Compared with weekdays and weekends, it is reasonable that the weekend label feature has a higher influence weight, and most of the other weights basically fall on the weather features.

[0163] It can be seen that after the attention mechanism is integrated, the dominant features at different time steps can be more conveniently and systematically summarized from the heat map.

[0164] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application but not to limit it, although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific embodiments of the present application can still be modified or replaced by the equivalent, without departing from the spirit and scope of the present application. Any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered within the protection scope of the claims of the present application.

Claims

1. A power distribution network voltage quality prediction method based on an improved CNN-LSTM neural network, characterized in that, The method comprises the following steps: Step one, select composite feature data, introduce Pearson correlation coefficient analysis to clarify the influence of different features on voltage, and construct an input matrix based on linear interpolation and sliding window method; Step two, construct a CNN-LSTM power distribution network quality prediction model based on convolutional neural network CNN and long short-term memory neural network LSTM; Step three, integrate attention mechanism into long short-term memory neural network LSTM, generate attention score by dot product of query matrix and hidden state based on hidden state sequence output by LSTM, and generate context vector as input for subsequent probability prediction by weighting; Step four, introduce Monte Carlo method to construct confidence interval, integrate L2 regularization and voltage limit constraint to improve loss function of CNN-LSTM power distribution network quality prediction model, and predict voltage quality by using the improved CNN-LSTM power distribution network quality prediction model.

2. The power distribution network voltage quality prediction method based on improved CNN-LSTM neural network according to claim 1, characterized in that, Step one specifically comprises: 1.1 Composite feature data selection: select voltage data V, weather data and calendar effect data, wherein the weather data adopts weather forecast data superimposed with normal distribution noise; 1.2 Pearson correlation analysis: calculate the Pearson correlation coefficient of the composite feature data and the voltage fluctuation, and select the features significantly related to the voltage fluctuation as the model input; 1.3 Data preprocessing: Missing value processing: the voltage data is linearly interpolated based on the mean value of the previous 3 hours, and the missing value of solar irradiance is set to 0; Abnormal value processing: delete the abnormal values with a difference of more than 15% from the neighborhood mean value and then re-insert them; Normalization processing: the Min-Max method is used to map the data to the [0, 1] interval, and the minimum and maximum parameters are saved for subsequent prediction result inverse transformation; 1.4 Sample generation: generate an input matrix based on a sliding window.

3. The power distribution network voltage quality prediction method based on improved CNN-LSTM neural network according to claim 2, characterized in that, The weather data includes temperature T, humidity H, wind speed W and solar irradiance I, and the calendar effect data includes weekend label R, summer label S and holiday label F; the voltage fluctuation is defined as the absolute difference between the last half sampling value and the average value of the historical voltage.

4. The power distribution network voltage quality prediction method based on improved CNN-LSTM neural network according to claim 2, characterized in that, The window size of the sliding window is 3 days, the dimension of the input matrix is 144*8*1, the input of the input matrix is the voltage, weather and calendar data of the previous three days, and the output is the fourth day voltage data.

5. The power distribution network voltage quality prediction method based on improved CNN-LSTM neural network according to claim 2, characterized in that, The calculation formula of the Pearson correlation coefficient is as follows: ; where x i , y i represent the characteristic value and the voltage fluctuation, respectively, and represent the characteristic value and the average of the voltage fluctuation, respectively; the voltage fluctuation is equal to the absolute value of the historical voltage data of the latter half of the sampling data minus the average thereof; when r = 1, it indicates that the characteristic and the voltage fluctuation are in complete positive correlation; when r = -1, it indicates that the characteristic and the voltage fluctuation are in complete negative correlation; when r = 0, it indicates that the characteristic and the voltage fluctuation have no correlation.

6. The power distribution network voltage quality prediction method based on improved CNN-LSTM neural network according to claim 1, characterized in that, The step two comprises: Build a CNN network to process high-dimensional features, which is used for multiple convolution and pooling on the input matrix constructed in step one to extract high-dimensional features and output reduced feature maps; when constructing the CNN network, set 5 batches, the samples of each batch are the same, and the dimension of the complete input matrix is 5*144*8*1; both convolution and pooling adopt Tensorflow architecture, the convolution kernel is a 3*3 dimensional matrix, the pooling kernel matrix is 2*2, and the Adam optimizer is used for parameter optimization; The LSTM network is built to capture time dependence, and the LSTM network includes four layers, the first layer and the second layer are 128 hidden units, and the third layer and the fourth layer are 64 hidden units; the output matrix dimension of each layer changes with the number of hidden units; after passing through the four-layer LSTM network, each hidden state in the output matrix contains the influence information of the corresponding time on the future voltage quality prediction.

7. The power distribution network voltage quality prediction method based on improved CNN-LSTM neural network according to claim 6, characterized in that, The transfer formula of the gating mechanism of the LSTM network is as follows: ; ; ; ; ; wherein, represents a convolution operation, represents a multiplication operation, represents a sigmoid activation function, tanh represents a hyperbolic tangent function, i t represents an input gate state, f t represents a forget gate state, C t represents a memory cell tensor, o t represents an output gate state, x t represents an input tensor at time t, h t represents a hidden state tensor; W f , W i , W C , W o are weights of the forget gate, input gate, state gate, and output gate, respectively, b f , b i , b C , b o are all respective biases.

8. The power distribution network voltage quality prediction method based on improved CNN-LSTM neural network according to claim 1, characterized in that, The specific steps of step three are as follows: 1) Calculate the attention weight The weight of each time step is determined by its attention score, and the calculation formula of the attention score is as follows: ; where h t denotes the LSTM output at the t-th time step, W h is a weight matrix, and b h is a bias. For the attention scores of all time steps, the softmax function is used to normalize them into the form of probability, so that the sum of the attention weights of all time steps is 1: ; wherein a t is the attention weight for the t-th time step, e t is the attention score for the corresponding time step, and T is the total length of the sequence. 2) Weighted summation After calculating the attention weight of each time step, the output of the LSTM network is weighted and summed to obtain the final context vector; the context vector contains information of all time steps, and the information weights of different time steps are different; the calculation formula of the context vector c is as follows: ; Finally, the context vector c after weighting by the attention mechanism is transmitted to the subsequent probability prediction network through the fully connected layer to continue the prediction task, and the output of the probability prediction network is the weighted result based on the importance of each time step in the time series.

9. The power distribution network voltage quality prediction method based on improved CNN-LSTM neural network according to claim 1, characterized in that, In step four, the Monte Carlo method is introduced to construct the confidence interval, and L2 regularization and voltage limit constraint are introduced to improve the loss function of the CNN-LSTM power distribution network quality prediction model, which includes: Monte Carlo probability prediction: enable MC Dropout in the LSTM inference stage, output prediction mean, variance and 95% confidence interval, when the confidence interval is narrow, the prediction result is more stable, the decision confidence is higher; on the contrary, when the confidence interval is wide, the related control measures should be handled with caution; the probability prediction network with MC Dropout is constructed as follows: for the same input x, the model performs N forward propagation in the inference stage, and different prediction values are obtained each time , i = 1, 2, …, N, the mean and variance of these samples are calculated to describe the uncertainty of the prediction, ; ; Using the properties of normal distribution, a 95% confidence interval is further constructed: ; Improved loss function: The loss function includes MSE error term, L2 regularization term and voltage limit penalty term, and the specific calculation formula of the loss function is as follows: ; wherein, represents a weight coefficient of the L2 regularization term, represents a weight coefficient of the voltage limit constraint, represents a voltage prediction value of the i-th sample, V i represents an actual voltage value of the i-th sample, N represents a total number of samples, V max and V min respectively represent upper and lower limit values of the voltage.

10. The power distribution network voltage quality prediction method based on improved CNN-LSTM neural network according to claim 1, characterized in that, In step four, the improved CNN-LSTM power distribution network quality prediction model is used to predict the voltage quality, which includes: According to step one, the actual power grid data and weather data are processed and the input matrix is constructed by using the sliding window; The processed feature data is input into the built CNN network for processing high-dimensional features, and the high-dimensional features are convolved and pooled multiple times to obtain the reduced feature matrix; The reduced feature matrix is input into the LSTM network to capture the time dependence between features to obtain the preliminary prediction value; The preliminary prediction value is input into the attention mechanism module coupled after the LSTM network to reflect the feature influence weight of the corresponding time step in real time; The Monte Carlo method is introduced to generate the mean, variance and 95% confidence interval of the prediction value, and L2 regularization and voltage limit constraint are introduced to improve the loss function, and the voltage quality prediction result is optimized.