Convolutional neural network short-term power load prediction method based on particle swarm optimization
By adopting the EEPSO-based CNN method in short-term power load prediction, the problems of low model optimization efficiency and insufficient prediction accuracy are solved, and more efficient model optimization and more accurate load prediction are achieved.
Patent Information
- Application Number
- CN202411973880.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has problems such as low model optimization efficiency and insufficient prediction accuracy in short-term power load prediction, especially in hyperparameter determination, input feature selection and network structure optimization.
The convolutional neural network (CNN) method based on enhanced elite particle swarm optimization (EEPSO) is used to determine the model input characteristics through Spearman correlation analysis and normalization processing, and combine the improved particle swarm optimization algorithm to automatically search and optimize the structural parameters and hyperparameters of CNN.
This greatly improves model performance and optimization efficiency, improves the accuracy and robustness of short-term power load prediction, and avoids the subjectivity and uncertainty of hyperparameter determination and feature selection in traditional methods.
Smart Images

Figure CN119940399A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power system load forecasting, and relates to a convolutional neural network short-term power load forecasting method, in particular to a convolutional neural network short-term power load forecasting method based on particle swarm optimization. Background Art
[0002] Power load forecasting is one of the key links in power system operation and planning. Its methods can be divided into short-term load forecasting (STLF), medium-term load forecasting (MTLF) and long-term load forecasting (LTLF). Short-term load forecasting plays a particularly important role in power systems and is widely used in dynamic economic dispatch, unit combination, demand response planning, load management, and the formulation of power market bidding strategies. For example, accurate short-term load forecasting can effectively reduce dispatching costs, optimize resource allocation, and ensure the safety of power grid operation. However, inaccurate forecasts may lead to power system operation problems. For example, too high a forecast load will increase the operating cost of backup power generation, while too low a forecast load may lead to insufficient power generation in the power grid and cause power supply failures.
[0003] Traditional load forecasting methods mostly use statistical models, such as the autoregressive moving average model (ARIMA) or the vector autoregressive model (VARMA), which predict future loads based on the historical characteristics of time series. However, such models are not good at dealing with complex nonlinear or chaotic relationships. In addition, deep learning technology has received widespread attention in the field of short-term load forecasting in recent years, especially convolutional neural networks (CNN) and long short-term memory networks (LSTM), which have shown strong capabilities in capturing nonlinear patterns in time series.
[0004] However, these methods still have the following shortcomings:
[0005] a) Hyperparameters (such as convolution kernel size, pooling window size, number of network layers, etc.) are usually determined through experience or trial and error, making it difficult to achieve the optimal model structure.
[0006] b) There is a lack of systematic analysis of the dimensionality and time lag selection of model input features.
[0007] c) The network structure optimization process usually fails to effectively balance local and global search capabilities, resulting in limited improvement in prediction accuracy.
[0008] In order to overcome the above challenges, this paper proposes a convolutional neural network (CNN) short-term load forecasting method based on enhanced elite particle swarm optimization (EEPSO).
[0009] After searching, no public documents of the prior art identical or similar to the present invention were found. Summary of the invention
[0010] The purpose of the present invention is to overcome the shortcomings of the prior art and propose a convolutional neural network (CNN) short-term load forecasting method based on enhanced elite particle swarm optimization (EEPSO), which can solve the problems of low model optimization efficiency and insufficient prediction accuracy in the prior art.
[0011] The present invention solves the practical problem by adopting the following technical solutions:
[0012] A convolutional neural network short-term power load forecasting method based on particle swarm optimization includes the following steps:
[0013] Step 1: Collect historical temperature and humidity data, hourly load information, and weekly load data of the city to be predicted, and perform Spearman correlation analysis and normalization on the collected data to finally obtain the input matrix of the CNN model;
[0014] Step 2: Based on the dimension of the input matrix of the CNN model obtained in step 1, a convolutional neural network CNN is constructed to obtain a sparse feature map;
[0015] Step 3: Based on the convolutional neural network CNN constructed in step 2, the improved enhanced elite particle swarm optimization algorithm EEPSO is used to obtain the optimal CNN parameters;
[0016] Step 4: Flatten the sparse feature map obtained in step 2 into a one-dimensional vector and connect it to the fully connected network; then train the entire model based on the optimal CNN parameters obtained in step 3 to obtain a high-precision, high-robust and fully optimized CNN model, and finally make a prediction based on the CNN model to obtain the short-term load forecast value corresponding to 24 hours.
[0017] Moreover, the specific steps of step 1 include:
[0018] (1) Collect historical temperature, humidity data, hourly load information, and weekly load data;
[0019] (2) Based on the data collected in step (1), considering various time lags of the historical time series, the Spearman correlation analysis is performed using the following formula to determine the correlation between the current value and the lagged observation value:
[0020]
[0021] Among them, S(t) is the data value at the current time point, S(t-λ) is the data value at the delayed time point, and μ1 or μ2 is the mean of S(t) and S(t-λ), respectively.
[0022] (3) Based on the data collected in step (1), the following formula is used for normalization processing to standardize the data to a fixed range to ensure that the data of different characteristics collected in step (1) can be analyzed at the same scale:
[0023]
[0024] Among them, X i Original data, min(a) and max(a) are the minimum and maximum values in the historical time series data, respectively.
[0025] (4) Based on the data correlation results determined in step (2) and the normalized data in step (3), the input matrix of the CNN model is obtained.
[0026] Furthermore, the step 2 comprises the following steps:
[0027] (1) Based on the dimension of the input matrix of the CNN model obtained in step (4) of step 1, the convolutional neural network is preliminarily constructed using the following formula:
[0028]
[0029] Where f is the activation function. Where w, k, y, and b represent the weight factor vector in the kernel, the kernel index, the input data vector, and the bias vector, respectively.
[0030] (2) Based on the dimension of the input matrix of the convolutional neural network initially constructed in step 2 (1) and the CNN model obtained in step 1 (4), determine the appropriate convolution kernel size, number, stride, and number of layers to ensure that the model strikes a balance between feature extraction and computational efficiency, and then obtain the feature map of the input data.
[0031] (3) Based on the output feature map of the convolutional layer in step 2 (2), the input of the pooling layer is determined, and the feature map after dimensionality reduction is output after the maximum pooling and average pooling operations, thereby enhancing the robustness of the features;
[0032] (4) Based on the pooling operation in step (3) of step 2, the reduced-dimensional feature map is output, and the sparse feature map is output after Dropout regularization.
[0033] Moreover, the specific method of enhancing the feature robustness of the feature map after the dimension reduction of the output of the maximum pooling and average pooling operation in step 2 (3) is:
[0034] First, it is necessary to perform feature state analysis before the pooling operation, and then perform the maximum pooling operation to select the maximum activation value from each pooling window as the output; finally, perform average pooling to calculate the average of all activation values in each pooling window as the output; finally, these two operations are spliced together to generate a feature map after dimensionality reduction.
[0035] Furthermore, the step 3 comprises the following steps:
[0036] (1) Based on the sparse feature map output after Dropout regularization in step 2 (4), define the search space to provide clear boundaries and parameter ranges for the optimization algorithm.
[0037] (2) Based on the definition of the search space in step 3 (1), input the historical hourly load, hourly temperature and humidity of the city to be predicted, and generate feasible particles in the space of unknown parameters.
[0038] (3) Based on the generation of feasible particles in step 3 (2), the root mean square error is determined as the objective function in the improved enhanced elite particle swarm optimization algorithm EEPSO to measure the prediction performance of the convolutional neural network (CNN) model. By minimizing the root mean square error, the structural parameters and hyperparameters of CNN are optimized to calculate the objective function value of each feasible particle.
[0039] (4) Based on step 3 (3), the objective function value of each feasible particle is obtained. If the objective function value of the current particle position is better than the historical record, the local optimal solution is updated and the chaotic descent inertia weight method is added. The formula for updating the feasible particle position is as follows:
[0040]
[0041] in The position of the pth particle at the κth iteration. The updated position of the pth particle at the κ+1th iteration, The velocity of the pth particle at the κ+1th iteration.
[0042] (5) Based on the update of a feasible particle position in step (4) of step 3, the particle gradually approaches the global optimal solution through multiple iterations of particle position updates. The optimization iteration uses the dynamic adjustment direction and amplitude of the position update to gradually search in the solution space and finally converge to the optimal solution, thereby obtaining the CNN parameters with the best performance.
[0043] Furthermore, the step 4 comprises the following steps:
[0044] (1) Flatten the sparse feature map obtained in step (4) of step 2 into a one-dimensional vector, and input the one-dimensional vector into a fully connected network for further processing;
[0045] (2) Based on the optimal CNN parameters obtained in step 3 and the processing results obtained in step 4 (1), the entire CNN model is trained to obtain a high-precision, high-robust and fully optimized CNN model.
[0046] (3) Based on step 4 (2), a high-precision, high-robustness and fully optimized CNN model is obtained to obtain the 24-hour short-term load forecast value.
[0047] Advantages and beneficial effects of the present invention:
[0048] 1. In view of the problem that hyperparameters (such as convolution kernel size, pooling window size, number of network layers, etc.) in traditional methods are usually determined by experience or trial and error, and it is difficult to achieve the optimal model structure, this invention adopts an improved enhanced elite particle swarm optimization algorithm, which greatly improves model performance and optimization efficiency through automated search and iterative optimization.
[0049] 2. In view of the lack of systematic analysis of the dimension and time lag selection of model input features in traditional methods, this paper proposes a method based on Spearman correlation analysis, which systematically selects and optimizes the dimension and time lag of model input features in combination with the time series characteristics of the data.
[0050] 3. Aiming at the problem that the local and global search capabilities cannot be effectively balanced in the process of network structure optimization, resulting in limited improvement in prediction accuracy, this paper proposes an improved enhanced elite particle swarm optimization algorithm (EEPSO), which introduces chaotic decreasing inertia weights and dynamically adjusts the particles' global search in the early stage and local search behavior in the later stage, thereby enhancing the optimization algorithm's ability to explore the entire solution space in the early stage and effectively improving the prediction accuracy.
[0051] 4. In view of the lack of sufficient discussion on the dimension of input data in traditional methods, the number of lagged observations (such as historical load or temperature) as model input is usually determined by heuristic methods. This paper proposes a method based on Spearman correlation analysis. Through Spearman correlation analysis, the correlation between current values and lagged observations (such as historical load, temperature, etc.) is systematically quantified, avoiding the subjectivity and uncertainty of traditional heuristic methods and ensuring that the selection of input features is scientific and reasonable. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A flowchart of the implementation of the convolutional neural network short-term power load forecasting based on particle swarm optimization of the present invention;
[0053] Figure 2 is an example diagram of the average / maximum pooling operation process of the present invention;
[0054] Figure 3 It is a processing flow chart of the present invention. DETAILED DESCRIPTION
[0055] The embodiments of the present invention are further described in detail below with reference to the accompanying drawings:
[0056] A convolutional neural network (CNN) short-term load forecasting method based on enhanced elite particle swarm optimization (EEPSO) is proposed. Figure 1 and Figure 3 As shown, the following steps are included:
[0057] Step 1: Collect the historical temperature, humidity data, hourly load information, and weekly load data of the city to be predicted, perform Spearman correlation analysis and normalization on the collected data, and finally obtain the input matrix of the CNN model.
[0058] The specific steps of step 1 include:
[0059] (1) Collect historical temperature, humidity data, hourly load information, and weekly load data;
[0060] In this embodiment, the working principle of step 1 (1) is:
[0061] The power system of the city is involved. Historical temperature and humidity data (T(t) and H(t)) are crucial for hourly load (L(t)) information. The above variables are important for accurately developing the CNN model because the weekly load profile has a great impact on habits, residents and businesses. Therefore, Monday, Tuesday, ..., Sunday (denoted as C(t)) and holidays (denoted as D(t)) are marked in two columns of exogenous variables. Therefore, the number of columns (denoted as c) of the 2D input data is 9; they give the data of L(t-1), T1(t), H1(t), T2(t), H2(t), T3(t), H3(t), C(t) and D(t).
[0062] (2) Based on the data collected in step (1), considering various time lags of the historical time series, the Spearman correlation analysis is performed using the following formula to determine the correlation between the current value and the lagged observations (such as historical load, temperature, etc.):
[0063]
[0064] Where S(t) is the data value at the current time point, S(t-λ) is the data value at the delayed time point, and μ1 or μ2 is the mean of S(t) and S(t-λ), respectively.
[0065] In this embodiment, the working principle of step (2) of step 1 is:
[0066] Spearman correlation analysis: The determination of the input data dimension involves correlation, using various time lags of historical time series. The present invention uses Spearman rank correlation for correlation analysis; it is a nonparametric measure of the monotonicity of the relationship between two data sets. Spearman correlation does not assume that both data sets are normally distributed. To implement this cross-correlation analysis, time shifts or lags must be considered. For example, if the time series S(t) and S(t-λ) both contain N scalar observations, the correlation coefficient ρ(λ) between S(t) and S(t-λ) is expressed as formula (4), where μ1 or μ2 is the mean of S(t) and S(t-λ), respectively. If the correlation coefficient is within [0.7, 0.9], the relationship between the two time series is strong. If the correlation coefficient is within [0.5, 0.7], the relationship between the two time series is moderate. Check the value of ρ(λ), starting from λ=0, until the value associated with ρ(λ)>0.5, if the correlation is moderate. Where ρ(λ) = 0.5 is the number of rows (r value) of CNN two-dimensional data. Since the historical hourly load is the most important value closely related to the future hourly load, only the correlation coefficient of L(t) is checked. The lag time of T1(t), H1(t), T2(t), H2(t), T3(t), H3(t), C(t), and D(t) is equal to the lag time of L(t).
[0067]
[0068] (3) Based on the data collected in step (1), the following formula is used for normalization processing to standardize the data to a fixed range to ensure that the data of different characteristics collected in step (1) can be analyzed at the same scale:
[0069]
[0070] X i Original data (actual observation value), min(a) and max(a) are the minimum and maximum values in the historical time series data, respectively.
[0071] In this embodiment, the working principle of step (3) of step 1 is:
[0072] Normalization processing: The value of each component in each time series vector is a={a1,a2,a3,……,a N}, where N is the number of samples of hourly load, hourly temperature, hourly humidity, etc., which are first normalized to [0,1] using the min-max normalization process in equation (2). The values of the calendar variable C(t) are set to sin(π / 2), sin(2π / 6), sin(π / 6), sin(0), -sin(π / 6), -sin(2π / 6) and -sin(π / 2) on Monday, Tuesday, ..., Sunday. The calendar variable D(t) is set to 0 and 1 on weekdays and holidays, respectively.
[0073]
[0074] (4) Based on the data correlation results determined in step (2) and the normalized data in step (3), the input matrix of the CNN model is obtained.
[0075] The working principle of step (4) of step 1 is:
[0076] Determination of the number of rows r: Spearman's rank correlation coefficient is used to analyze the correlation of time series data (such as historical load data). The specific method is to calculate the correlation coefficient ρ(λ) between the load L(t) and its lagged version L(t-λ). By increasing the lag time λ successively, find the maximum lag time point λ where the correlation ρ(λ) is greater than 0.5. This maximum lag time point λ determines the number of rows r of the input matrix, that is, the input data needs to contain r rows of historical time points.
[0077] Determination of the number of columns c: The data at each time point contains multiple features, specifically: load data, temperature of each city, humidity of each city, calendar variables (day of the week), and whether it is a holiday. These features together constitute c columns.
[0078] Final matrix construction: The shape of the input matrix is r×c, which serves as the input of the CNN.
[0079] Step 2: Based on the dimension of the input matrix of the CNN model obtained in step 1, a convolutional neural network (CNN) is constructed to obtain a sparse feature map;
[0080] Based on the dimension of the CNN input matrix obtained in step 1, the dimension of the input matrix directly determines the structural design of the CNN model, including the configuration of the convolutional layer, pooling layer, and fully connected layer. At the same time, the input matrix optimization provides high-quality feature input for the CNN model, ensuring the performance and prediction accuracy of the model. Finally, the feature map after dimensionality reduction is obtained.
[0081] The step 2 comprises the following steps:
[0082] (1) Based on the dimension of the input matrix of the CNN model obtained in step (4) of step 1, the convolutional neural network is preliminarily constructed using the following formula:
[0083]
[0084] Where f is the activation function. Where w, k, y, and b represent the weight factor vector in the kernel, the kernel index, the input data vector, and the bias vector, respectively.
[0085] In this embodiment, the working principle of step 2 (1) is:
[0086] Convolutional Layers and Kernels: Convolutional layers use kernels (filters) to generate activation (feature) maps from the input image. The kernel convolves the width and height of the input image (2D data) and uses inner products to generate activation maps. Different 2D kernels can detect different features. Let r and c be the indices of the feature map rows and columns respectively. For the κth layer, the convolution operation is denoted by It can be expressed as formula (3)
[0087]
[0088] Where f is the activation function. Where w, κ, s and b respectively represent the vector of weight factors in the kernel, the exponent of the kernel, the vector of input data and the vector of deviations. The present invention uses the Adam optimizer, and the number of convolutions and the number of kernels are determined by EEPSO.
[0089] Pooling layer: The pooling layer is after the convolutional layer and gradually reduces the spatial dimension. The two pooling operations are maximum pooling and average pooling. Average pooling uses all information to generate a mean in the pooling window. Maximum pooling identifies the maximum value in the pooling window. In the present invention, both maximum pooling and average pooling are used in series. The window size of the pooling layer is determined by EEPSO.
[0090] Fully connected neural network: The output of the pooling layer is converted from 2D data to 1D data and fed to a fully connected neural network (FCNN), which is a supervised multi-layer feedforward neural network. The present invention implements a hidden layer in FCNN. The number of neurons in the hidden layer is determined by EEPSO. The weight factors and biases in FCNN are obtained by Adam optimizer. In traditional methods, the number of neurons in the hidden layer is obtained by trial and error.
[0091] (2) Based on the dimension of the input matrix of the convolutional neural network initially constructed in step 2 (1) and the CNN model obtained in step 1 (4), determine the appropriate convolution kernel size, number, stride, and number of layers to ensure that the model strikes a balance between feature extraction and computational efficiency, and then obtain the feature map of the input data.
[0092] In this embodiment, the working principle of step 2 (2) is as follows:
[0093] Convolutional layer design: The convolutional layer extracts local features of the input data. The size and number of convolution kernels are determined by an optimization algorithm. Multiple convolutional layers are used to extract high-dimensional features.
[0094] (3) Based on the output feature map of the convolutional layer in step 2 (2), the input of the pooling layer is determined, and the feature robustness of the feature map enhanced by the reduced dimension output of the pooling maximum pooling and average pooling operations;
[0095] The specific method for enhancing the feature robustness of the feature map after the dimension reduction of the output of the maximum pooling and average pooling operation in step 2 (3) is:
[0096] First, it is necessary to perform feature state analysis before the pooling operation, and then perform the maximum pooling operation to select the maximum activation value from each pooling window as the output; finally, perform average pooling to calculate the average of all activation values in each pooling window as the output; finally, these two operations are spliced together to generate a feature map after dimensionality reduction.
[0097] In this embodiment, the working principle of step 2 (3) is as follows:
[0098] Applying max pooling and average pooling operations in parallel is as follows Figure 2 As shown in the figure, two pooling results are concatenated and fused. The pooling layer is used to reduce the feature dimension and enhance the robustness of the model. Figure 2 The operating principle is as follows:
[0099] Max Pooling:
[0100] Function: Select the maximum value from the pooling window.
[0101] Purpose: To extract the most significant features (e.g., peaks or spikes in load data).
[0102] Average Pooling:
[0103] Function: Calculate the average value from the pooling window.
[0104] Purpose: Smooth data and extract global feature information.
[0105] Parallel Operation:
[0106] In the present invention, the maximum pooling and average pooling operations are applied in parallel, that is, the two pooling operations are performed on the same input data. After the pooling is completed, the features of the maximum pooling and the features of the average pooling are obtained respectively.
[0107] Concatenation:
[0108] The results of maximum pooling and average pooling are concatenated to generate the final feature vector. The purpose is to fuse local significant features and global smooth features to enhance the feature expression ability of the model.
[0109] Optimization of pooling window size:
[0110] The size of the pooling window (e.g., 2×2 or 3×3) is determined by step 2 (2), thereby determining the pooling scale that best suits the model and data.
[0111] Application examples:
[0112] exist Figure 2 The detailed process of the pooling operation shows how to extract features from the input data through maximum pooling and average pooling respectively. Subsequently, the two features are combined through the splicing operation to form the final feature vector for the prediction task. This method can not only highlight the significant changes in the data, but also retain the global trend of the data, ultimately improving the accuracy and robustness of the prediction model.
[0113] (4) Based on the pooling operation in step (3) of step 2, the reduced-dimensional feature map is output, and the sparse feature map is output after Dropout regularization;
[0114] In this embodiment, the working principle of step (4) of step 2 is:
[0115] In order to prevent overfitting of multi-layer neural networks, regularization is essential. Dropout is an effective regularization method. The drop-out ratio controls the activation of neurons and represents the proportion of input units to be discarded. The present invention applies the dropout technique to the neurons of the input layer of FCNN, and the dropout ratio is determined by EEPSO.
[0116] Step 3: Based on the convolutional neural network (CNN) constructed in step 2, the improved enhanced elite particle swarm optimization algorithm (EEPSO) is used to obtain the optimal CNN parameters;
[0117] Based on the feature map after dimensionality reduction obtained at the end of step 2, it not only extracts key information and reduces redundancy, but also provides a more efficient objective function input for improving the particle swarm optimization algorithm, simplifies the search space, improves the optimization efficiency and the generalization ability of the results, and finally obtains the optimal CNN parameters.
[0118] The step 3 comprises the following steps:
[0119] (1) Based on the sparse feature map output after Dropout regularization in step 2 (4), the search space is defined to provide clear boundaries and parameter ranges for the optimization algorithm to effectively search for the best solution.
[0120] In this embodiment, the working principle of step 3 (1) is:
[0121] Each particle in the proposed EEPSO is associated with 9 unknowns representing the structural parameters and hyperparameters of CNN. The structural parameters of CNN are the number of convolutional layers and the number of neurons in the hidden layer in the fully connected layer. The hyperparameters of CNN are the number of kernels (filters) in each convolutional layer, the kernel size, the window size of the pooling layer, the spatial dropout ratio, and the dropout ratio.
[0122] (2) Based on the definition of the search space in step 3 (1), input the historical hourly load, hourly temperature and humidity of the city to be predicted, and generate feasible particles in the space of unknown parameters.
[0123] The specific method of step (2) of step 3 is:
[0124] Input the historical hourly load, hourly temperature and humidity of the city to be predicted, determine Sunday, Monday, ..., Saturday, weekdays and holidays; normalize the load data and meteorological data; calculate the Spearman correlation coefficient ρ(λ) between L(t) and L(t-λ). If ρ(λ)>0.5, then r=λ(number of rows in 2D data). Input the target, unknowns and population size to the EEPSO algorithm. Generate feasible particles in the nine-dimensional space of unknown parameters. Set κ=1
[0125] (3) Based on the generation of feasible particles in step 3 (2), the root mean square error is determined as the objective function in the improved enhanced elite particle swarm optimization algorithm (EEPSO) to measure the prediction performance of the convolutional neural network (CNN) model. By minimizing the root mean square error, the structural parameters and hyperparameters of CNN are optimized to calculate the objective function value of each feasible particle.
[0126] In this embodiment, the working principle of step 3 (3) is as follows:
[0127] In the EEPSO algorithm, the objective function (RMSE) runs through the entire search process. Starting from the initialization of the particles, the goal is to find the CNN structure and hyperparameter combination that minimizes the RMSE. In each iteration, the local optimal and global optimal particles (p b k est and g b k es ), and then update the speed and position of other particles according to these optimal particles and objective functions, and adjust parameters such as learning factors, so as to continuously optimize the CNN model structure and hyperparameters until the convergence conditions are met.
[0128] (4) Based on step 3 (3), the objective function value of each feasible particle is obtained. If the objective function value of the current particle position is better than the historical record, the local optimal solution is updated and the chaotic descent inertia weight method is added. The formula for updating the feasible particle position is as follows:
[0129]
[0130] in The position of the pth particle at the κth iteration. The updated position of the pth particle at the κ+1th iteration, The velocity of the pth particle at the κ+1th iteration.
[0131] In this embodiment, the working principle of step (4) of step 3 is as follows:
[0132] The mean search relies on the mean of the particles and the standard deviation of the distance between any pair of particles in the κth iteration. Let X κ m is the Ψ-dimensional vector of the average value of all particles. The last term in equation (4) is used to coordinate global and local searches, and the velocity of particle p is updated as follows:
[0133]
[0134] According to the chaotic descent inertia weight method, the inertia weight ω k It changes with the iteration index κ. In formula (5), ω1 and ω2 are the initial value and final value of the inertia weight respectively; MAX is the maximum number of iterations. The parameter of the chaotic descent inertia weight is set as: 0 =0.1, ω1=0.9, ω2=0.4, MAX=50.
[0135]
[0136] in
[0137] z κ =4×z κ-1 ×(1-z κ-1 ) (6)
[0138] Formula (6) describes a logical mapping that produces chaotic phenomena. The chaotic results are scattered in the range [0,1]. is a random number between 0 and 1. In equations (9) and (10), the learning factor and Increase / decrease the number of iterations exponentially ( Reduced from 2 to 1; From 1 to 2). In formulas (7) and (8), Consider The optimal target value (RMSE) after a good training of the proposed CNN model. τ is A multiplier, which can be set to 0.9×MAX. When the iterative process converges, the new learning factor in (9) Gradually decreases to zero, so that equation (9) becomes equation (10).
[0139]
[0140]
[0141]
[0142]
[0143] The particle position update formula (11) is: This update method introduces the chaotic decreasing inertia weight and the mean search strategy that balances the local search and global search capabilities, which enhances the optimization efficiency.
[0144]
[0145] (5) Based on the update of a feasible particle position in step (4) of step 3, through multiple iterations of particle position updates, the particle gradually approaches the global optimal solution. The optimization iteration uses the dynamic adjustment direction and amplitude of the position update to gradually search in the solution space and finally converge to the optimal solution, thereby obtaining the CNN parameters with the best performance.
[0146] In this embodiment, the working principle of step (5) of step 3 is:
[0147] Optimization iteration: Iterate the feasible particles generated in step 3 (2). The specific steps of the iteration process are as follows:
[0148] Step 1: Calculate all objective values (rmse) of all particles obtained by the Adam optimizer.
[0149] Step 2: Find the target value of all particles. and
[0150] Step 3: Calculate ω separately κ , and
[0151] Step 4: Evaluation σ κ and η κ If the particle is If the particle is outside the range, it will be clipped and copied Make a replacement.
[0152] The process of particle pruning / copying is as follows: Let the standard deviation of the distance between all particles in the κth iteration be σ k :right Particles outside the range are clipped; these clipped particles are copied replace. can be considered as a virtual particle, which is the average of all particles in the κth iteration. Such a value η κ Increased from 2 to 3 to cover 99.7% of all particles in range When all distances are Gaussian distributed.
[0153] Step 5: Calculate the speed of all particles according to formula (10).
[0154] Step 6: Update the positions of all particles according to equation (11).
[0155] Step 7: κ=κ+1
[0156] Step 8: If the iteration process meets the convergence criterion, stop and output the optimal CNN; otherwise, execute the first step.
[0157] Step 9: When k = MAX, stop iteration and output the optimal CNN model structure and hyperparameters.
[0158] Step 4: Flatten the sparse feature map obtained in step 2 into a one-dimensional vector and connect it to the fully connected network; then train the entire model based on the optimal CNN parameters obtained in step 3 to obtain a high-precision, high-robust and fully optimized CNN model, and finally make a prediction based on the CNN model to obtain the short-term load forecast value corresponding to 24 hours.
[0159] The step 4 comprises the following steps:
[0160] (1) Flatten the sparse feature map obtained in step (4) of step 2 into a one-dimensional vector, and input the one-dimensional vector into a fully connected network for further processing;
[0161] The fully connected network is the backend part of CNN, which is used to map the input features to the final output;
[0162] In this embodiment, the working principle of step 4 (1) is:
[0163] The two-dimensional features output by the pooling layer are flattened into one-dimensional features and input into the fully connected layer. The fully connected network includes a hidden layer, and the number of hidden neurons is determined by PSO optimization. The output of the pooling layer is converted from 2D data to 1D data and fed into a fully connected neural network, which is a supervised multi-layer feedforward neural network. The present invention implements a hidden layer in FCNN. The number of neurons in the hidden layer is determined by EEPSO. The weight factors and biases in FCNN are obtained by Adam optimizer.
[0164] (2) Based on the optimal CNN parameters obtained in step 3 and the processing results obtained in step 4 (1), the entire CNN model is trained to obtain a high-precision, high-robust and fully optimized CNN model.
[0165] In this embodiment, the working principle of step 4 (2) is as follows:
[0166] First, after generating feasible particles and determining the CNN structure and hyperparameters, the Adam optimizer is used to train the CNN model represented by each particle. In the convolution layer, the value of the convolution kernel is obtained by the Adam optimizer; in the fully connected layer, the weighting factor and bias are also obtained by the Adam optimizer. When calculating the target value corresponding to each particle, it is also calculated based on the predicted value and actual value obtained by the Adam optimizer training model. The entire training process aims to minimize the mean square error.
[0167] (3) Based on step 4 (2), a high-precision, high-robust and fully optimized CNN model is obtained to forecast the power load of a city for 24 hours. The short-term load forecast value for 24 hours is obtained.
[0168] In this embodiment, the working principle of step 4 (3) is as follows:
[0169] After model training is completed, the optimized CNN model is used to predict the power load for 24 hours. The final FCNN contains one hidden layer, and the number of neurons is determined by EEPSO. The output layer contains 24 neurons, corresponding to the predicted value for 24 hours.
[0170] It should be emphasized that the embodiments of the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific implementation modes. Any other implementation modes derived by those skilled in the art based on the technical solutions of the present invention also fall within the scope of protection of the present invention.
Claims
1. A convolutional neural network short-term power load forecasting method based on particle swarm optimization, characterized by: The following steps are involved: Step 1: Collect historical temperature and humidity data, hourly load information, and weekly load data of the city to be predicted, and perform Spearman correlation analysis and normalization on the collected data to finally obtain the input matrix of the CNN model; Step 2: Based on the dimension of the input matrix of the CNN model obtained in step 1, a convolutional neural network CNN is constructed to obtain a sparse feature map; Step 3: Based on the convolutional neural network CNN constructed in step 2, the improved enhanced elite particle swarm optimization algorithm EEPSO is used to obtain the optimal CNN parameters; Step 4: Flatten the sparse feature map obtained in step 2 into a one-dimensional vector and connect it to the fully connected network; then train the entire model based on the optimal CNN parameters obtained in step 3 to obtain a high-precision, high-robust and fully optimized CNN model, and finally make a prediction based on the CNN model to obtain the short-term load forecast value corresponding to 24 hours.
2. According to claim 1, a convolutional neural network short-term power load forecasting method based on particle swarm optimization is characterized by: The specific steps of step 1 include: (1) Collect historical temperature, humidity data, hourly load information, and weekly load data; (2) Based on the data collected in step (1), considering various time lags of the historical time series, the Spearman correlation analysis is performed using the following formula to determine the correlation between the current value and the lagged observation value: Among them, S(t) is the data value at the current time point, S(t-λ) is the data value at the delayed time point, and μ1 or μ2 is the mean of S(t) and S(t-λ) respectively; (3) Based on the data collected in step (1), the following formula is used for normalization processing to standardize the data to a fixed range to ensure that the data of different characteristics collected in step (1) can be analyzed at the same scale: Among them, X i Original data, min(a) and max(a) are the minimum and maximum values in the historical time series data, respectively; (4) Based on the data correlation results determined in step (2) and the normalized data in step (3), the input matrix of the CNN model is obtained.
3. The method for short-term power load forecasting based on convolutional neural network using particle swarm optimization according to claim 1 is characterized in that: The step 2 comprises the following steps: (1) Based on the dimension of the input matrix of the CNN model obtained in step (4) of step 1, the convolutional neural network is preliminarily constructed using the following formula: Where f is the activation function; where w, k, y, and b represent the weight factor vector in the kernel, the kernel index, the input data vector, and the bias vector, respectively; (2) Based on the dimension of the input matrix of the convolutional neural network initially constructed in step 2 (1) and the CNN model obtained in step 1 (4), determine the appropriate convolution kernel size, number, stride, and number of layers to ensure that the model strikes a balance between feature extraction and computational efficiency, and then obtain the feature map of the input data; (3) Based on the output feature map of the convolutional layer in step 2 (2), the input of the pooling layer is determined, and the feature map after dimensionality reduction is output after the maximum pooling and average pooling operations, thereby enhancing the robustness of the features; (4) Based on the pooling operation in step (3) of step 2, the reduced-dimensional feature map is output, and the sparse feature map is output after Dropout regularization.
4. The method for short-term power load forecasting based on convolutional neural network using particle swarm optimization according to claim 3 is characterized in that: The specific method for enhancing the feature robustness of the feature map after the dimension reduction of the output of the maximum pooling and average pooling operation in step 2 (3) is: First, it is necessary to perform feature state analysis before the pooling operation, and then perform the maximum pooling operation to select the maximum activation value from each pooling window as the output; finally, perform average pooling to calculate the average of all activation values in each pooling window as the output; finally, these two operations are spliced together to generate a feature map after dimensionality reduction.
5. The method for short-term power load forecasting based on convolutional neural network using particle swarm optimization according to claim 1 is characterized in that: The step 3 comprises the following steps: (1) Based on the sparse feature map output after Dropout regularization in step 2 (4), define the search space to provide clear boundaries and parameter ranges for the optimization algorithm; (2) Based on the definition of the search space in step 3 (1), input the historical hourly load, hourly temperature and humidity of the city to be predicted, and generate feasible particles in the space of unknown parameters; (3) Based on the generation of feasible particles in step 3 (2), the root mean square error is determined as the objective function in the improved enhanced elite particle swarm optimization algorithm EEPSO to measure the prediction performance of the convolutional neural network (CNN) model. By minimizing the root mean square error, the structural parameters and hyperparameters of CNN are optimized to calculate the objective function value of each feasible particle; (4) Based on step 3 (3), the objective function value of each feasible particle is obtained. If the objective function value of the current particle position is better than the historical record, the local optimal solution is updated and the chaotic descent inertia weight method is added. The formula for updating the feasible particle position is as follows: in The position of the pth particle at the κth iteration; The updated position of the pth particle at the κ+1th iteration, The velocity of the pth particle at the κ+1th iteration; (5) Based on the update of a feasible particle position in step (4) of step 3, the particle gradually approaches the global optimal solution through multiple iterations of particle position updates. The optimization iteration uses the dynamic adjustment direction and amplitude of the position update to gradually search in the solution space and finally converge to the optimal solution, thereby obtaining the CNN parameters with the best performance.
6. The method for short-term power load forecasting based on convolutional neural network using particle swarm optimization according to claim 1 is characterized in that: The step 4 comprises the following steps: (1) Flatten the sparse feature map obtained in step (4) of step 2 into a one-dimensional vector, and input the one-dimensional vector into a fully connected network for further processing; (2) Based on the optimal CNN parameters obtained in step 3 and the processing results obtained in step 4 (1), the entire CNN model is trained to obtain a high-precision, high-robust and fully optimized CNN model; (3) Based on step 4 (2), a high-precision, high-robustness and fully optimized CNN model is obtained to obtain the 24-hour short-term load forecast value.