Power load prediction method for electric energy meter data system and medium
By utilizing seasonal decomposition and an improved PSO optimization algorithm combined with a CNN-LSTM-Transformer neural network model, the power load forecasting method based on the electricity meter data system was improved, thus solving the problem of insufficient accuracy of traditional load forecasting methods under complex conditions and achieving high-precision forecasting of short-term power load.
Patent Information
- Application Number
- CN202511241530.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Traditional load forecasting methods lack sufficient accuracy under complex conditions and cannot meet the high-precision requirements of power load forecasting.
The power load forecasting method using the electricity meter data system is proposed. Through seasonal decomposition and deseasonalization, lag values and date attribute features are added. An improved PSO optimization algorithm is used to divide the data into independent and dependent variables. The training and test sets are optimized by combining the CNN-LSTM-Transformer neural network model.
It significantly improves the accuracy of short-term power load forecasting, increases model training speed and global search capability, enhances local search accuracy, and solves the problem of insufficient modeling of complex time series data by a single model.
Smart Images

Figure CN120806279B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric power data processing, in particular to a power load prediction method of an electric energy meter data system and a medium. BACKGROUND
[0002] The electric energy meter data system is an important part of the smart grid. The power load prediction is not only an important function of the electric energy meter data system, but also a core link of the smart grid safety dispatching and energy optimization management. Especially, the short-term prediction of the power load is crucial to the power generation dispatching and energy utilization, and directly affects the economic dispatching of the power system and the equipment operation efficiency. However, due to the complex factors affecting the load and the nonlinear and high volatility of the load sequence, the load prediction becomes extremely difficult.
[0003] The traditional load prediction method has limited prediction ability for such data, and cannot guarantee the prediction accuracy requirement under complex conditions. In order to improve the prediction accuracy, many scholars apply artificial intelligence technology to the power load prediction. Although the prediction method using neural network can comprehensively learn the data rules of the power load time series data and master the important transformation trend implied therein, the prediction accuracy still needs to be further improved, which is also the future development trend of the technology. SUMMARY
[0004] The technical problem to be solved by the present application is to provide a power load prediction method of an electric energy meter data system and a medium, which can significantly improve the prediction accuracy of the short-term power load.
[0005] In order to solve the above technical problems, the first technical solution adopted by the present application is:
[0006] The power load prediction method of the electric energy meter data system comprises:
[0007] S1: grouping the historical power load data according to the region ID to obtain each group of data sets;
[0008] S2: respectively performing seasonal decomposition on each group of data sets to obtain the corresponding feature items of each group of data sets;
[0009] S3: judging whether each group of data sets belongs to an additive model or a multiplicative model according to the corresponding feature items of each group of data sets, using a detrending algorithm corresponding to the belonging model to perform deseasonalization processing on the load data in each group of data sets, obtaining the deseasonalized load data of each group of data sets, and taking the deseasonalized load data as the corresponding feature item;
[0010] S4: adding the feature items of the lag value and the date attribute of each group of data sets including the load data;
[0011] S5: Normalizing all feature items corresponding to each group of data sets to obtain processed data;
[0012] S6: Dividing independent variables X and dependent variables Y of the processed data according to the improved PSO optimization algorithm, and further dividing to obtain a training set and a test set; wherein the improved PSO optimization algorithm updates the optimal particle according to a dynamic inertia weight and a dynamic learning factor, and the dynamic inertia weight and the dynamic learning factor are dynamically changed according to the Levy flight mechanism and the Gaussian disturbance mechanism;
[0013] S7: Creating an initial CNN-LSTM-Transformer neural network model;
[0014] S8: Training the initial CNN-LSTM-Transformer neural network model using the training set according to the improved PSO optimization algorithm until the final optimal particle is determined, and obtaining a CNN-LSTM-Transformer neural network model; wherein each training updates the hyperparameters of the model with the particle information updated by the improved PSO optimization algorithm in the previous training;
[0015] S9: Inputting the target test set processed by the S1 to the S5 into the CNN-LSTM-Transformer neural network model to obtain a short-term power load prediction result.
[0016] Optionally, the S6 further includes:
[0017] SS6: Converting the three-dimensional array originally corresponding to multiple time steps of the training set into a converted three-dimensional array corresponding to one time step.
[0018] Optionally, the conversion is specifically:
[0019] After sorting the three-dimensional array data originally corresponding to multiple time steps, it is fused into one time step.
[0020] Optionally, the process of the improved PSO optimization algorithm includes:
[0021] (1) Randomly initializing each particle;
[0022] (2) Evaluating each particle to obtain an optimal particle;
[0023] (3) Judging whether the optimal particle meets the end condition; if yes, ending the process and outputting the optimal particle; if not, executing (4);
[0024] (4) the inertia weight and the learning factor are dynamically changed through the Levy flight mechanism and the Gaussian disturbance mechanism, and the particle information and the speed and position of each particle are updated using the dynamic inertia weight and the dynamic learning factor;
[0025] (5) the fitness value of each updated particle is re-evaluated;
[0026] (6) the optimal particle is determined according to the fitness value of each particle, and step (3) is returned to be executed.
[0027] Optionally, in the (4), the Levy step generation algorithm used by the Levy flight mechanism and the Gaussian disturbance algorithm used by the Gaussian disturbance mechanism are respectively:
[0028] The position disturbance calculation formula in the Gaussian disturbance algorithm is: ,
[0029] wherein, is the position of particle i at t+1, is the position of particle i at t, is the velocity of particle i at t+1, is a disturbance intensity coefficient, is a Gaussian distributed noise, is a disturbance intensity which decays according to the number of iterations, and the decay formula is ;
[0030] The velocity disturbance calculation formula in the Gaussian disturbance algorithm is: ,
[0031] wherein, is the inertia weight, is the velocity of particle i at t, is the learning factor, is a uniformly distributed random number, is the individual historical optimal position of particle i, is the global optimal position of the particle group;
[0032] The Levy step generation algorithm is: ,
[0033] wherein, is a stability coefficient, and u is a random number conforming to a normal distribution.
[0034] Optionally, the S6 specifically comprises:
[0035] S61: determining the length of the independent variable X and the dependent variable Y according to the particle information of the optimal particle output by the improved PSO optimization algorithm;
[0036] S62: Dividing the normalized data into independent variables X and dependent variables Y according to the determined lengths of the independent variables X and the dependent variables Y;
[0037] S63: Dividing the divided processed data into a training set and a test set according to a preset ratio.
[0038] Optionally, the S1 specifically comprises:
[0039] S11: Obtaining historical power load data, the historical power load data comprising three columns of data of a region ID, a data time and load data;
[0040] S12: Grouping the historical power load data according to the region ID, and then sorting according to the data time to obtain each group of data sets.
[0041] Optionally, the initial CNN-LSTM-Transformer neural network model is created based on a Keras framework, and specifically comprises a multi-channel convolutional layer CNN, an LSTM network, a Transformer neural network layer and a fully connected layer connected in sequence.
[0042] The multi-channel convolutional layer CNN has a convolution kernel size and a receptive field size determined according to particle information updated by the improved PSO optimization algorithm, and is used to extract spatial features in time series data, such as similarity or difference features between regions, capture load mutations, identify load differences of date attributes, and retain local significant features.
[0043] The LSTM network is used to learn the rules of data through a gating mechanism including a forgetting gate, an input gate and an output gate, capture medium and short-term dependency relationships of all feature items in time series data, combine all feature items with LSTM memory cell information, and discard invalid information.
[0044] The Transformer neural network layer is used to calculate the weight of each feature item in time series data through a multi-head attention mechanism, capture long-distance dependency, and enhance feature expression capability through feedforward propagation.
[0045] Optionally, before the data output by the multi-channel convolutional layer CNN is input into the LSTM network, it further comprises:
[0046] After the three-dimensional array corresponding to multiple time steps output by the multi-channel convolutional layer CNN is converted into a three-dimensional array corresponding to one time step, it is input into the LSTM network.
[0047] Another technical solution provided by the application is:
[0048] A computer readable storage medium, having stored thereon a computer readable computer program, which, when executed from a processor, is capable of implementing all steps of the power load prediction method of the electric energy meter data system as described above.
[0049] The present application has the beneficial effects that: the present application first enhances the expression and stationarity of the load data by performing seasonal decomposition, deseasonalization processing on the power load data, and adding two characteristic items of hysteresis and date attribute; then, the training speed is greatly accelerated by converting the data shape and using the improved PSO optimization algorithm to obtain the training set and optimize the training data; then, the neural network model based on the combination of CNN, LSTM and Transformer is trained using the above training set according to the improved PSO optimization algorithm, the full-scale feature extraction of the load data from the micro to the macro is realized, the problem of insufficient modeling of single model for complex time series data is solved, and the prediction accuracy can be significantly improved; in particular, the improved PSO optimization algorithm is used to dynamically update the particles in the training set division and model training process, and the hyperparameters of the model are also dynamically updated using the dynamically updated particle information in the training process, which not only speeds up the model training convergence speed and increases the global search ability, but also significantly enhances the ability to jump out of the local optimum by combining the Levy flight mechanism and Gaussian disturbance, and improves the local search accuracy. Therefore, the power load prediction method of the electric energy meter data system provided by the present application can significantly improve the prediction accuracy of short-term power load data. BRIEF DESCRIPTION OF DRAWINGS
[0050] Figure 1 The flowchart of the power load prediction method of the electric energy meter data system provided by the first embodiment of the present application is shown in the figure;
[0051] Figure 2 The flowchart of the power load prediction method of the electric energy meter data system provided by the first embodiment of the present application is shown in the figure;
[0052] Figure 3 The flowchart of the power load prediction method of the electric energy meter data system provided by the first embodiment of the present application is shown in the figure;
[0053] Figure 4 The framework structure diagram of the CNN-LSTM-Transformer neural network model in the embodiment of the present application is shown in the figure;
[0054] Figure 5 The prediction flowchart of the model based on the CNN-LSTM-Transformer neural network model in the embodiment of the present application is shown in the figure; Figure 4
[0055] Figure 6 The data conversion flowchart in the second embodiment of the present application is shown in the figure;
[0056] Figure 7 A flowchart of the model training and prediction process in Embodiment Four of the present application is shown in the figure.
[0057] Figure 8 A diagram showing the effect of the power load short-term prediction of the power load prediction method of the electric energy meter data system provided by the present application is shown in the figure. DETAILED DESCRIPTION
[0058] To explain the technical content, achieved purposes and effects of the present application in detail, the following will be described in combination with the embodiments and the accompanying drawings.
[0059] Technical terms involved in the present application are explained as follows:
[0060] CNN, Convolutional Neural Network, is a kind of neural network used for processing data with grid topology, which realizes efficient feature learning and pattern recognition through core mechanisms such as local connection, weight sharing and hierarchical feature extraction. In recent years, it has shown significant advantages in time series prediction tasks. Through one-dimensional convolution kernel sliding, it can capture local patterns (such as periodic fluctuations and mutation points) in time series data. It is suitable for short-term dependence modeling, and for long sequence data, it needs to be used together with other neural networks.
[0061] LSTM, Long Short Term Memory, is a special kind of recurrent neural network, which solves the gradient vanishing / explosion problem of traditional recurrent neural networks in processing long sequence data through gating mechanism and cell state design, and can effectively capture long-term dependencies in time series data.
[0062] Transformer is a deep learning model based on attention mechanism, originally used for natural language processing. Its parallelization and global dependence modeling capabilities make it perform outstandingly in time series prediction. The high-order time series features such as the hidden cell state of LSTM are embedded through Transformer, and the position encoding is added to preserve the time order. The attention mechanism calculates the correlation weight between any two parts to dynamically capture local and global dependencies.
[0063] Seasonal decomposition, using the seasonal_decompose algorithm for seasonal decomposition, obtains three feature items including trend, seasonal and residual.
[0064] Embodiment One
[0065] The present embodiment provides a power load prediction method of an electric energy meter data system, as shown in the figure, at least comprising the following steps S1 to S9. Figure 1
[0066] The step S1 groups the historical power load data according to the region ID to obtain each group data set.
[0067] Specifically, each load data load in the historical power load data is associated with information such as region_id (region ID) and data_time (data time). In this embodiment, the load data and the corresponding region ID and data time are mainly taken as the analysis objects. Here, since the load is strongly associated with the region, and most of the prediction scenarios are also based on the specified region, the historical power load data obtained is grouped according to the region ID as the grouping basis to obtain each group data set.
[0068] In some preferred specific embodiments, the historical power load data is embodied in the form of a list with region ID, data time, and load data as table fields, which is more intuitive and convenient to operate. That is, the historical power load data includes three columns of region_id (region ID), data_time (data time), and load (load data).
[0069] In some preferred specific embodiments, after grouping the historical power load data according to the region ID, the data_time (data time) is sorted in ascending order (preferably) or descending order to obtain each group data set. This will achieve further optimization processing of each group data set, which is more conducive to reflecting the process and trend of load changes in each region, predicting the future, and regularity analysis.
[0070] The step S2 seasonally decomposes each group data set to obtain the characteristic items corresponding to each group data set.
[0071] Specifically, the seasonal_decompose algorithm is used to seasonally decompose each grouped data set to obtain three characteristic items including trend (trend), seasonal (seasonality), and residual (residual).
[0072] Here, by extracting the above three characteristic items of each grouped data, the time series variation trend and periodic variation of the load data can be better understood.
[0073] In some optional specific embodiments, the seasonal decomposition results of each group data set are displayed in a visual form to provide a way for users to intuitively observe the seasonal characteristics of each group data set, which is conducive to quickly observing abnormal data and processing.
[0074] The step S3 determines whether each group of data sets belongs to an additive model or a multiplicative model according to the corresponding characteristic item of each group of data sets, uses a detrending algorithm corresponding to the model to which the group of data sets belongs to perform deseasonalization processing on the load data in each group of data sets, obtains deseasonalized load data of each group of data sets, and takes the deseasonalized load data as the corresponding characteristic item.
[0075] The deseasonalization processing refers to determining whether the grouped data set belongs to an additive model or a multiplicative model according to whether the seasonal amplitude changes with time, and then removing the trend and the seasonality as additional features of the load data. Since the seasonal fluctuation affects the long-term trend, different deseasonalization methods are selected for different types of models (additive model-difference method, multiplicative model-proportional adjustment method), which are more targeted.
[0076] In some preferred embodiments, for each grouped data set, it is determined whether the grouped data set belongs to an additive model or a multiplicative model according to whether the decomposed data of the grouped data set satisfies Yt(t)=Trend(t)+Seasonal(t)+Residual(t) or satisfies Yt(t)=Trend(t)×Seasonal(t)×Residual(t); for the additive model, the deseasonalized data=original data-seasonal fluctuation; and for the multiplicative model, the deseasonalized data=original data / seasonal component.
[0077] Through the above steps, each group of data sets has four additional characteristic items: trend, seasonal, residual, and deseasonalized load. Preferably, the four newly added characteristic data are embodied in each group of data sets in the form of four columns of data.
[0078] It should be noted that, after the seasonal decomposition of each group of data sets, the difference method or the proportional adjustment method is used to remove the seasonal fluctuation; unlike the prior art which directly inputs the seasonal characteristic variable into the model for prediction, the embodiment removes the seasonal fluctuation, which helps to meet the stationarity assumption, improves the attention of the model to the non-seasonal pattern, and the subsequent use of LSTM also needs to combine the deseasonalization result for residual modeling.
[0079] The step S4 adds characteristic items of lag values and date attributes of the load data of each group of data sets.
[0080] In this embodiment, since each group of data sets belongs to time series data, the time series data usually has autocorrelation, that is, there is some correlation between the current value and the past value. Therefore, by introducing the lag value (the value of the same variable at a certain time in the past), the relationship can be explicitly included in the model; and adding the lag value can also reduce the influence of non-stationary components to some extent, thereby enhancing data stationarity.
[0081] In some preferred embodiments, the shift function provided by python (dataframe.shift(n)) helps to introduce the lag value. Specifically, each group of data sets adds lag_1, lag_2...lag_n columns to the load data of each row according to the pre-set n lag value stages, which represents the first value, the second value...the n value of the variable in the past, and enhances the dynamic interpretation of time series analysis by including its historical data. For example, by horizontally expanding each load data with the first 5 load data in the group according to time to the feature data under the feature item of "lag value" of this data.
[0082] Since there are many factors affecting power load, the date attributes such as whether it is a holiday, whether it is a weekday, and which day of the week have a great influence on power load. Therefore, by analyzing the data_time of the load data, the date attributes of whether it is a weekday and which day of the week can be obtained, and the "date attribute" is also added as a feature item.
[0083] The step S5 normalizes all feature items corresponding to each group of data sets to obtain processed data.
[0084] Since the dimensions of each data vary greatly, that is, each feature has different magnitude. For example, the value of load data may reach 1000 or more, while the value of region ID is between 1 and 100, and the value of residual is between -1 and 1. Therefore, a normalizer is created for each data to perform normalization processing. Normalization can eliminate the dimension difference of features, speed up the convergence speed of the model, enhance the training stability, and make the model easier to learn effective features.
[0085] In some preferred embodiments, the normalization calculation formula used in this embodiment is: wherein, is the normalized result of the ith independent variable, is the ith independent variable, is the minimum value of the independent variable, is the maximum value of the independent variable.
[0086] For example, Figure 2As shown, in the present embodiment, by using the above steps S1 to S5, the load data is processed by using various strategies, which can significantly enhance the information provided by the load data, help model training, and help improve the prediction accuracy of the model.
[0087] In the step S6, the normalized data is divided into independent variables X and dependent variables Y according to the improved PSO optimization algorithm, and the training set and the test set are obtained by further division.
[0088] The improved PSO optimization algorithm updates the optimal particle according to the dynamic inertia weight and the dynamic learning factor, and the dynamic inertia weight and the dynamic learning factor are dynamically changed according to the Levy flight mechanism and the Gaussian disturbance mechanism.
[0089] The power load prediction is to predict the load data in the future period of time through a period of historical load data, and therefore, after the load data characteristics are processed, the X (independent variable) and Y (dependent variable) decomposition needs to be performed for supervised model training. Different lengths of X and Y not only affect the model training speed, but also affect the final prediction result.
[0090] In the present embodiment, the length of X and Y is determined by the particle information of the optimal particle calculated by the improved PSO optimization algorithm, and the independent variable X and the dependent variable Y are divided. Since the optimal particle calculated by the improved PSO optimization algorithm is more accurate, the fitness of its particle information is higher, and therefore, the training set and the test set obtained by the division can be obviously optimized, and the adjustment difficulty of the training process can be obviously reduced.
[0091] In some preferred embodiments, the step S6 specifically includes the following steps:
[0092] S61: determining the length of the independent variable X and the dependent variable Y according to the particle information of the optimal particle output by the improved PSO optimization algorithm;
[0093] Here, the particle information is a parameter combination including X, Y, the number of neurons, the size of the convolution kernel, etc. The "X, Y" in it are the lengths of the independent variable X and the dependent variable Y, according to which the independent variable X and the dependent variable Y can be divided.
[0094] S62: dividing the normalized data into the independent variable X and the dependent variable Y according to the determined length of the independent variable X and the dependent variable Y;
[0095] S63: dividing the divided processed data into the training set and the test set according to a preset proportion.
[0096] In some preferred embodiments, as Figure 3The improved PSO optimization algorithm is specifically shown to include:
[0097] (1) randomly initializing each particle;
[0098] (2) evaluating each particle to obtain an optimal particle;
[0099] (3) determining whether the currently determined optimal particle meets an ending condition; if yes, ending the process and outputting the optimal particle; if no, performing (4);
[0100] (4) dynamically changing inertia weight w and learning factor by using a Levy flight mechanism and a Gaussian disturbance mechanism, and updating particle information, speed and position of each particle by using the dynamic inertia weight and the dynamic learning factor;
[0101] (5) reevaluating the fitness value of each updated particle;
[0102] (6) determining an optimal particle according to the fitness value of each particle and returning to perform (3).
[0103] Specifically, in the step (4), the Levy flight mechanism uses a Levy step generation algorithm, and the Gaussian disturbance mechanism uses a Gaussian disturbance algorithm.
[0104] Since the Gaussian disturbance can affect the position and speed in two dimensions, the disturbance is also divided into position disturbance and speed disturbance.
[0105] The position disturbance calculation formula in the Gaussian disturbance algorithm is: ,
[0106] wherein, is the position of particle i at t+1, is the position of particle i at t, is the speed of particle i at t+1, is a disturbance intensity coefficient, is a Gaussian distributed noise, is a disturbance intensity which decays according to the iteration number, and the decay formula is ;
[0107] The speed disturbance calculation formula in the Gaussian disturbance algorithm is: ,
[0108] wherein, is the inertia weight is the speed of particle i at t, is the learning factor, is a uniformly distributed random number, is the individual historical optimal position of particle i, is the global optimal position of the particle swarm;
[0109] The Levy step generation algorithm is: wherein, is a stability coefficient, controlling the heavy-tailed characteristics of the distribution, the smaller the value, the higher the probability of step length, the longer the step length, and the greater the parameter combination value span of exploration; u is a random number conforming to a normal distribution.
[0110] As a specific example, in the improved PSO optimization algorithm, a larger inertia weight w can be set initially to enhance global search, and gradually reduced in the later stage to improve local accuracy; the learning factors c1 and c2 can be dynamically adjusted by the number of iterations or the fitness value, for example, the individual experience is emphasized in the early stage (c1 is larger), and the group cooperation is emphasized in the later stage (c2 is larger), which can significantly improve the performance and effectively avoid the problems that may occur in the PSO algorithm of the prior art.
[0111] The improved PSO optimization algorithm used in this embodiment combines the Levy flight mechanism of generating Levy steps using the Mantegna algorithm, so it has strong long-distance jumping ability and can jump from one solution space to another solution space. When the particle state changes, the inertia weight w is affected and needs to be adaptively changed, so as to realize the dynamic change of the inertia weight w. Since random noise is introduced through the Gaussian disturbance mechanism to apply a small disturbance to the particle, resulting in variance adjustment, the learning factor also needs to be adaptively changed to improve the particle state, so as to realize the dynamic change of the learning factor .
[0112] Unlike the PSO algorithm of the prior art, which uses fixed inertia weight w and learning factors c1 and c2 to find the optimal particle, the velocity and position of the particle in the particle swarm are calculated, which can easily lead to low search efficiency, low search accuracy, and the problem of falling into a local optimal solution. The improved PSO optimization algorithm used in this embodiment realizes the dynamic change of the inertia weight w and the learning factors c1 and c2 through the Levy flight mechanism and the Gaussian disturbance, and updates the optimal particle accordingly, which enhances the ability to jump out of the local optimum and can effectively avoid the problem of falling into a local optimum, thereby significantly improving the local search accuracy and realizing global exploration.
[0113] At this point, the processing of the load data in the early stage of this embodiment is completed, and the following will start the building stage of the neural network model.
[0114] The S7 creates an initial CNN-LSTM-Transformer neural network model.
[0115] In combination with Figure 4 and Figure 5 For understanding, the created initial CNN-LSTM-Transformer neural network model takes a Sequential model based on Keras framework as the entry, and then sequentially connects a multi-channel convolutional layer CNN, an LSTM network, a Transformer neural network layer, and a fully connected layer.
[0116] Specifically, first, a Sequential model based on Keras framework is created, which provides a highly simplified interface for subsequent linearly stacked neural network layers. Then, a multi-channel convolutional layer CNN is built after the Sequential model interface in the Keras framework. Convolution along the time axis of time series data can extract spatial features: similarities or differences between regions, and capture load mutations, identify load differences between weekdays and weekends (based on the date attribute feature item), and also retain local significant features such as peak load. Among them, the convolution kernel size, receptive field size, etc. need to be set to adapt to different scales of local patterns, which are determined by the particle information updated by the improved PSO optimization algorithm. Then, an LSTM network is built. The data after convolution and pooling is calculated in the LSTM network. The gating mechanism of the LSTM's forget gate, input gate, and output gate can learn the rules of the data, capture the short-term and medium-term dependencies in the time series data, especially the trend, residual, lag value, etc. Feature information combined with LSTM memory unit information, discard unnecessary information, add new pattern information, and enhance the ability to capture dependencies. Then, a Transformer neural network layer is built. Through its multi-head attention mechanism, the weights of each part of the time series data are calculated, ignoring the time length limit, capturing long-distance dependencies, and enhancing the feature expression ability through feedforward propagation. Finally, a fully connected layer is built to output after the last layer of training. The pattern of time series data will be clear to the entire neural network model. In particular, the hyperparameters such as the number of network neurons and batch_size in the above created initial CNN-LSTM-Transformer neural network model are obtained by the improved PSO optimization algorithm provided in this embodiment.
[0117] In some preferred embodiments, the LSTM model is enhanced by a state reservation mechanism for modeling time series data, which reserves the cell state and hidden state (neuron weights, bias terms, etc.) after each batch data (small batches of data divided from the training set according to the batch size) training is completed, and passes these states to the next batch. Data enters the forget gate, input gate, and output gate, and then records the state. When each epoch data training is completed, the state between epochs is cleared according to the cleanState function in the callback method, so that each epoch data training is not affected by the previous round.
[0118] The step S8 trains the initial CNN-LSTM-Transformer neural network model using the training set according to the improved PSO optimization algorithm until the final optimal particle is determined, and obtains the CNN-LSTM-Transformer neural network model; wherein each training updates the hyperparameters of the model with the particle information updated by the improved PSO optimization algorithm in the previous training.
[0119] The hyperparameters include the receptive field of the convolutional network, the size of the convolutional kernel, the number of neurons of the LSTM, and the batch size for training, etc.
[0120] During the training process, when the end condition of the improved PSO optimization algorithm is met, the final optimal particle is obtained, and then the data is finally divided according to the optimal particle and the CNN-LSTM-Transformer neural network model is obtained. Finally, the model file is saved as the final prediction model.
[0121] The step S9 inputs the target test set obtained by processing the steps S1 to S5 into the CNN-LSTM-Transformer neural network model to obtain the short-term power load prediction result.
[0122] That is, to predict future power load data, the target test set needs to be obtained after processing steps S1 to S5, and then input into the CNN-LSTM-Transformer neural network model, i.e. the prediction model, to obtain the short-term power load prediction result.
[0123] In some preferred embodiments, in the process of model recursive prediction, each prediction result is part of the next prediction condition.
[0124] Specifically, the interval time granularity of the power load data is usually minute level, such as 15 minutes a point (a piece of power load data), and the power load prediction usually hopes to predict the power load data of the future 7 / 14 days, so as to complete such prediction, 672 / 1344 points need to be predicted. When creating a model for training, m points are usually used as independent variables, and n points starting from the m+1th point are used as dependent variables (n<50 will have better effect), and when using the model for prediction, it is considered that the last m points of the historical data are used to predict the next n points. In order to predict enough power load data, it is not enough to use a model once, and recursive prediction is needed, that is, the n points predicted in the first time are placed after m as part of the new historical data, and the n points are placed in the prediction result. Take the last m points from the new historical data to make the next prediction until the prediction result contains 672 / 1344 pieces of predicted power load data.
[0125] Embodiment Two
[0126] This embodiment is based on the above embodiment one, and by introducing "data conversion", the training time is greatly reduced on the basis of embodiment one, and the speed of model training and prediction is significantly improved.
[0127] The difference between this embodiment and embodiment one is that after obtaining the training set and the test set in the S6, the training set used for subsequent model training is also converted in data format, also known as shape conversion.
[0128] Specifically, the step S6 further includes:
[0129] SS6: converting the three-dimensional array (sample number, time step, feature number) corresponding to multiple time steps of the training set into the converted three-dimensional array corresponding to one time step.
[0130] Here, since the model of this embodiment uses LSTM neural network, it has requirements for the input format of data, which must be a three-dimensional array, and its general input format is (samples--sample number, time steps--time step, features--feature number). Therefore, it can be understood that the training set of this embodiment is a three-dimensional array in (samples--sample number, time steps--time step, features--feature number) data format before data format conversion, and the value of time steps--time step is greater than or equal to 2, and the value is usually large.
[0131] Using the input format described above when training the LSTM network will perform time steps of calculation and forward, backward propagation on the input time series data, which is very time-consuming. For example, a three-dimensional array with a shape of (1000, 100, 5), a sample size of 1000, a time step of 100, and a feature number of 5; when training in the neural network, it will be calculated in time steps, i.e. 100 times, including forward and backward propagation, which is time-consuming.
[0132] Here, based on the actual training, as long as the same data is obtained, the training effect should be consistent. By converting the data format of the data, the original three-dimensional array corresponding to multiple time steps is converted into processed data corresponding to only one time step; only one time step of calculation is needed, thereby greatly reducing the training time. For example, the three-dimensional array with a shape of (1000, 100, 5) described above is converted into a three-dimensional array with a shape of (1000, 1, 5).
[0133] In some preferred embodiments, the "conversion" in the above SS6 specifically includes:
[0134] The data corresponding to multiple time steps in the three-dimensional array is sorted and fused into one time step to obtain the processed data.
[0135] That is, all the data of time steps are sorted and fused into one time step.
[0136] In some preferred embodiments, as shown in Figure 6 The S6 and the SS6, respectively, specifically include:
[0137] The S6 specifically includes:
[0138] S61: Put every X rows of multi-dimensional data in the normalized structured two-dimensional data containing time, region ID, whether it is a working day, etc. features except for the power load feature (the power load value size also exists in the two-dimensional data, which is a separate column, which is part of the features needed for prediction, and also the column where the target data is located) into the independent variable (X) list, and put Y rows of power load features starting from the X+1 row into the dependent variable (Y) list; The "X" and "Y" in the "X rows" and "Y rows" are determined according to the improved PSO optimization algorithm.
[0139] S62: Convert the independent variable list and the dependent variable list into numpy arrays;
[0140] S63: Use the train_test_split method in the sklearn library to divide the training set and the test set in a ratio of 8:2; here, the division ratio can also be other numerical values, supporting flexible configuration.
[0141] The SS6 specifically comprises:
[0142] The SS6: the array data of the training set is sorted by time, and reshaped according to the specified shape (i.e. shape), so as to convert the three-dimensional array corresponding to multiple time steps into a three-dimensional array corresponding to one time step.
[0143] Optionally, the above-mentioned format conversion can be realized by using the reshape function of the numpy.array type in python, so as to obtain a shape corresponding to only one time step.
[0144] In some preferred embodiments, after the data is convolved and pooled by the multi-channel convolutional layer CNN, the data will be restored to a three-dimensional array corresponding to multiple time steps. Therefore, in the present embodiment, before the data output by the multi-channel convolutional layer CNN is input into the LSTM network, it further comprises:
[0145] The three-dimensional array corresponding to multiple time steps output by the multi-channel convolutional layer CNN is converted into a three-dimensional array corresponding to one time step, and then input into the LSTM network.
[0146] That is, the data output by the multi-channel convolutional layer CNN will be reshaped in shape before being input into the LSTM network. For example, a three-dimensional array with a shape of (10000, 100, 5) will be reshaped into a three-dimensional array with a shape of (10000, 1, 5).
[0147] The present embodiment can accelerate the training speed of the model by tens of times without affecting the training and prediction accuracy of the model by converting the data shape and fusing multiple time steps into one time step, and can also significantly improve the prediction speed.
[0148] Embodiment three
[0149] The present embodiment is further developed based on the above-mentioned embodiment one or embodiment two, and provides a computer readable storage medium having a computer readable computer program stored thereon, wherein the program can realize all steps contained in the power load prediction method of the electric energy meter data system according to the above-mentioned embodiment one or embodiment two when executed by a processor. The specific steps will not be described here, and details can be referred to the description of the above-mentioned embodiment one and embodiment two.
[0150] From the above description, it can be understood by those skilled in the art that all or part of the processes in the above technical solutions can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The program can include the processes of the above methods when executed. The program can also achieve the beneficial effects of the corresponding methods after being executed by a processor.
[0151] The storage medium can be a disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like.
[0152] Embodiment Four
[0153] This embodiment is based on any one of the above embodiments one to three and further expands the training and prediction process of the model through a preferred specific embodiment.
[0154] As shown in Figure 7 , the process includes:
[0155] S1 (not shown in the figure): A function for dividing independent variables, dependent variables, and dividing training sets and test sets is written in advance, and the input parameters of the function are the lengths of the independent variables and the dependent variables. Similarly, a function for creating a model and training is written in advance, and a model is created based on the Keras framework. First, there are two layers of conv1D, and the convolution kernel size and the receptive field are input parameters. One-dimensional convolution is performed on the tiled numpy array data to capture the features between regions. Then there are n layers of LSTM neural networks, and the gating mechanism will help to grasp the data regularity. n is an input parameter, and the number of LSTM neurons is also an input parameter. The information processed by LSTM is regarded as preliminary extracted information, which is input into the Transformer neural network. The Transformer global context capture ability is very strong, and combined with the multi-head attention mechanism, it can learn the feature relationship in different subspaces. In addition, early stopping strategy and learning rate dynamic adjustment callback strategy are added to optimize resource utilization and model training. The division function and the model creation and training function will be used in S3.
[0156] S2: Determine the upper and lower limits of various parameters that need to be put into the improved PSO optimization algorithm, call the improved PSO optimization algorithm, set the initial parameters of the algorithm itself, determine the fitness calculation method, which can be based on the coefficient of determination R-squared to calculate the fitness.
[0157] Compared with the prior art, the improved PSO optimization algorithm used in the embodiment has a larger search range and more accurate results due to the addition of the Levy flight mechanism and Gaussian disturbance; and the fitness calculation method is different, the R-squared, that is, the fitting degree of the model predicted data and the real data, is used as the basis, the default optimization direction, that is, the minimization direction, is used, the fitness is calculated by R-squared, the larger the R-squared, the smaller the fitness, and the PSO particle is better. Therefore, the optimal particle information obtained finally is more suitable for the data of the embodiment.
[0158] S3: Start training and improve the PSO optimization algorithm, determine the length of the dependent variable X and the independent variable Y through the data division information in the particle, call the division function to divide the data, and perform data format conversion, which not only reduces the number of experiments affected by the division of different length factors and independent variables on the model prediction ability, but also accelerates the model training speed by tens of times; then call the model creation and training function through the particle model creation and training information, set the model creation and training parameters, create a specific model, and set the training information, and start training using the divided and converted data.
[0159] S4: Calculate the fitness, update the particle, and iterate until the end condition is met to obtain the optimal particle; use the information of the optimal particle to divide the data and create the final model, and use the divided data to train the final model and save the final model.
[0160] S5: Recursively predict through the saved model, and use the prediction result of each time as part of the condition for the next prediction.
[0161] As shown in Figure 8 , it is an effect diagram for short-term power load prediction using the power load prediction method of the electric energy meter data system provided by the embodiment.
[0162] It is also the actual application of the embodiment of the present application in the electric energy meter system, wherein the blue curve represents the real power load data curve, and the green curve represents the power load prediction data curve obtained by predicting the model trained using the CNN-LSTM-Transformer power load prediction method. Figure 8 As can be seen, the data predicted by the power load prediction method of the electric energy meter data system is consistent with the trend of the real data, with a small difference, the determination coefficient R-squared of the two lines: 0.91 reaches an excellent level (when the value is closer to 1, it means that the model fitting effect is better; the value is closer to 0, the model effect is worse), and the root mean square error RMSE is 0.28, indicating that the deviation between the real value and the predicted value is also small. In summary, the power load prediction method of the electric energy meter data system provided by the embodiment is feasible, has high prediction accuracy, and can accurately predict the power load data.
[0163] Embodiment Five
[0164] The embodiment is further based on any of the above embodiments, and provides a CNN-LSTM-Transformer neural network model for short-term power load prediction.
[0165] The prediction model of the embodiment, i.e., the CNN-LSTM-Transformer neural network model, is structurally as shown in Figure 4 The model includes, as an entry, a Sequential model based on a Keras framework, and sequentially connected multi-channel convolutional layers CNN, an LSTM network, a Transformer neural network layer, and a fully connected layer.
[0166] The multi-channel convolutional layers CNN have kernel sizes and receptive field sizes determined according to particle information updated by the improved PSO optimization algorithm, and are used to extract spatial features in time series data, such as similarity or difference features between regions, capture load mutations, identify load differences of date attributes, and retain local significant features.
[0167] The LSTM network is used to learn the rules of data through a gating mechanism including a forgetting gate, an input gate, and an output gate, capture medium and short-term dependencies of all feature items in time series data, combine all feature items with LSTM memory cell information, and discard invalid information.
[0168] The Transformer neural network layer is used to calculate the weights of each feature item in time series data through a multi-head attention mechanism, capture long-distance dependencies, and enhance feature expression capability through feedforward propagation.
[0169] The CNN-LSTM-Transformer neural network model provided by the embodiment can achieve high-precision power load prediction function for power load data of different regions. The multi-channel convolutional layers CNN can extract spatial local patterns, the LSTM network can process time dependencies, the Transformer neural network layer can dynamically allocate weights for space-time features, and the model is trained according to the improved PSO optimization algorithm. Therefore, the combined network of the embodiment retains relatively complete data details, has strong adaptability to mutations, and is suitable for complex load data. In addition, the embodiment is used in combination with the power load prediction method provided in any of embodiments one to four, processes data by combining seasonal decomposition, increasing lag values, date attributes, and other strategies, enhances information provided by the data, and accelerates the model training speed by tens of times through converting data shapes when decomposing independent variables and dependent variables of the load data.
[0170] In summary, the power load prediction method and medium of the electric energy meter data system of the application combine CNN, LSTM and Transformer to realize full-scale feature extraction from micro to macro, solve the problem of insufficient modeling of single model for complex time series data, use the improved PSO optimization algorithm to no longer rely on experience to set parameters, reduce the experimental difficulty, significantly speed up the model training convergence speed, increase the global search ability, and improve the prediction accuracy by about 10%. Overall, the application provides an advanced and efficient power time series data prediction solution for the electric meter MDM system, provides a more reliable decision basis for power system dispatch, and provides a practical and innovative technology for the development of smart grids.
[0171] The above is only an embodiment of the application, and does not limit the patent scope of the application. Any equivalent transformation or direct or indirect application in related technical fields using the content of the application specification and drawings is also included in the patent protection scope of the application.
Claims
1. A method for power load forecasting of an electricity meter data system, characterized by, The method comprises the following steps: S1: grouping historical power load data according to area ID to obtain each group of data sets; S2: respectively performing seasonal decomposition on each group of data sets to obtain corresponding characteristic terms of each group of data sets; S3: determining whether each group of data sets belongs to an additive model or a multiplicative model according to the corresponding characteristic terms of each group of data sets, using a de-trending algorithm corresponding to the model to which each group of data sets belongs to de-seasonalize the load data in each group of data sets, and obtaining the de-seasonalized load data of each group of data sets as the corresponding characteristic term; S4: adding characteristic terms of lag values and date attributes of each group of data sets including load data; S5: normalizing all characteristic terms corresponding to each group of data sets; S6: dividing the normalized data into independent variables X and dependent variables Y according to the improved PSO optimization algorithm, and then dividing the normalized data into a training set and a test set; wherein the improved PSO optimization algorithm updates the optimal particle according to a dynamic inertia weight and a dynamic learning factor, and the dynamic inertia weight and the dynamic learning factor are dynamically changed according to the Levy flight mechanism and the Gaussian disturbance mechanism; S7: creating an initial CNN-LSTM-Transformer neural network model; S8: training the initial CNN-LSTM-Transformer neural network model using the training set according to the improved PSO optimization algorithm until the final optimal particle is determined, and obtaining a CNN-LSTM-Transformer neural network model; wherein the hyperparameters of the model are updated with the particle information updated by the improved PSO optimization algorithm in the last training each time; S9: inputting the target test set obtained by processing S1 to S5 into the CNN-LSTM-Transformer neural network model to obtain a short-term power load prediction result; The improved PSO optimization algorithm specifically comprises: (1) randomly initializing each particle; (2) evaluating each particle to obtain an optimal particle; (3) determining whether the optimal particle meets the end condition; if yes, ending the process and outputting the optimal particle; if no, executing (4); (4) dynamically changing the inertia weight and the learning factor through the Levy flight mechanism and the Gaussian disturbance mechanism, and updating the particle information, the speed and the position of each particle using the dynamic inertia weight and the dynamic learning factor; (5) re-evaluating the fitness value of each updated particle; (6) determining the optimal particle according to the fitness value of each particle and returning to execute (3); In the (4), the Levy step generation algorithm used by the Levy flight mechanism and the Gaussian disturbance algorithm used by the Gaussian disturbance mechanism are respectively: The position perturbation calculation formula in the Gaussian perturbation algorithm is: , wherein, is the position of particle i at t+1, is the position of particle i at t, is the velocity of particle i at t+1, is the perturbation intensity coefficient, is the Gaussian distributed noise, is the perturbation intensity which decays with the iteration number, the decay formula is ; The velocity perturbation calculation formula in the Gauss perturbation algorithm is: , wherein, is an inertia weight, is the velocity of particle i at t, is a learning factor, is a uniformly distributed random number, is the individual best position of particle i, is the global best position of the swarm. The Levy step generation algorithm is: wherein, is a stability coefficient and u is a random number following a normal distribution. is a stability coefficient and u is a random number following a normal distribution.
2. The power load forecasting method of an electric energy meter data system according to claim 1, wherein, The S6 further comprises: SS6: converting the original three-dimensional array corresponding to multiple time steps of the training set into a converted three-dimensional array corresponding to one time step.
3. The power load forecasting method of an electric energy meter data system according to claim 2, wherein, The conversion specifically comprises: sorting the data corresponding to multiple time steps in the original three-dimensional array corresponding to multiple time steps and merging them into one time step.
4. The power load forecasting method of an electric energy meter data system according to claim 1, wherein, The S6 specifically comprises: S61: Determine the lengths of the independent variable X and the dependent variable Y according to the particle information of the optimal particle output by the improved PSO optimization algorithm; S62: Divide the normalized data into the independent variable X and the dependent variable Y according to the lengths of the independent variable X and the dependent variable Y determined; S63: Divide the divided processed data into a training set and a test set according to a preset proportion.
5. The power load forecasting method of an electric energy meter data system according to claim 1, wherein, The S1 specifically includes: S11: Obtain historical power load data, which includes three columns of data of region ID, data time, and load data; S12: Group the historical power load data according to the region ID, and then sort each group of data sets according to the data time.
6. The power load forecasting method of an electric energy meter data system according to claim 1, wherein, The initial CNN-LSTM-Transformer neural network model is created based on the Keras framework, specifically including a multi-channel convolutional layer CNN, an LSTM network, a Transformer neural network layer, and a fully connected layer connected in sequence; The multi-channel convolutional layer CNN determines the convolution kernel size and receptive field size according to the particle information updated by the improved PSO optimization algorithm, and is used to extract spatial features in time series data, such as similarity or difference features between regions, capture load mutations, identify load differences of date attributes, and retain local significant features; The LSTM network is used to learn the rules of data through a gating mechanism including a forgetting gate, an input gate, and an output gate, capture the medium and short-term dependency relationships of all feature items in time series data, combine all feature items with LSTM memory cell information, and discard invalid information; The Transformer neural network layer is used to calculate the weight of each feature item in time series data through a multi-head attention mechanism, capture long-distance dependencies, and enhance feature expression ability through feedforward propagation.
7. The power load forecasting method of the electric energy meter data system according to claim 6, wherein, Before the data output by the multi-channel convolutional layer CNN is input into the LSTM network, it further includes: After the three-dimensional array corresponding to multiple time steps output by the multi-channel convolutional layer CNN is converted into a three-dimensional array corresponding to one time step, it is input into the LSTM network.
8. A computer readable storage medium having stored thereon a computer readable computer program, characterized in that, The program, when executed by a processor, can implement all steps included in the power load prediction method of the electric energy meter data system of any one of claims 1-7.
Citation Information
Patent Citations
Load prediction method based on PSO-CNN-LSTM model
CN114707750A
Power load prediction method based on NIWPSO + CNN + LSTM + Attention
CN119917842A