Power consumption time sequence prediction method and system based on self-attention mechanism, medium and processor
By combining a self-attention mechanism and a dual attention model with a convolutional neural network, the limitations of traditional power consumption prediction methods in terms of nonlinear features and long-term dependencies are overcome, achieving higher accuracy in power consumption prediction.
Patent Information
- Application Number
- CN202510793757.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional electricity consumption forecasting methods struggle to capture the nonlinear characteristics and long-term dependencies in electricity consumption data, resulting in limited forecast accuracy and an inability to distinguish the differences between different power units.
A time-series prediction method for power consumption based on a self-attention mechanism is adopted. Time and power unit features are extracted through a convolutional neural network, and feature fusion and optimization are performed by combining a dual attention mechanism. Backpropagation technology is used to optimize model parameters.
It improves the ability to capture nonlinear features and long-term dependencies, enhances prediction accuracy, is suitable for real-time prediction of large-scale power data, and reduces computational complexity and training time.
Smart Images

Figure CN120911734A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power consumption time series prediction, in particular to a power consumption time series prediction method, system, medium and processor based on a self-attention mechanism. BACKGROUND
[0002] With the acceleration of urbanization and the growth of energy demand, load forecasting of power systems becomes increasingly important. Accurate power consumption time series prediction helps optimize the scheduling of power grids, reduces energy waste, and improves power supply reliability. However, traditional power consumption prediction methods often rely on linear models such as time series analysis and regression analysis. These methods are difficult to capture the nonlinear characteristics and long-term dependencies of power consumption data, and the prediction accuracy is limited.
[0003] In recent years, deep learning technology has made breakthroughs in the field of power load forecasting. Methods based on recurrent neural networks and long short-term memory networks can effectively learn the time series characteristics of power consumption data, but have high computational complexity and long training time. At the same time, recurrent neural networks and long short-term memory networks have limited ability to capture long-term dependencies, making it difficult to predict power consumption over a longer period of time. Existing power load forecasting methods treat all power units as a whole for prediction, and cannot distinguish the differences between different power units, so the prediction accuracy needs to be improved. Therefore, it is necessary to develop a time series prediction method that can more accurately predict power consumption.
[0004] Therefore, a power consumption time series prediction method, system, medium and processor based on a self-attention mechanism are needed. SUMMARY
[0005] To solve the problems of high computational complexity, long training time, and limited ability to capture long-term dependencies in the prior art, the present application provides a power consumption time series prediction method, system, medium and processor based on a self-attention mechanism, which can achieve better results in terms of nonlinear feature and long-term dependency capture and prediction accuracy. The specific technical solutions are as follows:
[0006] A power consumption time series prediction method based on a self-attention mechanism, comprising:
[0007] S1: After obtaining the original data and performing data preprocessing, divide it into a training input set X and a training label set Y;
[0008] S2: Perform a first operation on the training input set X to obtain time features and power unit features;
[0009] S3: Calculate the fusion features F by double attention from the time features and power unit features;
[0010] S4: obtaining a prediction value Y' by performing a second operation on the fused feature F, and optimizing each calculation model parameter according to the prediction value Y';
[0011] S5: predicting a future power consumption value based on the optimized calculation model.
[0012] Further, in step S1, after obtaining the original data and performing data preprocessing, the original data is divided into a training input set X and a training label set Y, including the following steps:
[0013] S11: performing standardization processing on the original data;
[0014] S12: dividing the standardized data into a training input set and a training label set.
[0015] Further, in step S12, the standardized data is divided into a training input set and a training label set, including the following steps:
[0016] The standardized data is divided into individual data samples according to each power consumption unit and a key node of the power grid, and all data samples of the power consumption units and the key node of the power grid are collected to form a three-dimensional data set;
[0017] Each data sample in the three-dimensional data set D is divided into a plurality of time series data blocks d according to a time dimension t;
[0018] The data collected at the first t-1 time points in each time series data block d is divided into training input data d in , and all training input data d in is collected to form a training input set X;
[0019] The data collected at the tth time point in each time series data block d is divided into training labels d out , and all training labels d out is collected to form a training label set Y.
[0020] Further, in step S2, in the time feature and the power unit feature obtained by performing the first operation on the training input set X, the first operation adopts a convolution operation, including the following steps:
[0021] S21: performing two convolution operations on the standardized training input set X by using a convolution neural network (CNN) to change the dimension, and obtaining a time feature Tt and a power unit feature Te of the training data, respectively;
[0022] S22: for the time feature Tt, three feature vectors Q, K, and V of the time feature Tt are obtained by a convolution network, and the formula is as follows:
[0023] Q, K, V = ConV (Tt);
[0024] wherein, ConV represents a convolution operation, that is, the time feature Tt after convolution includes three feature vectors Q, K, and V;
[0025] S23: The power unit feature is also based on the same operation, and three feature vectors of the power unit feature are obtained through a convolution network.
[0026] Further, in step S3, the time feature and the power unit feature are calculated to obtain the fusion feature F through double attention, including the following steps:
[0027] S31: The three feature vectors of the time feature Tt are calculated through a self-attention mechanism to obtain the attention feature Attention(Q, K, V) of the time feature Tt, and the formula is as follows:
[0028]
[0029] wherein, p represents a feature dimension;
[0030] S32: The attention feature Attention(Q, K, V) obtained through calculation is calculated through a convolution neural network to obtain the updated time feature Tt';
[0031] S33: The three feature vectors of the power unit feature are calculated through the same operation to obtain the updated power unit feature Te';
[0032] S34: The updated time feature Tt' and the power unit feature Te' are fused through a convolution neural network to obtain the fusion feature F, and the fusion feature F is represented as:
[0033] F = ConV(Te' + Tt');
[0034] wherein, Tt' represents the updated time feature, Te' represents the updated power unit feature, and F represents the fusion feature;
[0035] S35: The fusion feature F is calculated again through a self-attention mechanism to obtain the updated fusion feature F.
[0036] Further, in step S4, the second operation adopts a convolution operation.
[0037] Further, in step S4, the predicted value Y' obtained by performing the second operation on the fusion feature F, and the predicted value Y' is used to optimize the parameters of each calculation model, including the following steps:
[0038] S41: The fusion feature F is calculated again through a convolution neural network to obtain the predicted value Y';
[0039] S42: calculate the loss value between the predicted value Y' and the training label d in the training label set Y, the formula is as follows: out
[0040]
[0041] Wherein, loss is the loss value;
[0042] S43: according to the loss value, the parameters of the above calculation model are updated by the back propagation technology, and after multiple iteration optimization, the calculation model with optimal parameters can be obtained.
[0043] A power consumption time series prediction system based on a self-attention mechanism, applied to the power consumption time series prediction method based on the self-attention mechanism described above, comprising:
[0044] A preparation module is used to obtain raw data and perform data preprocessing, and then divide into a training input set X and a training label set Y;
[0045] A first operation module is used to obtain time features and power unit features by performing a first operation on the training input set X;
[0046] A fusion module is used to obtain fusion features F by performing double attention calculation on the time features and the power unit features;
[0047] A second operation module is used to obtain a predicted value Y' by performing a second operation on the fusion features F, and optimize the parameters of each calculation model according to the predicted value Y';
[0048] A prediction module is used to predict future power consumption values based on the optimized calculation models.
[0049] A computer readable storage medium, the computer readable storage medium comprises a stored program, wherein when the program runs, the device where the computer readable storage medium is located is controlled to execute the power consumption time series prediction method based on the self-attention mechanism described above.
[0050] A processor for running a program, wherein the program runs to perform the power consumption time series prediction method based on the self-attention mechanism described above.
[0051] Compared with the prior art, the beneficial effects of the present application are:
[0052] 1. The nonlinear feature and the long-term dependence relationship have stronger capture ability
[0053] Traditional linear models such as time series analysis and regression analysis are difficult to capture the nonlinear characteristics and long-term dependencies in power consumption data. Recurrent neural networks (RNN) and long short-term memory networks (LSTM) can learn temporal features, but they have high computational complexity and limited long-term dependency capture capabilities.
[0054] The present solution can effectively mine nonlinear patterns and long-term dependencies across time dimensions through self-attention mechanisms, which can calculate the correlation weights between different time steps and power units in parallel, breaking through the limitations of traditional models.
[0055] 2. Double attention mechanism improves feature extraction comprehensiveness
[0056] Existing methods often predict power units as a whole, ignoring individual differences.
[0057] The present solution uses a double attention model:
[0058] Time dimension: mining temporal correlations (such as seasonality and periodicity) of data at different times through self-attention mechanisms;
[0059] Individual dimension: distinguishing the feature differences of different power units and grid nodes (such as industrial and residential power consumption patterns). The combination of the two can more comprehensively represent data features, solve the defects of traditional "one-size-fits-all" methods, and improve prediction relevance.
[0060] 3. Significant improvement in prediction accuracy
[0061] Convolution operation combined with attention mechanism: extract spatial features (such as power unit attributes) and temporal features through convolutional neural networks (CNN), then strengthen key feature weights through double attention mechanism, and reduce redundant information interference;
[0062] End-to-end optimization: calculate the difference between predicted values and true values through loss function, use backpropagation to iteratively optimize model parameters, and further improve prediction accuracy.
[0063] Experiments show that the present solution outperforms traditional RNN, LSTM, and other methods in terms of nonlinear feature capture and long-term prediction scenarios.
[0064] 4. Optimization of computational efficiency
[0065] Self-attention mechanisms can handle sequence data in parallel, avoiding the time-dependent computational bottleneck of recurrent networks, reducing training time costs, and making them more suitable for real-time prediction scenarios of large-scale power data.
[0066] 5. The power consumption time series prediction method based on the self-attention mechanism provided in the application can effectively capture the nonlinear characteristics and long-term dependencies in the power consumption data by using the self-attention mechanism, overcoming the limitations of traditional linear models; by using the double-attention model to learn in the time dimension and the individual dimension respectively, the time series characteristics of urban power consumption can be more comprehensively mined; the double-attention mechanism can more effectively extract data features, thereby improving the prediction accuracy, and through parameter updating training, the neural network parameters can be continuously optimized, further improving the prediction accuracy, and the application achieves better results in capturing nonlinear characteristics and long-term dependencies and prediction accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0067] In order to more clearly illustrate the specific embodiments of the application or the technical solutions in the prior art, the drawings needed in the specific embodiments or the prior art description will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual scale.
[0068] Figure 1 It is a flowchart of a power consumption time series prediction method based on a self-attention mechanism.
[0069] Figure 2 It is a structural schematic diagram of a power consumption time series prediction system based on a self-attention mechanism. DETAILED DESCRIPTION
[0070] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are some of the embodiments of the application, not all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0071] It should be understood that when used in the present application, the terms "include" and "contain" indicate the presence of described features, whole, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, whole, steps, operations, elements, components and / or sets thereof.
[0072] It should also be understood that the terms used in the specification of the application are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in the specification and the appended claims of the application, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0073] It should be further understood that the term "and / or" as used herein refers to any combination of one or more of the associated listed items, and all possible combinations, and includes these combinations.
[0074] Embodiment one
[0075] As Figure 1 The flowchart of a power consumption time series prediction method based on a self-attention mechanism is shown, including the following steps:
[0076] S1: After obtaining the original data and performing data preprocessing, divide it into training input set X and training label set Y.
[0077] Further, the original data includes the real-time measurement and recording of the current, voltage, power, power factor and other power consumption parameters of each power consumption unit in the city by the smart meter, and the temperature, humidity, vibration and other environmental parameters and electrical parameters monitored by the sensors on the key nodes of the power grid.
[0078] First, the power consumption parameters of each power consumption unit in the city measured and recorded by the smart meter, and the environmental parameters and electrical parameters monitored by the sensors on the key nodes of the power grid are standardized. The standardized data is divided. Assuming that a city has n power consumption units and key nodes of the power grid, the data collected in the entire city is
[0079] The data preprocessing includes the following steps:
[0080] S11: The data preprocessing includes standardizing the original data. Specifically, delete the error data, outliers and repeated records of the power consumption parameters of the power consumption unit in the original data generated in the collection process; for missing data, complete it by the mean filling method; then scale the data, the scaling range is in the range of 0 to 1.
[0081] S12: The data preprocessing includes dividing the standardized data into training input set and training label set.
[0082] The specific implementation steps are:
[0083] 1. The standardized data is divided into data samples according to each power consumption unit and key node of the power grid, such as based on the power consumption unit A, the power consumption parameters (i.e. power consumption parameters) generated during the operation of A are collected in time sequence as data samples D A ; all data samples of power consumption units and key nodes of the power grid are summarized to form a three-dimensional data set, represented as:
[0084] D∈R T×n×m ;
[0085] Wherein, t represents the total collected T time data, n represents the total number of power units, m represents the total number of power consumption parameters collected for each power unit A at each time, and the power consumption parameters include current, voltage, power, power factor and the like.
[0086] 2. Each data sample in the three-dimensional data set D is divided into a plurality of time series data blocks d according to the time dimension t.
[0087]
[0088] t is the time dimension or division window of the time series data block. Each data sample can be divided into T / t time series blocks in total.
[0089] 3. The data collected at the first t-1 time in each time series data block d is divided into training input data d in All training input data d in is summarized to form a training input set X, which is represented as:
[0090] d in ∈R (t-1)×n×m ;
[0091] 4. The data collected at the t time in each time series data block d is divided into training label d out All training labels d out is summarized to form a training label set Y, which is represented as:
[0092] d out ∈R 1×n×m ;
[0093] Finally, the divided time series data training input set X and the training label set Y are obtained.
[0094] In a time dimension, t is divided into a time series data block every 10 collection times, which is represented as:
[0095] d∈R 10×n×m
[0096] Then, a total of T / 10 time series blocks can be divided, and the first 9 elements d in ∈R 9 ×n×m in each time series data block are divided into the training input set X, and the 10th element d out ∈R 1×n×m is divided into the training label set Y.
[0097] S2: The time features and power unit features obtained by performing the first operation on the training input set X.
[0098] Further, the first operation can be a convolution operation, or a pooling operation, or other suitable feature extraction operation. In embodiments of the present application, the first operation is a convolution operation. Specifically, the first operation includes the following steps:
[0099] S21: Two convolution operations are performed on the standardized training input set X using a convolutional neural network (CNN) to change the dimensions, to obtain time features Tt and power unit features Te of the training data.
[0100] The time features are represented as:
[0101] Tt∈R t×m ;
[0102] The power unit features are represented as:
[0103] Te∈R n×m ;
[0104] S22: For the time features Tt, three feature vectors of the time features Tt are obtained through a convolution network, represented as:
[0105] Q,K,V=ConV(Tt);
[0106] wherein ConV represents a convolution operation, i.e., the time features Tt include three feature vectors Q, K, and V.
[0107] Specifically, for the time features Tt, three feature vectors Q, K, and V are obtained through a convolution network with a convolution kernel of 3 and a step size of 2.
[0108] S23: For the power unit features, three feature vectors of the power unit features are obtained through a convolution network based on the same operation.
[0109] S3: The time features and the power unit features are calculated to obtain fusion features F through double attention. Specifically, the following steps are included:
[0110] S31: One-way self-attention calculation is performed on the three feature vectors of the time features Tt using a self-attention mechanism to obtain attention features Attention(Q, K, V) of the time features Tt, wherein the attention features Attention(Q, K, V) are represented as:
[0111]
[0112] wherein p represents a feature dimension, and p takes a value of 768;
[0113] S32: The attention feature Attention(Q, K, V) is subjected to a convolution operation by using a convolutional neural network (CNN) to obtain an updated time feature Tt'.
[0114] S33: The same operation is performed on the three feature vectors of the power unit feature to obtain an updated power unit feature T e ′.
[0115] S34: The updated time feature T t ′ and the power unit feature T e ′ are subjected to feature fusion by using a convolutional neural network (CNN) to obtain a fusion feature F, which is represented as:
[0116] F = ConV(Te' + Tt');
[0117] wherein T t ′ represents the updated time feature, T e ′ represents the updated power unit feature, and F represents the fusion feature.
[0118] S35: The fusion feature F is subjected to attention calculation by using a self-attention mechanism again to obtain an updated fusion feature F.
[0119] It should be noted that the self-attention calculation is performed on the time feature and the power unit feature respectively in the one-time self-attention calculation stage, so as to respectively mine the correlation of the data in the time dimension and the individual dimension.
[0120] S4: A prediction value Y' is obtained by performing a second operation on the fusion feature F, and each calculation model parameter is optimized according to the prediction value Y'.
[0121] Further, the second operation can be a convolution operation, a fully connected layer, or other suitable output operation. In the embodiments of the present application, the second operation uses a convolution operation. Specifically, the following steps are included:
[0122] S41: The fusion feature F is subjected to a convolution operation by using a convolutional neural network (CNN) again to obtain a prediction value Y', which is represented as:
[0123] Y' ∈ R 1×n×m ;
[0124] wherein Y' represents the prediction value.
[0125] S42: A loss value between the prediction value Y' and a training label d out in a training label set Y is calculated, which is represented as:
[0126]
[0127] S43: According to the loss value, the parameters of the above calculation models are updated by a back propagation technique, and after multiple iterations and optimizations, the calculation models with optimal parameters are obtained.
[0128] S5: Based on the optimized calculation models, the future power consumption value is predicted.
[0129] Specifically, to predict the power consumption value at a time t', the data recorded by each power consumption unit of the city is taken out from the data D at the time t'-1 as the prediction input d in ∈R (t-1)×n×m The d in is input into the neural network model with optimal parameters, and the output value is the power consumption prediction value of each power consumption unit at a time t'.
[0130] The power consumption time series prediction method based on the self-attention mechanism provided in the application can effectively capture the nonlinear characteristics and long-term dependencies in the power consumption data, overcome the limitations of traditional linear models, and learn in the time dimension and individual dimension through the double attention model, so as to more comprehensively mine the time series characteristics of city power consumption. The double attention mechanism can more effectively extract data features, thereby improving the prediction accuracy. Through parameter updating and training, the neural network parameters can be continuously optimized, and the prediction accuracy is further improved. The application achieves better results in capturing nonlinear characteristics and long-term dependencies and prediction accuracy.
[0131] Embodiment Two
[0132] In an optional embodiment, the first operation can adopt a pooling operation. Specifically, a suitable pooling type can be selected according to the data characteristics and the prediction target, for example, the maximum pooling can extract the maximum value in each time series segment, the average pooling can extract the average value of each time series segment, the size and step of the pooling window are set, and the pooling operation is performed on each data sample d in the training input set X. Each pooling window outputs a feature vector, that is, the first feature.
[0133] In an optional embodiment, the second operation uses a fully connected layer, all elements in the second feature are connected into a one-dimensional vector, the flattened vector is input into one or more fully connected layers, a nonlinear activation function is applied after each fully connected layer to introduce a nonlinear transformation, and the output of the last fully connected layer is the prediction value.
[0134] Embodiment Three
[0135] As Figure 2 Fig. 1 shows a structural diagram of a power consumption time series prediction system based on a self-attention mechanism, which is applied to the power consumption time series prediction method based on a self-attention mechanism described above, and includes:
[0136] a preparation module for obtaining raw data and performing data preprocessing, and dividing the data into a training input set X and a training label set Y;
[0137] a first operation module for performing a first operation on the training input set X to obtain time features and power unit features;
[0138] a fusion module for calculating fusion features F by double attention from the time features and the power unit features;
[0139] a second operation module for performing a second operation on the fusion features F to obtain a prediction value Y', and optimizing each calculation model parameter according to the prediction value Y';
[0140] a prediction module for predicting future power consumption values based on the optimized calculation models.
[0141] Embodiment four
[0142] A computer-readable storage medium including a stored program, wherein the program controls the device where the computer-readable storage medium is located to perform the power consumption time series prediction method based on a self-attention mechanism described above when the program is running.
[0143] Embodiment five
[0144] A processor for running a program, wherein the program performs the power consumption time series prediction method based on a self-attention mechanism described above when the program is running.
[0145] The present application discloses a power consumption time series prediction method, system, medium and processor based on a self-attention mechanism, which relates to the technical field of power consumption time series prediction. The method includes obtaining raw data and dividing it into a training input set and a label set after preprocessing, extracting time and power unit features through convolution operation, calculating fusion features through double attention, obtaining prediction values through convolution operation and optimizing model parameters, and finally predicting future power consumption values based on the optimized model. The system includes preparation, first operation, fusion, second operation and prediction modules. The computer-readable storage medium and the processor respectively store and run the program for executing the method. The present application effectively captures the nonlinear features and long-term dependencies of the data through the self-attention mechanism, improves the feature extraction capability through the double attention model, and improves the prediction accuracy.
[0146] Those skilled in the art can understand that the units of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0147] In the embodiments provided by the present application, it should be understood that the division of units is only a logical functional division. In actual implementation, there can be another division manner. For example, multiple units can be combined into one unit, one unit can be split into multiple units, or some features can be ignored, etc.
[0148] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0149] When the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0150] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and they should be covered in the scope of the present application.
Claims
1. A power consumption time series prediction method based on a self-attention mechanism, characterized in that, The method comprises the following steps: S1: obtaining original data and performing data preprocessing, and then dividing into a training input set X and a training label set Y; S2: performing a first operation on the training input set X to obtain time characteristics and power unit characteristics; S3: calculating the fusion characteristics F of the time characteristics and the power unit characteristics through double attention; S4: obtaining a prediction value Y by performing a second operation on the fused features F ′ , according to the prediction value Y ′ optimizing parameters of each computing model; S5: predicting the future power consumption value based on the optimized calculation model.
2. The power consumption time series forecasting method based on self-attention mechanism according to claim 1, characterized in that, In step S1, after obtaining the original data and performing data preprocessing, the original data is divided into a training input set X and a training label set Y, which comprises the following steps: S11: standardizing the original data; S12: dividing the standardized data into a training input set and a training label set.
3. The power consumption time series forecasting method based on self-attention mechanism according to claim 1, characterized in that, In step S12, the standardized data is divided into a training input set and a training label set, which comprises the following steps: The standardized data is divided into data samples according to each power unit and a key node of a power grid, and all data samples of the power units and the key node of the power grid are collected to form a three-dimensional data set; Each data sample in the three-dimensional data set D is divided into a plurality of time sequence data blocks d according to a time dimension t; dividing the data collected at the first t-1 time instants in each time series data block d into training input data d in summing all the training input data d in forming a training input set X; dividing the data collected at the t-th time point in each time-series data block d into training labels d out , and aggregating all the training labels d out to form a training label set Y.
4. The power consumption time series forecasting method based on self-attention mechanism according to claim 1, characterized in that, In step S2, the first operation adopts a convolution operation, and specifically comprises the following steps: S21: performing twice convolution operation on the standardized training input set X through a convolutional neural network (CNN) to change the dimension, and obtaining time characteristics Tt and power unit characteristics Te of the training data respectively; S22: for the time characteristics Tt, a convolution network is used to obtain three feature vectors Q, K and V of the time characteristics Tt, and the formula is as follows: Q, K, V = ConV (Tt); Wherein, ConV represents a convolution operation, that is, the time characteristics Tt include three feature vectors Q, K and V after convolution; S23: for the power unit characteristics, a convolution network is also used to obtain three feature vectors of the power unit characteristics.
5. The self-attention mechanism based power consumption time series forecasting method according to claim 4, characterized in that, In step S3, the time characteristics and the power unit characteristics are calculated through double attention to obtain the fusion characteristics F, which comprises the following steps: S31: one-way self-attention calculation is performed on the three feature vectors of the time characteristics Tt through a self-attention mechanism to obtain the attention characteristics Attention (Q, K, V) of the time characteristics Tt, and the formula is as follows: Wherein, p represents the feature dimension; S32: Convolution operation is performed on the calculated attention feature Attention(Q, K, V) by using a convolutional neural network to obtain an updated time feature T t ′ ; S33: The same operation is performed on the three eigenvectors of the power unit feature, and the updated power unit feature T is obtained e ′ ; S34: adopting a convolutional neural network to update the time feature T t ′ , the power unit feature T e ′ performing feature fusion to obtain a fused feature F, and the fused feature F is represented as: F = ConV(Te ′ + Tt ′ ); wherein T t ′ denotes the updated time feature, T e ′ denotes the updated power consuming unit feature, F denotes a fused feature; S35: the fusion characteristics F are calculated again through the self-attention mechanism to obtain updated fusion characteristics F.
6. The power consumption time series forecasting method based on self-attention mechanism according to claim 5, characterized in that, In step S4, the second operation adopts a convolution operation.
7. The power consumption time series forecasting method based on self-attention mechanism according to claim 6, characterized in that, In step S4, the prediction value Y is obtained by performing a second operation on the fusion feature F ′ , according to the prediction value Y ′ Optimizing each calculation model parameter, comprising the following steps: S41: the fusion feature F is again subjected to convolution operation by a convolutional neural network to obtain a prediction value Y ′ ; S42: Calculate the predicted value Y ′ The loss value between the training label d out in the training label set Y and the predicted value Y is calculated as follows: Wherein, loss is a loss value; S43: the parameters of the above calculation model are updated through the back propagation technology according to the loss value, and after multiple iterations and optimization, the calculation model with the optimal parameters can be obtained.
8. A power consumption time series prediction system based on self-attention mechanism, characterized in that, The method is applied to the power consumption time sequence prediction method based on the self-attention mechanism in any one of claims 1 to 7, comprising: a preparation module for obtaining original data and performing data preprocessing, and then dividing into a training input set X and a training label set Y; a first operation module for performing a first operation on the training input set X to obtain time characteristics and power unit characteristics; a fusion module configured to obtain a fusion feature F by double attention calculation of the time feature and the power unit feature; a second operation module, configured to perform a second operation on the fused feature F to obtain a prediction value Y ′ , according to the prediction value Y ′ optimize parameters of each calculation model; a prediction module configured to predict a future power consumption value based on the optimized calculation models.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the power consumption time series prediction method based on the self-attention mechanism in any one of claims 1 to 7 when the program is running.
10. A processor, comprising: The processor is configured to run a program, wherein the program executes the power consumption time series prediction method based on the self-attention mechanism in any one of claims 1 to 7 when the program is running.
Citation Information
Cited By
Power consumption dynamic prediction method based on physical-semantic dual-channel modulation
CN121303468A
Power consumption dynamic prediction method based on physical-semantic dual-channel modulation
CN121303468B