Short-term power load prediction method and device
Through EEMD and MIC analysis combined with improved efficient multi-scale attention module and adaptive time-frequency integrated network, the problem of poor accuracy in the existing short-term power load prediction methods is solved, and the in-depth processing and accurate prediction of power load data is achieved.
Patent Information
- Application Number
- CN202510801982.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-08-26
AI Technical Summary
The existing short-term power load prediction methods cannot effectively explore the periodicity of power loads, resulting in unsatisfactory prediction accuracy.
The integrated empirical modal decomposition (EEMD) and maximum information coefficient (MIC) are used to analyze and extract the eigenmodal function components of load power data. Combined with an improved efficient multi-scale attention module and an adaptive time-frequency integrated network (ATFNet), the power system operation data is deeply processed to generate short-term power load prediction results.
The accuracy of short-term power load prediction is improved, the trend and periodic characteristics in the power load data are fully explored, the correlation between load power and various influencing factors is enhanced, and the accuracy of prediction is improved.
Smart Images

Figure CN120541618A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system dispatching, and in particular to a short-term power load forecasting method and device. Background Art
[0002] As the energy consumption side of the power system, electricity load limits the absorption of new energy to a certain extent. Accurate prediction of electricity load is conducive to the absorption and regulation of new energy power generation and promotes the further development of new power systems.
[0003] Power load is closely related to user behavior. Its composition is complex and diverse, influenced by various factors such as temperature, humidity, electricity prices, and time. These factors have varying degrees of influence on users, making user behavior both random and regular. Ultimately, power load sequences exhibit significant complexity, randomness, and nonlinearity, while also exhibiting temporal characteristics such as trends and periodicity.
[0004] Existing short-term power load forecasting methods primarily rely on various deep neural network models, such as recurrent neural networks (RNNs) represented by long short-term memory networks, convolutional networks, and time-convolutional networks. However, these networks only process load sequences from a temporal perspective and fail to further explore the cyclical nature of power loads, resulting in suboptimal forecasting accuracy. Summary of the Invention
[0005] The present invention provides a short-term power load forecasting method and device, which are used to solve the technical problem that the existing short-term power load forecasting method leads to unsatisfactory forecasting accuracy.
[0006] A first aspect of the present invention provides a short-term power load forecasting method, comprising:
[0007] Acquire the operating data of the power system to be tested, pre-process the operating data of the power system to be tested, and output the target power system operating data;
[0008] generating a target input tensor based on the target power system operation data;
[0009] A preset short-term power load forecasting model is used to perform forecasting based on the target input tensor to generate a short-term power load forecasting result.
[0010] Optionally, preprocessing the power system operation data to be tested and outputting target power system operation data includes:
[0011] performing an outlier elimination operation on the power system operation data to be tested to generate intermediate power system operation data;
[0012] The intermediate power system operation data is normalized to generate target power system operation data.
[0013] Optionally, the target power system operation data includes load power data, temperature data, humidity data, time data, and electricity price data; and generating a target input tensor based on the target power system operation data includes:
[0014] Performing collective empirical mode decomposition on the load power data to generate a plurality of intrinsic mode function components corresponding to the load power data;
[0015] Performing data gridding on the load power data, the temperature data, the humidity data, the time data, and the electricity price data using each of the intrinsic mode function components, and outputting a data grid corresponding to each of the intrinsic mode function components;
[0016] Calculating the mutual information corresponding to each of the data grids using the joint probability distribution and the marginal distribution corresponding to each of the data grids;
[0017] Normalizing the mutual information corresponding to each of the data grids, and outputting the normalized mutual information corresponding to each of the data grids;
[0018] Comparing the normalized mutual information corresponding to each of the data grids with a preset information threshold;
[0019] Taking any intrinsic mode function component corresponding to the normalized mutual information greater than the preset information threshold as the target intrinsic mode function component;
[0020] The target intrinsic mode function component, the load power data, the temperature data, the humidity data, the time data, and the electricity price data are combined to generate a target input tensor.
[0021] Optionally, the preset short-term power load forecasting model includes an improved efficient multi-scale attention module and an adaptive time-frequency integrated network; the step of using the preset short-term power load forecasting model to perform forecasting based on the target input tensor to generate a short-term power load forecast result includes:
[0022] Grouping the target input tensors to determine grouped tensors;
[0023] Using the improved efficient multi-scale attention module to perform enhanced feature extraction on the grouped tensors to generate an enhanced feature sequence;
[0024] Jointly modeling the enhanced feature sequence through the adaptive time-frequency integration network to generate a complex result;
[0025] Converting the complex number result to generate a real number result;
[0026] Denormalize the real number result to generate a short-term power load forecast result.
[0027] Optionally, the improved efficient multi-scale attention module includes a convolution layer, an average pooling layer, an average pooling layer in the horizontal direction, an average pooling layer in the vertical direction, a Softmax activation function layer, and a Sigmoid activation function layer; the improved efficient multi-scale attention module is used to perform enhanced feature extraction on the grouped tensors to generate an enhanced feature sequence, including:
[0028] Using the convolution layer, the average pooling layer in the horizontal direction, and the average pooling layer in the vertical direction to perform convolution operation, average pooling in the horizontal direction, and average pooling in the vertical direction on the grouped tensors, respectively, and output convolution features, horizontal average pooling features, and vertical average pooling features;
[0029] Using the average pooling layer and the Softmax activation function layer to perform average pooling and nonlinear mapping on the convolution features respectively, to generate a first average pooling feature and a first nonlinear mapping feature;
[0030] Splicing the horizontal axis average pooling feature and the vertical axis average pooling feature to generate a spliced feature;
[0031] Perform feature weighting on the splicing features through the Sigmoid activation function layer and output a fusion weight coefficient;
[0032] Multiplying the fusion weight coefficient and the grouped tensor to generate a fusion feature, performing group normalization on the fusion feature, and outputting a group normalized feature;
[0033] Using the average pooling layer and the Softmax activation function layer to perform average pooling and nonlinear mapping on the group normalized features, respectively, to generate a second average pooling feature and a second nonlinear mapping feature;
[0034] Performing matrix multiplication on the second average pooling feature and the first nonlinear mapping feature, and outputting a first matrix multiplication feature;
[0035] Performing matrix multiplication on the second nonlinear mapping feature and the first average pooling feature, and outputting a second matrix multiplication feature;
[0036] The first matrix multiplication feature and the second matrix multiplication feature are added to output an added feature, and the added feature and the grouped tensor are multiplied to output an enhanced feature sequence.
[0037] Optionally, the adaptive time-frequency integration network includes a frequency domain global dependency capture module, a time domain local dependency capture module, and a dominant harmonic series energy weighting module; and the enhanced feature sequence is jointly modeled by the adaptive time-frequency integration network to generate a complex result, including:
[0038] Using the frequency domain global dependency capture module to globally capture the enhanced feature sequence to generate global features;
[0039] Locally capturing the enhanced feature sequence through the time-domain local dependency capture module to generate local features;
[0040] Performing Fourier transform on the enhanced feature sequence, and outputting Fourier transform features and complex spectral coefficients corresponding to the Fourier transform features;
[0041] The dominant harmonic series energy weighting module is used to perform weighting according to the Fourier transform feature and the complex spectrum coefficient corresponding to the Fourier transform feature, and output a frequency domain weight;
[0042] A complex result is generated according to the frequency domain weight, the global feature, and the local feature.
[0043] Optionally, the frequency domain global dependency capture module includes a complex-valued spectrum attention network, a complex normalization module, a complex ReLU activation function layer, a random dropout layer, and a restored linear layer; the frequency domain global dependency capture module is used to globally capture the enhanced feature sequence to generate global features, including:
[0044] Performing an extended Fourier transform on the enhanced feature sequence and outputting an extended Fourier transform feature;
[0045] Using the extended Fourier transform features as input to a complex-valued spectral attention network to generate an attention output;
[0046] Adding the attention output to the enhanced feature sequence to determine a first added feature, and performing complex normalization on the first added feature using a complex normalization module to generate a first complex normalized feature;
[0047] Performing complex linear mapping on the first complex normalized feature, outputting a complex linear mapping feature, and using the complex linear mapping feature as an input of a complex ReLU activation function layer, outputting an intermediate feature;
[0048] A random dropout layer and a linear reduction layer are used to sequentially drop and linearly reduce the intermediate features to generate reduced features;
[0049] Randomly discarding the restored features through a random discard layer to generate random discard features;
[0050] Adding the randomly discarded feature and the first complex normalized feature to output a second added feature, and performing complex normalization on the second added feature through a complex normalization module to generate a second complex normalized feature;
[0051] Projecting the second complex normalized features to generate global features.
[0052] Optionally, the temporal local dependency capture module includes a self-attention mechanism module and a linear head; and locally capturing the enhanced feature sequence by the temporal local dependency capture module to generate local features includes:
[0053] Segmenting the enhanced feature sequence to generate a plurality of feature blocks;
[0054] Taking each of the feature blocks as the input of the self-attention mechanism-based module and outputting a self-attention output;
[0055] The self-attention output is linearly projected using the linear head to generate local features.
[0056] Optionally, the training process of the preset short-term power load forecasting model includes:
[0057] Acquiring a training power system operation data set, and preprocessing the training power system operation data set to generate an intermediate training power system operation data set;
[0058] generating a target training power system operation data set based on the intermediate training power system operation data set;
[0059] The target training power system operation data set is used to perform model training on the initial short-term power load forecasting model, and the trained preset short-term power load forecasting model is determined.
[0060] A second aspect of the present invention provides a short-term power load forecasting device, comprising:
[0061] An acquisition module is used to acquire the operating data of the power system to be tested, pre-process the operating data of the power system to be tested, and output the target power system operating data;
[0062] A generating module, configured to generate a target input tensor based on the target power system operation data;
[0063] The prediction module is used to use a preset short-term power load prediction model to make predictions based on the target input tensor to generate a short-term power load prediction result.
[0064] It can be seen from the above technical solutions that the present invention has the following advantages:
[0065] The above scheme of the present invention provides a short-term power load forecasting method. First, the operating data of the power system to be measured is obtained, and the operating data of the power system to be measured is preprocessed to output the target power system operating data; then, based on the target power system operating data, a target input tensor is generated; finally, a preset short-term power load forecasting model is used to make a forecast based on the target input tensor to generate a short-term power load forecasting result; based on the above scheme, the present invention uses a preset short-term power load forecasting model to make a forecast based on the generated target input tensor, which can deeply process the characteristics of the power system operating data itself, thereby improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0067] Figure 1 A flowchart of a short-term power load forecasting method provided in accordance with the first embodiment of the present invention;
[0068] Figure 2 The original load power curve diagram provided in the first embodiment of the present invention;
[0069] Figure 3 A schematic diagram of the results of EEMD decomposition of a load power curve provided in the first embodiment of the present invention;
[0070] Figure 4 This is a schematic diagram of the structure of the improved efficient multi-scale attention module provided in Example 1 of the present invention;
[0071] Figure 5 A schematic diagram of the structure of an adaptive time-frequency integration network provided in Example 1 of the present invention;
[0072] Figure 6 A schematic diagram of the structure of an adaptive time-frequency integration network provided in Example 1 of the present invention;
[0073] Figure 7 A flowchart of the steps of the training process of the preset short-term power load forecasting model provided in the second embodiment of the present invention;
[0074] Figure 8 A comparison chart of the power load forecast curves predicted by the preset short-term power load forecast model provided in Example 2 of the present invention, the BiLSTM method, and the CNN-LSTM method;
[0075] Figure 9 This is a structural block diagram of a short-term power load forecasting device provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0076] The embodiments of the present invention provide a short-term power load forecasting method and apparatus, which are used to solve the technical problem that existing short-term power load forecasting methods result in unsatisfactory forecasting accuracy.
[0077] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0078] See also Figure 1 , Figure 1 This is a flowchart of the steps of a short-term power load forecasting method provided in Example 1 of the present invention.
[0079] The present invention provides a short-term power load forecasting method, comprising:
[0080] Step 101: Acquire the operating data of the power system to be tested, pre-process the operating data of the power system to be tested, and output the target power system operating data.
[0081] Specifically, step 101 may include the following sub-steps S11-S12:
[0082] S11, performing an outlier elimination operation on the power system operation data to be measured to generate intermediate power system operation data;
[0083] S12. Normalize the intermediate power system operation data to generate target power system operation data.
[0084] It should be noted that the power system operation data is initially filtered out of outliers, and the maximum and minimum values are selected. The power system operation data is normalized using the Min-Max method to obtain the normalized data, which is the target power system operation data. The power system operation data includes load power data, temperature data, humidity data, time data, and electricity price data.
[0085] Furthermore, the specific operation of normalization using the Min-Max method is:
[0086] ;
[0087] in, is the normalized data; is the data before normalization; It is the minimum value in the power system operation data; It is the maximum value in the power system operation data.
[0088] Step 102: Generate a target input tensor based on the target power system operation data.
[0089] Specifically, step 102 may include the following sub-steps S21-:
[0090] Step S21: performing collective empirical mode decomposition on the load power data to generate multiple intrinsic mode function components corresponding to the load power data;
[0091] Please note that Figure 2-Figure 3 , performing ensemble empirical mode decomposition (EEMD) on the load power data in the target power system operation data, and obtaining multiple intrinsic mode data (i.e., intrinsic mode function components) of EEMD decomposition.
[0092] Specifically, firstly, the parameters such as the number of noise additions and the amplitude of white noise are set; then, in order to enhance the frequency component recognition ability of the signal, a white noise sequence that obeys the standard normal distribution is introduced into the original time series signal in each experiment. Suppose the original signal is , the white noise added in the i-th trial is , then the new signal after disturbance can be constructed In this way, a set of sample sequences containing random perturbations can be obtained for subsequent empirical mode decomposition. Perform EMD decomposition, construct local extreme value envelopes, extract IMF components, strip the extracted IMF (Intrinsic Mode Function) components from the signal, update the residual signal, and repeat the iterative extraction until the corresponding conditions are met to obtain a set of IMF components:
[0093] ;
[0094] in, represents the k-th order IMF component obtained by EMD decomposition of the i-th noisy signal; is the residual term remaining after signal decomposition; K is the number of IMF components finally decomposed.
[0095] Furthermore, repeat the above steps to calculate N noisy signals. A corresponding set of IMF components can be obtained, and the IMF components of the same order are averaged point by point to remove the influence of noise, that is: ,in, Indicates signal The k-th order IMF component of , that is, multiple intrinsic mode function components corresponding to the load power data.
[0096] Step S22: gridding the load power data, temperature data, humidity data, time data, and electricity price data using each intrinsic mode function component, and outputting a data grid corresponding to each intrinsic mode function component;
[0097] Step S23: Calculate the mutual information corresponding to each data grid using the joint probability distribution and marginal distribution corresponding to each data grid;
[0098] Step S24: normalize the mutual information corresponding to each data grid, and output the normalized mutual information corresponding to each data grid;
[0099] Step S25: Compare the normalized mutual information corresponding to each data grid with a preset information threshold;
[0100] Step S26: taking any intrinsic mode function component corresponding to the normalized mutual information greater than a preset information threshold as a target intrinsic mode function component;
[0101] Step S26: Combine the target intrinsic mode function components, load power data, temperature data, humidity data, time data, and electricity price data to generate a target input tensor.
[0102] It should be noted that the maximum information coefficient (MIC) is used to analyze the generalized correlation between the load power and its various modal data and various influencing factors, capture linear and nonlinear relationships, and select appropriate modal components as the input of the prediction model based on the correlation. Specifically, the load data and the modal data obtained after decomposition are gridded with temperature, humidity, time and electricity price respectively, that is, the load power data, temperature data, humidity data, time data and electricity price data are gridded using each intrinsic mode function component, and the data grid corresponding to each intrinsic mode function component is output, so as to calculate the mutual information of each grid. :
[0103] ;
[0104] in, is the joint probability distribution, and is the marginal distribution.
[0105] Furthermore, the mutual information of each grid is normalized to [0, 1] to eliminate the bias of complex grid structures on mutual information:
[0106] ;
[0107] in, is the normalized mutual information; The x in represents that the value range of the variable is divided into intervals, and the y in represents that the value range of the variable is divided into intervals.
[0108] Furthermore, the normalized mutual information matrix M after traversing all possible grid combinations is taken as the final MIC value, which captures the hidden relationship between the load power data and the modal data with more prominent periodicity after decomposition and the influencing factors.
[0109] Furthermore, based on the final MIC value, the corresponding modal decomposition quantity (i.e., the target intrinsic mode function component) is selected. The resulting multiple or single target intrinsic mode function components, load power data, temperature data, humidity data, time data, and electricity price data are combined to form a new feature quantity (i.e., the target input tensor), highlighting the periodicity of load power and strengthening the coupling between load power and various influencing factors.
[0110] Step 103: Use a preset short-term power load forecasting model to make a forecast based on the target input tensor to generate a short-term power load forecast result.
[0111] The preset short-term power load forecasting model includes the improved efficient multi-scale attention module (EMA) and the adaptive time-frequency integration network (ATFNet).
[0112] Specifically, step 103 may include the following sub-steps S31-S35:
[0113] Step S31: group the target input tensors and determine the grouped tensors;
[0114] It should be noted that the input tensor is reshaped into And divide G sub-features according to the number of channels to obtain a tensor of shape [C / G,H,W], that is, the grouped tensor, so that the feature space is segmented and the features in each subgroup can be calculated in parallel and independently modeled in the subsequent processing process, where C is the number of sample channels, which is set to the number of input features, H is the sample height, which is set to 1, and W is the sample width, which is set to the time step of the sample.
[0115] Step S32: Use the improved efficient multi-scale attention module to perform enhanced feature extraction on the grouped tensors to generate an enhanced feature sequence.
[0116] Please note that Figure 4 The improved efficient multi-scale attention module includes a convolution layer (i.e., a 3*3 convolution layer), an average pooling layer (Avgpool), an average pooling layer in the horizontal direction (X Avgpool), an average pooling layer in the vertical direction (Y Avgpool), a Softmax activation function layer, and a Sigmoid activation function layer. Figure 4 Group Norm in it is group normalization, and Matmul is matrix multiplication.
[0117] Furthermore, step S32 may include the following sub-step S321:
[0118] Step S321: Use a convolution layer, an average pooling layer in the horizontal direction, and an average pooling layer in the vertical direction to perform convolution operations, average pooling in the horizontal direction, and average pooling in the vertical direction on the grouped tensors, and output convolution features, horizontal average pooling features, and vertical average pooling features;
[0119] Step S322: Use an average pooling layer and a Softmax activation function layer to perform average pooling and nonlinear mapping on the convolution features, respectively, to generate a first average pooling feature and a first nonlinear mapping feature;
[0120] Step S323: concatenate the horizontal axis average pooling features and the vertical axis average pooling features to generate a concatenated feature;
[0121] Step S324: perform feature weighting on the splicing features through a Sigmoid activation function layer and output a fusion weight coefficient;
[0122] Step S325: Multiply the fusion weight coefficient and the grouped tensor to generate a fusion feature, and perform group normalization on the fusion feature to output the group normalized feature;
[0123] Step S326: Use an average pooling layer and a Softmax activation function layer to perform average pooling and nonlinear mapping on the group normalized features, respectively, to generate a second average pooling feature and a second nonlinear mapping feature;
[0124] Step S327: performing matrix multiplication on the second average pooling feature and the first nonlinear mapping feature, and outputting a first matrix multiplication feature;
[0125] Step S328: performing matrix multiplication on the second nonlinear mapping feature and the first average pooling feature, and outputting a second matrix multiplication feature;
[0126] Step S329: Add the first matrix multiplication feature and the second matrix multiplication feature, output the added feature, and multiply the added feature and the grouped tensor to output an enhanced feature sequence.
[0127] It should be noted that the present invention sets a parallel subnetwork (an improved efficient multi-scale attention module) as a front-end feature extractor to fully capture both global and local information. The first branch extracts global context information along the height and width directions respectively through 1D global average pooling, then concatenates the two dimensions and fuses and reduces the dimension through 1×1 convolution to obtain global features for the entire spatial area. Subsequently, two nonlinear Sigmoid functions are used to fit the two-dimensional binomial distribution on the linear convolution. The second branch uses 3×3 convolution to extract local neighborhood information from the original grouped features. The two subnetworks respectively output the feature weights after the softmax function and the corresponding local features.
[0128] Furthermore, matrix multiplication is used to perform dot multiplication of the weights of one branch with the spatial features of the other branch, and the two sets of results are added together to generate a set of adaptive attention weights across the height and width dimensions. After the attention weights are activated by Sigmoid, they are multiplied element-wise with the original grouped features and reshaped back to the original dimensions to obtain an enhanced feature sequence.
[0129] Step S33: Jointly model the enhanced feature sequence through an adaptive time-frequency integration network to generate a complex result.
[0130] Please note that Figure 5 The adaptive time-frequency integration network includes a frequency-domain global dependency capture module (F-Block, Frequency-domain Global Dependency Capture Block), a time-domain local dependency capture module (T-Block, Time-domain Local Dependency Capture Block), and a dominant harmonic series energy weighting module (i.e., a dominant harmonic energy weighting mechanism); the frequency-domain global dependency capture module includes a complex-valued spectrum attention network, a complex normalization module, a complex ReLU activation function layer, a random dropout layer, and a restored linear layer; the complex-valued spectrum attention network consists of a complex linear mapping layer, a random dropout layer, and a complex multi-head attention layer.
[0131] Specifically, step S33 may include the following sub-steps S331-:
[0132] Step S331: Use the frequency domain global dependency capture module to globally capture the enhanced feature sequence to generate global features.
[0133] Furthermore, step S331 may include the following sub-steps S3311-3318:
[0134] Step S3311: Perform extended Fourier transform on the enhanced feature sequence and output extended Fourier transform features;
[0135] Step S3312: Using the extended Fourier transform feature as input to the complex-valued spectral attention network to generate an attention output;
[0136] Step S3313: Add the attention output to the enhanced feature sequence to determine a first added feature, and use a complex normalization module to perform complex normalization on the first added feature to generate a first complex normalized feature;
[0137] Step S3314: Perform complex linear mapping on the first complex normalized feature, output the complex linear mapping feature, and use the complex linear mapping feature as the input of the complex ReLU activation function layer to output the intermediate feature;
[0138] Step S3315: Use a random drop layer and a linear reduction layer to perform random drop and linear reduction on the intermediate features in sequence to generate reduced features;
[0139] Step S3316: randomly discard the restored features through a random discard layer to generate random discard features;
[0140] Step S3317: Add the randomly discarded feature and the first complex normalized feature to output a second added feature, and perform complex normalization on the second added feature through a complex normalization module to generate a second complex normalized feature;
[0141] Step S3318: Project the second complex normalized features to generate global features.
[0142] It should be noted that, assuming the input time series (i.e., enhanced feature sequence) is , L is the length of the input time series, and an extended DFT (Discrete-Fourier Transform) is constructed. After the time series is input, the sequence length is extended from L to L+T by splicing a zero vector corresponding to the predicted length at the end, and then a real fast Fourier transform is applied to perform time-frequency conversion to obtain the extended DFT. , that is, the extended Fourier transform feature, is the length of the sequence after expansion, . The process can be expressed as:
[0143] ;
[0144] ;
[0145] ;
[0146] in, is the kth element in the result after extended Fourier transform; is the nth element in the result before extended Fourier transform, i.e. input; k is the frequency domain index; n is the time domain index; T is the prediction length; for The real part of for The imaginary part of .
[0147] Furthermore, the above formula yields a spectrum of length L + T. For real-valued time series, the conjugate symmetry of the output spectrum is an important property of the DFT. This property can be exploited to reduce computational cost by considering only the first half of the output spectrum.
[0148] Furthermore, the spectrum of each sample is normalized in the frequency domain to zero mean and unit amplitude to prevent certain frequency points from dominating the model training. Specifically, the extended Fourier transform features are further processed using a complex-valued spectral attention network, and the normalized complex spectrum is input into the first complex linear mapping layer (i.e., the complex linear mapping layer) in the complex-valued spectral attention network to generate the query, key, and value vectors of the multi-head attention: ,in, is the query vector; is the key vector; is the value vector; T is the transpose; is the multi-head attention function.
[0149] Furthermore, after being processed by the first random dropout layer in the complex-valued spectral attention network, it is input into the complex multi-head attention module to calculate the frequency domain correlation and obtain the attention output;
[0150] ;
[0151] ;
[0152] in, is the output of the h-th head in the multi-head attention mechanism; is the multi-head attention function; For splicing; is the output weight matrix; is the attention output; H is the number of heads in the multi-head attention mechanism.
[0153] This output is then passed through a second random dropout layer, added to the original input residual, and fed into the first complex normalization module to obtain the normalized result. This result is activated by a complex linear map and a complex ReLU (CReLU activation function) to generate intermediate data (intermediate features). This intermediate data then passes through a third random dropout layer and a dimension reduction linear layer to obtain the restored result (restored features). This intermediate data is then processed through a fourth random dropout layer. This is then added to the normalized residual and fed into the second complex normalization module to output the current coding layer result. By stacking multiple layers, this module can model frequency-domain dependencies and fill in spectral gaps caused by zero padding.
[0154] Furthermore, the second complex normalized feature is projected. The specific process is as follows: after frequency domain denormalization, iDFT (Inverse Discrete Fourier Transform) time series reconstruction and slicing, the global feature is obtained. .
[0155] Step S332: Locally capture the enhanced feature sequence through the time-domain local dependency capture module to generate local features.
[0156] It should be noted that the temporal local dependency capture module includes a self-attention mechanism module (Transformer) and a linear head (Linear Head).
[0157] Furthermore, step S332 may include the following sub-steps S3321-S3323:
[0158] Step S3321: Segment the enhanced feature sequence to generate multiple feature blocks;
[0159] Step S3322: take each feature block as input to the self-attention mechanism module and output the self-attention output;
[0160] Step S3323: Use the linear head to perform linear projection on the self-attention output to generate local features.
[0161] It should be noted that, the time series X (i.e., enhanced feature sequence) is input into the T-Block module constructed by the present invention, which is divided into N small blocks of sequence length p. Each small block is embedded into the encoder (composed of a self-attention mechanism module and a linear head), and linear projection is used to generate the output local features. .
[0162] Step S333: Perform Fourier transform on the enhanced feature sequence, and output the Fourier transform feature and the complex spectrum coefficient corresponding to the Fourier transform feature;
[0163] Step S334: Use the dominant harmonic series energy weighting module to perform weighting according to the Fourier transform characteristics and the complex spectrum coefficients corresponding to the Fourier transform characteristics, and output the frequency domain weight;
[0164] It should be noted that the dominant harmonic series energy weighting module is constructed to dynamically adjust the weights between the time-frequency domain modules (frequency domain weights, including the first frequency domain weight and the second frequency domain weight) according to the periodicity of the input sequence. The calculation process is as follows:
[0165] ;
[0166] ;
[0167] ;
[0168] ;
[0169] in, is the weight of the prediction result of the F-Block module (i.e. the first frequency domain weight), The weight of the T-Block module prediction result (i.e., the second frequency domain weight). It represents the complex spectrum coefficient at the i-th frequency component after performing N-point DFT on the signal frame weighted by the window function, that is, the complex spectrum coefficient corresponding to the Fourier transform feature; It represents the spectrum coefficient of the nth harmonic - that is, the DFT output of the frequency component corresponding to n times the fundamental frequency index k, that is, the Fourier transform feature; It represents the total energy of the dominant harmonic group (i.e. the set of frequencies related to the fundamental frequency and its integer multiples) in the spectrum; is the total energy of the entire spectrum.
[0170] Step S335: Generate a complex result based on the frequency domain weight, global features, and local features.
[0171] It should be noted that the plural result It can be expressed as: .
[0172] Step S34: convert the complex number result to generate a real number result;
[0173] Step S35: Denormalize the real number result to generate a short-term power load forecast result.
[0174] It should be noted that the final prediction result (i.e., the short-term power load forecast result) is obtained by converting the obtained complex number into a real number and performing inverse normalization.
[0175] For a comparison of technical performance, existing technologies can be used as a reference. Traditional power load forecasting primarily relies on mathematical statistical models, including time series techniques such as integrated moving average autoregressive and Markov chains. However, with the increasing complexity of user characteristics and the advancement of new power system construction, the nonlinearity and randomness of load curves have significantly increased. Traditional time series forecasting techniques are no longer able to meet current requirements due to model complexity and computational limitations when dealing with multiple coupled factors. With the rapid development of deep learning in recent years, various deep neural network models have been widely applied to load forecasting, such as recurrent neural networks (RNNs), represented by long short-term memory networks, convolutional networks, and time convolutional networks. However, these networks only process load series from a temporal perspective and do not further explore the periodicity of power load in the frequency domain. Furthermore, with high-dimensional, multi-factor coupled data, a single forecasting algorithm struggles to achieve ideal accuracy.
[0176] Fully exploring the trend and periodic characteristics contained in power load data and extracting the coupling relationship between various influencing factors and power load are of great significance for improving the accuracy of short-term power load forecasting. Based on this, the present invention provides a short-term power load forecasting method that deeply processes the various influencing factors and the characteristics of the power load to improve the forecast accuracy. Figure 6 , the present invention uses EEMD decomposition to highlight the periodicity of the original load data, strengthen the correlation between the load power and various influencing factors, and uses the maximum information coefficient to select the decomposed mode to reduce the interference of high-dimensional information. The EMA module is used to achieve the simultaneous extraction of local and global dependencies through feature grouping and spatial attention. In order to fully explore the inherent periodic laws and multi-factor coupling effects in power load data, the present invention adopts an adaptive time-frequency integrated network ATFNet. The network integrates time domain and frequency domain modules, and jointly models the feature-enhanced sequence from the two levels of local details and global trends, thereby accurately capturing periodic fluctuations and the coupling relationship between various factors, and effectively improving the prediction accuracy.
[0177] In an embodiment of the present invention, the present invention provides a short-term power load forecasting method. First, the operating data of the power system to be measured is obtained, and the operating data of the power system to be measured is preprocessed to output the target power system operating data; then, based on the target power system operating data, a target input tensor is generated; finally, a preset short-term power load forecasting model is used to make a forecast based on the target input tensor to generate a short-term power load forecasting result; based on the above scheme, the present invention uses a preset short-term power load forecasting model to make a forecast based on the generated target input tensor, which can deeply process the characteristics of the power system operating data itself, thereby improving the prediction accuracy.
[0178] For better explanation, refer to Figure 7, shows a flowchart of the steps of the training process of the preset short-term power load forecasting model provided by the second embodiment of the present invention, which may include the following steps:
[0179] Step 701: Acquire a training power system operation data set, and preprocess the training power system operation data set to generate an intermediate training power system operation data set.
[0180] Step 702: Generate a target training power system operation dataset based on the intermediate training power system operation dataset.
[0181] Step 703: Use the target training power system operation data set to perform model training on the initial short-term power load forecasting model to determine a trained preset short-term power load forecasting model.
[0182] It should be noted that the original power system operation data set (training power system operation data set) is obtained, including load power data, meteorological factors affecting load power (temperature, humidity), time factors (hours), and economic factors (electricity price). This raw data is preprocessed and divided into training sets, validation sets, and test sets.
[0183] Further, see Figure 8This paper demonstrates the effectiveness of the present invention through specific experiments. The 2020-2021 electricity load dataset for the New England region of the United States, operated by the New England Independent System Operator (NE-ISO), was used as the research subject. The experiments used traditional RNN (Recurrent Neural Network), BiLSTM (Bidirectional Long-Short Term Memory), CNNLSTM (Convolutional Neural Network - Long-Short Term Memory), and CNN-BiLSTM (Convolutional Neural Network - Bidirectional Long-Short Term Memory) algorithms as comparisons. The proposed EEMD-CNN-CBAM-BiLSTM (Ensemble Empirical Mode Decomposition - Convolutional Neural Network - Convolutional Block Attention Module - Bidirectional Long-Short Term Memory) model, a pre-built short-term power load forecasting model, was validated. During the verification process, the mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and mean squared percentage error (MSPE) were selected as evaluation criteria to achieve a comprehensive evaluation of the accuracy of this method.
[0184] Furthermore, as shown in Table 1, the error evaluation indicators of the EAM-ATFNet prediction model are all smaller than those of other comparison models, and the prediction performance of the EAM-ATFNet prediction model is more superior.
[0185] Table 1 Comparison and summary of evaluation indicators
[0186]
[0187] In an embodiment of the present invention, the present invention uses EEMD decomposition to highlight the periodicity of the original load data, strengthen the correlation between the load power and various influencing factors, and uses the maximum information coefficient to select the decomposed mode to reduce the interference of high-dimensional information. The EMA module is used to achieve the simultaneous extraction of local and global dependencies through feature grouping and spatial attention. In order to fully explore the inherent periodic laws and multi-factor coupling effects in power load data, the present invention uses an adaptive time-frequency integrated network ATFNet. The network integrates time domain and frequency domain modules, and jointly models the feature-enhanced sequence from the two levels of local details and global trends, thereby accurately capturing the periodic fluctuations and the coupling relationship between various factors, and effectively improving the prediction accuracy.
[0188] See also Figure 9 , Figure 9 This is a structural block diagram of a short-term power load forecasting device provided in Example 3 of the present invention.
[0189] The present invention provides a short-term power load forecasting device, comprising:
[0190] The acquisition module 901 is used to acquire the operating data of the power system to be tested, pre-process the operating data of the power system to be tested, and output the target power system operating data;
[0191] A generating module 902 is configured to generate a target input tensor based on target power system operation data;
[0192] The prediction module 903 is used to use a preset short-term power load prediction model to perform prediction according to the target input tensor to generate a short-term power load prediction result.
[0193] Furthermore, the acquisition module 901 is specifically configured to:
[0194] Perform outlier elimination on the power system operation data to be tested to generate intermediate power system operation data;
[0195] The intermediate power system operation data is normalized to generate the target power system operation data.
[0196] Furthermore, the target power system operation data includes load power data, temperature data, humidity data, time data, and electricity price data; the generation module 902 is specifically used to:
[0197] Performing collective empirical mode decomposition on the load power data to generate multiple intrinsic mode function components corresponding to the load power data;
[0198] Use each intrinsic mode function component to grid the load power data, temperature data, humidity data, time data, and electricity price data, and output the data grid corresponding to each intrinsic mode function component;
[0199] The joint probability distribution and marginal distribution corresponding to each data grid are used to calculate the mutual information corresponding to each data grid;
[0200] Normalize the mutual information corresponding to each data grid, and output the normalized mutual information corresponding to each data grid;
[0201] Compare the normalized mutual information corresponding to each data grid with the preset information threshold;
[0202] Any intrinsic mode function component corresponding to the normalized mutual information greater than a preset information threshold is used as the target intrinsic mode function component;
[0203] The target intrinsic mode function components, load power data, temperature data, humidity data, time data, and electricity price data are combined to generate a target input tensor.
[0204] Furthermore, the preset short-term power load forecasting model includes an improved efficient multi-scale attention module and an adaptive time-frequency integrated network; the forecasting module 903 includes:
[0205] The first submodule is used to group the target input tensors and determine the grouped tensors;
[0206] The second submodule is used to perform enhanced feature extraction on the grouped tensors using an improved efficient multi-scale attention module to generate an enhanced feature sequence;
[0207] The third submodule is used to jointly model the enhanced feature sequence through an adaptive time-frequency integration network to generate complex results;
[0208] The fourth submodule is used to convert the complex number result to generate a real number result;
[0209] The fifth submodule is used to perform denormalization on the real number results to generate short-term power load forecast results.
[0210] Furthermore, the improved efficient multi-scale attention module includes a convolution layer, an average pooling layer, an average pooling layer in the horizontal direction, an average pooling layer in the vertical direction, a Softmax activation function layer, and a Sigmoid activation function layer; the second submodule is specifically used to:
[0211] Use the convolution layer, the average pooling layer in the horizontal direction, and the average pooling layer in the vertical direction to perform convolution operations on the grouped tensors, average pooling in the horizontal direction, and average pooling in the vertical direction, respectively, and output convolution features, horizontal average pooling features, and vertical average pooling features;
[0212] The average pooling layer and the Softmax activation function layer are used to perform average pooling and nonlinear mapping on the convolution features respectively to generate the first average pooling feature and the first nonlinear mapping feature;
[0213] Concatenate the horizontal axis average pooling features and the vertical axis average pooling features to generate concatenated features;
[0214] The splicing features are weighted through the Sigmoid activation function layer, and the fusion weight coefficient is output;
[0215] Multiply the fusion weight coefficient and the grouped tensor to generate fusion features, and group normalize the fusion features to output group normalized features;
[0216] The average pooling layer and the Softmax activation function layer are used to perform average pooling and nonlinear mapping on the group normalized features respectively to generate the second average pooling feature and the second nonlinear mapping feature;
[0217] Perform matrix multiplication on the second average pooling feature and the first nonlinear mapping feature, and output the first matrix multiplication feature;
[0218] Performing matrix multiplication on the second nonlinear mapping feature and the first average pooling feature, and outputting a second matrix multiplication feature;
[0219] The first matrix multiplication feature and the second matrix multiplication feature are added to output the added feature, and the added feature and the grouped tensor are multiplied to output an enhanced feature sequence.
[0220] Furthermore, the adaptive time-frequency integration network includes a frequency domain global dependency capture module, a time domain local dependency capture module, and a dominant harmonic series energy weighting module; the third submodule includes:
[0221] The first unit is used to globally capture the enhanced feature sequence using a frequency domain global dependency capture module to generate global features;
[0222] The second unit is used to locally capture the enhanced feature sequence through a time domain local dependency capture module to generate local features;
[0223] The third unit is used to perform Fourier transform on the enhanced feature sequence, and output Fourier transform features and complex spectrum coefficients corresponding to the Fourier transform features;
[0224] The fourth unit is used to use the dominant harmonic series energy weighting module to perform weighting according to the Fourier transform characteristics and the complex spectrum coefficients corresponding to the Fourier transform characteristics, and output frequency domain weights;
[0225] The fifth unit is used to generate complex results according to frequency domain weights, global features, and local features.
[0226] Furthermore, the frequency domain global dependency capture module includes a complex-valued spectral attention network, a complex normalization module, a complex ReLU activation function layer, a random dropout layer, and a restoration linear layer; the first unit is specifically used to:
[0227] Perform extended Fourier transform on the enhanced feature sequence and output extended Fourier transform features;
[0228] The extended Fourier transform features are used as input to the complex-valued spectral attention network to generate attention output;
[0229] Adding the attention output to the enhanced feature sequence to determine a first added feature, and performing complex normalization on the first added feature using a complex normalization module to generate a first complex normalized feature;
[0230] Performing complex linear mapping on the first complex normalized feature, outputting the complex linear mapping feature, and using the complex linear mapping feature as the input of the complex ReLU activation function layer, outputting the intermediate feature;
[0231] Use random dropout layer and linear reduction layer to randomly drop and linearly reduce the intermediate features in turn to generate reduced features;
[0232] The restored features are randomly discarded through the random discard layer to generate random discard features;
[0233] Adding the randomly discarded feature and the first complex normalized feature to output a second added feature, and performing complex normalization on the second added feature through a complex normalization module to generate a second complex normalized feature;
[0234] Project the second complex normalized feature to generate the global feature.
[0235] Furthermore, the temporal local dependency capture module includes a self-attention mechanism module and a linear head; the second unit is specifically used to:
[0236] Split the enhanced feature sequence to generate multiple feature blocks;
[0237] Each feature block is used as the input of the self-attention mechanism module and outputs the self-attention output;
[0238] A linear head is used to linearly project the self-attention output to generate local features.
[0239] In an optional embodiment of the device, the device further comprises:
[0240] The first module is used to obtain a training power system operation data set, and preprocess the training power system operation data set to generate an intermediate training power system operation data set;
[0241] The second module is used to generate a target training power system operation dataset based on the intermediate training power system operation dataset;
[0242] The third module is used to perform model training on the initial short-term power load forecasting model using the target training power system operation data set, and determine the trained preset short-term power load forecasting model.
[0243] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, modules, sub-modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0244] An embodiment of the present invention further provides a computer device comprising a memory and a processor, wherein a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the short-term power load forecasting method as described in any of the above embodiments.
[0245] An embodiment of the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps of the short-term power load forecasting method as described in any of the above embodiments are implemented.
[0246] An embodiment of the present invention further provides a computer program product, including a computer program / instruction, which implements the steps of the short-term power load forecasting method as described in any of the above embodiments when the computer program / instruction is executed by a processor.
[0247] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0248] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0249] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A short-term power load forecasting method, characterized in that: include: Acquire the operating data of the power system to be tested, pre-process the operating data of the power system to be tested, and output the target power system operating data; generating a target input tensor based on the target power system operation data; A preset short-term power load forecasting model is used to perform forecasting based on the target input tensor to generate a short-term power load forecasting result.
2. The short-term power load forecasting method according to claim 1, characterized in that: The preprocessing of the power system operation data to be tested and outputting target power system operation data includes: performing an outlier elimination operation on the power system operation data to be tested to generate intermediate power system operation data; The intermediate power system operation data is normalized to generate target power system operation data.
3. The short-term power load forecasting method according to claim 1, characterized in that: The target power system operation data includes load power data, temperature data, humidity data, time data, and electricity price data; The generating a target input tensor based on the target power system operation data includes: Performing collective empirical mode decomposition on the load power data to generate a plurality of intrinsic mode function components corresponding to the load power data; Performing data gridding on the load power data, the temperature data, the humidity data, the time data, and the electricity price data using each of the intrinsic mode function components, and outputting a data grid corresponding to each of the intrinsic mode function components; Calculating the mutual information corresponding to each of the data grids using the joint probability distribution and the marginal distribution corresponding to each of the data grids; Normalizing the mutual information corresponding to each of the data grids, and outputting the normalized mutual information corresponding to each of the data grids; Comparing the normalized mutual information corresponding to each of the data grids with a preset information threshold; Taking any intrinsic mode function component corresponding to the normalized mutual information greater than the preset information threshold as the target intrinsic mode function component; The target intrinsic mode function component, the load power data, the temperature data, the humidity data, the time data, and the electricity price data are combined to generate a target input tensor.
4. The short-term power load forecasting method according to claim 1, characterized in that: The preset short-term power load forecasting model includes an improved efficient multi-scale attention module and an adaptive time-frequency integrated network; the preset short-term power load forecasting model is used to predict according to the target input tensor to generate a short-term power load forecast result, including: Grouping the target input tensors to determine grouped tensors; Using the improved efficient multi-scale attention module to perform enhanced feature extraction on the grouped tensors to generate an enhanced feature sequence; Jointly modeling the enhanced feature sequence through the adaptive time-frequency integration network to generate a complex result; Converting the complex number result to generate a real number result; Denormalize the real number result to generate a short-term power load forecast result.
5. The short-term power load forecasting method according to claim 4, characterized in that: The improved efficient multi-scale attention module includes a convolution layer, an average pooling layer, an average pooling layer in the horizontal direction, an average pooling layer in the vertical direction, a Softmax activation function layer, and a Sigmoid activation function layer; The improved efficient multi-scale attention module is used to perform enhanced feature extraction on the grouped tensors to generate an enhanced feature sequence, including: Using the convolution layer, the average pooling layer in the horizontal direction, and the average pooling layer in the vertical direction to perform convolution operation, average pooling in the horizontal direction, and average pooling in the vertical direction on the grouped tensors, respectively, and output convolution features, horizontal average pooling features, and vertical average pooling features; Using the average pooling layer and the Softmax activation function layer to perform average pooling and nonlinear mapping on the convolution features respectively, to generate a first average pooling feature and a first nonlinear mapping feature; Splicing the horizontal axis average pooling feature and the vertical axis average pooling feature to generate a spliced feature; Perform feature weighting on the splicing features through the Sigmoid activation function layer and output a fusion weight coefficient; Multiplying the fusion weight coefficient and the grouped tensor to generate a fusion feature, performing group normalization on the fusion feature, and outputting a group normalized feature; Using the average pooling layer and the Softmax activation function layer to perform average pooling and nonlinear mapping on the group normalized features, respectively, to generate a second average pooling feature and a second nonlinear mapping feature; Performing matrix multiplication on the second average pooling feature and the first nonlinear mapping feature, and outputting a first matrix multiplication feature; Performing matrix multiplication on the second nonlinear mapping feature and the first average pooling feature, and outputting a second matrix multiplication feature; The first matrix multiplication feature and the second matrix multiplication feature are added to output an added feature, and the added feature and the grouped tensor are multiplied to output an enhanced feature sequence.
6. The short-term power load forecasting method according to claim 4, characterized in that: The adaptive time-frequency integrated network includes a frequency domain global dependency capture module, a time domain local dependency capture module, and a dominant harmonic series energy weighting module; the enhanced feature sequence is jointly modeled by the adaptive time-frequency integrated network to generate a complex result, including: Using the frequency domain global dependency capture module to globally capture the enhanced feature sequence to generate global features; Locally capturing the enhanced feature sequence through the time-domain local dependency capture module to generate local features; Performing Fourier transform on the enhanced feature sequence, and outputting Fourier transform features and complex spectral coefficients corresponding to the Fourier transform features; The dominant harmonic series energy weighting module is used to perform weighting according to the Fourier transform feature and the complex spectrum coefficient corresponding to the Fourier transform feature, and output a frequency domain weight; A complex result is generated according to the frequency domain weight, the global feature, and the local feature.
7. The short-term power load forecasting method according to claim 6, characterized in that: The frequency domain global dependency capture module includes a complex-valued spectrum attention network, a complex normalization module, a complex ReLU activation function layer, a random dropout layer, and a restoration linear layer; the frequency domain global dependency capture module is used to globally capture the enhanced feature sequence to generate global features, including: Performing an extended Fourier transform on the enhanced feature sequence and outputting an extended Fourier transform feature; Using the extended Fourier transform features as input to a complex-valued spectral attention network to generate an attention output; Adding the attention output to the enhanced feature sequence to determine a first added feature, and performing complex normalization on the first added feature using a complex normalization module to generate a first complex normalized feature; Performing complex linear mapping on the first complex normalized feature, outputting a complex linear mapping feature, and using the complex linear mapping feature as an input of a complex ReLU activation function layer, outputting an intermediate feature; A random dropout layer and a linear reduction layer are used to sequentially drop and linearly reduce the intermediate features to generate reduced features; Randomly discarding the restored features through a random discard layer to generate random discard features; Adding the randomly discarded feature and the first complex normalized feature to output a second added feature, and performing complex normalization on the second added feature through a complex normalization module to generate a second complex normalized feature; Projecting the second complex normalized features to generate global features.
8. The short-term power load forecasting method according to claim 6, characterized in that: The temporal local dependency capture module includes a self-attention mechanism module and a linear head; the temporal local dependency capture module performs local capture on the enhanced feature sequence to generate local features, including: Segmenting the enhanced feature sequence to generate a plurality of feature blocks; Taking each of the feature blocks as the input of the self-attention mechanism-based module and outputting a self-attention output; The self-attention output is linearly projected using the linear head to generate local features.
9. The short-term power load forecasting method according to claim 1, characterized in that: The training process of the preset short-term power load forecasting model includes: Acquiring a training power system operation data set, and preprocessing the training power system operation data set to generate an intermediate training power system operation data set; generating a target training power system operation data set based on the intermediate training power system operation data set; The target training power system operation data set is used to perform model training on the initial short-term power load forecasting model, and the trained preset short-term power load forecasting model is determined.
10. A short-term power load forecasting device, characterized in that: include: An acquisition module is used to acquire the operating data of the power system to be tested, pre-process the operating data of the power system to be tested, and output the target power system operating data; A generating module, configured to generate a target input tensor based on the target power system operation data; The prediction module is used to use a preset short-term power load prediction model to make predictions based on the target input tensor to generate a short-term power load prediction result.