Multi-scale power load prediction method and system

By expanding the causal convolution network and multi-scale sparse attention mechanism, the power load data is decomposed, combined with probabilistic fragment sampling and dynamic pruning technology, the problem of high computational complexity of the Transformer model is solved, and efficient and accurate power load prediction is achieved.

CN120354984APending Publication Date: 2025-07-22ZHEJIANG GONGSHANG UNIVERSITY

Patent Information

Application Number
CN202510213955.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing power load prediction method based on the Transformer model has high computational complexity when processing long time series, making it difficult to effectively reduce the calculation amount while maintaining high prediction accuracy.

Method used

The expanded causal convolution network and multi-scale sparse attention mechanism are adopted, and the time series data is decomposed into multiple subsequences, combined with the probabilistic fragment sampling attention mechanism, dynamic pruning and multi-scale sparse attention modules are used to reduce the computational complexity and improve the prediction accuracy.

Benefits of technology

It realizes that while reducing the computational complexity, the accuracy and efficiency of power load prediction are improved, and that multivariate relationships in long time series can be captured quickly and efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354984A_ABST
    Figure CN120354984A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale power load prediction method and system, and the method comprises the steps: obtaining time series data related to power generation, and decomposing the time series data into a plurality of sub-series data, so as to capture key periodic features in the data; the time sequence data and the subsequences are sent into an encoder and a decoder for feature extraction, in the encoder, local important information is extracted through an expansion causal convolutional network, correlation features between the subsequences are fused into a subsequent attention module, in the encoder, an attention mechanism is sampled through a probabilistic fragment, and the time sequence data and the subsequences are extracted; randomly calculating the attention of a plurality of local blocks, performing dynamic pruning, capturing global dependency on coarse granularity through a multi-scale sparse attention mechanism, identifying a macroscopic mode of a sequence, and calculating a relationship between time steps on fine granularity to capture local dependency; and constructing and training a prediction model based on the encoder and the decoder, outputting a prediction result of the power load data, and performing result evaluation through dimensionality reduction and a loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power load forecasting, and particularly relates to a multi-scale power load forecasting method and system. Background Art

[0002] One prominent feature of electric energy is that it is difficult to store and is generated instantaneously when needed. Therefore, during the process from power generation to power transmission in the power system, synchronization needs to be maintained, so that the actual power consumption of electricity users is dynamically balanced with the supply of power companies, avoiding the phenomenon of waste of power resources. Therefore, any power grid company regards being able to provide suitable and stable electric energy to electricity users as a strategic goal. Therefore, predicting future changes based on historical load data is of great significance for power scheduling and distribution. Existing load forecasting methods based on deep learning represented by the Transformer model have been widely used in power load forecasting. Its self-attention mechanism can capture long-term dependencies in time series and automatically learn the relationships between elements in the sequence, and has good processing ability for non-linear data and high prediction accuracy. However, the computational complexity of the basic Transformer model lies in that it needs to calculate the attention scores between every pair of elements in the input sequence, which will lead to a sharp increase in the amount of calculation when dealing with long time series. Therefore, it is necessary to design an improved Transformer model to reduce the computational complexity while maintaining performance to meet the requirements of power load forecasting. Summary of the Invention

[0003] In order to solve the deficiencies of the prior art and achieve the purpose of reducing the calculation duration and improving the prediction accuracy, the present invention adopts the following technical solutions:

[0004] A multi-scale power load forecasting method includes the following steps:

[0005] Step 1: Obtain time series data related to power generation and decompose it into multiple subsequence data to capture key periodic features in the data;

[0006] Step 2: Send the obtained time series data and subsequences into an encoder and a decoder for feature extraction. In the encoder, dilated causal convolutional networks are used to extract local important information, and the correlation features between the subsequences are incorporated into the subsequent attention module. The value vectors at all positions in the sequence are weighted and summed, which is beneficial to extracting global and local sequence features, so as to obtain the output representation at the corresponding positions. In the encoder, through the probabilistic segment sampling attention mechanism, the attention of multiple local blocks is randomly calculated and dynamically pruned to reduce the computational complexity and highlight the most important local information. Then, through the multi-scale sparse attention mechanism, global dependencies are captured at a coarse granularity to identify the macroscopic patterns of the sequence, and the relationships between time steps are calculated at a fine granularity to capture local dependencies;

[0007] Step 3. Based on the encoder and decoder, construct and train a prediction model to output the prediction result of the power load data.

[0008] Further, in the first step, the historical time series data related to power generation is first cleared of outliers, and then the variable data of multiple power loads is normalized. The variational mode decomposition method is used to decompose multiple intrinsic mode functions so that each mode function has the smallest bandwidth.

[0009] Further, the variational mode decomposition method in the first step includes the following steps:

[0010] Step 1.1. Construct a Lagrangian function to handle the constrained optimization problem, decompose the time series f(t) of different scales into K mode functions so that each mode function has the smallest bandwidth, and the optimization problem formula is as follows:

[0011]

[0012] Among them, L(·) represents decomposing the time series f(t) of different scales into K mode functions, u k represents the mode function, w k represents the central frequency, α represents the smoothing parameter, α t represents the time derivative, t represents time, δ t represents the Dirac function, j represents the imaginary unit, represents the square of the L2 norm, and λ(t) represents the Lagrange multiplier;

[0013] Step 1.2. Solve the optimization problem by the alternating direction multiplier method and update the mode function and the central frequency. The formula is as follows:

[0014]

[0015] Among them, represents the (n + 1)-th iterative update of the k-th mode component of u(t) after Fourier transform, represents the n-th iterative update of the k-th mode component of u(t) after Fourier transform, represents the n-th iterative update of λ(t) after Fourier transform, represents f(t) after Fourier transform, ω represents the overall central frequency, ω k represents the central frequency of the k-th mode, and α represents the smoothing parameter;

[0016] Step 1.3. Update the Lagrange multiplier to gradually approximate the optimal solution of the optimization problem. The formula is as follows:

[0017]

[0018] Among them, λ n+1 (t) represents the center frequency of the (n + 1)-th update, and λ n (t) represents the Lagrange multiplier of the n-th update, and τ represents the step size parameter. represents the k-th mode function of the (n + 1)-th update;

[0019] Step 1.4, the end conditions for the update are as follows:

[0020]

[0021] Among them, ε represents the given determination accuracy.

[0022] Furthermore, in the encoder of the second step, the original normalized time series data is obtained, and the dilated causal convolutional network is used to extract the local important feature information in the standardized time history sequence;

[0023] In the dilated causal convolutional network, the convolution operation is designed to depend only on the current and past input information and is not interfered by future information, so that the convolution operation conforms to the causality of the time series; the causal convolution function for the input sequence X is defined as:

[0024]

[0025] Among them, X is the input sequence, s represents the input time series information, f represents the weight parameter of the convolution kernel, f(i) represents the weight value of the convolution kernel at the i-th position, q represents the convolution kernel size, which is used to set the receptive field of the network, and x s-i represents the located historical position information. When the sampling frequency is constant, by continuously deepening the network layer and expanding the convolution range, the local neighborhood feature information is extracted, and appropriate padding is performed at the network edge position to keep the output and input dimensions consistent.

[0026] Furthermore, in the encoder of the second step, the feature information extracted by the dilated causal convolutional network is sent to the multi-head sparse probabilistic self-attention module. The multi-head sparse probabilistic self-attention module depends on the sequence decomposition features, adds the subsequence similarity to extract the feature information, re-adjusts the calculation of the self-attention score item, and calculates the score by combining the sequence feature information;

[0027] When processing the information at a specific position in the time series, the correlation score between different positions in the input sequence is calculated through the self-attention mechanism, and the extracted features are fused into the self-attention mechanism for calculation. The optimized self-attention calculation is expressed as:

[0028]

[0029] Among them, A i represents the attention score at the i-th position, Q i represents the query vector at the i-th time step, K i represents the key vector at the i-th time step, V i represents the value vector at the i-th time step, Softmax represents the activation function, d represents the feature dimension, α’ represents the trainable parameter, R (i,j) represents the local correlation between different time steps i and j in the same subsequence;

[0030] By calculating the self-attention dot product of the input sequence features to obtain the feature information between time nodes, introducing calculating the covariance matrix of the k-th intrinsic mode function IMfs when the decomposition layer number is K.

[0031] After the similarity scores are normalized, a set of weights is obtained, which is used to weighted sum the value vectors at all positions in the sequence, facilitating the extraction of global and local sequence features, so as to obtain the output representation at the corresponding position.

[0032] Furthermore, in the decoder of the second step, first, the input data is subjected to a non-linear transformation, and features are extracted through a multi-layer perceptron to enhance the abstract representation of the model. The purpose of the multi-layer perceptron is to better capture the feature relationships between multi-dimensional variables; a probabilistic fragment sampling attention module is added after the multi-layer perceptron. By randomly calculating the attention of multiple local blocks and setting a score threshold, the attention below the threshold will be dynamically pruned to reduce the computational complexity and highlight the most important local information. Among them, through the probabilistic fragment sampling attention mechanism, the attention of different time steps inside the input data is calculated to capture the key information in the sequence. The processed data is further optimized through residual connection and layer normalization; a multi-scale sparse attention module is added after the probabilistic fragment sampling attention module. At the coarse-grained level, it quickly captures global dependencies and identifies the macroscopic patterns of the sequence. Sparsification is applied to the coarse-grained attention scores to reduce the computational complexity by screening important parts. The Factor amplification factor is used to amplify the weights of key time steps and reduce unnecessary calculations; at the fine-grained level, the relationships between time steps are finely calculated to capture local dependencies; finally, the coarse-grained attention scores and the fine-grained attention scores are added together; the multi-scale sparse attention mechanism module calculates the attention weights between the decoder input data and the encoder output, fuses the context information provided by the encoder, improves the understanding of the input sequence, enables the decoder to accurately capture the complex dependencies in the time series, and provides a high-quality prediction decoder module.

[0033] Further, in the probabilistic fragment sampling attention module of step two, multiple blocks of the sequence are randomly selected for local attention calculation, specifically as follows:

[0034] First, initialize a score matrix of all zeros:

[0035] scores←0∈R B×H×L×L ,

[0036] where B represents the batch size, H represents the number of heads, and L represents the sequence length. Initially, the score matrix is set to all zeros;

[0037] Then, randomly select multiple patches defined by randomly generated starting points, where each starting point defines a patch; within these patches, calculate the query, key, and value for each element in the input sequence, and generate local attention scores:

[0038] scores[b, h, i, j]=(q b,h,i ·k b,h,j ),

[0039] where q b,h,i and k b,h,i represent the query vector and key vector at positions i and j in the score matrix under batch b and head h, respectively. This local calculation significantly reduces the number of data points involved, reduces the demand for computing resources, and enhances the model's ability to capture key local temporal dependencies.

[0040] Introduce dynamic pruning technology to dynamically prune the score matrix; first calculate the absolute value and threshold, and then apply the threshold. By filtering out the lower scores and setting them to negative infinity, the attention scores are dynamically adjusted according to the set pruning rate, effectively reducing the computational complexity and alleviating the model burden.

[0041] Further, the multi-scale sparse attention module in step two aims to enhance the model's ability to capture global sequence information. This mechanism is achieved through two key steps, including coarse-grained region-level attention calculation and fine-grained token-token attention;

[0042] First, quickly capture global sequence information through rough region-level attention calculation, and calculate the attention scores at the coarse region level as follows;

[0043] coarse_scores bhls =queries blhe ×keys bshe ,

[0044] where queries blhe represents the query, keys bsheLet \(k\) denote the key, \(b\) denote the batch, \(l\) denote the length of the query sequence, \(h\) denote the number of queries and keys, \(e\) denote the feature dimension of the queries and keys, and \(s\) denote the length of the value sequence;

[0045] Then, using dynamic sparse attention, perform a product operation on the queries and keys, and apply the softmax activation function to calculate the initial routing scores:

[0046] routing_scores = softmax(coarse_scores) factor ,

[0047] where factor represents the amplification factor used to amplify the initial routing scores to improve sparsity. These scores are amplified to enhance sparsity and combined with the additional dense scores dense_scores to balance the distribution of the attention scores. The process of balancing attention is as follows:

[0048]

[0049] dense_scores = softmax(coarse_scores),

[0050] The attention scores are used to calculate the final attention weights, which are applied to the values to obtain the final output. The calculation process is as follows:

[0051]

[0052] where \(V\) blhd represents the weighted sum of the attention weights and values at each position in the sequence, sparse_scores bhls represents the sparse attention scores, \(l\) represents each position in the sequence, taking values from 1 to \(L\), and \(L\) represents the total sequence length.

[0053] Furthermore, the method further includes Step 4: Using t-SNE to reduce the dimension of the prediction results for visual expression, and evaluating the prediction model of the power load data in combination with the results obtained by calculating the loss function.

[0054] A multi-scale power load prediction system includes a sequence decomposition module, a feature extraction module, and a prediction module. According to the described multi-scale power load prediction method, the time series data related to power generation is sequentially decomposed, feature extracted, and the prediction model is constructed and trained.

[0055] The advantages and beneficial effects of the present invention are as follows:

[0056] In view of the problems that power load data is highly non-linear, strongly random and periodic, and it is difficult for previous models to capture the relationships between multi-variables in long time series, a prediction algorithm based on probabilistic segment sampling attention, multi-scale sparse attention mechanism and dilated convolutional causal network is proposed. In the encoder, a dilated causal convolutional network is added to extract local important information in the time series. Then, in the improved attention module, decomposed subsequence similarity is added to extract feature information, the attention score term calculation is readjusted, and the score is calculated by combining sequence feature information. In the decoder, a multi-layer perceptron is added to extract high-order non-linear features of the input sequence, and then a probabilistic sampling segment attention module is added to better focus on local feature information while reducing computational complexity. The strategies of random sampling and dynamic pruning are adopted to randomly calculate the attention of multiple local modules, and a score threshold is introduced to prune those below the threshold, reducing computational complexity and highlighting important information. Subsequently, a multi-scale sparse attention module is added, which combines global coarse-grained attention calculation and local fine-grained attention. In the coarse-grained aspect, it quickly captures global dependencies and identifies the macroscopic patterns of the sequence, and in the fine-grained aspect, it precisely calculates the relationships between time steps and captures local dependencies. The present invention can predict power load quickly and efficiently. Description of the Drawings

[0057] Figure 1 is the flowchart of the method of the embodiment of the present invention.

[0058] Figure 2 is the overall framework diagram of the embodiment of the present invention.

[0059] Figure 3 is the framework diagram of the multi-scale sparse probabilistic self-attention module in the embodiment of the present invention.

[0060] Figure 4 is the framework diagram of the probabilistic segment sampling attention module in the embodiment of the present invention.

[0061] Figure 5 is the framework diagram of the multi-scale sparse attention module in the embodiment of the present invention. Detailed Embodiments

[0062] The following will describe the detailed embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustration and explanation of the present invention, and are not intended to limit the present invention.

[0063] Aiming at the fact that power load data has high nonlinearity, strong randomness and periodicity, and previous models are difficult to capture the relationships among multi-variables in long time series, the present invention proposes a multi-scale power load forecasting method. First, the original data is normalized, and at the same time, the original data is decomposed into multiple subsequences. The original data and the subsequences are sent into an improved encoder and decoder for feature extraction of the data. Based on the dilated convolutional causal network in the encoder, the probabilistic segment sampling attention and multi-scale sparse attention mechanisms in the decoder, the power load is predicted to solve the problem that the large scale in the real world leads to long calculation time and reduced prediction accuracy, so as to quickly and efficiently predict the power load situation and provide help for subsequent power grid managers, such as Figure 1 , Figure 2 as shown, and specifically includes the following steps:

[0064] Step 1: Preprocessing of power data; by collecting historical power generation-related data of the power grid, a dataset for predicting the power grid load is obtained. First, the data is cleaned of outliers, and then the variable data of multiple power loads is normalized. At the same time, the variable data is decomposed using the variational mode decomposition algorithm to obtain multiple subsequence data. Among them, the normalized original power load data is decomposed into K intrinsic mode component sequences, where K>1, to ensure capturing the key periodic features in the data.

[0065] Among them, the variational mode decomposition algorithm includes the following steps:

[0066] Step 1.1: Initialize the center frequency, mode function, Lagrange multiplier and maximum number of iterations; decompose the time series of different scales into K said mode functions through the Lagrangian function, so that each mode function has the minimum bandwidth;

[0067] Construct a Lagrangian function to handle the constrained optimization problem. The goal is to decompose the time series of different scales f(t) into K mode functions, so that each mode function has the minimum bandwidth; the optimization problem is expressed as the following formula:

[0068]

[0069] where L(·) represents decomposing the time series of different scales f(t) into K mode functions, u k represents the mode function, w k represents the center frequency, α represents the smoothing parameter, α t represents the time derivative, t represents time, δ t represents the Dirac function, j represents the imaginary unit, represents the square of the L2 norm, f(t) represents the time series of different scales, and λ(t) represents the Lagrange multiplier.

[0070] Step 1.2: Solve the optimization problem by the alternating direction multiplier method and update the modal function and the central frequency. The formula is as follows:

[0071]

[0072]

[0073] Where, represents the (n + 1)-th iterative update of the k-th modal component of the Fourier-transformed u(t), represents the n-th iterative update of the k-th modal component of the Fourier-transformed u(t), represents the n-th iterative update of the Fourier-transformed λ(t), represents the Fourier-transformed f(t), ω represents the overall central frequency, ω k represents the central frequency of the k-th mode, and α represents the smoothing parameter.

[0074] Step 1.3: Update the Lagrange multiplier to gradually approximate the optimal solution of the optimization problem. The formula is as follows:

[0075]

[0076] Where, λ n+1 (t) represents the central frequency of the (n + 1)-th update, λ n (t) represents the Lagrange multiplier of the n-th update, τ represents the step size parameter, f(t) represents the time series of different scales, represents the modal function of the (n + 1)-th update.

[0077] Step 1.4: Repeat the update process in the above steps until the preset convergence condition or the maximum number of iterations is reached, and then output the final intrinsic mode function component sequence, that is, a series of modal function components.

[0078] The formula for the preset convergence condition or the maximum number of iterations is as follows:

[0079]

[0080] Where, ε represents the given determination accuracy.

[0081] Step 2: Feature extraction to efficiently and accurately predict the results; aiming at the highly nonlinear, strongly random and periodic characteristics of the power load data, and the difficulty of previous models in capturing the relationships between multi-variables in long time series, a prediction method based on probability fragment sampling attention, multi-scale sparse attention mechanism and dilated convolutional causal network is proposed.

[0082] First, a dilated causal convolutional network is added to the encoder to extract local important information in the time series. Then, the decomposed subsequence similarity is added to the improved attention module to extract feature information, the calculation of the attention score term is readjusted, and the score is calculated by combining the sequence feature information. Then, a multi-layer perceptron is added to the decoder to extract the high-order non-linear features of the input sequence, and a probabilistic sampling segment attention module is added to better focus on local feature information while reducing the computational complexity. In this module, the strategies of random sampling and dynamic pruning are adopted to randomly calculate the attention of multiple local modules. A score threshold is introduced, and values below this threshold indicate unimportance. Dynamic pruning technology is used for pruning to reduce the computational complexity and highlight important information. Subsequently, a multi-scale sparse attention module is added, which combines global coarse-grained attention calculation and local fine-grained attention. In the coarse-grained aspect, it quickly captures global dependencies and identifies the macroscopic patterns of the sequence. In the fine-grained aspect, it finely calculates the relationships between time steps and captures local dependencies. Through the above operations, the accuracy of power load prediction is significantly improved.

[0083] The specific steps are as follows:

[0084] Step 2.1: In the encoder, obtain the original normalized data, and first use the causal convolutional network to extract the local important feature information in the standardized time history sequence;

[0085] In the causal convolutional network, the convolution operation is designed to depend only on the current and past input information and is not interfered by future information, making the convolution operation conform to the causality of the time series. Let the convolution kernel size be denoted by k. The causal convolution function for the input sequence X is defined as:

[0086]

[0087] where X is the input sequence, s represents the input time series information, f represents the weight parameter of the convolution kernel, f(i) is the weight value of the convolution kernel at the i-th position, q represents the convolution kernel size and is used to set the receptive field of the network, and x s-i represents the located historical position information. When the sampling frequency is fixed, by continuously deepening the network layers, the convolution range can be expanded to extract local neighborhood feature information. Appropriate padding is performed at the network edge positions to keep the output and input dimensions consistent;

[0088] Step 2.2: In the encoder, send the extracted feature information into the multi-head sparse probabilistic self-attention module. The improved self-attention mechanism is designed based on the sequence decomposition features, and the subsequence similarity is added to extract the feature information. The calculation of the self-attention score term is readjusted, and the score is calculated by combining the sequence feature information, as Figure 3 shown. The specific steps are as follows:

[0089] When processing information at specific positions in a time series, the self-attention mechanism is achieved by calculating the correlation scores between different positions in the input sequence. The extracted features are fused into the self-attention mechanism for calculation. The optimized self-attention calculation is expressed as:

[0090]

[0091] where, A i represents the attention score at the i-th position, Q i represents the query vector at the i-th time step, K i represents the key vector at the i-th time step, V i represents the value vector at the i-th time step, Softmax represents the activation function, d represents the feature dimension, α’ represents the trainable parameter, and R (i,j) represents the local correlation between different time steps i and j in the same subsequence.

[0092] Through the first term the self-attention dot product calculation of the input sequence features is performed to obtain the feature information between time nodes. The second term is introduced to calculate the covariance matrix of the k-th intrinsic mode function IMfs when the decomposition layer number is K.

[0093] Specifically, the correlation features between the subsequences extracted from the decomposed time series are fused into the self-attention. After these similarity scores are normalized, a set of weights is obtained, which is used to weighted sum the value vectors at all positions in the sequence, facilitating the extraction of global and local sequence features, and thus obtaining the output representation at the corresponding positions.

[0094] Step 2.3. In the decoder, first, perform a non-linear transformation on the input data and extract features through a multi-layer perceptron to enhance the model's abstract representation. The purpose of the multi-layer perceptron is to better capture the feature relationships between multi-dimensional variables. After the multi-layer perceptron, add a probabilistic segment sampling attention module. By randomly calculating the attention of multiple local blocks, set a score threshold, and the attention below the threshold will be dynamically pruned to reduce the computational complexity and highlight the most important local information. Among them, through the probabilistic block sampling attention mechanism, calculate the attention for different time steps inside the input data to capture the key information in the sequence. The processed data is further optimized through residual connection and layer normalization. After the probabilistic random sampling segment module, add a multi-scale attention module. At the coarse-grained level, quickly capture the global dependencies and identify the macroscopic patterns of the sequence. Apply sparsification to the coarse-grained attention scores, reduce the computational complexity by screening the important parts, and use the Factor amplification factor to amplify the weights of the key time steps to reduce unnecessary calculations. At the fine-grained level, precisely calculate the relationships between time steps and capture local dependencies. Finally, sum the coarse-grained attention scores and the fine-grained attention scores. Through the multi-scale sparse attention mechanism, calculate the attention weights between the decoder input data and the encoder output, fuse the context information provided by the encoder, and improve the understanding of the input sequence. In this way, the decoder can accurately capture the complex dependencies in the time series and provide high-quality predictions for the decoder module; as Figure 4 、 Figure 5 shown, the specific implementation method is as follows:

[0095] Step 2.3.1. In the probabilistic block sampling attention module, randomly select multiple blocks of the sequence for local attention calculation. First, initialize a score matrix of all zeros:

[0096] scores←0∈R B×H×L×L

[0097] where B represents the batch size, H represents the number of heads, and L represents the sequence length. Initially, the score matrix is set to all zeros;

[0098] Then, randomly select multiple patches defined by randomly generated starting points, where each starting point defines a patch. For each element in the input sequence within these patches, calculate the query, key, and value, and generate local attention scores:

[0099] scores[b, h, i, j]=(q b,h,i ·k b,h,j )

[0100] where q b,h,i and k b,h,iDenote the query vector and key vector at positions i and j of the score matrix under batch b and head h respectively. This local calculation significantly reduces the number of data points involved, decreases the demand for computing resources, and enhances the model's ability to capture key local temporal dependencies;

[0101] Introduce a dynamic pruning technique to dynamically prune the score matrix; first calculate the absolute value and threshold, then apply the threshold, by filtering out the lower scores and setting them to negative infinity, the attention scores are dynamically adjusted according to the set pruning rate, effectively reducing the computational complexity and alleviating the model burden.

[0102] Step 2.3.2, The multi-scale sparse attention module aims to enhance the model's ability to capture global information of the sequence. This mechanism is achieved through two key steps, namely coarse-grained region-level attention calculation and fine-grained token-token attention.

[0103] First, through the rough region-level attention calculation, the model can quickly capture the global sequence information. Calculate the attention scores at the coarse region level as follows;

[0104] coarse_scores bhls =queries blhe ×keys bshe

[0105] where, queries blhe represents the query, keys bshe represents the key, b represents the batch, l represents the length of the query sequence, h represents the number of queries and keys, e represents the feature dimension of the query and key, and s represents the length of the value sequence;

[0106] Then, use dynamic sparse attention, perform a product operation on the queries and keys, and apply the softmax function to calculate the initial routing scores:

[0107] routing_scores=softmax(coarse_scores) tactor

[0108] where, the factor factor is used to amplify the initial routing scores to improve sparsity. These scores are amplified to enhance sparsity and combined with the additional dense scores dense_scores to balance the distribution of attention scores. The attention balancing process is as follows:

[0109]

[0110] dense_scores=softmax(coarse_scores)

[0111] These attention scores are used to calculate the final attention weights, which are then applied to the values to obtain the final output. The calculation process is as follows:

[0112]

[0113] Among them, V blhd represents the weighted sum of the attention weights and values at each position in the sequence, sparse_scores bhls represents the sparse attention scores, l represents each position in the sequence, and the value range of l is from 1 to L, where L represents the total sequence length.

[0114] Step 3: Use the improved encoder and decoder to train the prediction model, and complete the output of the prediction result through the fully connected layer. Each decoding layer generates a partial prediction result through linear projection, and the final predicted power load data result is obtained by gradually accumulating.

[0115] Step 4: Use t-SNE to perform dimensionality reduction on the prediction result for visual expression, and evaluate the prediction model by combining the results calculated by the loss function.

[0116] A multi-scale power load prediction system includes a sequence decomposition module, a feature extraction module, and a prediction module. According to a multi-scale power load prediction method, the decomposition, feature extraction, and construction and training of the prediction model of the time series data related to power generation are carried out in sequence.

[0117] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-scale power load forecasting method, characterized in that It includes the following steps: Step 1: Obtain time series data related to power generation and decompose it into multiple subsequence data to capture key periodic features in the data; Step 2: Send the obtained time series data and subsequences into an encoder and a decoder for feature extraction. In the encoder, dilated causal convolutional networks are used to extract local important information, and the correlation features between the subsequences are incorporated in the subsequent attention module. In the encoder, through the probabilistic segment sampling attention mechanism, the attention of multiple local blocks is randomly calculated and dynamically pruned, and then through the multi-scale sparse attention mechanism, global dependencies are captured at a coarse granularity to identify the macroscopic patterns of the sequence, and the relationships between time steps are calculated at a fine granularity to capture local dependencies; Step 3: Build and train a prediction model based on the encoder and decoder, and output the prediction result of the power load data.

2. The multi-scale power load forecasting method according to claim 1, wherein: In Step 1, first, the historical time series data related to power generation is cleared of outliers, and then the variable data of multiple power loads is normalized. The variational mode decomposition method is used to decompose multiple intrinsic mode functions so that each mode function has the smallest bandwidth.

3. A multi-scale power load forecasting method according to claim 2, characterized in that: The variational mode decomposition method in Step 1 includes the following steps: Step 1.1: Construct a Lagrangian function to handle the constrained optimization problem. Decompose the time series f(t) at different scales into K mode functions so that each mode function has the smallest bandwidth. The optimization problem formula is as follows: Among them, \(L(\cdot)\) represents decomposing the time series \(f(t)\) of different scales into \(K\) mode functions, \(u\) k represents the mode function, \(w\) k represents the central frequency, \(\alpha\) represents the smoothing parameter, \(\alpha\) t represents the time derivative, \(t\) represents time, \(\delta\) t represents the Dirac function, \(j\) represents the imaginary unit, represents the square of the \(L_2\) norm, \(\lambda(t)\) represents the Lagrange multiplier; Step 1.2: Solve the optimization problem by the alternating direction method of multipliers and update the mode function and the central frequency. The formula is as follows: Among them, represents the (n + 1)-th iterative update of the k-th modal component of u(t) after Fourier transform, represents the n-th iterative update of the k-th modal component of u(t) after Fourier transform, represents the n-th iterative update of λ(t) after Fourier transform, represents f(t) after Fourier transform, ω represents the overall center frequency, ω k represents the center frequency of the k-th mode, and α represents the smoothing parameter; Step 1.3: Update the Lagrange multiplier to gradually approximate the optimal solution of the optimization problem. The formula is as follows: where, λ n+1 (t) represents the central frequency of the (n + 1)-th update, λ n (t) represents the Lagrange multiplier of the n-th update, and τ represents the step size parameter, represents the k-th mode function of the (n + 1)-th update; Step 1.4: The end condition for updating is as follows: where ε represents the given determination accuracy.

4. A multi-scale power load forecasting method according to claim 1, characterized in that: In the encoder of Step 2, dilated causal convolutional networks are used to extract local important feature information from the time history sequence; In the dilated causal convolutional network, the convolution operation is designed to depend only on the current and past input information without being interfered by future information, so that the convolution operation conforms to the causality of the time series. The causal convolution function for the input sequence X is defined as: Among them, X is the input sequence, s represents the input time series information, f represents the weight parameter of the convolution kernel, f(i) represents the weight value of the convolution kernel at the i-th position, q represents the convolution kernel size, which is used to set the receptive field of the network, and x s-i represents the located historical position information. When the sampling frequency is constant, by continuously increasing the number of network layers and expanding the convolution range, local neighborhood feature information is extracted.

5. A multi-scale power load forecasting method according to claim 1, characterized in that: In the encoder of Step 2, the feature information extracted by the dilated causal convolutional network is sent into the multi-head sparse probabilistic self-attention module. The multi-head sparse probabilistic self-attention module depends on the sequence decomposition features, adds the subsequence similarity to extract feature information, re-adjusts the calculation of the self-attention score terms, and calculates the score by combining the sequence feature information; When processing the information at a specific position in the time series, the correlation scores between different positions in the input sequence are calculated through the self-attention mechanism, and the extracted features are fused into the self-attention mechanism for calculation. The optimized self-attention calculation is expressed as: Among them, A i represents the attention score at the i-th position, Q i represents the query vector at the i-th time step, K i represents the key vector at the i-th time step, V i represents the value vector at the i-th time step, Softmax represents the activation function, d represents the feature dimension, α’ represents the trainable parameter, R (i,j) represents the local correlation between different time steps i and j in the same subsequence; By performing self-attention dot product calculation on the input sequence features to obtain the feature information between time nodes, introducing calculating the covariance matrix of the k-th intrinsic mode function when the decomposition layer number is K; The similarity scores are normalized to obtain a set of weights, which are used to weighted sum the value vectors at all positions in the sequence to obtain the output representation at the corresponding position.

6. The multi-scale electric load forecasting method according to claim 1, wherein: In the decoder of Step 2, first, a nonlinear transformation is performed on the input data, and features are extracted through a multi-layer perceptron to capture the feature relationships between multi-dimensional variables; Add a probabilistic segment sampling attention module after the multi-layer perceptron. By randomly calculating the attention of multiple local blocks, set a score threshold, and the attention below the threshold will be dynamically pruned. Among them, through the probabilistic segment sampling attention mechanism, calculate the attention for different time steps inside the input data to capture the key information in the sequence. The processed data is further optimized through residual connection and layer normalization. Then add a multi-scale sparse attention module after the probabilistic segment sampling attention module to capture global dependencies at a coarse granularity, identify the macroscopic patterns of the sequence, apply sparsification to the coarse-grained attention scores, filter out the important parts, and use the amplification factor to amplify the weights of the key time steps. Calculate the relationship between time steps at a fine granularity to capture local dependencies. Finally, add the coarse-grained attention scores and the fine-grained attention scores together.

7. A multi-scale power load forecasting method according to claim 6, characterized in that: In the probabilistic segment sampling attention module in step two, randomly select multiple blocks of the sequence for local attention calculation, specifically as follows: First, initialize a score matrix of all zeros: scores ← 0 ∈ R B×H×L×L , Where B represents the batch size, H represents the number of heads, and L represents the sequence length. Initially, the score matrix is set to all zeros; Then, randomly select multiple patches defined by randomly generated starting points, where each starting point defines a patch. In these patches, calculate the query, key, and value for each element in the input sequence and generate local attention scores: scores[b, h, i, j] = (q b,h,i ·k b,h,j ), where q b,h,i and k b,h,i represent the query vector and the key vector, respectively, for the score matrix positions i and j under batch b and head h.

8. A multi-scale power load forecasting method according to claim 6, characterized in that: In the multi-scale sparse attention module in step two, first, quickly capture the global sequence information through rough region-level attention calculation, and calculate the attention scores at the coarse region level as follows; coarse_scores bhls = queries blhe × keys bshe , Among them, queries blhe represent queries, keys bshe represent keys, b represents batch, l represents the query sequence length, h represents the number of queries and keys, e represents the feature dimension of queries and keys, and s represents the length of the value sequence; Then, use dynamic sparse attention to perform a product operation on the query and the key, and apply the softmax activation function to calculate the initial routing scores: routing_scores = softmax(coarse_scores) factor , Where factor represents the amplification factor, the scores are amplified and combined with the additional dense scores densw_scores to balance the distribution of the attention scores. The process of balancing the attention is as follows: dense_score = softmax(coarse_scores), The attention scores are used to calculate the final attention weights and apply them to the values to obtain the final output. The calculation process is as follows: Among them, V blhd represents the weighted sum of the attention weights and values at each position in the sequence, and sparse_scores bhls represents the sparse attention scores. l represents each position in the sequence, taking values from 1 to L, and L represents the total sequence length.

9. A multi-scale power load forecasting method according to claim 1, characterized in that: The method further includes step four, reducing the dimension of the prediction result for visual expression, and evaluating the prediction model of the power load data by combining the results calculated by the loss function.

10. A multi-scale power load forecasting system, including a sequence decomposition module, a feature extraction module, and a prediction module, characterized in that, According to the multi-scale power load prediction method described in claim 1, decompose, extract features, and construct and train the prediction model for the time series data related to power generation in sequence.

Citation Information

Patent Citations

  • Load prediction method and system based on deep self-attention network

    CN113379164A

  • Adjustable load power multi-step prediction method based on improved TCN correction accumulative error

    CN114066052A

  • Power transformer load prediction method based on Transform model

    CN115622047A

  • Current load decomposition method based on deep convolution Seq2Seq-attention framework

    CN118094138A

  • Short-term load prediction method based on self-attention encoder and time convolutional network

    CN118554422A

Cited By

  • Time prediction method and device based on different channel attention mechanisms

    CN121350965A

  • Method and system for predicting residual life of battery cell

    CN121636965A