Power price prediction method and device, storage medium, electronic device and computer program product

By adopting a model based on a time-series fusion converter architecture in power price prediction, combining the gated residual network, variable selection network and multi-head attention mechanism, the problem of low price prediction accuracy in the long-term power spot market has been solved, and higher prediction accuracy and ability to capture market dynamics are achieved.

CN120218975APending Publication Date: 2025-06-27HUANENG CLEAN ENERGY RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510292514.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The accuracy of price predictions of the long-term electricity spot market has been low recently, mainly due to the influence of many factors such as the accumulation of forecast errors, the increase in uncertainties, market dynamic changes, economic and policy factors, and climate and environmental changes.

Method used

A model based on a time-series fusion converter architecture is adopted, combining a gated residual network, a variable selection network and a multi-head attention mechanism, and the power price at the target moment is predicted by obtaining historical meteorological data and power market operation data.

Benefits of technology

It improves the accuracy of forecasting power prices at the target moment, overcomes the difficulty of long-term power prices prediction, and enhances the model's ability to capture complex market dynamics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218975A_ABST
    Figure CN120218975A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power price prediction method and device, a storage medium, an electronic device and a computer program product, and relates to the field of computers.The electric power price prediction method comprises the steps that N sets of historical data at N moments are acquired, the ith group of historical data in the N groups of historical data comprises historical meteorological data and historical power market operation data corresponding to the ith moment in the N moments; the historical electricity market operation data comprises historical electricity price data, N is an integer greater than or equal to 2, and i is an integer greater than or equal to 1 and less than or equal to N; and on the basis of a target model, according to the N groups of historical data, predicting to obtain an electric power price at a target moment, the target model being a model based on a time sequence fusion converter architecture, and the target model having a gated residual network, a variable selection network and a multi-head attention mechanism. By adopting the technical scheme, the problem that the accuracy of predicting the electric power price is relatively low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular, to a method and device for predicting electricity prices, a storage medium, an electronic device, and a computer program product. Background Art

[0002] In the prediction of electricity spot prices, the longer the time span, the more likely the prediction errors will accumulate. The accumulation of such errors makes the deviation between the prediction result and the actual situation larger and larger, resulting in a significant reduction in the prediction accuracy. In addition, long-term price prediction needs to consider more uncertain factors, such as future policy changes, economic development trends, etc. The variability and unpredictability of these factors increase the difficulty of prediction and reduce the accuracy. Electricity spot trading is affected by various factors such as supply and demand relationships, energy policies, and environmental factors. These factors may change significantly in the long term, making it difficult for ordinary short-term prediction models to accurately capture future market dynamics. Therefore, due to the influence of various factors such as the accumulation of prediction errors, the increase in uncertain factors, the change of market dynamics, economic and policy factors, and climate and environmental changes, the prediction accuracy of the day-ahead price in the long-term electricity spot market is relatively low.

[0003] In view of the problem of relatively low accuracy in predicting electricity prices in the related art, no effective solution has been proposed yet.

[0004] Therefore, it is necessary to improve the related technology to overcome the defects in the related technology. Summary of the Invention

[0005] Embodiments of the present application provide a method and device for predicting electricity prices, a storage medium, an electronic device, and a computer program product, so as to at least solve the problem of relatively low accuracy in predicting electricity prices.

[0006] According to an aspect of the embodiments of the present application, a method for predicting electricity prices is provided, including: obtaining N sets of historical data at N moments, where the i-th set of historical data in the N sets of historical data includes: historical meteorological data and historical electricity market operation data corresponding to the i-th moment among the N moments; the historical electricity market operation data includes historical electricity price data, N is an integer greater than or equal to 2, and i is an integer greater than or equal to 1 and less than or equal to N; based on a target model, predicting the electricity price at a target moment according to the N sets of historical data, where the target model is a model based on a temporal fusion transformer architecture, and the target model has a gated residual network, a variable selection network, and a multi-head attention mechanism.

[0007] In an exemplary embodiment, the method further includes: in the process of predicting the electricity price at the target moment based on the target model according to the N sets of historical data, the following operations are performed through the gated residual network: processing the target vector and the target context vector through the following formula one to obtain the first data: η2 = ELU(W 2,ω α + W 3,ω c + b 2,ω ), formula one, where η2 is the first data, α is the target vector, c is the target context vector, ELU is the exponential linear unit activation function, W 2,ω is the second weight matrix, W 3,ω is the third weight matrix, and b 2,ω is the bias vector corresponding to the second weight matrix; the target vector μ x is the total number of features in a set of historical data corresponding to the t-th moment among the N moments, is the transformed input of the j-th feature in a set of historical data corresponding to the t-th moment, j is an integer greater than or equal to 1 and less than or equal to μ x , and the target context vector is the context vector calculated by the static covariate encoder in the target model; processing the first data through the following formula two to obtain the second data: η1 = W 1,ω η2 + b 1,ω , formula two; where η1 is the second data, W 1,ω is the first weight matrix, and b 1,ω is the bias vector corresponding to the first weight matrix; processing the target vector and the second data through the following formula three to obtain the third data: GRN ω (α, c) = LayerNorm(α + GLU ω (η1)) formula three; where GRN ω (α, c) is the third data, LayerNorm is the normalization process, GLU is the gated linear unit, ω is the weight sharing index, and the third data is to be used by the variable selection network.

[0008] In an exemplary embodiment, GLU ω (η1) = σ(W 4,ω η1 + b 4,ω ) ⊙ (W 5,ω η1 + b 5,ω ), GRN ω (η1)) = σ(W 4,ω η1 + b 4,ω ) ⊙ (W 5,ω η1 + b 5,ω ), where σ is the activation function, W4,ω is the fourth weight matrix, b 4,ω is the bias vector corresponding to the fourth weight matrix, W 5,ω is the fifth weight matrix, b 5,ω is the bias vector corresponding to the fifth weight matrix, and ⊙ is the element-wise product.

[0009] In an exemplary embodiment, the method further includes: in the process of predicting the electricity price at the target moment based on the target model according to the N sets of historical data, the following operations are performed through the variable selection network: processing the third data through the Softmax function to obtain the fourth data; obtaining the fifth data through the following formula four: wherein, is the fifth data, is obtained after non-linearly transforming , is the j-th element in the fourth data; the fourth data and the fifth data are to be used by the multi-head attention mechanism.

[0010] In an exemplary embodiment, the method further includes: in the process of predicting the electricity price at the target moment based on the target model according to the N sets of historical data, the following operations are performed through the multi-head attention mechanism: in the case of having m H attention heads, the attention output of the h-th attention head is obtained through the following formula five to obtain m H attention outputs of m H attention heads: H h is the attention output of the h-th attention head, Q is the query matrix, K is the key matrix, V is the value matrix, is the query weight matrix of the h-th attention head, is the key weight matrix of the h-th attention head, is the value weight matrix of the h-th attention head; the m H attention outputs are concatenated to obtain a concatenated vector; the concatenated vector is multiplied by the linear transformation weight matrix to obtain the multi-head attention output.

[0011] In an exemplary embodiment, the method further includes: in the process of predicting the electricity price at the target moment based on the target model according to the N sets of historical data, the following operations are also performed through the multi-head attention mechanism: calculating the shared value of the m H attention heads through the following formula six: wherein, W V is the weight array shared by different attention heads, where S is the shared value and H is the number of attention heads.

[0012] According to another aspect of the embodiments of the present application, there is also provided a device for predicting electricity prices, including: an acquisition module, configured to acquire N sets of historical data at N moments, where the i-th set of historical data in the N sets of historical data includes: historical meteorological data and historical electricity market operation data corresponding to the i-th moment among the N moments; the historical electricity market operation data includes historical electricity price data, N is an integer greater than or equal to 2, and i is an integer greater than or equal to 1 and less than or equal to N; a prediction module, configured to predict the electricity price at a target moment based on a target model according to the N sets of historical data, where the target model is a model based on a temporal fusion transformer architecture, and the target model has a gated residual network, a variable selection network, and a multi-head attention mechanism.

[0013] According to still another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and the program is set to execute the above-mentioned electricity price prediction method when running.

[0014] According to still another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is set to execute the above-mentioned electricity price prediction method through the computer program.

[0015] According to still another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program, and the computer program executes the above-mentioned electricity price prediction method when executed by a processor.

[0016] In the present application, the electricity price at the target moment is predicted according to N sets of historical data through the target model. Since the target model used is a model based on a temporal fusion transformer architecture and has a gated residual network, a variable selection network, and a multi-head attention mechanism, the accuracy of predicting the electricity price at the target moment is improved, thereby solving the problem of low accuracy in predicting electricity prices. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a hardware block diagram of a mobile terminal for a method of predicting electricity prices according to an embodiment of the present application;

[0020] Figure 2 is a flowchart of a method of predicting electricity prices according to an embodiment of the present application;

[0021] Figure 3 is an overall flowchart of a method of predicting electricity prices according to an embodiment of the present application;

[0022] Figure 4 is an overall architecture diagram of a method of predicting electricity prices according to an embodiment of the present application;

[0023] Figure 5 is a block diagram of a device for predicting electricity prices according to an embodiment of the present application. Detailed implementation manners

[0024] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0026] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware block diagram of a mobile terminal for a method of predicting electricity prices according to an embodiment of the present application. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1Only one processor 102 is shown (the processor 102 may include, but is not limited to, a processing device such as a microprocessor (MP) or a Field Programmable Gate Array (FPGA)), and a memory 104 for storing data. Among them, the mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown therein, or have a different configuration from Figure 1 that shown.

[0027] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the pneumatic imbalance detection method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0028] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the mobile terminal. In one instance, the transmission device 106 includes a Network Interface Controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0029] To solve the above problems, a method for predicting electricity prices is provided in this embodiment, including but not limited to being applied to the above-mentioned mobile terminal, Figure 2 is a flowchart of a method for predicting electricity prices according to an embodiment of the present application, as Figure 2 shown, and this process includes the following steps S202 - S204:

[0030] Step S202: Obtain N sets of historical data at N moments. Among them, the i-th set of historical data in the N sets of historical data includes: historical meteorological data and historical power market operation data corresponding to the i-th moment among the N moments; the historical power market operation data includes historical electricity price data. N is an integer greater than or equal to 2, and i is an integer greater than or equal to 1 and less than or equal to N.

[0031] Optionally, the i-th set of historical data in the N sets of historical data further includes: historical hydrological data corresponding to the i-th moment among the N moments. The historical meteorological data includes, but is not limited to: historical temperature, historical humidity, historical wind speed, historical wind direction, historical air pressure; the historical hydrological data includes, but is not limited to: historical local river flow, historical local river water level, historical local river water quality, historical local groundwater level. The historical power market operation data further includes, but is not limited to: historical market disclosure data (such as historical transaction information, historical generation plan and actual power generation).

[0032] Optionally, the N moments are N moments within a preset time period, where each moment is separated by a preset time interval.

[0033] Step S204: Based on the target model, predict the electricity price at the target moment according to the N sets of historical data. The target model is a model based on the temporal fusion transformer architecture, and the target model has a gated residual network, a variable selection network, and a multi-head attention mechanism.

[0034] Optionally, by predicting the price at the target moment, the prices at 96 moments per day for the next ten days can be predicted, where each moment is separated by 15 minutes.

[0035] It should be noted that a target model based on the Temporal-Fusion-Transformer (TFT) architecture is used for electricity price prediction. The target model has three key components: a Gated Residual Network (GRN), a variable selection network, and a multi-head attention mechanism. The gated residual network can adaptively adjust the contribution of historical data to the prediction, reduce information distortion, and maintain the deep learning ability of the model. The variable selection network dynamically calculates variable weights to screen out the data variables that have the most influence on electricity price prediction. The multi-head attention mechanism allows the model to process sequence information in parallel and capture the complex dependencies in electricity price fluctuations, especially long-term dependencies.

[0036] It should be noted that by introducing gating and residual connections, GRN improves the model's ability to capture non-linear relationships, reduces information loss, and maintains the depth and complexity of the model. The variable selection network can dynamically identify and adjust the weights of data variables to ensure that the most critical information is effectively utilized by the model, improving the accuracy of prediction. Through the multi-head attention mechanism, the model can process time series data from multiple perspectives, enhancing the ability to capture seasonal, periodic, and trend characteristics, making the prediction model more sensitive to the complex changes in time series.

[0037] In the above steps, through the target model, the electricity price at the target moment is predicted based on N sets of historical data. Since the target model used is a model based on the time series fusion transformer architecture and has a gated residual network, a variable selection network, and a multi-head attention mechanism, the accuracy of predicting the electricity price at the target moment is improved, thus solving the problem of low accuracy in predicting the electricity price.

[0038] In an exemplary embodiment, the method further includes the following steps: during the process of predicting the electricity price at the target moment based on the N sets of historical data by the target model, the following steps S11 - S13 are executed through the gated residual network:

[0039] Step S11: Process the target vector and the target context vector through the following formula to obtain the first data: η2 = ELU(W 2,ω α + W 3,ω c + b 2,ω ) Formula 1, where η2 is the first data, α is the target vector, c is the target context vector, ELU is the exponential linear unit activation function, W 2,ω is the second weight matrix, W 3,ω is the third weight matrix, and b 2,ω is the bias vector corresponding to the second weight matrix; the target vector μ x is the total number of features in a set of historical data corresponding to the t-th moment among the N moments, is the transformed input of the j-th feature in a set of historical data corresponding to the t-th moment, j is an integer greater than or equal to 1 and less than or equal to μ x , and the target context vector is the context vector calculated by the static covariate encoder in the target model;

[0040] It should be noted that the purpose of step S11 is to perform linear transformation and non-linear activation on the input data to prepare the input for subsequent gating operations. In GRN, W is a two-dimensional matrix with a dimension of dmodel×dmodel, where dmodel is the feature dimension of the model. The bias term b is one-dimensional and is added to the result of the matrix multiplication, adding the bias term b to each dimension. W and b are a general representation because GRN has multiple steps, and each step has a weight matrix and a bias matrix, where W 2,ω is the weight matrix for the second step, and W 3,ω is the weight matrix for the third step.

[0041] It should be noted that ELU can extract non-linear features in the input data, which is particularly crucial when dealing with complex and non-linear power market data. By combining with the context vector, GRN can consider the influence of static covariates on time series data and enhance the model's understanding of the environment and market conditions.

[0042] Optionally, α = Ξ t , c = c s .

[0043] Step S12: Process the first data through the following formula two to obtain the second data: η1 = W 1,ω , η2 + b 1,ω Formula two; where η1 is the second data, and W 1,ω is the first weight matrix, and b 1,ω is the bias vector corresponding to the first weight matrix;

[0044] It should be noted that W 1,ω is the weight matrix of the first step in GRN used to perform a linear transformation on η2, and its role is to weight the information transmitted from η2 calculated from the previous layer; b 1,ω is the bias vector corresponding to W 1,ω used to translate the result and is the bias matrix of the first layer linear transformation of the bias term b.

[0045] Optionally, W 1,ω and the corresponding bias vector b 1,ω together complete the linear transformation of the input η2.

[0046] It should be noted that through the weight matrix and the bias vector, the model can regulate the flow of information between the various layers of GRN to ensure that key features are fully expressed during the prediction process.

[0047] Step S13: Process the target vector and the second data through the following formula three to obtain the third data: GRN ω (α, c) = LayerNorm(α + GLUω (η1)) Formula III; where, GRN ω (α, c) is the third data, LayerNorm is the normalization process, GLU is the gated linear unit, ω is the weight sharing index, where the third data is to be used by the variable selection network.

[0048] It should be noted that GLU is a gated unit mechanism that combines a linear operation and a Sigmoid activation function, introducing gating selectivity and being suitable for scenarios that require efficient feature screening. GLU is different from ELU and is not just an activation function but a mechanism that combines information filtering and selection.

[0049] It should be noted that ω is used to share W 1,ω parameters across different positions or contexts.

[0050] It should be noted that through the gating mechanism, GRN can adaptively select which features are most critical for the prediction result, which enhances the model's adaptability and prediction accuracy; the normalization process and the use of the gated linear unit in the formula, combined with the residual connection, help the model learn deeper feature representations while maintaining the stability of training and preventing gradient vanishing or explosion.

[0051] It should be noted that the above steps make GRN an effective tool for dealing with the problem of predicting long-term time series of day-ahead prices in the electricity spot market in the TFT architecture, helping to overcome challenges such as data complexity, non-linear relationships, and long-term dependencies, thereby improving the reliability and accuracy of predictions.

[0052] In an exemplary embodiment, GLU ω (η1) = σ(W 4,ω η1 + b 4,ω ) ⊙ (W 5,ω η1 + b 5,ω ), where σ is the activation function, W 4,ω is the fourth weight matrix, b 4,ω is the bias vector corresponding to the fourth weight matrix, W 5,ω is the fifth weight matrix, b 5,ω is the bias vector corresponding to the fifth weight matrix, and ⊙ is the element-wise product.

[0053] It should be noted that ⊙ is the element-wise product (element-level Hadamard product), d model is the size of the hidden state and is unified throughout the architecture.

[0054] It should be noted that the weight matrix, weight sharing index, and bias vector are learned and adjusted during the training process to adapt to the dynamic changes in electricity market data. This enables the model to better handle the impacts of factors such as market rules, policy changes, and fluctuations in supply and demand relationships. The use of the Hadamard product, which combines the results of gating and linear transformation, helps optimize the gradient flow and avoid the common problems of vanishing gradients or exploding gradients when training deep networks. In this way, even when dealing with long-term time series data, the model can maintain good training performance and stability.

[0055] It should be noted that in the above steps, as the core mechanism in the GRN, GLU provides the TFT architecture with powerful non-linear modeling capabilities, an adaptive feature screening mechanism, and optimized training stability through its unique design and operation.

[0056] In an exemplary embodiment, the method further includes the following steps: during the process of predicting the electricity price at the target moment based on the target model and the N groups of historical data, the following steps S21 - S23 are executed through the variable selection network:

[0057] Step S21: Process the third data through the Softmax function to obtain the fourth data;

[0058] Optionally, the formula for the variable selection network to calculate the variable selection weight is: Where, is the third data, is the fourth data.

[0059] Step S22: Obtain the fifth data through the following formula four: Where, is the fifth data, is obtained after non-linearly transforming and is the j-th element in the fourth data;

[0060] Optionally, each generates a corresponding non-linear transformation through the corresponding GRN:

[0061] It should be noted that is the input feature weighted by the variable selection network and will be used for further processing and learning in the subsequent network layers. will be used as part of the input and fed into the multi-head attention layer to learn the long-term and short-term dependencies in the time series data; for each time step t, a will be calculated. If there are N time steps of input, N will be calculated. These It represents the input features weighted by the variable selection network at each time step t, which will be used in subsequent network layers for multi-step prediction.

[0062] Step S23: The fourth data and the fifth data are to be used by the multi-head attention mechanism.

[0063] It should be noted that the role of the variable selection network in the TFT architecture is to evaluate and select the most relevant input variables to optimize the prediction process. Through weight assignment by the Softmax function, as well as non-linear transformation and weighted fusion, the variable selection network can intelligently identify which variables are most important for predicting the electricity price at the target moment and highlight the features of these important variables. This not only improves the prediction ability of the model but also increases the interpretability of the model, making the basis for the prediction results clearer and helping electricity market participants better understand and respond to market dynamics.

[0064] In addition, by reducing the influence of redundant variables, the variable selection network can also reduce the complexity of the model, improve the training efficiency, reduce the risk of overfitting, and ensure that the model has higher reliability and generalization ability when dealing with the long-term time series prediction task of the day-ahead price in the electricity spot market. Finally, these optimized input data are passed to the multi-head attention mechanism to further extract the long-term dependencies in the sequence, laying the foundation for generating accurate prediction results.

[0065] In an exemplary embodiment, the method further includes the following steps: During the process of predicting the electricity price at the target moment based on the N sets of historical data according to the target model, the following steps S31 - S33 are performed by the multi-head attention mechanism:

[0066] Step S31: In the case of having m H attention heads, the attention output of the h-th attention head is obtained through the following formula five to obtain m H attention outputs of m H attention heads:

[0067] H h is the attention output of the h-th attention head, Q is the query matrix, K is the key matrix, V is the value matrix, is the query weight matrix of the h-th attention head, is the key weight matrix of the h-th attention head, is the value weight matrix of the h-th attention head;

[0068] It should be noted that the multi-head attention mechanism can extract features of input data from different perspectives by processing multiple attention heads in parallel, which is particularly important when dealing with diverse and complex datasets in the electricity market. Each attention head can focus on different aspects of the input data, such as seasonal patterns, short-term fluctuations, or long-term trends, thus capturing the time-series features of electricity prices more comprehensively and accurately.

[0069] Step S32: Concatenate the m H attention outputs to obtain a concatenated vector;

[0070] It should be noted that the multi-head attention mechanism allows the model to capture information from different subspaces, which helps the model learn richer representations. Each head can focus on different information. For example, some attention heads may focus on local patterns, while others may focus on global patterns. By concatenating the outputs of all attention heads together, the model can utilize these different pieces of information simultaneously.

[0071] Step S33: Multiply the concatenated vector by a linear transformation weight matrix to obtain the multi-head attention output.

[0072] Optionally, the working principle of the multi-head attention mechanism is as follows:

[0073] 1. Linear transformation: First, perform a linear transformation on Q, K, and V to obtain the query, key, and value for each attention head, which is achieved by multiplying by the corresponding weight matrices and ;

[0074] 2. Calculate attention: For each head, calculate the attention using the transformed query, key, and value. This is usually achieved through dot-product attention (Dot-Product Attention), and the formula is: Attention(Q, K, V) = A(Q, K)V, where

[0075] 3. Concatenate the outputs: Concatenate the attention outputs of all heads together to form a long vector: [H1,..., Hm H ;

[0076] 4. Linear transformation: Perform a linear transformation on the concatenated output to obtain the final multi-head attention output, which is achieved by multiplying by the weight matrix W H as follows: Multihead(Q, K, V) = [H1,..., Hm H W H , where Multihead(Q, K, V) is the final multi-head attention output.

[0077] Optionally, where \(R\) is the set of real numbers and \(n\) refers to the number of dimensions. Together, \(R^n\) n represents the \(n\)-dimensional real space, and \(d\) attention is an attention mechanism.

[0078] In an exemplary embodiment, the method further includes the following steps: In the process of predicting the electricity price at the target time based on the target model according to the \(N\) sets of historical data, the following steps are further performed through the multi-head attention mechanism: Calculate the shared value of the \(m\) H attention heads through the following formula six:

[0079]

[0080] where \(W\) V is the weight array shared by different attention heads, is the shared value, and \(H\) is the number of attention heads.

[0081] Optionally, represents the weight array shared by different heads, represents the final linear transformation result, \(H\) is the number of attention heads, i.e., \(H = m\) H .

[0082] It should be noted that the multi-head attention mechanism, the gated residual network, and the variable selection network are all interrelated components in the TFT. They jointly act on different parts of the target model to improve the performance and interpretability of multi-step time series prediction.

[0083] Optionally, the output of the GRN (weighted feature vector) is fed into the variable selection network for further processing. The GRN can be regarded as part of feature processing, which can help the model better understand and integrate input features; the output of the variable selection network (processed features) will be fed into the multi-head attention mechanism. Here, the multi-head attention mechanism will use these features to learn the long-term dependencies in the time series data; the output of the multi-head attention mechanism (merged multi-head features) will be used for subsequent prediction tasks, such as sequence-to-sequence prediction or quantization output; the GRN is responsible for screening important input features, the variable selection network is responsible for processing and integrating these features, and the multi-head attention mechanism is responsible for learning the complex relationships between features. These three components work together to enable the TFT to achieve high performance and interpretability in multi-step time series prediction tasks.

[0084] It should be noted that the GRN, variable selection network, and multi-head attention mechanism in the TFT architecture work together. They are respectively responsible for the non-linear processing and screening of features, the identification and weight adjustment of key variables, and the capture and information integration of complex relationships. The collaborative work of these components significantly improves the accuracy and interpretability of the model in processing the day-ahead price long-term time series prediction task in the electricity spot market. At the same time, it also improves the training efficiency and generalization ability of the model, providing more reliable and intelligent decision-making support for electricity market participants.

[0085] Obviously, the above-described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. To better understand the above method, the following will describe the above process in conjunction with embodiments, but it is not used to limit the technical solutions of the embodiments of the present invention. For example, Figure 3 Steps S301 to S312 in [Figure] schematically show the overall flowchart of a method for predicting electricity prices in an application embodiment. Specifically:

[0086] Offline stage (for the overall architecture diagram, refer to Figure 4 ):

[0087] Step S301: Offline data preprocessing. Collect historical electricity price data, historical market disclosure data, historical meteorological and hydrological data, and perform data cleaning and processing on the original data, including normalized data processing.

[0088] Step S302: Model construction. Build a prediction framework including a gated residual network, a variable selection network, and a multi-head attention mechanism through the TFT model.

[0089] Step S303: Model training. Use historical data to train the model and optimize the network parameters by minimizing the loss function.

[0090] Step S304: Model storage. Store the trained model parameters and structure to form an offline model.

[0091] Online stage:

[0092] Step S311: Collect historical electricity price data, operating day market disclosure data, meteorological and hydrological data, and perform data cleaning and processing on the original data, including normalized data processing.

[0093] Step S312: Input the TFT model stored in step S304 of the offline stage to obtain the 96-point day-ahead electricity price for the long term of 7 - 10 days in the future.

[0094] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present application.

[0095] In this embodiment, a device for predicting electricity prices is also provided. This device is used to implement the above embodiments and preferred implementation methods, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0096] Figure 5 is a structural block diagram of a device for predicting electricity prices according to an embodiment of the present application. This device includes:

[0097] An acquisition module 502, configured to acquire N sets of historical data at N moments. Among them, the i-th set of historical data in the N sets of historical data includes: historical meteorological data and historical electricity market operation data corresponding to the i-th moment among the N moments; the historical electricity market operation data includes historical electricity price data, N is an integer greater than or equal to 2, and i is an integer greater than or equal to 1 and less than or equal to N;

[0098] A prediction module 504, configured to predict the electricity price at a target moment based on a target model according to the N sets of historical data. Among them, the target model is a model based on a temporal fusion transformer architecture, and the target model has a gated residual network, a variable selection network, and a multi-head attention mechanism.

[0099] The above device predicts the electricity price at the target moment according to N sets of historical data through the target model. Since the target model used is a model based on a temporal fusion transformer architecture and has a gated residual network, a variable selection network, and a multi-head attention mechanism, the accuracy of predicting the electricity price at the target moment is improved, thereby solving the problem of low accuracy in predicting electricity prices.

[0100] In an exemplary embodiment, the above device further includes the following steps: a processing module, which, when predicting the electricity price at the target moment based on the target model according to the N sets of historical data, performs the following operations through the gated residual network: processing the target vector and the target context vector through the following formula one to obtain the first data: η2 = ELU(W 2,ω α + W 3,ω c + b 2,ω ), Formula One, where η2 is the first data, α is the target vector, C is the target context vector, ELU is the exponential linear unit activation function, W 2,ω is the second weight matrix, W 3,ω is the third weight matrix, and b 2,ω is the bias vector corresponding to the second weight matrix; the target vector μ x is the total number of features in a set of historical data corresponding to the t-th moment among the N moments, is the transformed input of the j-th feature in a set of historical data corresponding to the t-th moment, j is an integer greater than or equal to 1 and less than or equal to μ x , and the target context vector is the context vector calculated by the static covariate encoder in the target model; processing the first data through the following formula two to obtain the second data: η1 = W 1,ω η2 + b 1,ω , Formula Two; where η1 is the second data, W 1,ω is the first weight matrix, and b 1,ω is the bias vector corresponding to the first weight matrix; processing the target vector and the second data through the following formula three to obtain the third data: GRN ω (α, c) = LaterNorm(α + GLU ω (η1)), Formula Three; where GRN ω (α, c) is the third data, LayrNorm is the normalization process, GLU is the gated linear unit, ω is the weight sharing index, and the third data is to be used by the variable selection network.

[0101] In an exemplary embodiment, GLU ω (η1) = σ(W 4,ω η1 + b 4,ω ) ⊙ (W 5,ω η1 + b 5,ω ), where σ is the activation function, W 4,ω is the fourth weight matrix, b 4,ω is the bias vector corresponding to the fourth weight matrix, W 5,ω is the fifth weight matrix, and b5,ω is the bias vector corresponding to the fifth weight matrix, and ⊙ is the element-wise product

[0102] In an exemplary embodiment, the processing module is further configured to, in the process of predicting the electricity price at the target moment based on the target model according to the N sets of historical data, perform the following operations through the variable selection network: process the third data through the Softmax function to obtain the fourth data; obtain the fifth data through the following formula four: wherein, is the fifth data, is for obtained after a non-linear transformation of is the j-th element in the fourth data; the fourth data and the fifth data are to be used by the multi-head attention mechanism.

[0103] In an exemplary embodiment, the processing module is further configured to, in the process of predicting the electricity price at the target moment based on the target model according to the N sets of historical data, perform the following operations through the multi-head attention mechanism: in the case of having m H attention heads, obtain the attention output of the h-th attention head through the following formula five to obtain m H attention outputs of m H attention heads: H h is the attention output of the h-th attention head, Q is the query matrix, K is the key matrix, and V is the value matrix, is the query weight matrix of the h-th attention head, is the key weight matrix of the h-th attention head, is the value weight matrix of the h-th attention head; splice the m H attention outputs to obtain a spliced vector; multiply the spliced vector by the linear transformation weight matrix to obtain the multi-head attention output.

[0104] In an exemplary embodiment, the processing module is further configured to, in the process of predicting the electricity price at the target moment based on the target model according to the N sets of historical data, perform the following operations through the multi-head attention mechanism: calculate the shared value of the m H attention heads through the following formula six: wherein, W V is the weight array shared by different attention heads, is the shared value, and H is the number of attention heads.

[0105] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any of the above method embodiments when running.

[0106] Optionally, in this embodiment, the above storage medium may be configured to store a computer program for executing the following steps:

[0107] S1. Obtain N sets of historical data at N moments, where the i-th set of historical data in the N sets of historical data includes: historical meteorological data and historical power market operation data corresponding to the i-th moment among the N moments; the historical power market operation data includes historical electricity price data, N is an integer greater than or equal to 2, and i is an integer greater than or equal to 1 and less than or equal to N;

[0108] S2. Based on the target model, predict the electricity price at the target moment according to the N sets of historical data, where the target model is a model based on the architecture of a temporal fusion transformer, and the target model has a gated residual network, a variable selection network, and a multi-head attention mechanism.

[0109] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0110] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0111] An embodiment of the present application also provides a computer program product including a computer program, where the computer program executes the steps in any of the above method embodiments when executed by a processor.

[0112] An embodiment of the present application also provides an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.

[0113] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0114] S1. Obtain N sets of historical data at N moments. Among them, the i-th set of historical data in the N sets of historical data includes: the historical meteorological data and historical power market operation data corresponding to the i-th moment among the N moments; the historical power market operation data includes historical electricity price data, where N is an integer greater than or equal to 2, and i is an integer greater than or equal to 1 and less than or equal to N.

[0115] S2. Based on the target model, predict the electricity price at the target moment according to the N sets of historical data. Among them, the target model is a model based on the temporal fusion transformer architecture, and the target model has a gated residual network, a variable selection network, and a multi-head attention mechanism.

[0116] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0117] The specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0118] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.

[0119] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for predicting electricity prices, characterized in that: include: Acquire N groups of historical data at N moments, wherein the i-th group of historical data in the N groups of historical data includes: historical meteorological data and historical power market operation data corresponding to the i-th moment in the N moments; the historical power market operation data includes historical electricity price data, N is an integer greater than or equal to 2, and i is an integer greater than or equal to 1 and less than or equal to N; Based on the target model, the electricity price at the target moment is predicted according to the N groups of historical data, wherein the target model is a model based on a timing fusion converter architecture, and the target model has a gated residual network, a variable selection network and a multi-head attention mechanism.

2. The method according to claim 1, characterized in that The method further comprises: In the process of predicting the electricity price at the target time according to the N groups of historical data based on the target model, the following operations are performed by the gated residual network: The target vector and the target context vector are processed by the following formula to obtain the first data: η2 = ELU(W 2,ω α+W 3,ω c+b 2,ω ) Formula 1, where η2 is the first data, α is the target vector, c is the target context vector, ELU is the linear unit activation function, W 2,ω is the second weight matrix, W 3,ω is the third weight matrix, b 2,ω is the bias vector corresponding to the second weight matrix; the target vector μ x is the total number of features in a set of historical data corresponding to the tth moment among the N moments, is the transformation input of the jth feature in a set of historical data corresponding to the tth time, where j is greater than or equal to 1 and less than or equal to μ x An integer , wherein the target context vector is a context vector calculated by a static covariate encoder in the target model; The first data is processed by the following formula 2 to obtain the second data: η1=W W1ω η2+b 1,ω Formula 2: Wherein, η1 is the second data, W 1,ω is the first weight matrix, b 1,ω is the bias vector corresponding to the first weight matrix; The target vector and the second data are processed by the following formula 3 to obtain the third data: GRN ω (α,c)=LayerNorm(α+GLU ω (η1)) Formula 3; where GRN ω (α, c) is the third data, LayerNorm is normalization processing, GLU is a gated linear unit, and ω is a weight sharing index, wherein the third data is to be used by the variable selection network.

3. The method according to claim 2, characterized in that GLU ω (η1)=σ(W 4,ω η1+b 4,ω )☉(W 5,ω η1+b 5,ω ), where σ is the activation function, W 4,ω is the fourth weight matrix, b 4,ω is the bias vector corresponding to the fourth weight matrix, W 5,ω is the fifth weight matrix, b 5,ω is the bias vector corresponding to the fifth weight matrix, and ⊙ is the element-by-element product.

4. The method according to claim 2, characterized in that: The method further comprises: In the process of predicting the electricity price at the target time according to the N groups of historical data based on the target model, the following operations are performed through the variable selection network: The third data is processed by the Softmax function to obtain the fourth data; the fifth data is obtained by the following formula 4: in, is the fifth data, For After nonlinear transformation, we get: is the j-th element in the fourth data; The fourth data and the fifth data are to be used by the multi-head attention mechanism.

5. The method according to claim 1, characterized in that The method further comprises: Based on the target model, in the process of predicting the electricity price at the target time according to the N groups of historical data, the following operations are performed through the multi-head attention mechanism: In the presence of H In the case of h attention heads, the attention output of the hth attention head is obtained by the following formula 5 to obtain m H m attention heads H Attention output: H h is the attention output of the h-th attention head, Q is the query matrix, K is the key matrix, V is the value matrix, is the query weight matrix of the h-th attention head, is the key weight matrix of the h-th attention head, is the value weight matrix of the h-th attention head; The m H The attention outputs are concatenated to obtain the concatenated vector; Multiply the concatenated vector by the linear transformation weight matrix to obtain the multi-head attention output.

6. The method according to claim 5, characterized in that The method further comprises: Based on the target model, in the process of predicting the electricity price at the target time according to the N groups of historical data, the following operations are further performed through the multi-head attention mechanism: The m is calculated by the following formula 6 H The shared value of the attention heads: Among them, W V is the weight array shared by different attention heads, is the shared value, and H is the number of attention heads.

7. A device for predicting electricity prices, characterized in that: include: An acquisition module, used to acquire N groups of historical data at N moments, wherein the i-th group of historical data in the N groups of historical data includes: historical meteorological data and historical power market operation data corresponding to the i-th moment in the N moments; the historical power market operation data includes historical electricity price data, N is an integer greater than or equal to 2, and i is an integer greater than or equal to 1 and less than or equal to N; A prediction module is used to predict the electricity price at a target moment according to the N groups of historical data based on a target model, wherein the target model is a model based on a timing fusion converter architecture, and the target model has a gated residual network, a variable selection network and a multi-head attention mechanism.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 6 when executed.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 6 through the computer program.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.