Wind power generation power prediction method based on space-time diagram convolution and gating attention

By constructing a dynamic adjacency matrix and introducing spatiotemporal graph convolution and gated attention mechanism, the problem of insufficient spatial feature expression in wind power prediction is solved, and efficient and accurate wind power prediction is achieved to adapt to the complex dynamic characteristics of wind farms.

CN120597214APending Publication Date: 2025-09-05CHANGCHUN INST OF TECH

Patent Information

Application Number
CN202511094785.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing wind power prediction methods ignore the spatial coupling relationship between wind turbines, resulting in insufficient ability to express spatial features and difficulty in adapting to the dynamic and complex spatiotemporal characteristics of wind farm operation.

Method used

A method based on spatiotemporal graph convolution and gated attention is adopted. By obtaining the geographical location and historical wind power generation data of each wind turbine in the wind farm, a dynamic adjacency matrix is ​​constructed. Spatial features are extracted using a graph convolutional network, and an Informer encoder with a sparse attention mechanism is introduced for feature fusion to generate wind power prediction results for future time steps.

Benefits of technology

The accuracy and generalization ability of wind power generation prediction are improved, and it can adapt to complex dynamic changes under different meteorological conditions and wind turbine cluster sizes, reduce computational complexity and memory overhead, and shorten model inference latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597214A_ABST
    Figure CN120597214A_ABST
Patent Text Reader

Abstract

The invention relates to the field of new energy, and discloses a wind power generation power prediction method based on space-time diagram convolution and gating attention, and the method comprises the steps: obtaining the geographic position information, meteorological information and historical wind power generation power data of each fan in a wind power plant, and obtaining the normalized data; constructing a dynamic adjacency matrix based on the maximum information coefficient among the historical power data of each fan in the wind power plant, and generating a graph structure; a node set of the graph structure corresponds to each station in the wind power cluster, and an edge set is dynamically determined by a maximum information coefficient of historical power data between the stations; spatial feature extraction is carried out by using a graph convolutional network, and a graph structure learning module is introduced; and inputting the sequence output by the graph structure learning module into a gating circulation unit, introducing an Informer encoder based on a sparse attention mechanism, and generating a wind power prediction result of a future time step. According to the invention, high-precision prediction of the wind power generation power in a multi-fan scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of energy, and in particular to a wind power prediction method based on spatiotemporal graph convolution and gated attention. Background Art

[0002] Most current wind power prediction methods focus on time series modeling, often ignoring the potential spatial coupling relationship between wind turbines, or making static assumptions about the topological structure between wind turbines during modeling, resulting in insufficient ability to express spatial features and difficulty adapting to the dynamic and complex spatiotemporal characteristics during wind farm operation. Summary of the Invention

[0003] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a wind power prediction method based on spatiotemporal graph convolution and gated attention, comprising the following steps: Step 1: Obtain the geographical location information, meteorological information, and historical wind power data of each wind turbine in the wind farm, and perform normalization processing using the Min-Max normalization method to obtain normalized data; Step 2: construct a dynamic adjacency matrix based on the maximum information coefficient between the historical power data of each wind turbine in the wind farm, and generate a graph structure; the node set of the graph structure corresponds to each station in the wind farm cluster, and the edge set is dynamically determined by the maximum information coefficient of the historical power data between the stations; Step 3: Use graph convolutional networks to extract spatial features and introduce a graph structure learning module; Step 4: Input the sequence output by the graph structure learning module into the gated recurrent unit and introduce the Informer encoder based on the sparse attention mechanism; Step 5: The spatiotemporal features extracted and fused by multiple modules are input into the fully connected layer for mapping to generate the wind power prediction results for future time steps.

[0004] Furthermore, the acquisition of geographical location information, meteorological information, and historical wind power data of each wind turbine in the wind farm, and the normalization processing using the Min-Max normalization method to obtain normalized data include: Use Min-Max normalization to convert data of different dimensions to [0,1]: Where: is the normalized value, and Represent the maximum and minimum values ​​in the data respectively.

[0005] Furthermore, the method of constructing a dynamic adjacency matrix based on the maximum information coefficient between the historical power data of each wind turbine in the wind farm and generating a graph structure includes: discretizing and gridding the power data of each station in a two-dimensional space to achieve normalized calculation of the maximum mutual information value between variables, using the following formula: Where: Parameter B is 0.6 power of the total number of samples, the constraint condition Indicates the limit on the number of grid divisions; Based on the correlation characteristics of multiple meteorological elements and station power, the constructed graph is as follows: The node set Corresponding wind power cluster N Stations, side gatherings The adjacency matrix generated by the dynamic determination of the maximum information coefficient of the historical power data between stations is for: Where: The diagonal element 1 represents the node autocorrelation, and the off-diagonal element Reflects the intensity of nonlinear spatiotemporal correlation between stations.

[0006] Furthermore, the use of graph convolutional networks for spatial feature extraction and the introduction of graph structure learning modules include: picture The Laplace matrix of for: in is the Laplace matrix, is the node degree matrix, which is used to represent the number of edges connected to each node. It is the adjacency matrix of the graph, which is used to describe the connection relationship between nodes, and the regularized Laplace matrix for: Where: is the identity matrix, is the degree matrix, is the adjacency matrix, is the eigenvector matrix, , yes The eigenvalue diagonal moment of .

[0007] Furthermore, the order Chebyshev polynomials are used to approximate the graph convolution operation, including: is a positive definite matrix, which can be considered as a basis for the spectral space: Where: and are the representations of the signal in the node domain and spectral domain respectively.

[0008] According to the node domain signal on the graph and , combined with the convolution theorem to obtain the graph convolution operator: Where: For Hadamard, It is the convolution operation of the spectral graph; Use convolution kernel replace , the original Hadamard product multiplication will become a matrix multiplication: use Chebyshev polynomials of order are used to approximate the graph convolution operation: Where: , , are the Chebyshev polynomial coefficients, yes Chebyshev polynomials of order, , yes The maximum eigenvalue of The GCN layered propagation formula is: Where: Represents node self-connection, , yes The degree matrix of is the activation function, is the feature matrix, It is The learning parameters of the layer.

[0009] Furthermore, the spatiotemporal graph convolution module integrates the one-dimensional causal convolution kernel with the gated linear unit activation mechanism, which is expressed as: Where: and are the gating weight matrix and bias term respectively.

[0010] Furthermore, the gated recurrent unit is: Where: is the update gate output; is the reset gate output; Represented as candidate hidden states; It is a hidden state; 、 、 ; is the parameter weight between the input layer and the hidden layer; 、 、 is the parameter weight between hidden layers; 、 、 is the offset.

[0011] Furthermore, the sparse attention mechanism filters the query vector by using a sparsity evaluation function, which is composed of the difference between the logarithmic sum exponential term and the arithmetic mean term: in, is the key vector sequence length, Characterize the input feature dimension; The attention output is implemented through sparse matrix operations: Where: is the filtered sparse query matrix, whose dimension satisfies .

[0012] Furthermore, the sparse attention mechanism of the Informer encoder includes a distillation operation. Layer to The feature distillation process of the layer is: Where: Indicates the The feature tensor of the layer processed by the multi-head ProbSparse self-attention module, is a one-dimensional convolution operation with a kernel width of 3. is the exponential linear unit activation function, The feature dimension is further compressed by downsampling.

[0013] Furthermore, the Informer decoder preprocesses the target sequence using a zero-filling strategy. The decoder input feature matrix is ​​constructed by concatenating the start token and the target placeholder, including: The decoder input feature matrix is ​​constructed by concatenating the start token and the target placeholder: Where: Indicates the starting identifier containing historical temporal semantics, The target sequence occupancy matrix is ​​initialized to all zeros, and Respectively represent the encoding window length and the prediction sequence step length, Defined as the model hidden layer feature dimension.

[0014] The beneficial effects of the present invention are: by introducing the maximum information coefficient (MIC) to construct a dynamic adjacency matrix, the nonlinear spatiotemporal correlation between wind farms can be accurately quantified, overcoming the limitations of traditional static topology assumptions on spatial feature expression, so that the graph structure can adapt to the complex dynamic characteristics of wind farms, laying a more practical foundation for subsequent feature extraction; at the same time, the spatiotemporal graph convolution module combines spectral domain graph convolution with one-dimensional causal convolution to achieve a deep fusion of spatial local features and time series features, and the gated attention module extracts local temporal features through GRU and combines it with the sparse attention mechanism to capture long-range dependencies. The synergistic effect of multiple modules significantly improves the accuracy of power prediction.

[0015] Min-Max normalization is used in the data preprocessing stage to unify the data to the range of [0,1], accelerating the convergence of the model. In the gated attention module, the sparse attention mechanism uses probability-driven screening of key time steps, reducing redundant operations in attention calculations and lowering the computational complexity and memory overhead of the model. The distillation operation of the Informer encoder further compresses the feature dimensions, and the zero-value filling strategy of the decoder uses generative single-step prediction instead of dynamic recursive decoding, which significantly shortens the model inference latency and enables the system to efficiently process long-sequence time series data. At the same time, it enhances the numerical stability of links such as LSE operations.

[0016] The construction method of the dynamic adjacency matrix enables the system to flexibly adapt to the dynamic changes in spatiotemporal characteristics during the operation of the wind farm without relying on fixed topological structure assumptions; the gated linear unit (GLU) of the spatiotemporal graph convolution module dynamically adjusts the activation strength of the feature channel through a dual-path gating mechanism, and the gated recurrent unit (GRU) can effectively handle the dependencies in the time series. Combined with the sparse attention's ability to focus on key information, the model can maintain good prediction performance under different meteorological conditions and different wind turbine cluster sizes, and has stronger generalization capabilities.

[0017] High-precision wind power generation prediction results can provide a scientific basis for power system scheduling planning, energy storage configuration, and wind power grid stability control, helping to reduce wind curtailment rates, improve the utilization efficiency of wind energy resources, and promote the economic, safe, and stable operation of new energy power systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Schematic diagram of the flow of wind power prediction method based on spatiotemporal graph convolution and gated attention. DETAILED DESCRIPTION

[0019] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the following.

[0020] The features and performance of the present invention are further described in detail below with reference to the embodiments.

[0021] like Figure 1 As shown in the figure, the wind power prediction method based on spatiotemporal graph convolution and gated attention includes the following steps: Step 1: Obtain the geographical location information, meteorological information, and historical wind power data of each wind turbine in the wind farm, and perform normalization processing using the Min-Max normalization method to obtain normalized data; Step 2: construct a dynamic adjacency matrix based on the maximum information coefficient between the historical power data of each wind turbine in the wind farm, and generate a graph structure; the node set of the graph structure corresponds to each station in the wind farm cluster, and the edge set is dynamically determined by the maximum information coefficient of the historical power data between the stations; Step 3: Use graph convolutional networks to extract spatial features and introduce a graph structure learning module; Step 4: Input the sequence output by the graph structure learning module into the gated recurrent unit and introduce the Informer encoder based on the sparse attention mechanism; Step 5: The spatiotemporal features extracted and fused by multiple modules are input into the fully connected layer for mapping to generate the wind power prediction results for future time steps.

[0022] The acquisition of geographical location information, meteorological information and historical wind power data of each wind turbine in the wind farm, and normalization using the Min-Max normalization method to obtain normalized data includes: Use Min-Max normalization to convert data of different dimensions to [0,1]: Where: is the normalized value, and Represent the maximum and minimum values ​​in the data respectively.

[0023] The method of constructing a dynamic adjacency matrix based on the maximum information coefficient between the historical power data of each wind turbine in the wind farm and generating a graph structure includes: discretizing and gridding the power data of each station in a two-dimensional space to achieve normalized calculation of the maximum mutual information value between variables, using the following formula: Where: Parameter B is 0.6 power of the total number of samples, the constraint condition Indicates the limit on the number of grid divisions; Based on the correlation characteristics of multiple meteorological elements and station power, the constructed graph is as follows: The node set Corresponding wind power cluster N Stations, side gatherings The adjacency matrix generated by the dynamic determination of the maximum information coefficient of the historical power data between stations is for: Where: The diagonal element 1 represents the node autocorrelation, and the off-diagonal element Reflects the intensity of nonlinear spatiotemporal correlation between stations.

[0024] The method of using graph convolutional networks to extract spatial features and introducing a graph structure learning module includes: picture The Laplace matrix of for: in is the Laplace matrix, is the node degree matrix, which is used to represent the number of edges connected to each node. It is the adjacency matrix of the graph, which is used to describe the connection relationship between nodes, and the regularized Laplace matrix for: Where: is the identity matrix, is the degree matrix, is the adjacency matrix, is the eigenvector matrix, , yes The eigenvalue diagonal moment of .

[0025] The graph convolution operation is approximated using Chebyshev polynomials of order 1, including: is a positive definite matrix, which can be considered as a basis for the spectral space: Where: and are the representations of the signal in the node domain and spectral domain respectively.

[0026] According to the node domain signal on the graph and , combined with the convolution theorem to obtain the graph convolution operator: Where: For Hadamard, It is the convolution operation of the spectral graph; Use convolution kernel replace , the original Hadamard product multiplication will become a matrix multiplication: use Chebyshev polynomials of order are used to approximate the graph convolution operation: Where: , , are the Chebyshev polynomial coefficients, yes Chebyshev polynomials of order, , yes The maximum eigenvalue of The GCN layered propagation formula is: Where: Represents node self-connection, , yes The degree matrix of is the activation function, is the feature matrix, It is The learning parameters of the layer.

[0027] The spatiotemporal graph convolution module combines a one-dimensional causal convolution kernel with a gated linear unit activation mechanism, which is represented as: Where: and are the gating weight matrix and bias term respectively.

[0028] The gated recurrent unit is: Where: is the update gate output; is the reset gate output; Represented as candidate hidden states; It is a hidden state; 、 、 ; is the parameter weight between the input layer and the hidden layer; 、 、 is the parameter weight between hidden layers; 、 、 is the offset.

[0029] The sparse attention mechanism is to filter the query vector by a sparsity evaluation function, which is composed of the difference between the logarithmic sum exponential term and the arithmetic mean term: in, is the key vector sequence length, Characterize the input feature dimension; The attention output is implemented through sparse matrix operations: Where: is the filtered sparse query matrix, whose dimension satisfies .

[0030] The sparse attention mechanism of the Informer encoder includes a distillation operation, Layer to The feature distillation process of the layer is: Where: Indicates the The feature tensor of the layer processed by the multi-head ProbSparse self-attention module, is a one-dimensional convolution operation with a kernel width of 3. is the exponential linear unit activation function, The feature dimension is further compressed by downsampling.

[0031] The informer decoder preprocesses the target sequence using a zero-filling strategy. The decoder input feature matrix is ​​constructed by concatenating the start token and the target placeholder, including: The decoder input feature matrix is ​​constructed by concatenating the start token and the target placeholder: Where: Indicates the starting identifier containing historical temporal semantics, The target sequence occupancy matrix is ​​initialized to all zeros, and Respectively represent the encoding window length and the prediction sequence step length, Defined as the model hidden layer feature dimension.

[0032] Specifically, in order to maintain the wind data values ​​at the same time within a certain range and further improve the prediction accuracy and convergence speed of the model, Min-Max normalization is used to convert data of different dimensions to [0,1]: Where: is the normalized value, and Represent the maximum and minimum values ​​in the data respectively.

[0033] In the field of meteorological data analysis for power systems, discretized datasets based on numerical weather prediction (NWP) provide multi-scale meteorological elements for modeling the spatiotemporal correlation between stations. To this end, the maximum information coefficient (MIC) is introduced to quantify the nonlinear spatiotemporal correlation between station power. Its theoretical basis can be expressed as the mutual information metric: Where: mutual information Indicates station and The statistical correlation of For the station The information entropy of Condition Entropy, which represents the joint probability density function, and is the marginal probability distribution.

[0034] In the graph structure construction method based on the maximum information coefficient (MIC), the power data of each station is discretized and gridded in a two-dimensional space to achieve the normalized calculation of the maximum mutual information value between variables. This process can be formally described as: Where: parameter B is 0.6 power of the total number of samples, and the constraint condition is Indicates the limit on the number of mesh divisions.

[0035] Based on the correlation characteristics of multiple meteorological elements and station power, the constructed meteorological map model can be represented as , where the node set Corresponding to the stations in the wind power cluster, the edge collection The MIC coefficient of the historical power data between stations is dynamically determined. The adjacency matrix generated based on this is It can be expressed as: Where: The diagonal element 1 represents the node autocorrelation, and the off-diagonal element Reflects the intensity of nonlinear spatiotemporal correlation between stations.

[0036] Spatiotemporal Graph Convolution Module Using the convolution theorem on graphs to define graph convolution from the spectral domain is called spectral method, which can realize the extraction of local structural features and learn the network hidden layer representation. Laplacian Matrix for ,in is the Laplace matrix, is the node degree matrix, which is used to represent the number of edges connected to each node, that is, the number of edges connected to the node. Is the adjacency matrix of the graph, which is used to describe the connection relationship between nodes. The regularized Laplacian Matrix is: Where: is the identity matrix, is the degree matrix, is the adjacency matrix, is the eigenvector matrix, , yes The diagonal matrix of eigenvalues.

[0037] is a positive definite matrix, which can be considered as a basis for the spectral space: Where: and are the representations of the signal in the node domain and spectral domain respectively.

[0038] According to the node domain signal on the graph and , combined with the convolution theorem to obtain the graph convolution operator: Where: For Hadamard, It is the lineage graph convolution operation.

[0039] Use convolution kernel replace , the original Hadamard product multiplication will become a matrix multiplication: use Chebyshev polynomials of order are used to approximate the graph convolution operation: Where: , , are the Chebyshev polynomial coefficients, yes Chebyshev polynomials of order, , yes The maximum eigenvalue of .

[0040] Finally, the GCN layer-by-layer propagation formula is: Where: Represents node self-connection, , yes The degree matrix of is a nonlinear activation function, is the feature matrix, are the learning parameters of the layer.

[0041] The temporal modeling module of the Spatiotemporal Graph Convolutional Network (STGCN) constructs a feature extraction framework with temporal causal constraints by fusing a 1D causal convolution kernel with a gated linear unit (GLU) activation mechanism. The 1D causal convolution kernel ensures the unidirectionality of temporal information transmission by limiting the direction of the receptive field of the convolution operation, thus preventing future information leakage. The GLU function dynamically adjusts the activation strength of the feature channel through a dual-path gating mechanism. Its mathematical form can be expressed as: Where: and are the gating weight matrix and bias term respectively.

[0042] Gated Attention Module First, the STGCN output sequence is fed into the GRU unit to extract local temporal features. The ProbSparse Attention mechanism in the Informer is then introduced to select a small number of key time steps for global modeling through probabilistic driving, effectively reducing computational complexity and enhancing the model's ability to capture long-range dependencies. The calculation formula for the Gated Recurrent Unit (GRU) neural network is: Where: is the update gate output; is the reset gate output; Represented as candidate hidden states; It is a hidden state; 、 、 ; is the parameter weight between the input layer and the hidden layer; 、 、 is the parameter weight between hidden layers; 、 、 is the offset.

[0043] The sparse attention mechanism introduces the Kullback-Leibler divergence to quantify the sparsity of the query vector, thereby optimizing the focusing efficiency of the attention distribution. Define the query vector With key vector The associated probability distribution is , whose sparsity is improved by contrasting the uniform distribution The greater the distribution difference, the higher the query vector activity. It is composed of the difference between the logarithmic sum exponential term (Log-Sum-Exp, LSE) and the arithmetic mean term: Where: is the key vector sequence length, Characterizes the input feature dimension.

[0044] Its original calculation process exists Memory overhead and numerical stability risks of LSE operations. To this end, an approximate optimization solution is proposed: when When the value is high, the attention probability distribution shows a long tail characteristic, and the sparsity score is preferred. Finally, the attention output is realized by the sparse matrix operation shown in the following formula: Where: is the filtered sparse query matrix, whose dimension satisfies .

[0045] The distillation operation in the sparse attention mechanism effectively reduces the redundancy of the encoder feature map through the feature screening strategy. The core of this operation is to strengthen the weight distribution of the dominant attention features and optimize the feature focusing ability through hierarchical transmission. Layer to The feature distillation process of the layer can be expressed as: Where: Indicates the The feature tensor of the layer processed by the multi-head ProbSparse self-attention module, is a one-dimensional convolution operation with a kernel width of 3. is the exponential linear unit activation function, The feature dimension is further compressed by downsampling.

[0046] In the decoder architecture design, a zero-padding strategy is used to preprocess the target sequence to improve the processing efficiency of long time series data. This mechanism replaces traditional dynamic recursive decoding with a generative single-step prediction paradigm, significantly reducing model inference latency.

[0047] The decoder input feature matrix is ​​constructed by concatenating the start token and the target placeholder: Where: Indicates the starting identifier containing historical temporal semantics, The target sequence occupancy matrix is ​​initialized to all zeros, and Respectively represent the encoding window length and the prediction sequence step length, Defined as the model hidden layer feature dimension.

[0048] A wind power prediction system based on spatiotemporal graph convolution and gated attention, which applies a wind power prediction method based on spatiotemporal graph convolution and gated attention, including a data input module, a data preprocessing module, a graph structure construction module, a spatiotemporal graph convolution module, a gated attention module, and a prediction output module; The data input module is used to input meteorological information, wind turbine geographical information and historical wind power generation data; The data preprocessing module is used to perform normalization processing on the input data; The graph structure construction module is used to construct a dynamic connection matrix based on input data to generate a graph structure, wherein the node set of the graph structure corresponds to each station in the wind power cluster, and the edge set is dynamically determined by the maximum information coefficient (MIC) of historical power data between stations; The spatiotemporal graph convolution module is used to extract spatial features using a graph convolutional network and includes a graph structure learning submodule; The gated attention module is used to receive the sequence output by the spatiotemporal graph convolution module and input it into the gated recurrent unit (GRU), and includes an informer encoder based on the sparse attention mechanism; The prediction output module is used to map the spatiotemporal features extracted and fused by multiple modules through a fully connected layer to generate wind power prediction results for future time steps.

Claims

1. A wind power prediction method based on spatiotemporal graph convolution and gated attention, characterized by: The steps include: Step 1: Obtain the geographical location information, meteorological information, and historical wind power data of each wind turbine in the wind farm, and perform normalization processing using the Min-Max normalization method to obtain normalized data; Step 2: construct a dynamic adjacency matrix based on the maximum information coefficient between the historical power data of each wind turbine in the wind farm, and generate a graph structure; the node set of the graph structure corresponds to each station in the wind farm cluster, and the edge set is dynamically determined by the maximum information coefficient of the historical power data between the stations; Step 3: Use graph convolutional networks to extract spatial features and introduce a graph structure learning module; Step 4: Input the sequence output by the graph structure learning module into the gated recurrent unit and introduce the Informer encoder based on the sparse attention mechanism; Step 5: The spatiotemporal features extracted and fused by multiple modules are input into the fully connected layer for mapping to generate the wind power prediction results for future time steps.

2. The wind power prediction method based on spatiotemporal graph convolution and gated attention according to claim 1 is characterized in that: The acquisition of geographical location information, meteorological information and historical wind power data of each wind turbine in the wind farm, and normalization using the Min-Max normalization method to obtain normalized data includes: Use Min-Max normalization to convert data of different dimensions to [0,1]: Where: is the normalized value, and Represent the maximum and minimum values ​​in the data respectively.

3. The wind power prediction method based on spatiotemporal graph convolution and gated attention according to claim 1 is characterized in that: The method of constructing a dynamic adjacency matrix based on the maximum information coefficient between the historical power data of each wind turbine in the wind farm and generating a graph structure includes: discretizing and gridding the power data of each station in a two-dimensional space to achieve normalized calculation of the maximum mutual information value between variables, using the following formula: Where: Parameter B is 0.6 power of the total number of samples, the constraint condition Indicates the limit on the number of grid divisions; Based on the correlation characteristics of multiple meteorological elements and station power, the constructed graph is as follows: The node set Corresponding wind power cluster N Stations, side gatherings The adjacency matrix generated by the dynamic determination of the maximum information coefficient of the historical power data between stations is for: Where: The diagonal element 1 represents the node autocorrelation, and the off-diagonal element Reflects the intensity of nonlinear spatiotemporal correlation between stations.

4. The wind power prediction method based on spatiotemporal graph convolution and gated attention according to claim 3 is characterized in that: The method of using graph convolutional networks to extract spatial features and introducing a graph structure learning module includes: picture The Laplace matrix of for: in is the Laplace matrix, is the node degree matrix, which is used to represent the number of edges connected to each node. It is the adjacency matrix of the graph, which is used to describe the connection relationship between nodes, and the regularized Laplace matrix for: Where: is the identity matrix, is the degree matrix, is the adjacency matrix, is the eigenvector matrix, , yes The eigenvalue diagonal moment of .

5. The wind power prediction method based on spatiotemporal graph convolution and gated attention according to claim 4 is characterized in that: The graph convolution operation is approximated using Chebyshev polynomials of order 1, including: is a positive definite matrix, which can be considered as a basis for the spectral space: Where: and are the representations of the signal in the node domain and spectral domain respectively According to the node domain signal on the graph and , combined with the convolution theorem to obtain the graph convolution operator: Where: For Hadamard, It is the convolution operation of the spectral graph; Use convolution kernel replace , the original Hadamard product multiplication will become a matrix multiplication: use Chebyshev polynomials of order are used to approximate the graph convolution operation: Where: , , are the Chebyshev polynomial coefficients, yes Chebyshev polynomials of order, , yes The maximum eigenvalue of The GCN layered propagation formula is: Where: Represents node self-connection, , yes The degree matrix of is the activation function, is the feature matrix, It is The learning parameters of the layer.

6. The wind power prediction method based on spatiotemporal graph convolution and gated attention according to claim 1, characterized in that: The spatiotemporal graph convolution module combines a one-dimensional causal convolution kernel with a gated linear unit activation mechanism, which is represented as: Where: and are the gating weight matrix and bias term respectively.

7. The wind power prediction method based on spatiotemporal graph convolution and gated attention according to claim 6 is characterized in that: The gated recurrent unit is: Where: is the update gate output; is the reset gate output; Represented as candidate hidden states; It is a hidden state; 、 、 ; is the parameter weight between the input layer and the hidden layer; 、 、 is the parameter weight between hidden layers; 、 、 is the offset.

8. The wind power prediction method based on spatiotemporal graph convolution and gated attention according to claim 1 is characterized in that: The sparse attention mechanism is to filter the query vector by a sparsity evaluation function, which is composed of the difference between the logarithmic sum exponential term and the arithmetic mean term: in, is the key vector sequence length, Characterize the input feature dimension; The attention output is implemented through sparse matrix operations: Where: is the filtered sparse query matrix, whose dimension satisfies .

9. The wind power prediction method based on spatiotemporal graph convolution and gated attention according to claim 1, characterized in that: The sparse attention mechanism of the Informer encoder includes a distillation operation, Layer to The feature distillation process of the layer is: Where: Indicates the The feature tensor of the layer processed by the multi-head ProbSparse self-attention module, is a one-dimensional convolution operation with a kernel width of 3. is the exponential linear unit activation function, The feature dimension is further compressed by downsampling.

10. The wind power prediction method based on spatiotemporal graph convolution and gated attention according to claim 1, characterized in that: The informer decoder preprocesses the target sequence using a zero-filling strategy. The decoder input feature matrix is ​​constructed by concatenating the start token and the target placeholder, including: The decoder input feature matrix is ​​constructed by concatenating the start token and the target placeholder: Where: Indicates the starting identifier containing historical temporal semantics, The target sequence occupancy matrix is ​​initialized to all zeros, and Respectively represent the encoding window length and the prediction sequence step length, Defined as the model hidden layer feature dimension.

Citation Information

Patent Citations

  • Unsupervised depth multi-scale SAR image change detection method based on spatial frequency domain

    CN116778207A

  • Multi-wind power plant power prediction method based on attention space-time synchronization graph convolutional network

    CN117498296A

  • Traffic flow prediction method based on dynamic sparse graph convolution GRU

    CN119694142A

  • Multi-node fan cluster wind power prediction method, computer program product and storage medium

    CN120124007A

Cited By

  • Charging pile order quantity prediction method and system based on evolution graph and attention mechanism

    CN120931325A

  • Charging pile order quantity prediction method and system based on evolutionary graph and attention mechanism

    CN120931325B

  • Robot visual control method and system based on spatio-temporal context perception

    CN120985680A

  • Typhoon multi-tide station water level prediction method based on dynamic attention map circulation network

    CN120995023A

  • Photovoltaic power generation prediction method and system based on uncertainty graph convolution

    CN121052461A