Multi-node fan cluster wind power prediction method, computer program product and storage medium
Through the spatial and temporal parallel model framework, combined with the multi-head cross-attention mechanism of the spatial and temporal branch modules, the problem of space-time dependence integration in multi-node offshore wind power prediction is solved, and the prediction accuracy and efficiency are significantly improved.
Patent Information
- Application Number
- CN202510185876.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to capture and integrate multi-level space-time dependence in parallel in multi-node offshore wind power power prediction, resulting in insufficient prediction accuracy and efficiency.
The space-time parallel model framework is adopted, including the spatial branch module and the time branch module, and multi-grained interaction is achieved through the space-time fusion module. The spatial branch module uses Euclidean distance and Pearson correlation coefficient to construct spatial adjacency graphs and functional relationship graphs, and combines Chebishev polynomial approximation optimization graph convolution networks. The time branch module uses a multi-resolution convolution kernel-gating module and a recursive time module to extract time dependencies and integrates it through a multi-head cross-attention mechanism.
Parallel capture and integration of multi-level space-time dependence is realized, which significantly improves the accuracy and efficiency of multi-node wind power power prediction.
Smart Images

Figure CN120124007A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system prediction, and particularly relates to a method for predicting wind power of a multi-node wind turbine cluster, a computer program product, and a storage medium. Background Art
[0002] With the continuous depletion of fossil fuels and the increasingly acute environmental problems caused by them, the utilization of renewable energy has become a key topic of concern in the global energy field. As a renewable energy source, wind energy has played an important role in the global energy transition due to its cleanliness, sustainability, and cost-effectiveness. Firstly, since offshore wind farms are far from land, they cause less interference to human life and the ecological environment. In addition, compared with onshore wind energy, offshore wind energy has fewer power generation calm periods, so it has longer power generation time and higher power generation efficiency. However, the offshore power generation power is restricted by the inherent randomness and intermittency of wind energy and the more complex and changeable offshore wind power environment. At the same time, with the rapid growth of wind power penetration, the large-scale grid connection of wind power poses great challenges to the dispatching and planning of the power system. Therefore, improving the accuracy of offshore wind power prediction not only helps to optimize the stability of the power system but also can improve the wind power consumption rate.
[0003] The research on wind power prediction technology can be divided into short-term prediction (0 - 24h) and long-term prediction (1 - 7d) according to the time scale. Short-term prediction helps the real-time dispatching of the power grid, such as fan control and load tracking. Long-term prediction focuses more on the long-term maintenance and planning of the power grid, and the accuracy of the prediction results will decrease as the complexity of the data increases. To improve the accuracy of short-term wind power prediction, many scholars have done a lot of relevant research. The existing prediction methods can be roughly divided into four categories: physical-based methods, traditional statistical methods, machine learning methods, and deep learning methods.
[0004] Physical methods predict the power generation of wind turbines by simulating and deducing the physical processes in the real world. Among them, numerical weather prediction (NWP) plays a crucial role in physical methods. Physical methods perform well in long-term prediction, but they strongly rely on a large amount of weather forecast data. For example, there are many problems in the simulation of meteorological data for offshore wind power by the NWP model, which leads to extremely complex calculations of physical methods based on NWP data and greatly limits their application in short-term offshore wind power prediction. Traditional statistical methods establish the mapping relationship between predicted values and measured values based on historical data itself. They analyze the distribution, trend, and correlation of historical data to establish a certain correlation between wind power detection variables. Common traditional statistical methods include autoregressive (AR), autoregressive moving average (ARMA), and autoregressive integrated moving average (ARIMA). Compared with physical methods, statistical models do not require complex variables and have better prediction effects in short-term prediction. However, statistical methods often rely on fixed assumptions, which limits their prediction performance when dealing with non-linear and high-dimensional data. Different from physical methods and traditional statistical methods, machine learning and deep learning use data as the driving force to establish prediction models. Machine learning captures the complex pattern relationship between input and output data through non-linear functions, showing the advantages of self-learning, self-organization, and self-adaptation. Common deep learning methods include extreme learning machine (ELM), multi-layer perceptron (MLP), and support vector machine (SVR). Although machine learning methods can effectively capture the non-linear characteristics of offshore wind power data, they may have defects when dealing with extreme or complex information. Deep learning methods generally use deep neural networks, which can effectively solve the shortcomings of machine learning methods.
[0005] In real scenarios, when multiple wind turbines generate wind power, they usually have both temporal and spatial correlations. Existing studies have shown that when considering the spatio-temporal correlation of wind power data, the wind power prediction error is significantly reduced. To combine spatio-temporal features, relevant studies have introduced the convolutional neural network (CNN), which is good at extracting local spatial features, and the basic strategy is to use it in combination with the recurrent neural network family.
[0006] Although the method of combining convolutional neural networks and recurrent neural networks and their variant networks has high flexibility in extracting spatial domain information and capturing temporal characteristics. However, the convolutional operation of CNN is based on a fixed connection pattern and can only process regular grid-like structure data (commonly two-dimensional images, one-dimensional time series), which greatly limits the ability of spatio-temporal models to extract spatial features between multiple nodes.
[0007] To break through the technical bottleneck of traditional convolution, graph neural networks (GNNs) map the constructed graphs (where nodes are regarded as different variables and edges represent the relationships between variables) to Euclidean space, utilize the structural features between nodes for information propagation, and perform relational reasoning based on associations, thereby capturing the spatial dependence relationships between multiple nodes. As a general graph learning structure of GNNs, graph convolutional networks (GCNs) perform excellently in processing non-Euclidean data. However, in the application of wind power prediction, there is still room for further optimization and improvement. Although existing methods have made some progress in the multi-node offshore wind power prediction task, there are still some limitations.
[0008] For different wind turbines, the historical wind power data and some other input variables may have different periodic characteristics. Local time features have a significant impact on prediction performance, while existing methods often perform poorly in adapting to multiple periodic characteristics; the actual distribution of wind turbines has unstructured and non-Euclidean characteristics and does not have a predefined and explicit graphical structure; in addition, the propagation of wind in the spatial domain has a certain degree of persistence, and the spatio-temporal information between adjacent wind turbines is dynamically highly correlated. However, existing methods lack the exploration of implicit laws when constructing the spatial relationship matrix and are not sufficient to capture all dependencies; at this time, the spatial correlation features between variables will affect the time features of power generation, indicating that both jointly affect the future power generation. However, the spatial and time features extracted by existing methods are completely separated. Therefore, it is very important to extract and fuse the key information of spatio-temporal features in parallel. Summary of the Invention
[0009] Aiming at the above deficiencies of the existing technology, the technical problem to be solved by the present invention is: how to provide a prediction method for the wind power of a multi-node wind turbine cluster that can capture and integrate multi-level spatio-temporal dependencies in parallel and improve the accuracy and efficiency of multi-node wind power prediction.
[0010] To solve the above technical problems, the present invention adopts the following technical solutions:
[0011] A prediction method for the wind power of a multi-node wind turbine cluster, comprising:
[0012] Obtain the geographical coordinate data and historical power data of the wind turbine cluster, and perform normalization processing on the data;
[0013] Construct a spatio-temporal parallel model framework, which includes a spatial branch module, a time branch module, and a spatio-temporal fusion module;
[0014] The spatial branch module measures the spatial proximity between wind turbines using the Euclidean distance based on geographical coordinate data to obtain the spatial adjacency graph SAG. Based on historical power data, it linearly quantifies the functional similarity relationship between wind turbines using the Pearson correlation coefficient to obtain the functional relationship graph FRG. It approximates and optimizes the graph convolutional network using Chebyshev polynomials, and jointly aggregates the cross-graph feature correlation between the spatial adjacency graph SAG and the functional relationship graph FRG through a multi-head cross-attention mechanism to output multi-graph fusion features.
[0015] The temporal branch module uses parallel multi-resolution convolutional kernel-gated modules and a recurrent temporal module to extract the short-term and long-term temporal dependencies of the wind power sequence, and integrates them through a multi-head cross-attention mechanism to output temporal sequence features.
[0016] The spatial branch module and the temporal branch module work in parallel. Through the spatio-temporal fusion module, a multi-head cross-attention mechanism is used to fuse the multi-graph fusion features and the temporal sequence features, realizing multi-granularity interaction between spatio-temporal features and outputting the prediction result of wind power.
[0017] As an optimization, the construction of the spatial adjacency graph SAG includes: based on the geographical coordinate data of the wind turbine cluster, obtaining the geographical coordinate matrix of the wind turbine cluster, and calculating the Euclidean distance D between any pair of wind turbines in the geographical coordinate matrix ij =||C i -C j || 2 ,where C i and C j represent the coordinate vectors of the i-th and j-th wind turbines respectively, and ||·|| 2 represents quantifying the geometric distance between two wind turbines through the definition of the second norm.
[0018] As an optimization, the construction of the functional relationship graph FRG includes: based on the historical power data of the wind turbine cluster, linearly quantifying the functional similarity relationship between wind turbines using the Pearson correlation coefficient to obtain the functional relationship matrix of the wind turbine cluster. The element R in the functional relationship matrix ij represents the spatial similarity between any two wind turbines i and j: where P i and P j represent the historical power vectors of the i-th and j-th wind turbines respectively, is the power sequence after mean removal, is the mean of the i-th wind turbine power sequence, and ||·|| 2 represents the Euclidean norm.
[0019] As an optimization, in the process of approximating and optimizing the graph convolutional network using Chebyshev polynomials, the first kind of Chebyshev polynomial T is introduced k(x) = cos(karccosx), where |x| ≤ 1, replaces the convolution kernel of the spectral graph convolution: T k (·) is the Chebyshev polynomial of order k; β k is the parameter updated by training iteration; is obtained by performing a transfer transformation on Λ to get a rescaled eigenvalue diagonal matrix, resulting in a graph convolutional network: Then, the matrix operation is put into the Chebyshev polynomial to obtain: After performing Laplacian normalization, the first-kind Chebyshev polynomial approximation of the graph convolution is finally expressed as: where represents the symmetric normalized Laplacian matrix After re-adjusting the normalization, λ max is the largest eigenvalue of and y sd represents the spatial feature based on the spatial adjacency graph SAG, and y sc represents the spatial feature based on the functional relationship graph FRG, and both are flattened by time step and node into The output y s ′ d of the SAG module provides the query Q s = y s ′ d W h Q The output y s ′ c of the FRG module provides the key K s = y s ′ c W h K and the value V s = y s ′ c W h V to obtain the linear transformation output Y that dynamically fuses the features of the two Spa = Output SpatialBranch = Concat(head 1 , head 2 ,..., head h )W t O .
[0020] As an optimization, the multi-resolution convolutional kernel-gating module utilizes convolutional kernels of different resolutions to simulate dependencies at different time scales in the time series, combines a gating mechanism to control the information flow through the network layer, and then weights and combines the convolutional results of different resolutions through the gating mechanism to obtain multi-scale time modality features;
[0021] The recurrent time module is used to retain the context information of long time steps and output the hidden state y of the entire sequence tg =[h 1 ,h 1 ,...,h t , as a supplement to extracting long-term dependencies, h t is the current hidden state at time t, and T is the length of the input sequence;
[0022] The time features captured by the multi-resolution convolutional kernel-gating module and the recurrent time module are flattened by time step and node dimension into The multi-resolution convolutional kernel-gating module provides the query Q t =y t ′ c W h Q The recurrent time module provides the key K t =y t ′ g W h K and the value V t =y t ′ g W h V , and the multi-head cross-attention mechanism is used to integrate multi-time modality information to obtain a linearly transformed output of the fused features: Y Temp =Output TemporalBranch =Concat(head 1 ,head 2 ,...,head h )W t O .
[0023] As an optimization, the spatio-temporal fusion module realizes multi-granularity interaction between spatio-temporal features by using the fused multi-graph fusion features from the spatial branch module and the time series features from the time branch module :
[0024] A computer program product includes a computer program, and when the computer program is executed by a computer, the method described in any one of the above is implemented.
[0025] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a computer, the method described in any one of the above is implemented.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] (1) The spatial branch module and the temporal branch module work in parallel, independently extracting spatial and temporal features respectively, and performing preliminary feature fusion through the attention mechanism. After capturing the dependency relationship between the two sets of different-level features, the model processes the features from the spatial branch module and the temporal branch module in parallel to achieve multi-granularity interaction and deep feature fusion;
[0028] (2) The temporal branch module focuses on the transformation of data over time, uses multi-resolution convolutional kernels to adapt to time-series data of different frequencies, and dynamically controls the flow of information through a gating mechanism. The multi-resolution convolutional kernel-gating module and the recurrent temporal module of the framework structure of the temporal branch module are executed in parallel, aiming to extract short-term and long-term temporal dependencies at multiple levels, thereby enhancing the feature selection ability and receptive field of the model, and further improving the model's processing ability for sequence data;
[0029] (3) The spatial branch module is responsible for propagating information between nodes, introduces Chebyshev polynomial approximation to optimize the graph convolution operation, and captures the physical relationship in space and the functional similarity by constructing a multi-view graph structure, so as to simultaneously capture the physical relationship in space and the functional similarity to obtain the optimal lateral interdependence relationship between different variables, thereby enhancing the model's modeling ability for complex spatial dependency relationships and the computational efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a convolution operation diagram used by the resolution convolutional kernel-gating unit in the present invention;
[0031] Figure 2 It is an overall framework structure diagram of the spatio-temporal parallel model framework DBANN in the present invention;
[0032] Figure 3 It is a comparison diagram of the offshore wind power prediction results between DBANN and the baseline model in the present invention (node 1);
[0033] Figure 4 It is a comparison diagram of the offshore wind power prediction results between DBANN and the baseline model in the present invention (node 2);
[0034] Figure 5 It is a comparison diagram of evaluation indexes of various configurations of the multi-resolution temporal convolutional kernel in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0036] The method for predicting the wind power of a multi-node wind turbine cluster in this specific embodiment includes:
[0037] Obtain the geographical coordinate data and historical power data of the wind turbine cluster, and perform normalization processing on the data;
[0038] Construct a spatio-temporal parallel model framework, which includes a spatial branch module, a temporal branch module, and a spatio-temporal fusion module;
[0039] The spatial branch module uses the Euclidean distance to measure the spatial proximity between wind turbines based on the geographical coordinate data to obtain a spatial adjacency graph SAG. Based on the historical power data, the Pearson correlation coefficient is used to linearly quantify the functional similarity relationship between wind turbines to obtain a functional relationship graph FRG. The Chebyshev polynomial approximation is used to optimize the graph convolutional network, and the multi-head cross-attention mechanism is combined to aggregate the cross-graph feature correlation between the spatial adjacency graph SAG and the functional relationship graph FRG, and output multi-graph fusion features;
[0040] The temporal branch module uses a parallel multi-resolution convolutional kernel-gated module and a recursive temporal module to extract the short-term and long-term temporal dependencies of the wind power sequence, and integrates them through the multi-head cross-attention mechanism to output temporal sequence features;
[0041] The spatial branch module and the temporal branch module work in parallel. The spatio-temporal fusion module uses the multi-head cross-attention mechanism to fuse the multi-graph fusion features and the temporal sequence features, realizes multi-granularity interaction between spatio-temporal features, and outputs the prediction result of the wind power.
[0042] In this specific embodiment, the construction of the spatial adjacency graph SAG includes: based on the geographical coordinate data of the wind turbine cluster, obtain the geographical coordinate matrix of the wind turbine cluster, and calculate the Euclidean distance D between any pair of wind turbines in the geographical coordinate matrix ij = ||C i - C j || 2 , C iand C j respectively represent the coordinate vectors of the i-th and j-th wind turbines, and ||·|| 2 represents quantifying the geometric distance between two wind turbines through the definition of the second norm.
[0043] In this specific embodiment, the construction of the function relationship graph FRG includes: based on the historical power data of the wind turbine cluster, using the Pearson correlation coefficient to linearly quantify the function-based similarity relationship between wind turbines, obtaining the function relationship matrix of the wind turbine cluster. In the function relationship matrix, the element R i represents the spatial similarity between any two wind turbines i and j: where P i and P j respectively represent the historical power vectors of the i-th and j-th wind turbines, is the power sequence after removing the mean value, is the mean value of the power sequence of the i-th wind turbine, and ||·|| 2 represents the Euclidean norm.
[0044] In this specific embodiment, in the process of optimizing the graph convolutional network by the Chebyshev polynomial approximation, the first kind of Chebyshev polynomial T k (x) = cos(karccosx), |x| ≤ 1 is introduced to replace the convolution kernel of the spectral graph convolution: T k (·) is the k-th order Chebyshev polynomial; β k is the parameter updated by training iteration; is to perform a transfer transformation on Λ to obtain a rescaled eigenvalue diagonal matrix, and obtain the graph convolutional network: Then, the matrix operation is put into the Chebyshev polynomial to obtain: After performing Laplacian normalization, the first kind of Chebyshev polynomial approximation graph convolution is finally expressed as: where represents the symmetric normalized Laplacian matrix After re-adjusting the normalization, λ max is 's largest eigenvalue, K is the number of consecutive screening operations in the model; by aggregating and transforming the k-hop neighborhood historical information of the nodes to smooth the node signals, we get: and y sd represents the spatial feature based on the spatial adjacency graph SAG, y sc represents the spatial feature based on the function relationship graph FRG, and both are flattened by the time step and the node dimension into The output y s ' d of the SAG module provides the query Qs = y s ′ d W h Q , the output y of the FRG module s ′ c provides the key K s = y s ′ c W h K and the sum value V s = y s ′ c W h V , obtaining the linear transformation output Y that dynamically fuses the features of the two Spa = Output SpatialBranch = Concat(head 1 , head 2 ,..., head h )W t O .
[0045] In this specific embodiment, the multi-resolution convolution kernel-gating module simulates the dependencies of different time scales in the time series using convolution kernels of different resolutions, combines the gating mechanism to control the information flow through the network layer, and weights and combines the convolution results of different resolutions through the gating mechanism to obtain multi-scale time modality features;
[0046] The recursive time module is used to retain the context information of long time steps and outputs the hidden state y of the entire sequence tg = [h 1 , h 1 ,..., h t , as a supplement to extracting long-term dependencies, h t is the current hidden state at time t, and T is the length of the input sequence;
[0047] The time features captured by the multi-resolution convolution kernel-gating module and the recursive time module are flattened by time step and node dimension into The multi-resolution convolution kernel-gating module provides the query Q t = y t ′ c W h Q , the recursive time module provides the key K t = y t ′ g W h K and the sum value V t = y t ′ g Wh V , the multi - head cross - attention mechanism is used to integrate multi - temporal modal information to obtain the linear transformation output of the fused features: Y Temp = Output TemporalBranch = Concat(head 1 , head 2 ,..., head h )W t O .
[0048] In this specific embodiment, the spatio - temporal fusion module realizes the multi - granularity interaction between spatio - temporal features for the fused multi - map fusion features from the spatial branch module and the time - series features from the time - branch module :
[0049] A computer program product includes a computer program, and when the computer program is executed by a computer, it implements the method described in any one of the above.
[0050] A computer - readable storage medium stores a computer program, and when the computer program is executed by a computer, it implements the method described in any one of the above.
[0051] Affected by specific terrain environments and atmospheric movements, the wind speed within the propagation path has spatial correlation. Theoretically, for wind turbines located within the same wind zone, due to similar wind - field environmental factors such as terrain and meteorological factors (such as wind speed, wind direction, air pressure, etc.), and considering the wake effect, there is a certain degree of interaction between wind turbines with close geographical locations, and their output powers often show similar fluctuations. That is to say, the wind power outputs between wind turbines within a specific geographical area are not independent but spatially correlated. Therefore, it is necessary to establish the data correlation between different wind turbines. Moreover, considering that offshore wind power data has high volatility and its performance is more complex. For this reason, we have established two spatial - association graph structures, one is the spatial adjacency graph SAG, and the other is the functional - relationship graph FRG.
[0052] The former uses the Euclidean distance to measure the spatial proximity between wind turbines to obtain multiple elements d ij to construct the spatial - adjacency graph matrix The wind - farm geographical - coordinate matrix contains n wind - turbine coordinate vectors c i =(long i , lat i ), 1 ≤ i ≤ n. The elements in matrix D represent the Euclidean distance between any pair of wind turbines: D ij = ||C i - C j ||2 , where ||·|| 2 represents quantifying the geometric distance between two nodes through the definition of the second norm. The diagonal elements D of the matrix ii = 0 because the physical distance of each wind turbine to itself is absolutely 0.
[0053] The latter linearly quantifies the function-based similarity relationship among n wind turbines using the Pearson correlation coefficient to obtain the function relationship graph matrix Historical power contains the historical power vectors of n wind turbines In the function relationship graph matrix R, the element R ij represents the spatial similarity between any two wind turbines i and j:
[0054]
[0055] where is the power sequence after mean removal,
[0056] is the mean of the power sequence of the i-th turbine, and ||·‖ 2 represents the Euclidean norm.
[0057] If R ij = 1, it represents a perfect positive correlation; R ij = -1, it represents a perfect negative correlation; R ij = 0 indicates no linear correlation. The function relationship graph matrix R is a symmetric matrix, and the diagonal element R ii = 1 because the correlation of each wind turbine with itself is 1.
[0058] The spatial adjacency graph matrix most intuitively provides the constraints in physical space, and the function relationship graph matrix reveals the dynamic association of wind power sequences, enabling the graph structure to take into account both geographical topological features and wind farm dynamics features to more accurately and comprehensively capture the spatial features between wind turbines.
[0059] To make full use of the spatial topological features between wind turbines and alleviate the high cost caused by directly performing eigenvalue decomposition on the Laplacian matrix, the first kind of Chebyshev polynomial T k (x) = cos(karccosx), |x| ≤ 1 is introduced to replace the convolution kernel of spectral graph convolution: where T k (·) is the Chebyshev polynomial of order k; β k is the corresponding coefficient (i.e., the parameter updated by training iteration); is the re-scaled eigenvalue diagonal matrix obtained by performing a shift transformation on Λ.
[0060] Accordingly, the graph convolution formula is Then put the matrix operation into the Chebyshev polynomial To match the domain of the Chebyshev polynomial, perform Laplacian normalization again. Finally, the first-kind Chebyshev polynomial approximation of graph convolution is expressed as: Where represents the symmetric normalized Laplacian matrix After re-adjusting the normalization, λ max is the largest eigenvalue of. By standardization, the structural characteristics of the graph are retained, and high-order graph convolution in polynomial form is realized. Applying a graph convolution stack with a first-order approximation in the vertical direction can obtain an effect similar to that of K-localized convolution in the horizontal direction, and all convolutions aggregate the information of the K-1 order neighborhood from the central node. K is the number of consecutive screening operations or convolutional layers in the model, which is set to 3 in this embodiment. In addition, due to the order of approximation being limited to 1, it is efficient for large-scale graphs.
[0061] By aggregating and transforming the k-hop neighborhood historical information of nodes to smooth the node signals, the advantage is that its filter can be in the spatial domain and supports multi-dimensional input. Thus, we get: Where y sd represents the spatial feature based on SAG, y sc represents the spatial feature based on FRG, and both For adapting to the multi-head cross-attention mechanism, they are flattened by time step and node dimension into The output y s ′ d of the SAG mode provides the query Q s = y s ′ d W h Q , the output y s ′ c of the FRG mode provides the key K s = y s ′ c W h K and the value V s = y s ′ c W h V , and obtain the linear transformation output Y that dynamically fuses the features of both Spa = Output SpatialBranch = Concat(head 1 , head 2 ,..., head h )W tO .
[0062] The graph convolution module combines the multi-head cross-attention mechanism to aggregate multi-graph information and extract multi-hop historical information of wind farms. The Chebyshev convolution avoids the high computational cost of directly calculating the Laplace eigendecomposition, while making the convolution kernel have spatial localization, thereby avoiding the global aliasing effect. Its polynomial order K can control the receptive field range from the first-order neighborhood information (1-hop neighborhood) to the k-hop field to update the node features, so as to effectively model the multi-scale field features of the nodes in the graph. The query, key and value matrices of the attention mechanism are cross-defined by the output of different graph structures, thereby capturing the cross-graph feature association between SAG and FRG. The multi-head mechanism uses multiple heads to process the feature relationships of different subspaces in parallel, which improves the computational efficiency of the model. Moreover, the k-layer Chebyshev convolution can be seamlessly combined with time series modeling to form a synergistic mechanism between the graph convolution layer and the time series layer. Finally, the spatial branch module outputs multi-graph fusion features suitable for downstream spatiotemporal sequence prediction tasks.
[0063] In the time dimension, the output power of wind turbines often exhibits different dependencies at different time scales. It is necessary to capture both the long-term (reflecting trend information) and short-term (reflecting local features) time dependencies of wind power sequences to improve prediction accuracy. Although the RNN family can effectively extract long-term time dependencies, the extraction of local time dependencies is still insufficient. To this end, we propose a parallel time branch, including a multi-resolution convolution kernel-gating module to more comprehensively and accurately characterize local time patterns, and a recursive time module to capture the contextual information of long-term time steps in parallel to supplement the long-term time patterns. Figure 1 shown.
[0064] Temporal convolution often uses 1-D small convolution kernels to capture local time dependencies, but the intrinsic natural frequency of actual wind power data may have different patterns and noisy data. Therefore, considering the conventional characteristics of the time dimension, two dilated starting convolutions and dilated causal convolutions are used in the temporal convolution module, and in the calculation process of the dilated causal convolution, the expansion rate increases with the depth of the network. In this embodiment, kernels with multiple resolutions 1*2, 1*3, 1*4 are designed to enable the network to calculate in parallel to discover the time dependencies of different ranges of input variables.
[0065] First, input data x t The convolution operation generates the following three receptive fields:
[0066]
[0067] where {K 1 ,K 2,K 3} represents three groups of kernel sizes, is the convolutional kernel corresponding to the i-th scale, i ∈ {1, 2, 3}, and * represents the convolution operation.
[0068] Subsequently, the three outputs are controlled to be of the same length and concatenated in the channel dimension to form the output of the DIC module, that is, the new feature map containing all scale information At the same time, a gating mechanism is combined to control the information flow through the network layer, and the convolution results of different scales are weighted and combined through the gating mechanism to obtain multi-scale time-modal features ⊙ represents element-wise multiplication (Hadamard product).
[0069] The multi-resolution convolutional kernels {k 1 ,k 2 ,k 3} simulate the dependencies of different time scales in the time series. At the same time, the activation function and the gating mechanism provide flexible control of the dynamically changing signals, enabling the model to adapt to diverse and multi-resolution local time series patterns.
[0070] GRU is a lightweight recurrent neural network that can retain the context information of long time steps, is good at modeling complex temporal dependencies, and outputs the hidden state y of the entire sequence tg = [h 1 ,h 1 ,...,h t ], as a supplement to extracting long-term dependencies, h t is the current hidden state at time t, T is the length of the input sequence, and this enhanced representation improves the robustness of the model to temporal patterns. Similarly, the temporal features captured by the two temporal Blocks are also flattened by time steps and nodes into The multi-resolution convolutional kernel - gating module provides the query Q t = y′ tc W h Q The recursive time module provides the key K t = y′ tg W h K and the value V t = y′ tg W h V Using multi-head cross-attention to further integrate multi-time-modal information to obtain the linear transformation output Y of the deeply fused features Temp = Output TemporalBranch = Concat(head 1 ,head 2 ,...,headh )W t O 。
[0071] For the high volatility and multi - variation characteristics of the time dimension of offshore wind power data, this subsection designs a time branch to achieve deep fusion of multi - modal time features, including a multi - resolution convolution kernel - gated module, a recursive time module, and an attention mechanism, which can effectively capture multi - scale temporal dynamics and enhance the interaction of spatio - temporal features, and finally output a hierarchical multi - perspective time dependence suitable for downstream spatio - temporal sequence prediction tasks.
[0072] In this embodiment, a scheme of using a multi - head cross - attention mechanism (concatenated according to the node dimension n, keeping the feature dimension f unchanged) is designed to fuse two sets of features from the parallel time branch and space branch, namely the time - series features (from the time branch) and the spatial features (from the space branch), so as to achieve multi - granularity interaction between spatio - temporal features: As Figure 2 shown.
[0073] To successfully output the deeply fused spatio - temporal feature Y, in the output layer, it is first converted into a one - dimensional vector through a flattening operation , and then input into the first fully - connected layer to map it to the hidden space, and further mapped through the second fully - connected layer: where each dense layer has its weight matrix W and bias vector b.
[0074] Then, a reshaping operation is used to obtain the spatio - temporal feature Y with the feature dimension f = 1:
[0075] Finally, a formal definition of the end - to - end model constructed in this embodiment from the input X to the output Y is given: Y = P(X; Θ), where P(·) is the model function including multiple deep - learning operations such as convolution, concatenation, flattening, and reshaping, and Θ is the set of all trainable parameters.
[0076] Considering the complexity of offshore wind power, we select a certain domestic wind farm as the case study object, and the cluster of wind turbines from No. 1 to No. 40 are fully labeled.
[0077] The collection time of the wind power dataset is from 08:00:00 on January 7, 2022 to 07:45:00 on August 31, 2022, with a time resolution of about 15 minutes, and a total of 906,240 records. The dataset is further divided into three parts according to the ratio of 80:20: the first 724,992 samples are used as the training set (during the model training process, 20% of the samples in the training set are used as the validation set), and the last 181,248 samples are used as the test set. The wind farm attribute information is shown in Table 1.
[0078] Attribute Total number of days Total electric power Average electric power Maximum electric power Minimum electric power Standard deviation of electric power Value 237(d) 73664930 (kW) 81.2863 (kW) 493.1800 (kW) 0.0000 (kW) 83.4109 (kW)
[0079] Table 1
[0080] During the training phase, the training loss and validation loss curves of wind power prediction based on the spatio-temporal parallel model framework (DBANN) decrease and tend to level off as the number of training epochs increases (up to 30 epochs). The model converges well and there is no obvious overfitting. In addition, to ensure the fairness of testing and the stability of the validation model, the average value of the evaluation metrics from 10 repeated tests is used as the final evaluation metric. All floating-point numbers in the test experiment results of this embodiment are kept to 4 decimal places.
[0081] Regression evaluation metrics can effectively evaluate the error between the prediction results and the actual values, and can intuitively show the prediction accuracy of the model. Considering the data characteristics of wind power, root mean square error (MSE), mean absolute error (MAE), mean absolute percentage error (MAPE), coefficient of determination R 2 and time consumption are used as evaluation metrics.
[0082] To prove the effectiveness of the proposed model DBANN, it is compared with traditional single models. DBANN is compared with other prediction methods, including classical traditional models and advanced spatio-temporal models. A brief introduction to these models is as follows:
[0083] ARIMA: Autoregressive Integrated Moving Average model. It analyzes and detects the long-term memory characteristics of wind power sequences by means of the autocorrelation function (ACF) and partial autocorrelation function (PACF), and then uses the model to predict the linear components of wind power sequences.
[0084] GRU: Gated Recurrent Unit, a variant of the Recurrent Neural Network (RNN). The two gate structures of GRU effectively control the flow of information: the Reset Gate selectively resets historical information, and the Update Gate controls the update and retention of information.
[0085] CNN-LSTM: A hybrid architecture that combines the feature extraction ability of Convolutional Neural Network (CNN) and the advantage of Long Short-Term Memory Network (LSTM) in processing sequence data, and is used to process wind power sequence data with spatio-temporal characteristics. It first uses CNN to extract spatial features from regularly gridded wind farm data, and then uses LSTM to capture the changing trends in time series.
[0086] STGCN: Spatio-Temporal Graph Convolutional Network, which uses Graph Convolution (GCN) and Temporal Convolution (TCN) to extract the spatial and temporal features of data simultaneously. GCN mines the correlation features of data in the spatial dimension, and TCN captures the features of data changing over time, thereby obtaining spatio-temporal information.
[0087] Graph Wavenet: A graph neural network architecture that can capture the inherent hidden spatial dependencies of data and is then applied to spatio-temporal graph modeling to mine the deep relationships of data in the spatio-temporal dimension.
[0088] ASTGCN: An attention-based spatio-temporal graph convolutional network that combines spatio-temporal graph convolution and spatio-temporal attention mechanism to synchronously capture the complex spatio-temporal features in data.
[0089] To verify the effectiveness of DBANN, systematic experiments were designed from multiple perspectives, including key hyperparameter analysis experiments, comparison experiments with different baseline methods, ablation experiments of the model, and serial and parallel experiments of spatio-temporal branches. To ensure the reliability and robustness of the experimental results, the results of each experiment were obtained from the statistical analysis of the average values of the evaluation indicators for the prediction of 3 actual nodes after the model ran independently 10 times. The comparison of the predicted values of DBANN, the true values of wind power, and the predicted values of some spatio-temporal models in the present invention is as Figure 3 (Node 1) and Figure 4 (Node 2) are shown.
[0090] Based on the prediction results, this embodiment systematically analyzes and optimizes the key hyperparameters of the proposed model DBANN through a trial-and-error method to ensure that the model can fully exert its performance. Tables 2 and 3 show the evaluation indicators of the proposed model DBANN under different hyperparameter settings.
[0091]
[0092] Table 2
[0093]
[0094] Table 3
[0095] The specific analysis is as follows:
[0096] The influence of the number of attention heads. When the number of attention heads increases from 2 to 8, the model performance is gradually optimized, and RMSE, MAE, and MAPE all reach the lowest values (10.9067, 6.6767, and 13.5900 respectively), and R 2 reaches the highest value (0.9865). When the number of attention heads further increases to 16 or 32, the performance decreases and the time consumption increases significantly. Therefore, when the number of attention heads is 8, the model reaches the best balance between performance and efficiency.
[0097] Effect of batch size. When the batch size for training is 128, the RMSE, MAE, and MAPE are 10.9067, 6.6767, and 13.59% respectively, and R 2 reaches 0.9865, and the time consumption is only 13 seconds. When the batch size increases to 256, the performance slightly decreases. The ideal choice for the batch size is 128.
[0098] Effect of number of training epochs. When the number of training epochs is 30, the performance is the best (RMSE is 10.9067, MAE is 6.6767), while when the number of training epochs is greater than 60, the performance tends to be stable or even slightly decreases, indicating that a higher number of epochs may lead to overfitting.
[0099] Effect of time window step size. When the time window step size is 4, the RMSE, MAE, and MAPE are the lowest (10.9067, 6.6767, and 13.5900), and R 2 is the highest (0.9865), and the time consumption is moderate (13 seconds). When the step size is greater than 6, the performance decreases and the time increases significantly. An overly long time window may introduce unnecessary noise information and affect the model performance.
[0100] Effect of learning rate. Both too small or too large learning rates (such as 0.0001 or 0.005) result in a decrease in performance. When the learning rate is 0.001, the model performance is the best (RMSE is 10.9067, MAE is 6.6767, MAPE is 13.59%, and R 2 is 0.9865), and the time is only 13 seconds.
[0101] To verify the accuracy of the proposed model (DBANN) of the present invention, a comparative experiment was conducted by comparing it with multiple existing prediction models such as ARIMA, SVR, GRU, TCN, CNN-LSTM, STGCN, Graph WaveNet, and ASTGCN on the same dataset and parameter settings. The offshore wind power prediction results of different advanced spatio-temporal prediction models are shown in Table 4:
[0102] The present invention ARIMA SVR GRU TCN CNN-LSTM STGCN Graph WaveNet ASTGCN RMSE 10.9067 11.4133 13.7108 12.3500 13.5600 11.4600 11.2867 17.0533 21.2467 MAE 6.6767 6.8600 10.8102 7.5667 8.4700 7.1067 6.9000 10.8367 11.1833 MAPE 13.5900 14.9933 16.8732 17.0200 17.0467 14.9467 13.8567 23.6700 28.8200 <![CDATA[R 2 > 0.9865 0.9812 0.9773 0.9682 0.9704 0.9850 0.9753 0.9143 0.9191
[0103] Table 4
[0104] The necessity of considering spatio-temporal features that simultaneously incorporate complex spatial domain information and long-term and short-term temporal domain information. In this embodiment, multiple graphical structures are constructed in the spatial dimension to consider multi-hop spatial domain associations at multiple levels. At the same time, long-term and short-term dependencies are comprehensively modeled in the temporal dimension. This enables the model to not only focus on spatio-temporal associations between neighboring nodes but also capture the potential influence of more distant nodes on the target node. Compared with traditional spatio-temporal models such as STGCN, CNN-LSTM, Graph WaveNet, and ASTGCN, the RMSE of DBANN is reduced from 21.2467 to 10.9067, the MAE is reduced from 11.1833 to 6.6767, the MAPE is reduced from 28.82% to 13.59%, and the R 2 value is increased from 0.9143 to 0.9865. These improvements are mainly attributed to the model's comprehensive consideration of complex spatio-temporal information, thus significantly enhancing the effectiveness of wind power spatio-temporal feature extraction.
[0105] The proposed model DBANN in the present invention effectively avoids the problems of information loss and insufficient model expressiveness that may be caused by separately extracting temporal or spatial features. For example, ARIMA and SVR, due to only focusing on time series modeling and ignoring the spatial associations between nodes, result in MAPEs of 14.9933% and 16.8732% respectively, while the MAPE of the proposed model is significantly reduced to 13.59%, indicating that synchronous modeling of spatio-temporal features can effectively improve prediction accuracy.
[0106] In this embodiment, spatio-temporal dependency matrices are constructed from multiple perspectives, and at the same time, a feature fusion mechanism is combined to fully integrate the potential correlations in the temporal and spatial dimensions. Compared with Graph WaveNet and ASTGCN, the MAPE of DBANN is reduced from 23.67% and 28.82% to 13.59% respectively, and the R 2 is increased from 0.9143 and 0.9191 to 0.9865 respectively, indicating that effective fusion of spatio-temporal features can effectively improve the robustness of prediction.
[0107] Therefore, overall, DBANN has the best prediction effect on offshore wind power, and it has significant advantages in both prediction accuracy and fitting ability, specifically manifested as the comprehensive superiority of DBANN in multiple performance indicators such as RMSE (10.9067), MAE (6.6767), MAPE (13.59%), and R 2 (0.9865).
[0108] To verify the impact of key modules in the proposed model on the wind power prediction performance, ablation experiments were conducted. The experiments were carried out under the same dataset and parameter settings, and the contributions of the following modules to the model performance were mainly analyzed: multi-resolution temporal convolutional kernels, SAG, FRG, and the attention mechanism in the spatio-temporal fusion module.
[0109] To verify the capture performance of different features by the multi-scale and multi-resolution convolution kernel configuration ({2, 3, 4}) in the temporal fusion modeling, 10 groups of resolution convolution kernel configurations are designed and classified into three configuration schemes according to the types of convolution kernel sizes in the configuration: single resolution (the same convolution kernel size), two resolutions (combining different convolution kernel sizes by the exhaustive method), and three resolutions (the multi-scale convolution kernel configuration adopted by the DBANN of the present invention).
[0110] The prediction results of the offshore wind power with different resolution convolution kernel configurations are shown in Table 5:
[0111]
[0112] Table 5
[0113] The single resolution setting scheme includes three configurations of {2, 2, 2}, {3, 3, 3}, and {4, 4, 4}, where the size of the convolution kernel remains the same in all layers. Judging from the results, the {2, 2, 2} configuration performs relatively better, with the average RMSE and MAE being 11.7733 and 7.2833 respectively, which is better than the {4, 4, 4} and {3, 3, 3} configurations. However, since the single resolution can only capture single-time scale features and has limited modeling ability for complex time series, the overall prediction effect is limited.
[0114] The two-resolution setting covers six configurations of {2, 2, 3}, {2, 2, 4}, {3, 3, 2}, {3, 3, 4}, {4, 4, 2}, and {4, 4, 3}. Through comparison, it is found that mixing convolution kernel configurations with different resolutions can significantly improve the model's ability to capture multi-scale features of time series. The average RMSE decreases from 11.7733 (configuration {2, 2, 2}) of the single resolution to 11.0233, and the MAE and MAPE also decrease from 7.2833 and 14.7567% to 6.7733 and 13.9233% respectively. This shows that introducing a combination of convolution kernels with different time scales can effectively improve the overall prediction accuracy of the model. However, due to only including two resolutions, its feature fusion ability still has certain limitations.
[0115] The three-resolution setting adopts the {2, 3, 4} configuration (the scheme used by the DBANN), and more balanced optimization is achieved in the short-term and long-term time dependence modeling. The experimental results show that this configuration is better than other resolution combinations in all indicators. The average RMSE is further reduced to 10.9067, the MAE and MAPE are 6.6767 and 13.59% respectively, R 2The value is as high as 0.9865, which is an improvement compared to 0.98 for a single resolution and 0.9839 for two resolutions. In terms of node performance, the RMSE and MAE of the three-resolution modes are only 3.65 and 2.38 respectively, indicating that while capturing local features, the three-resolution configurations can better integrate global information at different time scales.
[0116] Through experimental analysis of the configuration scheme of multi-resolution convolutional kernels, it can be seen that the combination of three resolutions (configuration {2, 3, 4}) shows the best comprehensive performance in the ablation experiment, indicating that it can more efficiently capture diverse time-dependent features when modeling complex time series data. As Figure 5 shown.
[0117] To verify the influence of the graph structure in the spatial branch on the wind power prediction performance, analyze the contributions of two spatial graph structures, SAG and FRG, to the model performance.
[0118] The wind power prediction results of offshore wind farms applying different spatial graph structures are shown in Table 6:
[0119]
[0120] Table 6
[0121] The graph structure modeled in this embodiment combines the geographical proximity (SAG) and spatial functional correlation (FRG) of wind farms, and can capture the complex spatial correlations between different nodes. Under this configuration, the average RMSE, MAE, MAPE, and R 2 of the model are 10.9067, 6.6767, 13.59%, and 0.9865 respectively, showing the best performance.
[0122] Remove the SAG module and only retain the FRG based on spatial functional correlation. The experiment shows that the average RMSE and MAE increase to 11.1833 and 6.8867, the MAPE reaches 13.6633%, and R 2 drops to 0.9786. Although FRG can more accurately describe the functional relationships between wind farms, its performance is also limited to a certain extent when used alone without the support of geographical information.
[0123] Remove the FRG module and only retain the SAG based on geographical proximity. Under this configuration, the average RMSE and MAE of the model increase to 11.58 and 7.0933 respectively, the MAPE increases to 14.4467%, and R 2 drops to 0.9704. The results show that although SAG can capture the geographical adjacency between wind farms, the lack of the ability to model spatial functional relationships leads to a decline in performance.
[0124] The experimental results show that the two graph structures (SAG and FRG) each have their own advantages in capturing the spatial correlation of wind farms. SAG provides basic spatial correlation information through geographical proximity, while FRG further enhances the accuracy of spatial modeling through functional correlation. The model DBANN proposed in this embodiment combines the advantages of both and performs best in terms of prediction performance, which indirectly proves the importance and effectiveness of the multi-level graph structure design.
[0125] To verify the impact of the feature fusion mechanism in the spatio-temporal fusion module on the wind power prediction performance, ablation experiments with three different configurations were conducted, namely: the original model, the model without the cross-attention mechanism (W / O Cross-Attention), and the model without the attention mechanism (W / O Attention).
[0126] The wind power prediction results of offshore wind farms applying different spatio-temporal fusion mechanisms are shown in Table 7:
[0127]
[0128] Table 7
[0129] The original model includes a complete spatio-temporal fusion module, which uses the multi-head cross-attention mechanism to integrate two sets of features from parallel branches to achieve multi-granularity interaction between spatio-temporal features. Under this configuration, the average RMSE, MAE, MAPE, and R of the model 2 are 10.9067, 6.6767, 13.59%, and 0.9865 respectively, showing the best performance.
[0130] In the spatio-temporal fusion module that only retains the multi-head attention mechanism, the average RMSE and MAE of the model increase to 12.5250 and 7.8500 respectively, the MAPE increases to 16.2267%, and R 2 drops to 0.9740, and the performance deteriorates.
[0131] In the configuration where the attention mechanism is removed and spatio-temporal features are integrated only through simple feature concatenation, all error metrics increase significantly (the average RMSE and MAE increase to 14.3667 and 9.2700, and the MAPE increases to 18.7433%), and the goodness of fit of the model decreases significantly (R 2 drops to 0.9578).
[0132] The experimental results show that the proposed spatio-temporal fusion module can comprehensively and effectively fuse time and space features, showing the best performance in the wind power prediction task, further verifying the importance and effectiveness of the cross multi-head attention mechanism in this module.
[0133] The key structure of the proposed model DBANN consists of two branches, and each branch has two sub-branches. String-parallel experiments can verify the impact of the structure settings in the model on the model performance. For this purpose, 7 groups of experimental schemes are designed and divided into three structures according to the connection methods and interaction modes of the sub-branches in the two branches: parallel branch structure, cross-branch structure, and cross-sub-branch structure.
[0134] In the parallel branch structure (i.e., Configuration 1), the time branch and the space branch operate independently in parallel. Branch 1 includes sub-branch LRB and sub-branch MGB, and Branch 2 includes sub-branch SR and sub-branch SA.
[0135] There are cross-connections between the time branch and the space branch in the cross-branch structure, including two configurations. In Configuration 2, Branch 1 includes sub-branch LRB and sub-branch SA, and Branch 2 includes sub-branch MGB and sub-branch SR; in Configuration 3, Branch 1 includes sub-branch LRB and sub-branch SR, and Branch 2 includes sub-branch MGB and sub-branch SA.
[0136] In the cross-sub-branch structure, there are not only cross-connections between the time branch and the space branch, but also each branch has only 1 sub-branch, including four configurations. In Configuration 4, Branch 1 has sub-branch LRB and Branch 2 has sub-branch SA; in Configuration 5, Branch 1 has sub-branch LRB and Branch 2 has sub-branch SR; in Configuration 6, Branch 1 has sub-branch MGB and Branch 2 has sub-branch SA; in Configuration 7, Branch 1 has sub-branch MGB and Branch 2 has sub-branch SR.
[0137] Through the analysis of indicators such as RMSE, MAE, MAPE, and R 2 It can be seen that the parallel branch structure performs optimally in all indicators. Its RMSE value is 10.9067, MAE value is 6.6767, MAPE value is 13.59%, and R 2 value is 0.9865, indicating that the difference between the model prediction value and the true value is small, the prediction error is small, the prediction accuracy is high, and the data fitting degree is high under this structure, and the model performance is the best. In the cross-sub-branch structure, the R 2 value of Configuration 5 is relatively high but still lower than that of the parallel branch structure, and Configurations 4, 6, and 7 perform poorly in terms of RMSE, MAE, MAPE, and R 2 indicators. The values of each indicator of each configuration in the cross-branch structure are greater than those of the parallel branch structure, and its performance is inferior to that of the parallel branch structure in all aspects.
[0138] The prediction results of the offshore wind power using different branch structure configurations are shown in Table 8:
[0139]
[0140] Table 8
[0141] In summary, the parallel branch structure adopted in this embodiment performs excellently. When processing wind power spatiotemporal prediction tasks, this structure can accurately capture the complex spatiotemporal characteristics in the data, thereby achieving a more accurate prediction of wind power.
[0142] In this embodiment, DBANN fully considers the complex spatiotemporal characteristics of offshore wind power data and cleverly balances the correlation between the time domain and spatial dimensions. It designs multiple key components based on multi-level spatiotemporal perception, including a time branch based on hierarchical multi-view time dependency, a spatial branch based on spatial dependency containing multi-image information, and a spatiotemporal fusion module based on a multi-head cross-attention machine. These designs enable the DBANN model to perform well in terms of prediction accuracy and effectiveness.
[0143] In the case study experiment based on a real relevant offshore wind power data set, this embodiment carried out comparative experiments, ablation experiments, and serial-parallel experiments with eight other baseline methods, fully verifying the effectiveness and advancement of the model proposed in this embodiment. The comparative experimental results show that DBANN performs well in offshore wind power prediction, and all comprehensive evaluation indicators are better than other baseline methods. The ablation experiment results further reveal the important contribution of key modules and components in DBANN to the improvement of model performance. The serial-parallel experimental results show that the connection method and interaction mode of each component in the DBANN structure are reasonably designed, which effectively improves the prediction performance of the model.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit the technical solution. Those skilled in the art should understand that those modifications or equivalent substitutions of the technical solution of the present invention that do not depart from the purpose and scope of the technical solution should be included in the scope of the claims of the present invention.
Claims
1. A method for predicting wind power of a multi-node wind turbine cluster, characterized in that: include: Obtain the geographic coordinate data and historical power data of the wind turbine cluster and normalize the data; Construct a spatiotemporal parallel model framework, which includes a spatial branch module, a temporal branch module, and a spatiotemporal fusion module; The spatial branch module uses the Euclidean distance to measure the spatial proximity between wind turbines based on the geographic coordinate data to obtain the spatial adjacency graph SAG. Based on the historical power data, the Pearson correlation coefficient is used to linearly quantify the functional similarity relationship between wind turbines to obtain the functional relationship graph FRG. The Chebyshev polynomial approximation is used to optimize the graph convolutional network, and the multi-head cross-attention mechanism is combined to aggregate the cross-graph feature associations between the spatial adjacency graph SAG and the functional relationship graph FRG, and output multi-graph fusion features. The time branch module uses parallel multi-resolution convolution kernel-gating module and recursive time module to extract the short-term and long-term time dependencies of wind power series, integrates them through multi-head cross attention mechanism, and outputs time series features; The spatial branch module and the temporal branch module work in parallel. The spatiotemporal fusion module uses a multi-head cross-attention mechanism to fuse multi-image fusion features and time series features, realize multi-granularity interaction between spatiotemporal features, and output the prediction results of wind power.
2. The method for predicting wind power of a multi-node wind turbine cluster according to claim 1, characterized in that: The construction of the spatial adjacency graph SAG includes: obtaining a geographic coordinate matrix of the wind turbine cluster based on the geographic coordinate data of the wind turbine cluster, and calculating the Euclidean distance D between any pair of wind turbines in the geographic coordinate matrix. ij =||C i -C j ||2, C i and C j denote the coordinate vectors of the i-th and j-th wind turbines respectively, and ||·||2 denotes the quantification of the geometric distance between the two wind turbines through the definition of the second norm.
3. The method for predicting wind power of a multi-node wind turbine cluster according to claim 1, characterized in that: The construction of the functional relationship graph FRG includes: based on the historical power data of the wind turbine cluster, using the Pearson correlation coefficient to linearly quantify the functional similarity relationship between the wind turbines, and obtaining the functional relationship matrix of the wind turbine cluster. In the functional relationship matrix, the element R ij Represents the spatial similarity between any two wind turbines i and j: Where P i and P j denote the historical power vectors of the i-th and j-th wind turbines respectively, is the power series after removing the mean, is the mean of the i-th wind turbine power sequence, and ||·||2 represents the Euclidean norm.
4. The method for predicting wind power of a multi-node wind turbine cluster according to claim 1, characterized in that: In the process of optimizing graph convolutional networks with Chebyshev polynomial approximation, the first kind of Chebyshev polynomial T is introduced. k (x) = cos(karccosx), |x| ≤ 1 replaces the convolution kernel of the spectral convolution: T k (·) is a Chebyshev polynomial of order k; β k is the parameter updated during training iterations; The rescaled eigenvalue diagonal matrix is obtained by transferring Λ to obtain the graph convolution network: Then put the matrix operation into the Chebyshev polynomial to get: After Laplace normalization, the final first-class Chebyshev polynomial approximate graph convolution is expressed as: in Represents the symmetric normalized Laplacian matrix After re-normalization, λ max yes The maximum eigenvalue of , K is the number of consecutive screening operations in the model; the signal of the node is smoothed by aggregating and transforming the k-hop neighborhood history information of the node, and we get: and y sd represents the spatial feature based on the spatial adjacency graph SAG, y sc Represents the spatial features based on the function relationship graph FRG, and both Flattened by time step and node dimension The output y of the SAG module s ' d Provide Query Q s =y s ' d W h Q , the output y of the FRG module s ' c Provide key K s =y s ' c W h K Sum value V s =y s ' c W h V , and obtain the linear transformation output Y that dynamically integrates the two features Spa =Output SpatialBranch =Concat(head1,head2,...,head h )W t O .
5. The method for predicting wind power of a multi-node wind turbine cluster according to claim 1, characterized in that: The multi-resolution convolution kernel-gating module uses convolution kernels of different resolutions to simulate the dependencies of different time scales in the time series, combines the gating mechanism to control the information flow through the network layer, and weights the convolution results of different resolutions through the gating mechanism to obtain multi-scale time modal features; The recursive time module is used to retain the context information of the long time step and output the hidden state y of the entire sequence tg =[h1,h1,...,h t ], as a supplement to extracting long-term dependencies, h t is the current hidden state at time t, and T is the length of the input sequence; The temporal features captured by the multi-resolution convolution kernel-gated module and the recursive temporal module are flattened by the time step and node dimensions respectively as The multi-resolution convolution kernel-gating module provides the query Q t =y t ' c W h Q , the recursive time module provides the key K t =y t ' g W h K Sum value V t =y t ' g W h V , use the multi-head cross attention mechanism to integrate multi-temporal modal information and obtain the linear transformation output of the fused features: Y Temp =Output TemporalBranch =Concat(head1,head2,...,head h )W t O .
6. The method for predicting wind power of a multi-node wind turbine cluster according to claim 1, characterized in that: The spatiotemporal fusion module combines the fused multi-image features from the spatial branch module and time series features from the time branch module Realize multi-granularity interaction between spatiotemporal features:
7. A computer program product, characterized in that: The method comprises a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a computer, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Wind power generation power prediction method based on space-time diagram convolution and gating attention
CN120597214A
New energy time-sharing multi-day power generation capacity prediction system and method based on large model
CN120952244A
A new energy time-sharing multi-day power generation capacity prediction system and method based on a large model
CN120952244B