A Multidimensional Wind Power Prediction Method Based on Wind Speed Climbing Identification and Matching
By employing the Informer algorithm, which combines a spatiotemporal graph convolutional filling model, dynamic quantile standardization, and multidimensional data fusion with a wind speed similarity matching mechanism, the problem of accurately predicting wind power ramping events is solved, thereby improving the accuracy and adaptability of wind power prediction.
Patent Information
- Application Number
- CN202511173881.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing wind power prediction methods cannot accurately predict wind power ramping events in sudden weather conditions, leading to power imbalance. Furthermore, existing models have poor universality and cannot deeply mine the effective information between wind speed and meteorological elements.
The model employs a spatiotemporal graph convolutional filling model to handle missing data, dynamic quantile standardization, and coupled anomaly detection. It combines reinforcement learning and the Informer algorithm with multidimensional data fusion, uses a wind speed similarity matching mechanism for wind power prediction, and incorporates a physical information mutation event classifier and a dynamic adaptive wind speed mutation identification algorithm to improve the model's feature representation and generalization ability.
It improves the accuracy and responsiveness of wind power forecasting, adapts to sudden wind speed changes, enhances the model's feature representation and generalization ability, and is suitable for short-term forecasting under various complex meteorological conditions.
Smart Images

Figure CN120728586B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wind power data prediction, and in particular relates to a multi-dimensional wind power prediction method based on wind speed ramp identification and matching. Background Technology
[0002] In wind power forecasting, wind speed and power data in sudden weather events such as cold waves and thunderstorms typically exhibit strong randomness and abrupt changes. When wind power fluctuates dramatically within a short period, and the magnitude and rate of change exceed the regulation capacity of the power system's control strategies, power imbalances occur. Existing conventional wind power forecasting methods cannot provide accurate and reliable predictions in such situations; this is known as a wind power ramping event. While existing research on ramping events has yielded significant results, issues such as unclear ramping definitions and limited data dimensions remain to be addressed. If the system could predict future power values based on current wind power data, and thus forecast wind power during abrupt events, grid dispatching could prepare for emergencies in advance, better ensuring the safety of grid operations.
[0003] In recent years, research on wind power ramping events has mainly focused on two aspects. One type of research aims to construct criteria for identifying ramping events, such as using wind power change rate, standard deviation sliding window, and abrupt change detection methods to accurately capture the start and end boundaries of ramping sections. However, there is currently no unified standard for defining ramping events in the wind power field, and existing definitions are difficult to apply directly to complex practical problems. Therefore, determining a suitable definition based on different plant environmental conditions and operating conditions is an urgent problem to be solved. The second type of research focuses on deep learning and intelligent algorithms. With the development of algorithms, wind power ramping prediction models based on network structures such as Long Short-Term Memory (LSTM), Convolutional Neural Networks (CNN), and Transformer networks with self-attention mechanisms have gradually become research hotspots, improving the response capability for periods of drastic wind power changes in short-term predictions. However, many existing prediction models suffer from overly specific targeting and poor universality. Moreover, abrupt wind speed changes involve multiple factors and have complex nonlinear characteristics. Relying solely on wind speed data and a single neural network model cannot deeply mine the effective information between wind speed and various meteorological elements along high-speed rail lines to obtain more accurate predictions.
[0004] Therefore, how to achieve the best possible prediction results while considering factors such as ramp characteristics and multimodal meteorological factors still requires further research. Effectively establishing wind power prediction models to mine the characteristic information between meteorological elements, ramp events, and wind power to obtain more accurate prediction results is an important problem that needs to be solved. However, a review of relevant literature revealed no such findings. Summary of the Invention
[0005] Objective of the Invention: The technical problem to be solved by this invention is to address the shortcomings of existing technologies by providing a multi-dimensional wind power prediction method based on wind speed ramp identification and matching, comprising the following steps:
[0006] Step 1: Preprocess the wind speed and power data of the wind farm: use a spatiotemporal graph convolution filling model to handle data loss caused by sensor failure; use dynamic quantile standardization to standardize the original data; use coupled anomaly detection to detect anomalies in the raw wind speed and power data; finally, use the novel mode decomposition algorithm VMD-I.
[0007] Step 2: Use reinforcement learning to optimize the dynamic window width, add spatiotemporal factors and jointly analyze and screen the extreme point rate, add a physical information-based mutation event classifier to form a dynamic adaptive wind speed mutation identification algorithm, and use the dynamic adaptive wind speed mutation identification algorithm to denoise the wind speed data.
[0008] Step 3: Expand the dimensions of meteorological factors and use a wind speed time period matching algorithm to match wind speed data with historical wind speeds.
[0009] Step 4: Based on the Informer algorithm with attention mechanism, which integrates multidimensional data fusion, predict the future wind power generation capacity.
[0010] Beneficial Effects: This invention introduces a wind speed similarity matching mechanism to fully exploit the structural similarity between historical wind speed segments and current wind speed sequences, improving the adaptability of power prediction to abrupt wind speed changes. By combining multimodal information such as wind speed statistical features, wind speed matching difference features, power response features, and meteorological element features, a deep fusion approach guides the main prediction network, significantly enhancing the model's feature representation and generalization capabilities. This invention employs an improved Informer main network to achieve parallel long-sequence modeling, possessing good prediction accuracy and inference efficiency. It can effectively identify the impact of wind speed anomalies on wind power, improving the accuracy of response to wind power ramp-up events. Furthermore, this method has a flexible structure, strong scalability, and industrial deployment adaptability, making it suitable for short-term wind power prediction scenarios under various complex meteorological conditions. Attached Figure Description
[0011] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.
[0012] Figure 1 This is a flowchart illustrating the specific implementation of the present invention.
[0013] Figure 2 A schematic diagram of the sudden weather identification and prediction network architecture designed for this invention.
[0014] Figure 3 This is a schematic diagram of the wind speed statistical feature network.
[0015] Figure 4 This is a schematic diagram of a network for matching wind speed differences.
[0016] Figure 5 This is a schematic diagram of the power response characteristic network.
[0017] Figure 6 This is a schematic diagram of the meteorological element characteristic network.
[0018] Figure 7 A schematic diagram of the wind speed similarity coefficient algorithm designed for this invention.
[0019] Figure 8 The preprocessing method designed for this invention. Detailed Implementation
[0020] like Figure 1 , Figure 2 As shown in the figure, this embodiment of the invention provides a multi-dimensional wind power prediction method based on wind speed ramp identification and matching, including the following steps:
[0021] Step 1: Preprocess the wind speed and power data of the Jiangxi Baijiashe wind farm, such as... Figure 8 As shown: a spatiotemporal graph convolutional filling model (STGC-Impute) is used to handle data loss caused by sensor failure; dynamic quantile normalization is used to standardize the original data; coupled anomaly detection (CAD) is used to detect anomalies in the original data of sudden wind speed and power; finally, a novel mode decomposition algorithm, VMD-I, is used to address the shortcomings of traditional mode decomposition algorithms in generating non-physical modes (such as noise-dominated IMF components) in scenarios of sudden wind speed changes.
[0022] Step 2: Use reinforcement learning to optimize the dynamic window width, add spatiotemporal factors and jointly analyze and screen the extreme point rate, add a physical information-based mutation event classifier to form a dynamic adaptive wind speed mutation identification algorithm, and use the dynamic adaptive wind speed mutation identification algorithm to denoise the wind speed data.
[0023] Step 3: Expand the dimensions of meteorological factors and use a wind speed time period matching algorithm to match wind speed data with historical wind speeds.
[0024] Step 4: Based on the Informer algorithm with attention mechanism (Informer is an algorithm designed for long time series forecasting) of multi-dimensional data fusion, predict the future wind power generation.
[0025] Step 1 includes:
[0026] Step 1-1: Collect wind speed and wind power data from the wind farm. Divide the data into past data and historical data. The past data consists of wind speed and power data from the wind farm over the past year, while the historical data consists of wind speed and power data from the wind farm over the past two years, from three years ago to one year ago.
[0027] Collect multimodal meteorological data, including numerical weather prediction (NWP) data for the wind field location;
[0028] Steps 1-2 involve preprocessing past and historical data, including spatiotemporal graph convolutional filling model, dynamic quantile normalization, coupling anomaly detection, and VMI-I mode decomposition.
[0029] Step 1-2-1 involves inputting the wind speed and power data from the wind farm into a spatiotemporal graph convolutional filling model for processing, including the following steps:
[0030] Step 1-2-1-1, input the following data: wind speed V, power P, unit coordinates ,wind direction and missing mask (1 indicates valid, 0 indicates missing); where Represents the space of real numbers; Let N represent the two-dimensional spatial coordinates of the i-th unit; N represents the total number of units; T represents the number of time steps, for example, if there is a time point every ten minutes within a day, then T=144.
[0031] Step 1-2-1-2: Construct a dynamic graph structure and build node features. The feature vector of each node (unit) contains both static and dynamic information.
[0032] ,
[0033] in It is the input feature vector used for node i at time t in wind power forecasting; Let be the wind speed of the i-th unit at time t; Let be the power generation capacity of the i-th unit at time t; Let be the altitude of the location of the i-th unit; For the t-th wind turbine unit, there is the model or category. To convert the wind direction angle into a cosine value, which is used to represent wind direction characteristics; To convert the wind direction angle into a sine value, which is used to represent wind direction characteristics;
[0034] Next, the edge weights are dynamically updated to reflect spatiotemporal correlation:
[0035] ,
[0036] in Indicates edge weight; Indicates the spatial attenuation radius (default 3km). Let be the wind direction angle of the i-th unit at time t, with the wind direction term ranging from [0-1], where 1 indicates that the winds are in the same direction; exp represents the natural exponential function.
[0037] Steps 1-2-1-3, designing the spatiotemporal graph convolutional filling model, include the following steps:
[0038] Step 1-2-1-3-1: Establish a Spatial Graph Convolutional Network (ST-GCN). The input feature matrix of the Spatial Graph Convolutional Network is... , The input feature matrix is the l-th layer; The input feature dimension is ; the spatial graph convolutional network outputs a feature matrix of . , The output feature matrix of the l-th layer, To define the output feature dimension, a dynamic adjacency matrix is constructed based on the input and output feature matrices. For time step t, the adjacency matrix... The element is defined as:
[0039] ,
[0040] in, The maximum neighborhood radius (typically taken as 5km); Representing the adjacency matrix The element in the i-th row and j-th column; Let be the wind direction angle of the j-th unit at time t;
[0041] Next, we perform degree matrix normalization. for:
[0042] ,
[0043] in This represents the graph adjacency weight between node i and node j at time step t; This represents the sum of connection weights of node i, i.e., the degree of node i. Represents a diagonal matrix function; It is a matrix with non-zero elements only on its diagonal;
[0044] Perform symmetric normalization:
[0045] ,
[0046] in This is the normalized adjacency matrix;
[0047] Finally, graph convolution is performed, where the mathematical expression for a single-layer spatial graph convolution is:
[0048] ,
[0049] in, To train the weight matrix for the lesson, For the jump connection weight matrix, The activation function is LeakyReLU (negative slope = 0.01).
[0050] Step 1-2-1-3-2 introduces a temporal attention mechanism. Temporal attention is a key component of STGC-Impute, specifically designed to: capture long-term temporal dependencies in wind speed and power data (such as daily cycles and weather system transits), dynamically weight important time nodes (such as sudden changes during the passage of a typhoon eye), and address the fixed field-of-view limitation of traditional recurrent neural networks (RNNs) / convolutional neural networks (CNNs). This includes the following steps:
[0051] Step 1-2-1-3-2-1, Perform time attention scoring: For each node Define the query vector as The key vector is The value vector is , where X is the input feature tensor of the entire model; B is the batch size, i.e. the number of input samples at one time; d is the projection dimension (Embedding Dim), which maps the original features to d dimensions for attention calculation; It is a learnable projection matrix that maps input features to a vector of query Q, key K, and value V; the attention weights are calculated using the following formula:
[0052] ,
[0053] in This indicates that the nth node is in the nth position. The time step for the first Attention score at each time step; n is the node number (wind turbine number). These are the indices for the current time step and other time steps, respectively. For the nth node in the nth position The query vector at each time step; For the nth node in the nth position The key vector at each time step; F represents the original feature dimension, i.e., the input feature dimension;
[0054] For all Perform softmax weight normalization:
[0055] ,
[0056] in For node n in the th... The time step for the first Attention weights for each time step; This indicates that the nth node is in the nth position. Each time step scores the attention of the k-th time step;
[0057] Finally, the attention-weighted output is:
[0058] ,
[0059] in Indicates that node n is in time Temporal feature representation after attention-weighted fusion; For value vectors, Is node n in time Embedded representation on; Indicates the nth node at time Regarding time Attention weights;
[0060] Step 1-2-1-3-2-2, the overall temporal attention matrix is represented as follows:
[0061] ,
[0062] ,
[0063] Where A is the attention matrix of each node at each time step to other time steps; Z is the output of the time attention mechanism, representing the weighted representation of each node at each time step;
[0064] Step 1-2-1-3-3, further model the time dependency using a gated loop unit:
[0065] ,
[0066] ,
[0067] ,
[0068] ,
[0069] in, This is the input for the t-th time step; This is the hidden state from the previous moment; For those who wish to update, decide how much historical information to retain; To reset the gate, control the degree to which historical states affect the current state; This is the candidate hidden state; This represents the final hidden state at the current time step. These are learnable linear transformation weights; For the Sigmoid function; It is the hyperbolic tangent activation function; The product of Hadamard; This involves concatenating vectors.
[0070] Steps 1-2-1-4: Establish the physical constraint loss function and jointly optimize intermediate parameters. and To achieve (value in this embodiment) The spatiotemporal graph convolutional filling model (STGC-Impute) is physically constrained to achieve the following: using a physical prior (wind speed-power curve) as auxiliary supervision to enhance the physical rationality of the model's predictions, and reducing the power curve loss. for:
[0071] ,
[0072] in It is the power of node i at time t predicted by the spatiotemporal graph convolutional padding model STGC-Impute. Let i be the wind speed at time t. It is a function of wind speed and wind power (power curve function) obtained from experience or fitting. This represents a set of target points for prediction where wind speed is known, power is missing, or is uncertain. Optional forms include: three-segment linear, logistic function, or third-order polynomial fitting of the power curve;
[0073] The total loss function of the spatiotemporal graph convolutional padding model STGC-Impute for:
[0074] ,
[0075] in This represents the weight hyperparameter, used to control the strength of the power curve constraint; It is the standard reconstruction loss, calculated using the following formula:
[0076] ,
[0077] in The set of locations that can be supervised by the true values (training set non-missing points); These are the predicted values from the spatiotemporal graph convolutional filling model. This represents the actual value from the wind speed data;
[0078] Step 1-2-2, Dynamic Quantile Standardization: Traditional standardization methods (such as Min-Max and Z-Score) in wind power scenarios suffer from the problems of using fixed parameters, neglecting the non-stationary and abrupt changes in wind speed. Feature distortion leads to sensitivity to extreme values, causing the standardized data distribution to deviate from physical laws. Furthermore, the joint distribution relationship between wind speed and power is not fully utilized due to multimodal fragmentation. In this invention, to improve the normalization robustness of wind speed data and multimodal meteorological data under abrupt events (such as wind power ramp-up events), a dynamic quantile standardization method based on a sliding window is proposed. This method constructs a local window for each time step in the time series, calculates the quantiles (such as 25% and 75%) in the local distribution, and normalizes the original values, thereby avoiding the influence of global extreme points on the normalization results and improving the model's adaptability to abnormal and drastic fluctuations in wind speed. The steps include:
[0079] Step 1-2-2-1: Construct a sliding window centered on the current time point for the wind speed data at each moment. ,in This represents a local time window centered at the current time t, where w is the window length. This indicates a round-down operation; Indicates the offset from the current time t towards the future. Wind speed values at each time step;
[0080] Step 1-2-2-2, calculate the quantiles within the window:
[0081] ,
[0082] ,
[0083] ,
[0084] in , , These represent the 25th percentile, median (50th percentile), and 75th percentile of the data within the current window, respectively. A quantile function is a function or operation used to extract values at specified quantile positions from a set of data.
[0085] Steps 1-2-2-3: Calculate the quantile interval: ,in, Indicates the interquartile range. It reflects the distribution range of the middle 50% of the data, that is, the dispersion of the middle 50% of the data, and can be used for anomaly-resistant normalization;
[0086] Step 1-2-2-4: Analyze the wind speed data at the current moment. For standardization, choose any of the following formulas:
[0087] Median symmetric normalization: ;
[0088] Min-Max normalization: ;
[0089] Logarithmic compression normalization: ;
[0090] in, This represents the standardized wind speed or meteorological variable value. The smoothing constant introduced to prevent division by zero errors is generally set to a value of [value missing]. ;
[0091] Step 1-2-2-5, Anomaly Detection Rules: If the following conditions are met: or If the data point at time t is determined to be a local outlier, this rule is derived from the outlier identification method defined in the box plot and is used to identify sudden wind speed events within a local time window.
[0092] Steps 1-2-3: The improved variational mode decomposition (VMD-I) algorithm is used to preprocess the wind speed time series.
[0093] Steps 1-2-3 include the following steps:
[0094] Step 1-2-3-1: Use the original VMD algorithm to process the wind speed sequence. Decomposed into K modal functions ;
[0095] Steps 1-2-3-2, for each mode function Calculate the following indicators:
[0096] ,
[0097] ,
[0098] ,
[0099] in Spectral entropy reflects the degree of frequency disorder in a mode; the smaller the value, the purer the mode. Energy percentage indicates the proportion of this mode in the overall signal; a larger percentage indicates that the mode is more important. To reconstruct the correlation coefficient, which measures the correlation between the mode and the original wind speed, the larger the value, the more the main trend is preserved; Indicates the proportion of energy in the frequency component; Represents the j-th mode function; corr represents the correlation coefficient function;
[0100] Steps 1-2-3-3: Calculate the modal validity score and reconstructed signal output:
[0101] ,
[0102] ,
[0103] ,
[0104] ,
[0105] ,
[0106] in Non-negative weights; This indicates the validity of the k-th mode; The threshold is typically set to [value]. , used to filter the set of valid modes; To reconstruct the signal and obtain the noise-suppressed main wind speed signal; This represents the set of modal indices that satisfy the validity score being greater than or equal to the threshold.
[0107] Step 2 includes:
[0108] Step 2-1, Extracting Extreme Values: For each sub-signal IMF, calculate and extract the maximum value. Minimum value And all extreme points, all extreme points form a set. , This represents the nth extracted extreme point, where n represents the total number of extreme points.
[0109] Step 2-2: Remove the pseudo-extreme points in each sub-signal IMF (pseudo-extreme points refer to points where the distance between two extreme points is too small, which often occurs in the fluctuating wind speed range).
[0110] Given the excessive density of extreme points during the climbing phase, false poles are removed:
[0111] Calculate dynamic window width based on adaptive coefficients :
[0112] ,
[0113] in The maximum wind speed in the sub-signal IMF. The minimum wind speed in the sub-signal IMF. This represents the adaptive coefficient, and its value is generally 1. ;
[0114] Steps 2-3: Screening for extreme points:
[0115] For each extreme point, calculate the Euclidean distance between the extreme point and its nearest neighbor. If the Euclidean distance is greater than the dynamic window width... ,satisfy If the extreme point is retained, it will be counted as a valid extreme point. Otherwise, discard the extreme point; where Represents the k-th extreme point and the (k+1)th extreme point Euclidean distance;
[0116] Finally, all the extreme points that satisfy the conditions are retained, resulting in the set of effective extreme points. ,in This represents the f-th effective extreme point;
[0117] Steps 2-4: Based on the selected extreme point set output by the adaptive pole selection model, calculate the pole rate according to the number of selected poles. This enables the selection of sub-signal data from the VMD mode decomposition results for the identification of sudden wind speed events, and the calculation of the pole rate of each sub-signal (IMF) component. :
[0118] ,
[0119] in Indicates the first Extreme rate of wind speed data This represents the total number of reconstructed extreme points in the preprocessed wind speed data. Indicates the first The total number of reconstructed extreme points for each wind speed data point is used to filter the wind speed data based on the extreme point rate of each data point, ensuring that the data meets the requirements. Wind speed data with values less than or equal to a threshold of 0.35 are retained, resulting in the retained wind speed dataset. ;
[0120] Steps 2-5, wind speed reconstruction:
[0121] The retained wind speed dataset The wind speed data is reconstructed by overlaying, and the reconstructed wind speed data is denoted as... :
[0122] ,
[0123] Where sum refers to the superposition function; The reconstructed wind speed data; This represents the retained wind speed dataset.
[0124] Step 3 includes proposing a wind vector suitable for sudden wind power events. This vector is composed of four types of features: statistical features of wind speed segments, difference features of matching segments, power response features, and weather forecast elements. Multimodal, multi-dimensional, and comprehensive extraction of corresponding wind segment features is used as input to the algorithm to improve prediction accuracy.
[0125] Step 3-1: Construct statistical indicators for the selected wind segments within the current time window, including the following steps:
[0126] Step 3-1-1: Calculate the gradeability coefficient for each wind speed range.
[0127] ,
[0128] in This represents the magnitude of wind speed variation, where 'a' represents the sequence number of all time points within the time period 'c', and 'c' represents the time period of extreme points. Numerically, it is the product of the number of time series points 'a' and the fixed time step 'Δt'. This represents the slope coefficient within time period c. The ramp factor represents the ramp factor at each point, which is numerically equal to the ramp factor at each time series point. The derivative;
[0129] Step 3-1-2: Construct the wind speed segment for the current time window, and calculate the mean wind speed, mean acceleration, average intensity, standard deviation, coefficient of variation, and upward trend ratio for the wind speed segment.
[0130] ,
[0131] ,
[0132] ,
[0133] ,
[0134] ,
[0135] ,
[0136] in is a wind speed sequence of length L; Accel is the mean acceleration of the wind speed segment, which reflects the absolute value of the second difference of the wind speed sequence, describes the degree of wind speed change, and is used to detect whether there are drastic wind speed fluctuations; M is the average intensity value of the wind speed segment. The average wind speed of the wind speed range. Standard deviation represents the degree of dispersion or fluctuation in wind speed; The coefficient of variation is used to measure the relative dispersion of wind speed. The ratio of upward trends. This is an indicator function, which takes the value 1 if the wind speed at the next moment is greater than the wind speed at the current moment (indicating an upward trend), and takes the value 0 otherwise (indicating no upward trend).
[0137] Step 3-1-3: Organize the past wind speed segments to be matched. and Historical wind speed ranges matched after wind speed matching algorithm The steps include:
[0138] Step 3-1-3-1: Calculate the wind speed intensity difference. :
[0139] ,
[0140] in The wind speed sampling interval is typically set to 15 minutes. Indicates the first Historical wind speed data at any given time. Indicates the first Historical wind speed data at any given time, wind speed intensity difference Normalization interval set to ;
[0141] Step 3-1-3-2, calculate the wind speed trend difference. :
[0142] ,
[0143] in, Indicates the first Historical wind speed data at any given time This represents the past wind speed data at time a+1; the resulting wind speed trend difference... Perform normalization processing, and set the normalization interval to... ;
[0144] Step 3-1-3-3, as follows Figure 7 As shown, calculate the wind speed similarity coefficient:
[0145] ,
[0146] in Taking into account both the difference in wind speed amplitude and the difference in wind speed trend, the final similarity metric is determined.
[0147] Step 3-1-4: Integrate the statistical features to form a wind speed statistical feature vector. :
[0148] ,
[0149] like Figure 3 As shown, then The established Statistical Feature Embedding Network (SFEN) module for wind speed is then fed in:
[0150] ,
[0151] ,
[0152] ,
[0153] ,
[0154] ,
[0155] ,
[0156] ,
[0157] in It is the input tensor after tensor reshaping; It is an activation function used to correct linear units; This represents a one-dimensional convolution operation; This is the output of the convolutional layer, i.e., the number of features after convolution; This is a one-dimensional max pooling operation; As a result of pooling, each channel becomes a feature of length 2; For flattening operation; The flattened vector; This is the weight matrix of the fully connected layer, with an output dimension of 64. For bias terms; This is the output of the first fully connected layer; This represents the final embedded wind speed statistical features; This represents the weight matrix of the second fully connected layer. This represents the shape reshaping operation, where s represents the original statistical feature vector. Indicates the size of the convolutional pooling kernel. Indicates the number of output channels;
[0158] Wind speed statistical feature vector The vectors are processed through tensor reshaping, convolutional layers, and pooling, flattened into vectors, and then fed into two fully connected layers to obtain the final output. The output combines wind speed statistical feature embeddings with convolutional and pooling features, resulting in stronger feature learning and abstraction capabilities.
[0159] Step 3-2, design the wind speed matching difference feature vector, including the following steps:
[0160] Step 3-2-1: Construct a set of historical candidate segments, including historical wind speed data. Divide the data into sliding windows of length L, and extract wind speed segments of length L. ; This represents the Nth sliding window segment in the historical wind speed data. ;
[0161] Step 3-2-2: Construct the wind speed matching difference feature vector:
[0162] ,
[0163] Where S is the average slope and D is the amplitude change, the calculation formula is:
[0164] ,
[0165] ;
[0166] Step 3-2-3: Perform sliding extraction on the historical wind speed sequence to obtain historical wind speed segments. ,calculate Corresponding feature vector And calculate the similarity metric:
[0167] ,
[0168] in , These are weighting coefficients. ; Indicates current wind speed segment With historical fragments The overall similarity score indicates that the smaller the value, the more similar the two people are. Used to measure the similarity in shape between two time series; historical wind speed segments obtained by similarity matching of wind speeds in the current time period;
[0169] For the feature extraction function of composite slope climbing and intensity, select Extract the power sequence from the smallest matching segment. ;
[0170] Step 3-2-4: Design the power response feature vector. To model the changing behavior of the power sequence in the matching segment and its response characteristics to wind speed fluctuations, calculate the power response feature vector Feat(y), which is used to quantify dynamic information such as average power level, rate of change, and number of abrupt changes. Also, design the fused mode vector, defined as:
[0171] ,
[0172] ,
[0173] ,
[0174] ,
[0175] ,
[0176] in Average power; This represents the maximum power change rate. The mean of the slope of change; The ratio of fluctuation amplitude; The number of mutations. The mutation threshold is typically set at 10% of the maximum rate of change. The loss function; This represents the wind power value corresponding to the i-th time step, which comes from the power time series corresponding to the matched historical wind speed segments; This represents the time interval between two adjacent time steps; This represents the maximum value in the wind power segment y, reflecting the upper limit of power fluctuation; This represents the minimum value in the wind power segment y, reflecting the lower limit of power fluctuation; This indicates an indicator function; a value of 1 indicates that the condition is true, and a value of 0 indicates that the condition is false.
[0177] Step 3-3: Fuse to obtain feature vectors :
[0178] ,
[0179] Multilayer Perceptron (MLP) is used to upscale vectors to a unified vector dimension. :
[0180] ,
[0181] in The power response features after unifying the vector dimensions are fed into the main network as one of the modality fusion inputs.
[0182] Steps 3-4 involve expanding the dimensions and weighting the meteorological factors, and constructing a feature matrix by combining the derived features.
[0183] Steps 3-4 include the following steps:
[0184] Step 3-4-1: Collect meteorological data including wind speed, cloud cover, air pressure, temperature, humidity, and irradiance;
[0185] Step 3-4-2: Perform data preprocessing, using interpolation and mean to fill missing values, converting data of different dimensions to the same scale, and performing Z-score normalization on each column of variables.
[0186] ,
[0187] ,
[0188] ,
[0189] in, This represents the normalized value. This represents the t-th data in the i-th row; This represents the sample mean of the variable in the i-th row; Represents the sample standard deviation of the variable in the i-th row;
[0190] After processing, the normalized matrix Norm is obtained:
[0191] ,
[0192] in This represents the standardized value of the nth variable at time step t;
[0193] Step 3-4-3: Use a correlation matrix to analyze the correlation between NWP meteorological factors and wind power at the Baijiashe wind farm, and use the Pearson correlation coefficient to measure the correlation between elements of the multivariate matrix.
[0194] ,
[0195] ,
[0196] Where R represents the correlation coefficient matrix; This represents the Pearson correlation coefficient between variable i and variable j;
[0197] Pearson correlation analysis was used to quantify the statistical linear relationship between various meteorological factors and wind power (or gradeability coefficient). Wind speed, irradiance, and temperature, which have the highest correlation with wind power, were selected as the meteorological element feature vectors, specifically represented as follows:
[0198] ,
[0199] in Represents a meteorological element matrix. , , Let i represent the wind speed, temperature, and radiation at time i, respectively.
[0200] Step 4 includes:
[0201] Step 4-1: Design a feature embedding subnetwork. The feature embedding subnetwork adopts a multimodal feature fusion prediction method based on wind speed similarity matching vector. In view of the problem that wind power sudden change event features are obvious and traditional single input is difficult to model, four types of feature vectors are introduced: wind speed statistical feature vector Stat, wind speed matching difference feature vector Diff, power response feature vector Resp, and meteorological element feature vector NWP.
[0202] The feature embedding subnetwork includes a wind speed statistical feature subnetwork, a wind speed matching difference feature network, a power response feature network, and a meteorological element feature network;
[0203] A wind speed statistical feature subnetwork is established. This subnetwork utilizes the statistical features (mean, variance, skewness, kurtosis) of the wind speed time series to construct feature vectors. A nonlinear feature mapping is implemented based on a feedforward neural network structure with two or more layers. Finally, a unified-dimensional deep wind speed statistical representation is output, expressed as:
[0204] First layer of feedforward neural network:
[0205] ,
[0206] in This represents the input wind speed statistical feature vector; This represents the weight matrix of the first layer; Indicates the bias term; Indicates the activation function; This is the hidden vector output from the first layer; The output dimension of the first hidden layer, i.e., the first feedforward neural network maps the original input (such as wind speed statistical features) to a... In the dimensional representation space;
[0207] Second layer feedforward neural network:
[0208] ,
[0209] in This represents the weight matrix of the second layer. This is the second-level bias term; This is the nonlinear output of the second layer; This represents the output dimension of the second hidden layer, which is the final output dimension of the entire wind speed statistical feature subnetwork.
[0210] Execution layer normalization:
[0211] ,
[0212] in For vectors The mean; Representing vectors The variance; It is a very small constant to prevent division by zero, generally 0. ; These are learnable scaling and offset parameters; normalization is used to improve convergence speed and alleviate the vanishing or exploding gradient problem. Represents a vector The result after the normalization operation at the execution layer;
[0213] Output wind speed statistical feature embedding representation :
[0214] ,
[0215] As part of the fusion input in the main model Informer, such as splicing or attention injection;
[0216] Step 4-2: Establish a wind speed matching difference feature network, the specific structure of which is as follows: Figure 4 As shown, the wind speed matching difference feature network uses a two-layer one-dimensional convolutional neural network structure to extract local difference change patterns, and obtains a global representation through global average pooling. The wind speed matching difference feature vector is then used. This represents the differences in characteristics and power between wind speed segments and similar wind speed segments, allowing Diff-Net to better learn the mapping relationship between power differences caused by wind speed differences:
[0217] Input vector:
[0218] ,
[0219] in ; To match the power sequence of historical wind speed ranges; Differential power; For wind speed matching, use the difference feature vector; This represents the wind power response sequence corresponding to the historical wind speed segment that best matches the current moment.
[0220] The first layer of the convolutional network extracts local patterns by performing convolution on the wind speed matching difference feature vectors:
[0221] ,
[0222] in It is an intermediate representation after the first convolutional layer. This represents the first layer of one-dimensional convolution. This represents the number of output channels for the first convolutional layer. Represents the nonlinear activation function ReLU.
[0223] The second convolutional network further extracts deep local structural information:
[0224] ,
[0225] in, This represents the number of output channels for the second convolutional layer. It is the intermediate representation after the second convolutional layer. This represents the second layer of one-dimensional convolution;
[0226] Perform global average pooling:
[0227] ,
[0228] in This is the output feature vector after global average pooling; This represents the output features of the second convolutional network at time step t;
[0229] The output mapping layer takes the wind speed matching difference feature vector as input to the fully connected layer and maps it to a d-dimensional embedding vector. :
[0230] ,
[0231] in This is the weight matrix; For bias terms;
[0232] Step 4-3: Establish the power response characteristic network, the specific structure of which is as follows: Figure 5 As shown, based on the wind speed matching difference feature network, the power response feature network integrates the power response features into a rate response feature vector. :
[0233] ;
[0234] Power response feature vector Input power response feature network:
[0235] ,
[0236] ,
[0237] ,
[0238] in This is a learnable third-layer weight matrix; This is the third-level bias term; The final output power response embedding vector can be used to concatenate with other sub-features and input into the main network.
[0239] Final output embedding vector As a power response feature vector, it can characterize the sensitivity of power to changes in wind speed, the severity of fluctuations, and the frequency of abrupt changes, enabling the model to better identify the effectiveness of matching power and its predictive contribution.
[0240] Step 4-4: Construct a meteorological element feature network, the specific structure of which is as follows: Figure 6 As shown, this method aims to fully utilize multidimensional meteorological elements (temperature, humidity, air pressure, etc.) from numerical weather prediction (NWP) systems. The meteorological element feature network is used to extract important weather background features affecting wind speed and power response, and includes the following steps:
[0241] Step 4-4-1, Organize the input features:
[0242] ,
[0243] in This represents the input multidimensional meteorological element data matrix;
[0244] Step 4-4-2: The input passes through two convolutional layers, and global average pooling is performed on the time dimension.
[0245] ,
[0246] ,
[0247] ,
[0248] in, For global average pooling; This represents a global representation of meteorological characteristics obtained by compressing the time dimension.
[0249] Step 4-4-3, project as modality embedding vector:
[0250] ,
[0251] ,
[0252] Where W and b represent the learnable weight matrix and bias, respectively; The weights are obtained after Softmax normalization. ; This represents the attention weight at time point t. This represents the meteorological information obtained by incorporating the attention mechanism; These are the parameters for linear transformation; This is the final meteorological feature vector, used to concatenate the outputs of other sub-networks; In mathematical or engineering expressions, it is often used to introduce the definition of variables or additional descriptions.
[0253] Steps 4-5 involve building the main network of the prediction model, including the following steps:
[0254] Step 4-5-1, Prepare the Informer main model, including the following steps:
[0255] Step 4-5-1-1, Prepare the input event sequence ;
[0256] Steps 4-5-1-2 involve passing the vectors to the encoder and decoder in the Informer. The Informer uses an Encoder-Decoder architecture, similar to the Transformer; this includes the following steps:
[0257] Step 4-5-1-2-1, Vector input encoder: The input first undergoes linear mapping and positional encoding to form an embedding. :
[0258] ,
[0259] Including the historical time steps of Step;
[0260] The encoder in Informer consists of two or more encoder blocks, each of which contains a sparse probabilistic multi-head attention mechanism and a feedforward neural network (FFN).
[0261] The proposed sparse probabilistic multi-head attention mechanism replaces the standard multi-head attention mechanism. It employs a sparse multi-head attention mechanism, selecting only query vectors with high information content for attention calculation. The feedforward neural network (FFN) calculation formula is as follows:
[0262] ,
[0263] in This represents a feedforward neural network; x represents the input vector.
[0264] And combine residuals and layer normalization:
[0265] ,
[0266] in This indicates the final output of the current encoder; Representation layer normalization.
[0267] The encoder adds a sequence compression module (such as max pooling, convolution, or average pooling) after each encoder block to achieve hierarchical abstraction.
[0268] ,
[0269] in This represents the feature representation output after sequence compression. Indicates max pooling; Represents one-dimensional convolution; or represents the choice of neural network.
[0270] Step 4-5-1-2-2: The Informer decoder is used to convert the temporal features extracted by the encoder into prediction results for two or more future time steps. It adopts an improved mechanism based on the Transformer architecture, supports parallel prediction, and includes an input concatenation module (providing necessary input information for the decoder), multi-layer decoder blocks, and a linear mapping layer.
[0271] The decoder receives a sequence composed of historical true values and placeholder prediction vectors, with a length of [length missing]. ,in This is the known part. This is the part to be predicted;
[0272] The multi-layer decoder block includes a self-attention module, a cross-attention module, and a feedforward neural network;
[0273] The self-attention module is used to model the self-dependencies of sequences within the decoder;
[0274] The cross-attention module is used to align the encoder output with the decoder input to achieve information flow transmission;
[0275] The feedforward neural network is used to extract nonlinear features, and residual connections and layer normalization are used to enhance training stability.
[0276] The linear mapping layer is used to restore the final output of the decoder to the target dimension through a linear transformation, thereby achieving the predicted output:
[0277] ,
[0278] in This represents the final prediction result vector; Represents a linear mapping layer;
[0279] Step 4-5-2, introduce a gating mechanism:
[0280] ,
[0281] This mechanism enables the model to automatically adjust the fusion strength of modal information according to the current time step, avoiding information interference caused by strong injection and improving robustness;
[0282] in This represents the encoder input after fusing modal information; This represents the original encoder input feature sequence; This represents the weight matrix in the gating mechanism; Represents the modality fusion feature vector; This represents the bias parameter in the gating mechanism;
[0283] Step 4-5-3, Layered Injection: The modality fusion vector is injected into each layer of the Informer encoder, as shown below:
[0284] ,
[0285] in This refers to the Informer encoder module; This represents the modal fusion vector injected into the l-th layer;
[0286] Each floor uses a different projection:
[0287] ,
[0288] in This indicates the independent fully connected projection layer used by layer l; This represents the fused modal feature vector.
[0289] Step 4-5-4, Modality selection attention mechanism: For each modality feature, attention weights are constructed and then weighted and combined.
[0290] ,
[0291] in represents the attention weight of the i-th modality; w represents the learnable attention projection vector; Represents the weight matrix in learnable attention; This represents the embedding vector of the i-th mode; This represents the embedding vector of the j-th mode.
[0292] Final fusion vector:
[0293] ,
[0294] This mechanism allows the model to automatically determine the most critical modal source in the current context during prediction;
[0295] Step 4-5-5 involves performing sub-network fusion representation, including the following steps:
[0296] Step 4-5-5-1, Feature Concatenation and Fusion: The input wind speed sequence is processed through the wind speed statistical feature sub-network, the wind speed matching difference feature network, the power response feature network, and the meteorological element feature network, respectively, to obtain the output vectors of the four sub-networks (each dimension is...). ):
[0297] ,
[0298] in This is the statistical feature vector of wind speed; For wind speed matching, use the difference feature vector; This is the power response feature vector; For meteorological element feature vectors;
[0299] The feature vectors are concatenated along their feature dimensions to construct a multimodal fusion feature vector:
[0300] ,
[0301] Linearly mapped to a vector with the same embedding dimension as the original encoder, with dimension 1. :
[0302] ,
[0303] in Represents the modal fusion vector. , For trainable parameters, , ; This represents the dimension of each modality embedding vector;
[0304] Step 4-5-5-2, Embedding Projection Mapping: In order to match the modality vector dimension with the encoder's main network input, a fully connected layer is needed to map the fused feature vector to the embedding dimension d (usually consistent with the embedding dimension of the Informer).
[0305] Step 4-5-5-3, Broadcast and Fusion Injection.
[0306] Step 4-5-5-3 includes: merging the modality fusion vector Repeated replication to construct a time-consistent mode matrix :
[0307] ,
[0308] in This means that the vector will be copied T times along the time dimension;
[0309] modality matrix With encoder input embedding By adding element by element, we obtain the injected modulated input representation:
[0310] ,
[0311] Finally, the input sequence It serves as input to the Informer encoder for subsequent modeling.
[0312] This method also includes step 5: prior processing and power prediction, which includes the following steps:
[0313] Step 5-1: Obtain the predicted NWP meteorological variable values for the current time period from the Jiangxi Baijiashe Meteorological Station, including:
[0314] ,
[0315] Where C is the number of NWP feature dimensions. A two-dimensional input sequence composed of meteorological variables;
[0316] This represents the temperature sequence from time t-T+1 to time t;
[0317] This represents the pressure sequence from time t−T+1 to time t.
[0318] This represents the humidity sequence from time t−T+1 to time t;
[0319] This represents the wind speed height layer characteristics from time t−T+1 to time t (which may be the Z wind component or elevation).
[0320] Step 5-1-1, Input and Processing, including normalization of all variables (min-max or z-score); unifying the time resolution and aligning with wind speed and power sequences; and filling missing points using linear interpolation or mean values.
[0321] Step 5-1-2, extract structural features :
[0322] ,
[0323] Structural features It will be added to the main network, and together with the outputs of the wind speed statistical feature subnetwork, the wind speed matching difference feature network, and the power response feature network, it will constitute the structure modulation input;
[0324] Step 5-2, prior similarity matching of wind speed sequences, includes the following steps:
[0325] Step 5-2-1: Current wind speed sequence sampling. Extract the wind speed time sequence for the current moment from the real-time observation data of the wind farm. ;
[0326] Step 5-2-2: Extract wind speed segment sequences in the form of sliding windows from historical wind speed records to form a historical wind speed library:
[0327] ,
[0328] This represents the Tth (last) wind speed value in the i-th segment of the wind speed sequence, which is the end wind speed value of that segment; Represents a historical wind speed database;
[0329] Step 5-2-3: Before power prediction, perform similarity matching between the wind speed sequence to be predicted and historical wind speed segments. This is done quickly using a wind speed matching difference feature network to perform wind speed matching for similar segments.
[0330] ,
[0331] ,
[0332] ,
[0333] in, and These represent the structural feature vector of the current segment and the structural feature vector of the i-th historical segment, respectively; To minimize the similarity metric fragment; arg is the function that takes the minimum value;
[0334] Step 5-3: Construct a matching difference sequence by comparing the current wind speed sequence with the most similar historical wind segment point by point and calculating the difference sequence. :
[0335] ,
[0336] in, This represents the wind speed value at the last time step of the most similar historical wind speed segment.
[0337] Matching differential feature extraction (Diff-Net): Extracting differential sequences The input is fed into a wind speed matching difference feature network (Diff-Net) for encoding, and the output is a fixed-dimensional vector:
[0338] ,
[0339] in The wind speed matching difference feature vector output by the wind speed matching difference sub-network is used for subsequent modality fusion. Indicates network naming; This indicates that the difference vector is input into the Diff-Net network structure for forward computation.
[0340] Step 5-4: Obtain the power response characteristics of the wind speed range, including the following steps:
[0341] Step 5-4-1: Obtain the power response of the matched segment and calculate the response characteristics;
[0342] Step 5-4-2: Mapping historical wind speed segments to power segments, obtaining historical wind speed segments through similarity retrieval. Search Historical wind power series for the corresponding time period :
[0343] ,
[0344] in The historical wind power output at time point T;
[0345] Step 5-4-3, construct the response feature vector, and... Considering the historical response curves as a subnetwork of wind speed response characteristics, the following five types of structural statistical features are extracted from them: average power Maximum power change rate Mean slope of change ; fluctuation range ratio Number of mutations This constitutes the power response feature vector (Power Response Embedding).
[0346] Unifying vector dimensions through feedforward neural network MLP:
[0347] ,
[0348] ,
[0349] in Embedded vectors for power response features; The output vector of the power response feature network (dimension: );
[0350] Step 5-5 involves inputting the Informer model, which combines four subnetworks: wind speed statistical feature subnetwork, wind speed matching difference feature network, power response feature network, and meteorological element feature network. This includes the following steps:
[0351] Step 5-5-1, the input feature dimension includes the following four parts:
[0352] ,
[0353] ,
[0354] ,
[0355] ,
[0356] in Embed vectors for wind speed statistical features; The wind speed matching difference feature vector is output by Diff-Net (wind speed matching difference sub-network) and is used to capture the difference between wind speed and similar segments; Output from Resp-Net (Power Response Feature Network) is used to characterize the response to power changes; Input the raw meteorological element sequence from the Numerical Weather Prediction System (NWP);
[0357] Step 5-5-2, Modality Fusion and Injection, fuses the output feature vectors of the sub-networks into... And guide the encoder input through projection, repetition, gating, or interlayer injection:
[0358] ;
[0359] Step 5-5-3, Informer network prediction: The modulated wind speed embedding sequence is input, and after passing through the encoder and decoder structure, the future wind power is predicted.
[0360] ,
[0361] in Let be the wind power at each predicted time in step t+1; To predict wind power series;
[0362] Steps 5-6: Network Training: A loss function with physical information is used to train a neural network consisting of an Informer main network, a wind speed statistical feature sub-network, a wind speed matching difference feature network, a power response feature network, and a meteorological element feature network. Data loss and physical loss are optimized to obtain optimized prediction results. Besides the error in the wind speed prediction stage, the conversion from wind speed prediction to power prediction is another source of error. Starting from practical engineering applications, this paper analyzes a real-world case of the wind power prediction system currently used in the target wind farm using power prediction evaluation indicators, studying the shortcomings of current wind power prediction systems in the wind power field. Based on commonly used error indicators for wind power prediction, the newly released national standard of the People's Republic of China is added, introducing new assessment indicators for wind farms. The technical requirements for wind power or photovoltaic power prediction systems in Jiangxi Province, where the target is located, are selected as the performance indicator, namely the root mean square error (RMSE).
[0363] ;
[0364] Mean Absolute Error :
[0365] ;
[0366] accuracy :
[0367] ;
[0368] Power qualification rate :
[0369] ,
[0370] ,
[0371] in This represents the predicted wind power output at time i. The measured value of wind power at time i; For the first The sum of the online capacity of the target wind farms at any given time; is the power point qualification label at time i, where 1 indicates that the prediction of point i is qualified and 0 indicates that the prediction of point i is unqualified; time is the prediction duration.
[0372] This embodiment uses a wind farm site in Baijiashe area of Jiangxi Province and meteorological monitoring stations within a 20-kilometer radius as the research object to establish a prediction model and use historical data to predict the wind power output for the next fifteen minutes.
[0373] Taking a wind farm site in Baijiashe area of Jiangxi Province as an application example, the site is located in an area with complex terrain, strong time-varying wind speed, and obvious wind speed upswing and sudden drops. Combining 10-minute resolution wind speed data collected by the wind farm's main control system and meteorological data such as temperature, humidity, and air pressure provided by multiple meteorological monitoring stations within a 20-kilometer radius, the wind power prediction method based on wind speed similarity matching proposed in this invention is applied to achieve short-term prediction of wind power within the next hour (6 time steps).
[0374] In the experiment, typical seasonal wind speed data from the past year were selected for modeling and prediction validation. The results show that the mean absolute error (MAE) of the proposed method decreased by approximately 15.7% in predicting sudden wind speed changes, and the F1 score for predicting sudden power changes (such as sharp increases or decreases exceeding 30% of the rated power) increased by 22.3%. Compared with the standard Long Short-Term Memory (LSTM) network method, the conventional CNN-Transformer combination method, and traditional statistical regression models (such as SVR), this invention has superior anomaly response capability and fitting accuracy.
[0375] Its advantages are mainly reflected in the following aspects:
[0376] Structural matching-driven prediction enhancement: By matching prior wind speed similarity, the model can be provided with structurally similar historical wind power response segments, enabling structural analogy and prediction compensation for future wind speed change trends, which is particularly suitable for scenarios with multiple sudden wind speed changes.
[0377] Deep fusion of multimodal features: This method introduces four heterogeneous feature sources: wind speed statistics, structural differences, power response, and NWP meteorological factors. This effectively avoids the problem of insufficient robustness of the model to external disturbances under a single wind speed input and enhances the interpretability and correlation between time series features.
[0378] The improved Informer architecture has the ability to model long sequences: Compared with conventional Transformer or LSTM methods, the improved Informer main network still has strong feature extraction and parallel computing capabilities when dealing with prediction tasks with long windows (such as more than 12 steps), which significantly improves prediction efficiency and performance.
[0379] It has strong generalization ability for practical deployment: the model shows stable prediction performance under multiple seasons and multiple wind farm site data, and has strong cross-time and cross-regional migration ability, which can provide effective prediction support for wind farm power generation scheduling, active power control and grid connection.
[0380] In summary, this invention not only improves the accuracy of wind power prediction, especially during periods of rapid wind speed change, but also has good potential for engineering applications.
[0381] This invention provides a multi-dimensional wind power prediction method based on wind speed ramp identification and matching. This method uses wind speed statistical features, wind speed structure difference features, power response features, and numerical weather forecast information as multi-source inputs, fusing them to construct a unified modal embedding vector, which is integrated into the Informer main network framework to achieve high-precision short-term wind power prediction. The specific implementation can be flexibly adjusted according to engineering deployment requirements, possessing good scalability and engineering implementation capabilities.
[0382] The method of this invention is applicable to various industrial-grade application scenarios such as intelligent operation and scheduling platforms for wind farms, wind power prediction service systems, and new energy grid-connected power regulation systems. It can significantly improve the accuracy of wind power prediction and system robustness, and is of great significance for improving wind power absorption capacity and grid connection stability.
Claims
1. A multidimensional wind power prediction method based on wind speed ramp identification and matching, characterized in that, Includes the following steps: Step 1: Preprocess the wind speed and power data of the wind farm: Use a spatiotemporal graph convolution filling model to handle data loss caused by sensor failure; The original data was standardized using dynamic quantile standardization; Coupled anomaly detection is used to detect anomalies in the raw data of sudden wind speed and power; finally, a novel mode decomposition algorithm, VMD-I, is used. Step 2: Use reinforcement learning to optimize the dynamic window width, add spatiotemporal factors and jointly analyze and screen the extreme point rate, add a physical information-based mutation event classifier to form a dynamic adaptive wind speed mutation identification algorithm, and use the dynamic adaptive wind speed mutation identification algorithm to denoise the wind speed data. Step 3: Expand the dimensions of meteorological factors and use a wind speed time period matching algorithm to match wind speed data with historical wind speeds. Step 4: Based on the Informer algorithm with attention mechanism for multi-dimensional data fusion, predict the future wind power generation capacity; Step 2 includes: Step 2-1, Extracting Extreme Values: For each sub-signal IMF, calculate and extract the maximum value. Minimum value And all extreme points, all extreme points form a set. , This represents the nth extreme point extracted; Step 2-2: Remove the pseudo-extreme points in each sub-signal IMF; Given the excessive density of extreme points during the climbing phase, false poles are removed: Calculate dynamic window width based on adaptive coefficients : , in The maximum wind speed in the sub-signal IMF. The minimum wind speed in the sub-signal IMF. Indicates the adaptive coefficient; Steps 2-3, Filtering Extremes: For each extreme point, calculate the Euclidean distance between the extreme point and its nearest neighbor. If the Euclidean distance is greater than the dynamic window width... ,satisfy If the extreme point is retained, it will be counted as a valid extreme point. Otherwise, discard the extreme point; where Represents the k-th extreme point and the (k+1)th extreme point Euclidean distance; Finally, all the extreme points that satisfy the conditions are retained, resulting in the set of effective extreme points. ,in This represents the f-th effective extreme point; Steps 2-4: Based on the selected extreme point set output by the adaptive pole selection model, calculate the pole rate according to the number of selected poles. This enables the selection of sub-signal data from the VMD mode decomposition results for the identification of sudden wind speed events, and the calculation of the pole rate of each sub-signal IMF component. : , in Indicates the first Extreme rate of wind speed data This represents the total number of reconstructed extreme points in the preprocessed wind speed data. Indicates the first The total number of reconstructed extreme points for each wind speed data point is used to filter the wind speed data based on the extreme point rate of each data point, ensuring that the data meets the requirements. Wind speed data less than or equal to a threshold are retained, resulting in a retained wind speed dataset. ; Steps 2-5, wind speed reconstruction: The retained wind speed dataset The wind speed data is reconstructed by overlaying, and the reconstructed wind speed data is denoted as... : , Where sum refers to the superposition function; The reconstructed wind speed data; This represents the retained wind speed dataset.
2. The method according to claim 1, characterized in that, Step 1 includes: Step 1-1: Collect wind speed and wind power data from the wind farm, and divide the data into past data and historical data; Collect multimodal meteorological data, including numerical weather prediction data for the wind field location; Steps 1-2 involve preprocessing past and historical data, including spatiotemporal graph convolutional filling model, dynamic quantile normalization, coupling anomaly detection, and VMI-I mode decomposition. Step 1-2-1 involves inputting the wind speed and power data from the wind farm into a spatiotemporal graph convolutional filling model for processing, including the following steps: Step 1-2-1-1, input the following data: wind speed V, power P, unit coordinates ,wind direction and missing mask ;in Represents the space of real numbers; Represents the two-dimensional spatial coordinates of the i-th unit; N represents the total number of units; T represents the number of time steps; Step 1-2-1-2: Construct a dynamic graph structure and build node features. The feature vector of each node contains both static and dynamic information. , in It is the input feature vector used for node i at time t in wind power forecasting; Let be the wind speed of the i-th unit at time t; Let be the power generation capacity of the i-th unit at time t; Let be the altitude of the location of the i-th unit; For the t-th wind turbine unit, there is the model or category. To convert the wind direction angle to a cosine value; To convert the wind direction angle to a sine value; Next, the edge weights are dynamically updated to reflect spatiotemporal correlation: , in Indicates edge weight; Indicates the spatial attenuation radius. Let be the wind direction angle of the i-th unit at time t; exp represents the natural exponential function; Steps 1-2-1-3, designing the spatiotemporal graph convolutional filling model, include the following steps: Step 1-2-1-3-1: Establish a spatial graph convolutional network. The input feature matrix of the spatial graph convolutional network is... , The input feature matrix is the l-th layer; The input feature dimension is ; the spatial graph convolutional network outputs a feature matrix of . , The output feature matrix of the l-th layer, To define the output feature dimension, a dynamic adjacency matrix is constructed based on the input and output feature matrices. For time step t, the adjacency matrix... The element is defined as: , in, The maximum neighborhood radius; Representing the adjacency matrix The element in the i-th row and j-th column; Let be the wind direction angle of the j-th unit at time t; Next, we perform degree matrix normalization. for: , in This represents the graph adjacency weight between node i and node j at time step t; This represents the sum of the connection weights of node i; Represents a diagonal matrix function; It is a matrix with non-zero elements only on its diagonal; Perform symmetric normalization: , in This is the normalized adjacency matrix; Finally, graph convolution is performed, where the mathematical expression for a single-layer spatial graph convolution is: , in, To train the weight matrix for the lesson, For the jump connection weight matrix, Use the LeakyReLU activation function; Step 1-2-1-3-2 introduces a time attention mechanism, including the following steps: Step 1-2-1-3-2-1, Perform time attention scoring: For each node Define the query vector as The key vector is The value vector is , where X is the input feature tensor of the entire model; B is the batch size; and d is the projection dimension; It is a learnable projection matrix that maps input features to a vector of query Q, key K, and value V; the attention weights are calculated using the following formula: , in This indicates that the nth node is in the nth position. The time step for the first Attention score at each time step; n is the node number; These are the indices for the current time step and other time steps, respectively. For the nth node in the nth position The query vector at each time step; For the nth node in the nth position The key vector at each time step; F represents the original feature dimension; For all Perform softmax weight normalization: , in For node n in the th... The time step for the first Attention weights for each time step; This indicates that the nth node is in the nth position. Each time step scores the attention of the k-th time step; Finally, the attention-weighted output is: , in Indicates that node n is in time Temporal feature representation after attention-weighted fusion; For value vectors, Is node n in time Embedded representation on; Indicates the nth node at time Regarding time Attention weights; Step 1-2-1-3-2-2, the overall temporal attention matrix is represented as follows: , , Where A is the attention matrix of each node at each time step to other time steps; Z is the output of the time attention mechanism; Step 1-2-1-3-3, further model the time dependency using a gated loop unit: , , , , in, This is the input for the t-th time step; This is the hidden state from the previous moment; For the updates; To reset the door; This is the candidate hidden state; This represents the final hidden state at the current time step. These are learnable linear transformation weights; Use the Sigmoid activation function; It is the hyperbolic tangent activation function; The product of Hadamard; This involves concatenating vectors. Steps 1-2-1-4: Establish the physical constraint loss function and jointly optimize intermediate parameters. and To implement the physical constraints of the spatiotemporal graph convolution filling model, power curve loss. for: , in It is the power of node i at time t predicted by the spatiotemporal graph convolutional filling model. Let i be the wind speed at time t. It is a function of wind speed and wind power obtained from experience or fitting. This represents a set of target points for prediction where wind speed is known, power is missing, or is uncertain. Total loss function of spatiotemporal graph convolutional filling model for: , in This represents the weight hyperparameter; It is the standard reconstruction loss, calculated using the following formula: , in The set of locations that can be monitored by real values; These are the predicted values from the spatiotemporal graph convolutional filling model. This represents the actual value from the wind speed data; Step 1-2-2, dynamic quantile standardization, includes the following steps: Step 1-2-2-1: Construct a sliding window centered on the current time point for the wind speed data at each moment. ,in This represents a local time window centered at the current time t, where w is the window length. This indicates a round-down operation; Indicates the offset from the current time t towards the future. Wind speed values at each time step; Step 1-2-2-2, calculate the quantiles within the window: , , , in , , These represent the 25th percentile, median, and 75th percentile of the data within the current window, respectively. Represents the quantile function; Steps 1-2-2-3: Calculate the quantile interval: ,in, Indicates the interquartile range; Step 1-2-2-4: Analyze the wind speed data at the current moment. For standardization, choose any of the following formulas: Median symmetric normalization: ; Min-Max normalization: ; Logarithmic compression normalization: ; in, This represents the standardized wind speed or meteorological variable value. A smoothing constant introduced to prevent division by zero errors; Step 1-2-2-5, Anomaly Detection Rules: If the following conditions are met: or If , then the data point at time t is determined to be a local outlier. Steps 1-2-3 involve preprocessing the wind speed time series using the improved Variational Mode Decomposition (VMD-I) algorithm.
3. The method according to claim 2, characterized in that, Steps 1-2-3 include the following steps: Step 1-2-3-1: Use the original VMD algorithm to process the wind speed sequence. Decomposed into K modal functions ; Steps 1-2-3-2, for each mode function Calculate the following indicators: , , , in For spectral entropy; This refers to the percentage of energy used. To reconstruct the correlation coefficient; Indicates the proportion of energy in the frequency component; Represents the j-th mode function; corr represents the correlation coefficient function; Steps 1-2-3-3: Calculate the modal validity score and reconstructed signal output: , , , , , in Non-negative weights; This indicates the validity of the k-th mode; For threshold; To reconstruct the signal and obtain the noise-suppressed main wind speed signal; This represents the set of modal indices that satisfy the validity score being greater than or equal to the threshold.
4. The method according to claim 3, characterized in that, Step 3 includes: Step 3-1: Construct statistical indicators for the selected wind segments within the current time window, including the following steps: Step 3-1-1: Calculate the gradeability coefficient for each wind speed range. , in This represents the magnitude of wind speed variation, where 'a' represents the sequence number of all time points within the time period 'c', and 'c' represents the time period of extreme points. Numerically, it is the product of the number of time series points 'a' and the fixed time step 'Δt'. This represents the slope coefficient within time period c. The ramp factor represents the ramp factor at each point, which is numerically equal to the ramp factor at each time series point. The derivative; Step 3-1-2: Construct the wind speed segment for the current time window, and calculate the mean wind speed, mean acceleration, average intensity, standard deviation, coefficient of variation, and upward trend ratio for the wind speed segment. , , , , , , in Let L be the wind speed sequence of length L; Accel is the mean acceleration of the wind speed segment; M is the average intensity of the wind speed segment. The average wind speed of the wind speed range. Standard deviation; The coefficient of variation; The ratio of upward trends. For indicator functions; Step 3-1-3: Organize the past wind speed segments to be matched. and Historical wind speed ranges matched after wind speed matching algorithm The steps include: Step 3-1-3-1: Calculate the wind speed intensity difference. : , in Indicates the wind speed sampling interval. Indicates the first Historical wind speed data at any given time. Indicates the first Historical wind speed data at any given time, wind speed intensity difference Normalization interval set to ; Step 3-1-3-2, calculate the wind speed trend difference. : , in, Indicates the first Historical wind speed data at any given time This represents the past wind speed data at time a+1; the resulting wind speed trend difference... Perform normalization processing, and set the normalization interval to... ; Step 3-1-3-3, calculate the wind speed similarity coefficient: , in Taking into account both the magnitude and trend differences in wind speed; Step 3-1-4: Integrate the statistical features to form a wind speed statistical feature vector. : , Then The established wind speed statistical features are embedded into the network module: , , , , , , , in It is the input tensor after tensor reshaping; For activation functions; This represents a one-dimensional convolution operation; This is the output of the convolutional layer; This is a one-dimensional max pooling operation; This is the result of pooling; For flattening operation; The flattened vector; This is the weight matrix of the fully connected layer; For bias terms; This is the output of the first fully connected layer; This represents the final embedded wind speed statistical features; This represents the weight matrix of the second fully connected layer. This represents the shape reshaping operation, where s represents the original statistical feature vector. Indicates the size of the convolutional pooling kernel. Indicates the number of output channels; Wind speed statistical feature vector The vectors are processed through tensor reshaping, convolutional layers, and pooling, flattened into vectors, and then fed into two fully connected layers to obtain the final output. ; Step 3-2, design the wind speed matching difference feature vector, including the following steps: Step 3-2-1: Construct a set of historical candidate segments, including historical wind speed data. Divide the data into sliding windows of length L, and extract wind speed segments of length L. ; This represents the Nth sliding window segment in the historical wind speed data; x(t) represents the wind speed sequence. Step 3-2-2: Construct the wind speed matching difference feature vector: , Where S is the average slope and D is the amplitude change, the calculation formula is: , ; Step 3-2-3: Perform sliding extraction on the historical wind speed sequence to obtain historical wind speed segments. ,calculate Corresponding feature vector And calculate the similarity metric: , in , These are weighting coefficients. ; Indicates current wind speed segment With historical fragments The overall similarity score; Used to measure the similarity in shape between two time series; historical wind speed segments obtained by similarity matching of wind speeds in the current time period; For the feature extraction function of composite slope climbing and intensity, select Extract the power sequence from the smallest matching segment. ; Step 3-2-4: Design the power response feature vector, calculate the power response feature vector Feat(y), and design the fused mode vector, defined as: , , , , , in Average power; This represents the maximum power change rate. The mean of the slope of change; The ratio of fluctuation amplitude; The number of mutations. The mutation threshold, The loss function; This represents the wind power value corresponding to the i-th time step; This represents the time interval between two adjacent time steps; This represents the maximum value in the wind power segment y; This represents the minimum value in the wind power segment y; Step 3-3: Fuse to obtain feature vectors : , Multilayer Perceptron (MLP) is used to upscale vectors to a unified vector dimension. : , in The power response features after unifying the vector dimensions; Steps 3-4 involve expanding the dimensions and weighting the meteorological factors, and constructing a feature matrix by combining the derived features.
5. The method according to claim 4, characterized in that, Steps 3-4 include the following steps: Step 3-4-1: Collect meteorological data including wind speed, cloud cover, air pressure, temperature, humidity, and irradiance; Step 3-4-2: Perform data preprocessing, using interpolation and mean to fill missing values, converting data of different dimensions to the same scale, and performing Z-score normalization on each column of variables. , , , in, This represents the normalized value. This represents the t-th data in the i-th row; This represents the sample mean of the variable in the i-th row; Represents the sample standard deviation of the variable in the i-th row; After processing, the normalized matrix Norm is obtained: , in This represents the standardized value of the nth variable at time step t; Step 3-4-3: Use a correlation matrix to analyze the correlation between NWP meteorological factors and wind power of wind farms, and use the Pearson correlation coefficient to measure the correlation between elements of the multivariate matrix. , , Where R represents the correlation coefficient matrix; This represents the Pearson correlation coefficient between variable i and variable j; Pearson correlation analysis was used to quantify the statistical linear relationship between various meteorological factors and wind power. Wind speed, irradiance, and temperature, which have the highest correlation with wind power, were selected as the meteorological element feature vectors, specifically represented as follows: , in Represents a meteorological element matrix. , , Let i represent the wind speed, temperature, and radiation at time i, respectively.
6. The method according to claim 5, characterized in that, Step 4 includes: Step 4-1: Design a feature embedding subnetwork. The feature embedding subnetwork adopts a multimodal feature fusion prediction method based on wind speed similarity matching vector. In view of the problem that wind power sudden change event features are obvious and traditional single input is difficult to model, four types of feature vectors are introduced: wind speed statistical feature vector Stat, wind speed matching difference feature vector Diff, power response feature vector Resp, and meteorological element feature vector NWP. The feature embedding subnetwork includes a wind speed statistical feature subnetwork, a wind speed matching difference feature network, a power response feature network, and a meteorological element feature network; A wind speed statistical feature subnetwork is established. This subnetwork constructs feature vectors using the statistical features of wind speed time series data. A nonlinear feature mapping is achieved based on a feedforward neural network structure with two or more layers. Finally, a unified-dimensional deep wind speed statistical representation is output, expressed as: First layer of feedforward neural network: , in This represents the input wind speed statistical feature vector; This represents the weight matrix of the first layer; Indicates the bias term; Indicates the activation function; This is the hidden vector output from the first layer; This refers to the output dimension of the first hidden layer. Second layer feedforward neural network: , in This represents the weight matrix of the second layer. This is the second-level bias term; This is the nonlinear output of the second layer; Indicates the output dimension of the second hidden layer; Execution layer normalization: , in For vectors The mean; Representing vectors The variance; It is a constant; These are learnable scaling and offset parameters; Represents a vector The result after the normalization operation at the execution layer; Output wind speed statistical feature embedding representation : , Step 4-2: Establish a wind speed matching difference feature network. This network employs a two-layer one-dimensional convolutional neural network structure to extract local difference change patterns and obtains a global representation through global average pooling. The wind speed matching difference feature vector is then used. : Input vector: , in ; To match the power sequence of historical wind speed ranges; Differential power; For wind speed matching, use the difference feature vector; This represents the wind power response sequence corresponding to the most matching historical wind speed segment at the current moment; The first layer of the convolutional network extracts local patterns by performing convolution on the wind speed matching difference feature vectors: , in It is an intermediate representation after the first convolutional layer. This represents the first layer of one-dimensional convolution. This represents the number of output channels for the first convolutional layer. The second convolutional network further extracts deep local structural information: , in, This represents the number of output channels for the second convolutional layer. It is the intermediate representation after the second convolutional layer. This represents the second layer of one-dimensional convolution; Perform global average pooling: , in This is the output feature vector after global average pooling; This represents the output features of the second convolutional network at time step t; The output mapping layer takes the wind speed matching difference feature vector as input to the fully connected layer and maps it to a d-dimensional embedding vector. : , in This is the weight matrix; For bias terms; Step 4-3: Establish a power response feature network based on the wind speed matching difference feature network. This power response feature network integrates the power response features into a rate response feature vector. : ; Power response feature vector Input power response feature network: , , , in This is a learnable third-layer weight matrix; This is the third-level bias term; The final output power response embedding vector; Final output embedding vector As a power response feature vector; Step 4-4: Construct a meteorological element feature network. This network is used to extract important weather background features that affect wind speed and power response, and includes the following steps: Step 4-4-1, Organize the input features: , in This represents the input multidimensional meteorological element data matrix; Step 4-4-2: The input passes through two convolutional layers, and global average pooling is performed on the time dimension. , , , in, For global average pooling; This represents a global representation of meteorological characteristics obtained by compressing the time dimension. Step 4-4-3, project as modality embedding vector: , , Where W and b represent the learnable weight matrix and bias, respectively; The weights are obtained after Softmax normalization. ; This represents the attention weight at time point t. This represents the meteorological information obtained by incorporating the attention mechanism; These are the parameters for linear transformation; This is the final meteorological feature vector; In mathematical or engineering expressions, it is used to introduce the definition or additional description of variables; Steps 4-5 involve building the main network of the prediction model, including the following steps: Step 4-5-1, Prepare the Informer main model, including the following steps: Step 4-5-1-1, Prepare the input event sequence ; Steps 4-5-1-2 involve passing the vectors to the encoder and decoder in the Informer; this includes the following steps: Step 4-5-1-2-1, Vector input encoder: The input first undergoes linear mapping and positional encoding to form an embedding. : , Including the historical time steps of Step; The encoder in Informer consists of two or more encoder blocks, each of which contains a sparse probabilistic multi-head attention mechanism and a feedforward neural network (FFN). The sparse probabilistic multi-head attention mechanism is used to replace the standard multi-head attention mechanism; the calculation formula for the feedforward neural network FFN is: , in This represents a feedforward neural network; x represents the input vector. And combine residuals and layer normalization: , in This indicates the final output of the current encoder; Representation layer normalization; The encoder adds a sequence compression module after each encoder block to achieve hierarchical abstraction: , in This represents the feature representation output after sequence compression. Indicates max pooling; Represents one-dimensional convolution; or represents the choice of neural network. Step 4-5-1-2-2: The Informer decoder is used to convert the temporal features extracted by the encoder into prediction results for two or more future time steps. It adopts an improved mechanism based on the Transformer architecture, supports parallel prediction, and includes an input concatenation module, a multi-layer decoder block, and a linear mapping layer. The decoder receives a sequence composed of historical true values and placeholder prediction vectors, with a length of [length missing]. ,in This is the known part. This is the part to be predicted; The multi-layer decoder block includes a self-attention module, a cross-attention module, and a feedforward neural network; The self-attention module is used to model the self-dependencies of sequences within the decoder; The cross-attention module is used to align the encoder output with the decoder input to achieve information flow transmission; The feedforward neural network is used to extract nonlinear features, and residual connections and layer normalization are used to enhance training stability. The linear mapping layer is used to restore the final output of the decoder to the target dimension through a linear transformation, thereby achieving the predicted output: , in This represents the final prediction result vector; Represents a linear mapping layer; Step 4-5-2, introduce a gating mechanism: , in, This represents the encoder input after fusing modal information; This represents the original encoder input feature sequence; This represents the weight matrix in the gating mechanism; Represents the modality fusion feature vector; This represents the bias parameter in the gating mechanism; Step 4-5-3, Layered Injection: The modality fusion vector is injected into each layer of the Informer encoder, as shown below: , in This refers to the Informer encoder module; This represents the modal fusion vector injected into the l-th layer; Each floor uses a different projection: , in This indicates the independent fully connected projection layer used by layer l; This represents the fused modal feature vector; Step 4-5-4, Modality selection attention mechanism: For each modality feature, attention weights are constructed and then weighted and combined. , in represents the attention weight of the i-th modality; w represents the learnable attention projection vector; Represents the weight matrix in learnable attention; This represents the embedding vector of the i-th mode; This represents the embedding vector of the j-th mode; Final fusion vector: ; Step 4-5-5 involves performing sub-network fusion representation, including the following steps: Step 4-5-5-1, Feature Concatenation and Fusion: The input wind speed sequence is passed through the wind speed statistical feature sub-network, the wind speed matching difference feature network, the power response feature network, and the meteorological element feature network, respectively, to obtain the output vectors of the four sub-networks: , in This is the statistical feature vector of wind speed; For wind speed matching, use the difference feature vector; This is the power response feature vector; For meteorological element feature vectors; The feature vectors are concatenated along their feature dimensions to construct a multimodal fusion feature vector: , Linearly mapped to a vector with the same embedding dimension as the original encoder. , in Represents the modal fusion vector. , For trainable parameters, , ; This represents the dimension of each modality embedding vector; Step 4-5-5-2, Embedding Projection Mapping: In order to match the modality vector dimension with the encoder main network input, a fully connected layer is needed to map the fused feature vector to the embedding dimension d; Step 4-5-5-3, Broadcast and Fusion Injection.
7. The method according to claim 6, characterized in that, Step 4-5-5-3 includes: merging the modality fusion vector Repeated replication to construct a time-consistent mode matrix : , in This means that the vector will be copied T times along the time dimension; modality matrix With encoder input embedding By adding element by element, we obtain the injected modulated input representation: , Finally, the input sequence It serves as input to the Informer encoder for subsequent modeling.
8. An electronic device, characterized in that, It includes a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 7.
9. A storage medium, characterized in that, It stores a computer program or instructions that, when run on a computer, perform the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Support vector machine wind power prediction device
CN114727560A
Wind power plant climbing event prediction method based on data enhancement
CN117117968A