Bayesian fusion-based wind farm ultra-short-term power prediction method and system

By employing a graph Bayesian fusion-based method for ultra-short-term wind power prediction, and utilizing multi-source data standardization and graph convolution propagation techniques, the problem of high accuracy and high efficiency in ultra-short-term wind power prediction is solved, achieving robust prediction under complex environments.

CN122495321APending Publication Date: 2026-07-31NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF INFORMATION SCI & TECH
Filing Date
2026-04-24
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing wind power forecasting methods struggle to simultaneously meet the demands for high accuracy and efficiency when dealing with ultra-short-term forecasts. In particular, their forecasting accuracy is limited in complex environments, and the calculation process is time-consuming.

Method used

A graph Bayesian fusion-based method for predicting ultra-short-term power in wind farms is adopted. This method collects and standardizes multi-source monitoring data, constructs a wind turbine graph structure, performs graph convolution propagation, extracts fused feature vectors, and calculates wind speed probability prediction and power prediction deviation to achieve ultra-short-term power prediction.

Benefits of technology

It improves the accuracy and stability of ultra-short-term power prediction for wind farms, enhances the prediction accuracy and spatial generalization ability under complex wind field conditions, provides a confidence interval reference for wind speed prediction, and ensures the reliability and practicality of the prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122495321A_ABST
    Figure CN122495321A_ABST
Patent Text Reader

Abstract

This invention relates to the field of wind power prediction, and more particularly to a method and system for ultra-short-term power prediction of wind farms based on graph Bayesian fusion. The method includes the following steps: collecting multi-source monitoring data of wind turbines in a wind farm; standardizing the multi-source monitoring data to obtain a standard time-series dataset; performing continuous time-series segmentation on the standard time-series dataset to construct a standardized sample tensor set; identifying the location coordinates of the wind turbines in the wind farm; constructing a wind turbine graph structure based on the location coordinates; injecting the standardized sample tensor set into the wind turbine graph structure for graph convolution propagation to obtain a fused feature vector; and performing wind speed probability prediction based on the fused feature vector to obtain the wind speed prediction result. This invention improves the accuracy and efficiency of wind farm power prediction and optimizes the reliability and practicality of the prediction model in actual wind farm operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wind power prediction, and in particular to a method and system for ultra-short-term power prediction of wind farms based on graph Bayesian fusion. Background Technology

[0002] Compared to traditional short-term or medium-to-long-term power forecasting, ultra-short-term power forecasting typically refers to wind power forecasting for the next few minutes to hours. It is of significant importance for real-time grid dispatching, rapid response to load fluctuations, and the optimization of wind farm operations. Wind power is significantly affected by meteorological conditions (such as wind speed, wind direction, and temperature) and topographic factors, exhibiting rapid changes and nonlinear characteristics, leading to considerable uncertainty and complexity in ultra-short-term forecasting. Existing wind power forecasting methods mainly include physical model-based methods, statistical methods, and machine learning methods. Physical models rely on wind farm meteorological data and turbine characteristics, providing basic trends in wind power output, but their prediction accuracy is low when facing rapid changes in local wind fields. Statistical methods use historical power data for time-series modeling, which can capture certain patterns, but their prediction accuracy remains limited in complex environments, and the computation process is time-consuming and inefficient. Traditional machine learning methods, while capable of handling nonlinear problems, often require extensive feature engineering and manual intervention, and their real-time prediction efficiency is also insufficient, failing to meet the dual requirements of high accuracy and high efficiency for ultra-short-term wind farm power forecasting. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention proposes a method and system for ultra-short-term power prediction of wind farms based on graph Bayesian fusion, thereby solving at least one of the aforementioned technical problems.

[0004] To achieve the above objectives, this invention provides a method for ultra-short-term power prediction of wind farms based on graph Bayesian fusion, comprising the following steps: Step S1: Collect multi-source monitoring data of wind turbines in the wind farm, and standardize the multi-source monitoring data to obtain a standard time series dataset; Step S2: Perform continuous temporal segmentation on the standard time series dataset to construct a standardized sample tensor set; Step S3: Identify the location coordinates of the wind turbines in the wind farm; construct a wind turbine diagram structure based on the location coordinates; Step S4: Inject the standardized sample tensor set into the wind turbine graph structure and perform graph convolution propagation to obtain the fused feature vector; Step S5: Perform wind speed probability prediction based on the fused feature vector to obtain the wind speed prediction result; Step S6: Calculate the power prediction deviation and perform directional training correction on the wind speed prediction results to execute the ultra-short-term power prediction operation.

[0005] This specification provides a wind farm ultra-short-term power prediction system based on graph Bayesian fusion, used to execute the wind farm ultra-short-term power prediction method based on graph Bayesian fusion as described above, including: The preprocessing module is used to collect multi-source monitoring data of wind turbines in wind farms, and to standardize the multi-source monitoring data to obtain a standard time series dataset. The segmentation module is used to perform continuous temporal segmentation on the standard time series dataset and construct a standardized sample tensor set. The graph creation module is used to identify the location coordinates of wind turbines in a wind farm; and to construct a wind turbine graph structure based on the location coordinates. The convolution module is used to inject the standardized sample tensor set into the wind turbine graph structure for graph convolution propagation to obtain the fused feature vector; The prediction module is used to perform wind speed probability prediction based on the fused feature vector to obtain the wind speed prediction result; The training correction module is used to calculate the power prediction deviation and perform targeted training correction on the wind speed prediction results in order to perform ultra-short-term power prediction operations.

[0006] The specific steps of step S1 are as follows: Collect multi-source monitoring data from wind turbines in wind farms; the multi-source monitoring data includes raw SCADA observation data, reanalysis meteorological data, and topographic raster data; Extract the fan number and time coordinates from the multi-source monitoring data; Based on the wind turbine number and time coordinates, a standardized data table is obtained by uniform standardization. Standard preprocessing is performed on the standardized data table to obtain a standard time series dataset.

[0007] Furthermore, the standard preprocessing specifically includes: Identify wind speed and power fields based on standardized data tables; Perform physical constraint range checks on the wind speed and power fields, mark outliers and remove them; Detect the temporal continuity of the wind speed and power fields; Based on the time continuity, abnormal fields are distinguished to obtain short-term missing measurement segments and long missing measurement segments that disrupt continuity. Linear interpolation is used to repair short-term missing measurement segments; the long-term missing measurement segments are truncated. Physical constraint range checks and time continuity verifications are performed on wind speed and power fields. Outliers are removed and short-term missing data segments are repaired using linear interpolation. Long missing data segments that disrupt continuity are directly truncated.

[0008] Furthermore, step S2 specifically involves the following steps: Set both the fixed input duration and the prediction duration as sliding windows; The standard time-series dataset is continuously segmented into time sequences according to the sliding window to obtain supervised learning samples. Robust normalization is performed on the dynamic input channel and target power channel of the supervised learning samples respectively, and the normalizer parameters are saved; Time series are partitioned based on normalizer parameters to construct a standardized sample tensor set; the standardized sample tensor set includes a training set, a validation set, a test set, and a time series K-fold set.

[0009] Furthermore, step S3 specifically involves the following steps: Identify the location coordinates of wind turbines in a wind farm; Based on the location coordinates, perform wind turbine layout analysis and construct a node set; Based on the location coordinates and the distance between computer groups, the angle between the wind turbine and the main wind direction is identified; The side set is determined based on the distance between the units and the angle between the fan and the main wind direction; Construct the wind turbine graph structure based on the set of nodes and the set of edges.

[0010] Furthermore, the specific steps of step S4 are as follows: One-dimensional convolution, dilated convolution, and residual stacking are performed on the standardized sample tensor set to extract local change trends and multi-scale temporal features. By coupling local change trends and multi-scale temporal features, time-coded features are obtained. The time-encoded features are injected into the wind turbine graph structure and graph convolution propagation is performed to obtain the wind turbine graph propagation features; Multi-head attention cross-modal fusion is performed on the propagation features and time-coded features of the wind turbine diagram, followed by attention enhancement and normalization to obtain the fused feature vector.

[0011] Furthermore, the specific steps of step S5 are as follows: A time series folding model is constructed by training the model based on fused feature vectors. Wind speed probability prediction is performed based on a time series model, generating multiple prediction results; The wind speed prediction result is obtained by performing second-order moment fusion on multiple prediction results.

[0012] Furthermore, the model training calculation specifically includes: The mean prediction head and variance prediction head calculate and output the fused feature vector to obtain the predicted mean, predicted standard deviation, and optional mixture distribution parameters; The first stage of training involves optimizing the MSE loss for the predicted mean. The second phase of training involves calibrating the predicted standard deviation using Gaussian negative log-likelihood loss. The optimal model parameters are output based on two-stage training. Construct a time series fold model based on the optimal model parameters.

[0013] Furthermore, the specific steps of step S6 are as follows: The latest raw SCADA observation data is collected based on the multi-source monitoring data; Based on the wind speed prediction results, the unit numbers are precisely aligned using the latest SCADA raw observation data, and the power prediction deviation is calculated to obtain error indicators and probability indicators. Based on the error index and probability index, a prediction deviation assessment is performed to obtain a prediction assessment report; Based on the prediction and evaluation report, the time series folding model is trained and corrected in a targeted manner, and an optimized prediction model is output to perform ultra-short-term power prediction tasks.

[0014] A graph Bayesian fusion-based ultra-short-term power prediction system for wind farms, used to execute the graph Bayesian fusion-based ultra-short-term power prediction method for wind farms as described in claim 1, comprising: The preprocessing module is used to collect multi-source monitoring data of wind turbines in wind farms, and to standardize the multi-source monitoring data to obtain a standard time series dataset. The segmentation module is used to perform continuous temporal segmentation on the standard time series dataset and construct a standardized sample tensor set. The graph creation module is used to identify the location coordinates of wind turbines in a wind farm; and to construct a wind turbine graph structure based on the location coordinates. The convolution module is used to inject the standardized sample tensor set into the wind turbine graph structure for graph convolution propagation to obtain the fused feature vector; The prediction module is used to perform wind speed probability prediction based on the fused feature vector to obtain the wind speed prediction result; The training and correction module is used to calculate power prediction bias and perform targeted training corrections on the wind speed prediction results in order to perform ultra-short-term power prediction operations. The beneficial effects of this invention are as follows: By collecting multi-source monitoring data and standardizing it uniformly, differences in sources and dimensions can be eliminated, improving data consistency and continuity, providing high-quality input for subsequent model training, and enhancing the stability and robustness of predictions. Continuous temporal segmentation constructs sample tensors, encoding historical dynamic information into trainable inputs, while normalization reduces the impact of outliers, making the model more accurate and robust in short-term predictions. Through wind turbine location identification and graph structure construction, the spatial relationships between wind turbines and the influence of wakes are introduced into the model, enabling information sharing between neighboring turbines and improving prediction accuracy and spatial generalization ability under complex wind field conditions. Graph convolutional propagation fuses temporal and spatial features, capturing the dynamic changes of the wind turbine itself and its neighbors, making the model more accurate under nonlinear fluctuations and local anomalies, and enhancing feature representation capabilities. Probabilistic prediction based on fused feature vectors simultaneously obtains the mean wind speed and uncertainty, improving point prediction accuracy while providing confidence interval references, which is helpful for extreme value fluctuation prediction and risk management. By evaluating power prediction bias and correcting it through targeted training, the model can be adaptively optimized, improving the accuracy and stability of power prediction and ensuring the reliability and practicality of ultra-short-term prediction in actual operation. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the steps of a wind farm ultra-short-term power prediction method based on graph Bayesian fusion according to the present invention. Figure 2 This is a detailed flowchart illustrating the implementation steps of step S1. Figure 3 This is a detailed flowchart illustrating the implementation steps of step S2; Figure 4 A flowchart illustrating the standard preprocessing process; Figure 5 This is a schematic diagram showing the geographical distribution of wind turbines; Figure 6 This is a schematic diagram for predicting wind turbine power. Figure 7 A comparison chart of the predictive capabilities of multiple models; Figure 8 A schematic diagram illustrating the training objective function; Figure 9 A diagram showing the comparison of error metrics for different models; Figure 10 This is a schematic diagram of multi-fold weighted integration; Figure 11 This is a schematic diagram of the architecture of a time series folding model. Detailed Implementation

[0016] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0017] This application provides a method and system for ultra-short-term power prediction of wind farms based on graph Bayesian fusion. The execution entities of the method and system include, but are not limited to, the following: mechanical equipment, data processing platform, cloud server node, network upload device, etc., which can be considered as general computing nodes of this application. The data processing platform includes, but is not limited to, at least one of the following: audio and image management system, information management system, and cloud data management system.

[0018] Please see Figures 1 to 11 This invention provides a method for ultra-short-term power prediction of wind farms based on graph Bayesian fusion, comprising the following steps: Step S1: Collect multi-source monitoring data of wind turbines in the wind farm, and standardize the multi-source monitoring data to obtain a standard time series dataset; Step S2: Perform continuous temporal segmentation on the standard time series dataset to construct a standardized sample tensor set; Step S3: Identify the location coordinates of the wind turbines in the wind farm; construct a wind turbine diagram structure based on the location coordinates; Step S4: Inject the standardized sample tensor set into the wind turbine graph structure and perform graph convolution propagation to obtain the fused feature vector; Step S5: Perform wind speed probability prediction based on the fused feature vector to obtain the wind speed prediction result; Step S6: Calculate the power prediction deviation and perform directional training correction on the wind speed prediction results to execute the ultra-short-term power prediction operation.

[0019] In the embodiments of the present invention, see Figure 1 The diagram below illustrates the steps of a wind farm ultra-short-term power prediction method based on graph Bayesian fusion according to the present invention. In this example, the steps of the wind farm ultra-short-term power prediction method based on graph Bayesian fusion include: Step S1: Collect multi-source monitoring data of wind turbines in the wind farm, and standardize the multi-source monitoring data to obtain a standard time series dataset; In this embodiment, multi-source data acquisition and unified preprocessing are performed on each wind turbine within the wind farm to form a high-quality standard time-series dataset. Data sources include raw SCADA operational data (sampling period 1 s to 10 s), reanalysis meteorological data (such as 10 m / 100 m wind speed, temperature, and air pressure, with a time resolution of 1 h), and terrain raster data (resolution 30 m, including elevation and roughness). After SCADA data is acquired in real-time via an industrial protocol, it is aggregated in 10-minute time windows to extract the mean, standard deviation, and extreme values. Meteorological data is mapped to wind turbine locations through spatial interpolation (such as bilinear interpolation) and upscaled to a 10-minute resolution through temporal interpolation (such as linear or spline interpolation). Terrain data is directly matched to each wind turbine node as a static feature. In the preprocessing stage, physical constraints are first checked (e.g., wind speed limited to 0–30 m / s, power limited to ±5% of rated value), and outliers are identified. Then, temporal continuity is checked; short-term missing data (≤3 time steps) are repaired using linear interpolation, and long missing data segments are truncated. Next, standardization processing (e.g., Z-score or quantile normalization) is performed on each feature, and the time index and turbine number are unified to construct a structured data table.

[0020] Step S2: Perform continuous temporal segmentation on the standard time series dataset to construct a standardized sample tensor set; In this embodiment, a sliding window mechanism is used to perform continuous temporal segmentation to construct the standardized sample tensor set required for supervised learning. The input window length L and prediction step size H are set, for example, L=6 (past 60 minutes) and H=2 (future 20 minutes), and the window is moved using a sliding step size stride=1. For any time point t, the data from its previous L time steps are extracted as input X(i,t−L+1:t), and the power or wind speed for the next H steps is used as the label Y(i,t+1:t+H). In a multi-wind turbine scenario, the data from all wind turbine nodes within the same time window are stacked to form a three-dimensional input tensor X∈R^(N×L×F), where N is the number of wind turbines (e.g., 50) and F is the feature dimension. Subsequently, robust normalization is performed on the samples, the median and interquartile range (IQR) are calculated based on the training set, and a uniform transformation is performed on all samples, while the normalization parameters are saved for subsequent use. To ensure the reliability of model evaluation, the samples are divided into training set (70%), validation set (15%) and test set (15%) in chronological order, and a time series K-fold set (e.g., K=5) is constructed for cross-validation.

[0021] Step S3: Identify the location coordinates of the wind turbines in the wind farm; construct a wind turbine diagram structure based on the location coordinates; In this embodiment, a sliding window mechanism is used to convert continuous time series into supervised learning samples. First, the input window length L and prediction step size H are set, for example, L=12 (representing the past 120 minutes) and H=3 (predicting the next 30 minutes). Then, the data for each wind turbine is chronologically segmented with a step size of 1, generating a large number of input-output sample pairs. The input data dimension is (number of wind turbines N × time step L × number of features F, for example, 50 × 12 × 30), and the output is the target variable (such as wind speed or power) for the next H steps. During the segmentation process, temporal continuity must be maintained, missing segments must be avoided, and samples with a missing percentage exceeding 5% must be filtered out. Robust normalization is performed on the samples (based on the median and IQR calculated from the training set), and the normalization parameters are saved for subsequent inference. The dataset is divided chronologically into a training set (70%), a validation set (15%), and a test set (15%), and a time series K-fold (e.g., K=5) cross-validation set is constructed.

[0022] Step S4: Inject the standardized sample tensor set into the wind turbine graph structure and perform graph convolution propagation to obtain the fused feature vector; In this embodiment, the standardized sample tensor set is fused with the wind turbine graph structure, and spatial-temporal coupled features are extracted through graph convolution propagation. For each time window sample, its input tensor X∈R^(N×L×F) is first encoded into a temporal feature representation H_t∈R^(N×L×C) by a temporal feature extraction module (such as one-dimensional convolution or dilated convolution), where C is the number of feature channels (e.g., 64). Subsequently, the node features of each time step are mapped to the graph structure, and graph convolution operations are performed based on the adjacency matrix A. For example, the GCN method is used, which weights and aggregates the features of neighboring nodes through a normalized adjacency matrix. During the calculation, two layers of graph convolution can be stacked, each followed by an activation function (ReLU) and a normalization operation (BatchNorm), and Dropout (e.g., 0.2) is set to prevent overfitting. To fuse temporal and spatial information, an attention mechanism (e.g., multi-head attention) can be introduced after graph convolution to achieve adaptive weighting of features across time and space.

[0023] Step S5: Perform wind speed probability prediction based on the fused feature vector to obtain the wind speed prediction result; In this embodiment, a dual-branch output structure is designed, including a mean prediction head and a variance prediction head. The mean head outputs the predicted mean μ(i,t+h) through a fully connected network (such as a two-layer MLP with 128 hidden layers), and the variance head outputs the standard deviation σ(i,t+h), which is guaranteed to be positive through the Softplus function. During training, a two-stage optimization strategy is adopted: the first stage uses MSE loss to optimize the mean prediction, so that the model accurately fits the expected wind speed; the second stage uses Gaussian negative log-likelihood (NLL) loss to calibrate the variance, so that the predicted distribution can reasonably reflect the uncertainty. Training uses the Adam optimizer (learning rate 1e-3→5e-4), batch size 64, training epochs of approximately 50-80, and combines an Early Stopping strategy. To improve model stability, time series K-fold training (e.g., K=5) is used to obtain multiple sub-models outputting multiple prediction results (μ_k, σ_k²).

[0024] Step S6: Calculate the power prediction deviation and perform directional training correction on the wind speed prediction results to execute the ultra-short-term power prediction operation.

[0025] In this embodiment, the predicted wind speed mean μ is input into the wind turbine power curve model (or empirical function) and converted into a predicted power value P_pred. This predicted power is then compared with the actual observed power P_real, and error indices such as MAE, RMSE, and relative error MAPE are calculated. At the probabilistic level, the prediction variance σ² is combined to calculate NLL and confidence interval coverage (e.g., 95% interval coverage) to assess the reasonableness of uncertainty. Subsequently, the error results are analyzed to identify areas of concentrated error (e.g., specific wind turbines or specific time periods), and a prediction evaluation report is generated. During the model correction phase, targeted training strategies are adopted, such as increasing the weight of high-error samples (e.g., weighting coefficient 1.5) or making local adjustments to the structure of specific wind turbine subgraphs. Simultaneously, the model undergoes incremental training (learning rate 1e-4, training for 10-20 rounds) to adapt to the latest data distribution changes. Within the time series model framework, poorly performing sub-models can be replaced or retrained.

[0026] In this embodiment, see Figure 2 The diagram below illustrates the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Collect multi-source monitoring data from wind turbines in wind farms; the multi-source monitoring data includes raw SCADA observation data, reanalysis meteorological data, and topographic raster data; Extract the fan number and time coordinates from the multi-source monitoring data; Based on the wind turbine number and time coordinates, a standardized data table is obtained by uniform standardization. Standard preprocessing is performed on the standardized data table to obtain a standard time series dataset.

[0027] In this embodiment, the SCADA system collects real-time operating status data through the wind turbine's built-in sensors. The sampling period is typically set to 1 s to 10 s. Key variables collected include wind speed (0–25 m / s), active power (0–rated power, such as a 2 MW unit), engine speed (5–20 rpm), pitch angle (0–90°), and nacelle temperature. The data is uploaded to the central server via an industrial communication protocol (such as OPC-UA), and a buffer queue (e.g., 1000 entries) is set up to prevent data loss due to momentary communication interruptions. For the meteorological data analysis, a data source such as ERA5 is selected to obtain meteorological variables such as wind speed, temperature, and air pressure at a height of 100 m, with a time resolution of 1 hour, which are periodically retrieved via an API interface. Subsequently, a bilinear interpolation method is used to map the meteorological grid (0.25° × 0.25°) to the central area of ​​the wind farm, and cubic spline interpolation is used to improve the time resolution to the 10-minute level. The terrain raster data was imported into the DEM model (30 m resolution) via a GIS system, extracting elevation, slope (0–30°), aspect, and surface roughness parameters for each wind turbine location. The three types of data were then integrated into a raw data pool through a unified data access interface (such as Kafka data streams or the time-series database InfluxDB). SCADA (Supervisory and Data Acquisition System) records raw observations such as wind turbine operating status, wind speed, and power. ECMWF / ERA5 (Reanalysis Meteorological Data Source) provides external meteorological drivers for the turbine locations.

[0028] The wind turbine equipment ID field (e.g., WTG_ID) is parsed from the raw SCADA data to establish a unique set of wind turbine numbers (e.g., WTG_01 to WTG_50), and a wind turbine topology mapping table is constructed for subsequent graph structure modeling. For meteorological data, which exists in a grid format, spatial matching methods (e.g., nearest neighbor search or KD Tree) are used to map each meteorological grid point to its corresponding wind turbine location, generating a "wind turbine-meteorological node" relationship. For time processing, a unified standard time system (e.g., UTC+8) is adopted, and timestamp formats from different data sources (Unix timestamps, string timestamps, etc.) are parsed and converted, with a unified time granularity Δt = 10 min. For high-frequency SCADA data, a time window aggregation method is used to calculate the mean, maximum, and standard deviation within each 10-minute window; for meteorological data, a linear interpolation method is used for time upsampling to align it with the SCADA data. A time matching tolerance (e.g., ±30 s) is set, and data exceeding the threshold is discarded to avoid mismatch.

[0029] For continuous variables (such as wind speed, power, and temperature), the Z-score normalization method is used to calculate the mean μ and standard deviation σ based on the training set (the first 70% of the time period), and all data are mapped to a standard normal distribution. For non-normally distributed features (such as the skewed distribution of power in low wind speed areas), logarithmic transformation or Box-Cox transformation is used for distribution correction. Due to the periodicity of wind direction variables, trigonometric function encoding is used to convert the original angle θ into two components, sin(θ) and cos(θ), thus avoiding abrupt angle changes. Topographic raster data, as static features, is mapped to the [0,1] interval using the Min-Max normalization method; meteorological variables are processed using either Z-score or quantile normalization methods according to their statistical characteristics. At the data fusion level, using (wind turbine number, time) as the primary key, SCADA, meteorological, and topographic features are horizontally concatenated to form a unified structured data table, with each record containing approximately 20–40 dimensions of features. For missing data, set a missing rate threshold (e.g., 10%), use K-nearest neighbor interpolation (K=5) or time series linear interpolation to complete low-missing data, and remove high-missing samples.

[0030] Let the fan assembly be The default implementation of the project . No. Typhoon machines at all times The original observations are recorded as follows:

[0031] in, Indicates wind speed. Indicates active power. This represents the static terrain vector of the wind turbine. This represents the timestamp. The meaning of this formula is that the original observation at any given time not only includes dynamic quantities, but also the terrain conditions with the camera position unchanged, which facilitates subsequent multi-source fusion.

[0032] The project first aggregates minute-level SCADA data into hourly-level data, using the following aggregation method:

[0033] in, This indicates the number of original sampling points included within one hour. Its engineering significance lies in eliminating the sampling frequency differences between different data sources by using a uniform hourly granularity.

[0034] Rule constraints are constructed based on the physical characteristics of the wind turbine. For example, when the wind speed is below 3 m / s, the power should be close to 0. If the deviation exceeds 20% of the rated power, it is judged as an anomaly. Statistical methods (such as the 3σ criterion) and unsupervised algorithms (such as Isolation Forest) are introduced to identify outliers in the multidimensional feature space, with the anomaly ratio typically controlled between 2% and 5%. Secondly, noise suppression is applied to the time series data using moving average filtering (window length 3–5) or Savitzky-Golay filtering to smooth the data and reduce the impact of high-frequency disturbances. In terms of sample construction, a sliding time window mechanism is used to generate supervised learning data. The input sequence length is set to L=6 (i.e., data from the past 60 minutes), and the prediction step size is H=1–3 (predicting power for the next 10–30 minutes). Each sample is represented as X(i,t−L+1:t) and labeled as Y(i,t+H). Temporal feature encoding (such as sin / cos representation of hourly periods) is also introduced to enhance the model's ability to characterize periodicity. Finally, the data is divided into training set (70%), validation set (15%), and test set (15%) according to time order to ensure that there is no time leakage problem, thus forming a standard time series dataset.

[0035] In this embodiment, the standard preprocessing specifically includes: Identify wind speed and power fields based on standardized data tables; Perform physical constraint range checks on the wind speed and power fields, mark outliers and remove them; Detect the temporal continuity of the wind speed and power fields; Based on the time continuity, abnormal fields are distinguished to obtain short-term missing measurement segments and long missing measurement segments that disrupt continuity. Linear interpolation is used to repair short-term missing measurement segments; the long-term missing measurement segments are truncated. Physical constraint range checks and time continuity verifications are performed on wind speed and power fields. Outliers are removed and short-term missing data segments are repaired using linear interpolation. Long missing data segments that disrupt continuity are directly truncated.

[0036] In this embodiment, wind speed and power fields are identified and extracted from a standardized data table. Based on the physical operating mechanism of wind turbines, a constraint rule system is established for the wind speed and power fields to identify abnormal data that clearly violate physical laws. For wind speed variables, a reasonable range is typically set at 0–30 m / s (the upper limit can be relaxed to 35 m / s to account for extreme weather redundancy). Data below 0 or above the upper limit is directly judged as abnormal. For power variables, the range is set at 0–2000 kW based on the rated capacity of the wind turbine (e.g., a 2 MW unit), with an allowable measurement error range of ±5%, i.e., data outside the range of [-100 kW, 2100 kW] is considered abnormal. At the same time, wind speed-power relationship constraints (i.e., power curve constraints) are introduced. For example, when the wind speed is lower than the cut-in wind speed (usually 3 m / s), the power should be close to 0 (less than 5% of the rated power). If there is a significant deviation (e.g., wind speed of 2 m / s but power reaching 500 kW), it is marked as a physical conflict anomaly. During implementation, data records can be scanned one by one based on the rule engine, and an exception flag field (flag=1 indicates an exception) can be generated.

[0037] A complete time index sequence is constructed based on a unified time step (e.g., Δt = 10 min), and the data for each wind turbine is time-aligned. The difference between adjacent timestamps, Δt_i = t_i - t_{i-1}, is calculated to determine if there are any time interval anomalies. When Δt_i > Δt (e.g., exceeding a tolerance of 600 s ± 30 s), a data gap is identified at that location. The location of the missing segment is further confirmed by combining this with field missingness (NaN value). In the implementation, a time tolerance threshold of ±30 s can be set to accommodate slight delays in actual data acquisition systems. For continuous detection results, the start and end times of each missing segment are recorded, and the length of the missing segment (e.g., the number of missing points N) is calculated. In experiments, short-term communication interruptions commonly result in the loss of 1-3 time steps, while equipment failures may lead to long-term interruptions exceeding 10 time steps.

[0038] A threshold N_th for the length of missing data is set (e.g., N_th=3, corresponding to a 30-minute time span). When the number of consecutive missing points N ≤ N_th, it is defined as a short-term missing data segment; when N > N_th, it is defined as a long-term missing data segment. The selection of this threshold is usually based on experimental experience and the sensitivity of the prediction model to continuity. For example, in ultra-short-term prediction (10–30 min) tasks, data missing exceeding 30 minutes will significantly affect the integrity of the temporal features, and therefore needs to be handled separately. In the implementation process, each missing data interval is traversed and statistically analyzed, and a missing data type label (type=short or type=long) is generated. To improve classification accuracy, the trend of data changes before and after the missing data can also be used for auxiliary judgment. For example, if the wind speed changes drastically before and after the missing data (more than 5 m / s), it can be marked as an unstable interval even if there are few missing data points.

[0039] For short-term missing data segments, linear interpolation is used for data reconstruction, which involves linear fitting based on valid data points at both ends of the missing data interval. For example, for a missing point x(t_k), its value can be estimated through the linear relationship between the previous time step x(t_{k-1}) and the next valid time step x(t_{k+m}), with the interpolation formula being a time-weighted average. This method has high accuracy over short periods and is suitable for scenarios where wind speed and power changes are relatively stable. The interpolation error is typically controlled within 5%. For long-term missing data segments, due to their large span and the inability to guarantee the linear assumption, truncation is used. This involves directly deleting the missing data segment and its adjacent range (e.g., one time step before and after) to avoid boundary effects interfering with the model. After truncation, the time series index needs to be reconstructed, and the continuity of the remaining data must be ensured.

[0040] All labeled outliers are uniformly processed, converted into missing values, and the missing segments are identified based on the temporal continuity detection results. Then, linear interpolation repair and data truncation are performed according to short-term and long-term missing value classification strategies, respectively. After processing, data consistency is checked, such as re-checking for values ​​exceeding physical limits and whether time interval anomalies still exist. Secondary verification thresholds can be set (e.g., wind speed limit 30 m / s, power limit ±5%) to ensure data quality. The final output data sequence should meet the following conditions: no obvious physical anomalies, continuous time series (except for truncated portions), and a missing value rate of less than 5%. This high-quality dataset will serve as a key input to the graph Bayesian fusion model, ensuring that the model is not affected by abnormal noise when learning the correlation structure and temporal dependencies between wind turbines, thereby improving the accuracy and stability of ultra-short-term power prediction.

[0041] Set retention criteria for each record:

[0042] in, The function represents the indicative function, taking the value 1 when the conditions are met and 0 when they are not. This formula indicates that the sample will proceed to the next step only if the wind speed, power, and non-empty conditions are simultaneously met.

[0043] For repairable missing sections, linear interpolation is used:

[0044] This formula represents the addition of values ​​proportional to time between two valid endpoints. If the missing measurement time span is too long, the project implementation will not add values ​​but will directly delete the corresponding window to prevent semantic misalignment of subsequent supervision samples.

[0045] In this embodiment, see Figure 3 The diagram below illustrates the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Set both the fixed input duration and the prediction duration as sliding windows; The standard time-series dataset is continuously segmented into time sequences according to the sliding window to obtain supervised learning samples. Robust normalization is performed on the dynamic input channel and target power channel of the supervised learning samples respectively, and the normalizer parameters are saved; Time series are partitioned based on normalizer parameters to construct a standardized sample tensor set; the standardized sample tensor set includes a training set, a validation set, a test set, and a time series K-fold set.

[0046] In this embodiment, based on the timescale characteristics of the wind farm's ultra-short-term power prediction task, the input duration (historical observation length) and prediction duration (future prediction span) of the sliding window are set. In engineering practice, the selection of the input duration L and prediction duration H typically requires a trade-off between "information sufficiency" and "model complexity." For example, in a data system with a 10-minute time resolution, the input window length L can be set to 6–12 minutes, corresponding to historical information from the past 60–120 minutes; the prediction duration H can be set to 1–3 minutes, corresponding to the power prediction target for the next 10–30 minutes. In experiments, L=6 and H=2 are often used as the baseline configuration, i.e., using data from the past hour to predict the power for the next 20 minutes. The sliding window moves using a stride, typically set to 1 (i.e., sliding once every 10 minutes) to maximize sample utilization; in scenarios with limited computational resources, stride=2 or 3 can be set to reduce the number of samples. To ensure window effectiveness, boundary constraints must also be set, i.e., samples are only generated when the length of the continuous time series is greater than L+H. For multi-wind turbine systems, windows can be constructed synchronously in the spatial dimension, meaning that each time window contains data from all wind turbine nodes, thus providing a unified input tensor for subsequent graph structure modeling.

[0047] Using (wind turbine number i, time t) as an index, the multivariate time series of each wind turbine is traversed. For any time point t, the data of the previous L time steps are extracted to form the input sequence X(i,t−L+1:t), and the power values ​​of the next H time steps are extracted as the prediction target Y(i,t+1:t+H). In multi-wind turbine scenarios, the data of all wind turbines within the same time window can be stacked to form a three-dimensional tensor structure X∈R^(N×L×F), where N is the number of wind turbines (e.g., 50), L is the time length (e.g., 6), and F is the feature dimension (e.g., 20-dimensional); the target tensor Y∈R^(N×H). During the segmentation process, it is necessary to strictly ensure temporal continuity, that is, only the data segments that have undergone prior processing (without long missing measurement segments) are sliced ​​to avoid constructing samples across breakpoints. To enhance the model's ability to capture dynamic changes, differential features (e.g., Δwind speed, Δpower) or time-encoded features (e.g., hourly sin / cos) can be added to the input features.

[0048] Robust normalization is performed on both the input features and the target variable. Unlike traditional Z-score normalization, robust normalization scales based on quantile statistics (such as the median and interquartile range, IQR), making it more suitable for wind power data exhibiting peaked and long-tailed distributions. For each feature channel f in the input tensor X, its median Med_f and interquartile range IQR_f (i.e., Q3−Q1) on the training set are calculated, and the data is converted to (x−Med_f) / IQR_f. For the target power channel Y, its Med_y and IQR_y are also calculated and normalized. The quantile range can be set to 25%–75%, and a smoothing term ε is added for features with excessively small IQR values ​​(e.g., less than 1e-3) to avoid numerical instability. It is important to emphasize that the normalization parameters are calculated only on the training set and saved as normalizer parameters (including Med and IQR vectors), which are then uniformly applied to the validation and test sets to avoid data leakage. In multi-fan scenarios, you can choose "global normalization" (all fans share parameters) or "local normalization" (each fan calculates parameters independently). Experiments show that local normalization is more effective when there are large differences between fans.

[0049] All samples are divided chronologically, typically using a ratio of 70% training set, 15% validation set, and 15% test set. For example, 4300 samples can be divided into three subsets: 3010, 645, and 645, ensuring fairness and generalization ability in model evaluation. To enhance the model's robustness across different time periods, a Time Series K-Fold cross-validation mechanism is introduced. For instance, setting K=5 divides the time series into five consecutive intervals, selecting one interval as the validation set each time and using the rest as the training set, while maintaining the chronological order. In terms of tensor construction, the normalized input data is organized into a four-dimensional tensor X∈R^(B×N×L×F), where B is the batch size of the samples; the target data is Y∈R^(B×N×H). To adapt to the graph Bayesian model structure, adjacency matrix or graph structure encoding information can also be added. The final output dataset includes training tensors, validation tensors, test tensors, and a K-fold training subset set, along with normalizer parameter files (such as Med and IQR vectors).

[0050] Given input window length and predicted length Construct supervised samples: ; in, It is the model input. It is the target sequence at H future time points. In engineering implementation, hourly data is input with 8-dimensional features by default, of which the first 6 dimensions are dynamic channels and the last 2 dimensions are static channels.

[0051] To ensure the temporal semantic continuity of window samples, the following requirements must be met: ; in, This is a fixed time interval. This constraint means that the window cannot contain skipped times, out-of-order sequences, or long gaps; otherwise, the entire window segment will be deleted.

[0052] To reduce the impact of outliers, robust normalization is used in the dynamic channel:

[0053] In the above formula, The median is used, and the denominator is the interquartile range. This method is more robust to extreme wind conditions than mean-variance normalization. The project saves both the input and output scalers for online denormalization.

[0054] In this embodiment, step S3 includes the following steps: Identify the location coordinates of wind turbines in a wind farm; Based on the location coordinates, perform wind turbine layout analysis and construct a node set; Based on the location coordinates and the distance between computer groups, the angle between the wind turbine and the main wind direction is identified; The side set is determined based on the distance between the units and the angle between the fan and the main wind direction; Construct the wind turbine graph structure based on the set of nodes and the set of edges.

[0055] In this embodiment, the precise geographical location information of each wind turbine within the wind farm is obtained, typically expressed in latitude and longitude, and uniformly using the WGS84 coordinate system as the spatial reference standard. Data sources can include the wind farm GIS system, SCADA configuration files, or engineering design drawings. Each wind turbine has a unique number (e.g., WTG_01 to WTG_50) along with its geographic coordinate information. To improve calculation accuracy, latitude and longitude coordinates can be converted to planar projected coordinates (e.g., UTM coordinate system) for subsequent distance calculations and spatial analysis. In practical processing, the conversion accuracy can be controlled within meter-level error (≤1 m). Elevation information for each wind turbine can also be extracted from terrain raster data as an auxiliary spatial feature. To ensure data consistency, the coordinate data needs to be validated, for example, checking for duplicate coordinates (two wind turbines coinciding) or significant offsets (e.g., exceeding the wind farm boundary). The wind turbine coordinates are visualized (e.g., a two-dimensional planar scatter distribution) to observe their arrangement characteristics, such as whether they are regular grid-like, row-column distributed, or irregularly distributed. Furthermore, clustering analysis methods (such as K-means or DBSCAN) can be used to spatially group wind turbines to identify locally dense regions or subarray structures. For example, K=5 can be set for region division, or a DBSCAN neighborhood radius ε=500 m can be set to identify local clusters. This analysis is helpful for subsequent edge weight design. During node construction, each node not only includes its spatial location but can also be appended with static attributes (such as altitude and surface roughness) and dynamic attributes (such as current wind speed and power), thus forming a multi-dimensional node feature vector. In the graph Bayesian model, nodes are usually represented as random variables whose states change over time. The number of nodes is the same as the number of wind turbines (e.g., N=50), and the node feature dimension can be set to F=30 (including meteorological and operational variables).

[0056] Based on the planar coordinates (X, Y) of the wind turbines, the distance d_ij between any two wind turbines is calculated using the Euclidean distance formula, i.e., d_ij = √[(x_i − x_j)² + (y_i − y_j)²]. In actual wind farms, the spacing between wind turbines is typically 300–800 meters, so the calculation results can be used to determine the range of mutual influence between units. To improve calculation efficiency, a distance matrix (N×N) can be constructed, for example, a 50×50 matrix can be generated for 50 wind turbines. Next, the dominant wind direction needs to be determined. This parameter can be obtained through statistical analysis of historical wind speed data, for example, by selecting the mode or weighted average of the wind direction data during the training period (e.g., the dominant wind direction is 270°, i.e., westerly). Subsequently, for any two wind turbines i and j, the direction angle θ_ij of the line connecting i to j is calculated and compared with the dominant wind direction θ_w to obtain the included angle Δθ = |θ_ij − θ_w|. This included angle is used to determine the upwind and downwind relationship between wind turbines: when Δθ is small (e.g., <30°), it means that j is located in the downwind region of i and may be affected by the wake.

[0057] Potential connections are filtered based on a distance threshold, for example, setting a distance threshold d_th = 1500 m. If the distance between two wind turbines is less than this value, a spatial association is considered to exist; otherwise, no connection is established. Next, edges are further filtered using wind direction angle information, for example, only retaining connections that satisfy Δθ < 30° to highlight the influence of the wake propagation direction. For wind turbine pairs (i, j) that meet the condition, a directed edge i→j is established (indicating that i has an influence on j). The edge weight can be calculated using a multi-factor fusion method, for example, defining the weight w_ij = exp(−d_ij² / σ²) × cos(Δθ), where σ is 1000 m to control the degree of distance attenuation, and cos(Δθ) to reflect wind direction consistency. The weight is maximum when Δθ is close to 0°, and approaches 0 when it is close to 90°. This method can construct a weighted directed graph structure. To avoid the graph being too dense, the maximum number of connections for each node can be limited (e.g., using the K-nearest neighbor method, K=5), retaining only the 5 closest neighbors. The final set of edges is E={(v_i, v_j, w_ij)}, which is usually on the order of N×K (e.g., 50×5=250 edges).

[0058] Construct an adjacency matrix A (N×N), where A_ij = w_ij represents the connection weight from node i to node j, with a value of 0 if there is no connection. For directed graphs, A is usually an asymmetric matrix; however, it can be symmetricized in some models (e.g., A = (A + Aᵀ) / 2). Subsequently, the adjacency matrix can be normalized, for example, by row normalization (making the sum of each row equal to 1) or symmetric normalization (D^(-1 / 2) AD^(-1 / 2)) to improve the numerical stability of graph neural networks or graph Bayesian models. In the graph Bayesian fusion model, this graph structure will serve as a priori dependencies to describe the conditional probability propagation paths between wind turbines. The node feature matrix (N×F) also needs to be combined with time-series data to form a spatiotemporal graph structure input (e.g., T×N×F). In the experimental setup, this graph structure significantly improves the model's ability to model spatial correlations; for example, in wind fields with wake effects, the prediction error (RMSE) can be reduced by approximately 5%–12%.

[0059] In this embodiment, step S4 includes the following steps: One-dimensional convolution, dilated convolution, and residual stacking are performed on the standardized sample tensor set to extract local change trends and multi-scale temporal features. By coupling local change trends and multi-scale temporal features, time-coded features are obtained. The time-encoded features are injected into the wind turbine graph structure and graph convolution propagation is performed to obtain the wind turbine graph propagation features; Multi-head attention cross-modal fusion is performed on the propagation features and time-coded features of the wind turbine diagram, followed by attention enhancement and normalization to obtain the fused feature vector.

[0060] In this embodiment, a standardized sample tensor is used as input, typically with the structure (Batch, number of wind turbines N, time step L, feature dimension F), for example (32, 50, 12, 30). A one-dimensional convolution (1D-CNN) operation is performed on the time dimension to capture local trends over a short period. The data from each wind turbine is treated as an independent time series, and convolution is performed along the time axis. The kernel size is typically set to k=3 or k=5, the stride is 1, and padding uses a "same" method to maintain a constant time length. The number of output channels can be set to 64 or 128 to enhance feature representation. Based on this, dilated convolution is introduced to expand the receptive field without increasing the number of parameters. For example, multiple dilated convolutions with dilation rates d=2 and d=4 are used to enable the model to capture trends over longer time spans (e.g., 30–60 minutes). Subsequently, ordinary convolutions and dilated convolutions are combined to form multi-scale convolutional blocks, which are then stacked through residual connections. For example, every two convolutional layers form a residual unit with an output of F(x)+x, in order to alleviate the gradient vanishing problem in deep networks.

[0061] Features from different convolutional layers (such as outputs from ordinary convolutions and outputs from dilated convolutions with different dilation rates) are concatenated or weighted and fused. For example, features from three scales are concatenated along the channel dimension to form higher-dimensional features (e.g., expanding from 128 dimensions to 384 dimensions). Then, a 1×1 convolution (i.e., pointwise convolution) is used for dimensionality reduction and feature recombination, compressing the high-dimensional features to the target dimension (e.g., 128 dimensions) while achieving a linear combination of information from different scales. A gating mechanism can be introduced, such as using a structure similar to GLU (Gated Linear Unit), generating weights through the sigmoid function to adaptively weight features at different scales, thereby enhancing the impact of key temporal patterns (e.g., sudden changes in wind speed) on the overall features. To further enhance temporal representation capabilities, positional encoding can be superimposed, for example, using sine and cosine functions to encode time steps, enabling the model to perceive temporal sequence information. The final temporal encoded feature dimension is (Batch, N, L, D), where D is typically set to 128 or 256.

[0062] Temporally encoded features are used as node feature inputs (each time step corresponds to a graph signal), and graph convolution is performed in conjunction with the adjacency matrix A (N×N). Common methods are spectral domain graph convolution or spatial domain graph convolution, for example, using a normalized adjacency matrix  = D^(-1 / 2) AD^(-1 / 2), where D is the degree matrix, and then feature propagation is performed: H' = σ(ÂHW), where H is the input feature (N×D), W is the learnable weight matrix (D×D'), and σ is the activation function (such as ReLU). In practice, to preserve the temporal dimension, graph convolution can be performed independently for each time step, or a spatiotemporal graph convolution (ST-GCN) structure can be used to jointly model temporal and spatial convolutions. For example, the number of graph convolution layers can be set to 2–3, with an output dimension of 128 for each layer, and Dropout (e.g., 0.2) can be added between layers to prevent overfitting. The propagation effect can be enhanced by combining edge weights (based on distance and wind direction), so that the information of the upwind wind turbines can be more effectively transmitted to the downwind wind turbines.

[0063] A multi-head attention mechanism is employed for cross-modal fusion, where temporal encoded features are used as the query, and graph propagation features are used as the key and value (or vice versa). Attention calculation enables dynamic interaction between the two types of features. The attention calculation formula is: Attention(Q,K,V)=softmax(QKᵀ / √d)V, where d is the feature dimension (e.g., 128). The multi-head mechanism is typically set to h=4 or h=8 heads, with each head learning association patterns in different subspaces, thereby improving the model's expressive power. The fused features are then concatenated and linearly transformed to restore the original dimension (e.g., 128). Subsequently, attention enhancement mechanisms are introduced, such as self-attention, to further strengthen the weight distribution of key time steps or key wind turbine nodes, making the model focus more on regions that contribute significantly to prediction. Simultaneously, layer normalization and residual connections are added at the output stage (i.e., the output is Attention(x)+x) to improve training stability and prevent feature degradation.

[0064] In one specific implementation, a one-dimensional convolutional encoding is performed on the time series of each sample. Let the convolutional layer be... Layer output is Then it can be written as: ; in, Indicates nonlinear activation. Indicates batch normalization, This represents the void ratio. The formula illustrates that the project doesn't simply use ordinary convolutions, but rather expands the receptive field by using different void ratios, covering a longer time range with fewer layers.

[0065] The improved CNN encoder in the project contains 4 convolutional layers, with the second and third layers using dilated convolutions with dilation rates of 2 and 4, respectively; residual connections are also configured.

[0066] This expression represents the addition of the current layer output to the bypass features to reduce deep training degradation.

[0067] After pooling and full connection, the temporal encoding features are obtained:

[0068] in, Indicates batch size, Indicates the number of nodes. This indicates the hidden dimension. This feature will be fed into the graph branch and the fusion branch for further processing.

[0069] The graph encoder of this invention adopts a structure with GAT as the primary component and residual connections as a secondary component. Its attention coefficient can be expressed as: ; in, It is a linear transformation matrix. For attention parameters, For nodes The neighborhood set, This indicates vector concatenation. The text here means that the system does not use all neighbor information equally, but instead automatically learns which neighbor is more important to the current wind turbine.

[0070] The node update process is as follows:

[0071] in, This represents the activation function. The project implementation uses a three-layer GATConv, with residuals and normalization added after the second and third layers.

[0072] If graph convolution is used, it can also be written as:

[0073] This formula describes global propagation using a normalized adjacency matrix. In other words, the spatial propagation concept of this invention is not limited to a specific type of graph operator; GAT and GCN are equivalent substitutions.

[0074] Let the CNN branch output be The graph branch output is Then cross-modal fusion can be written as:

[0075] in, This represents multi-head attention. It means using time-branch features as the query and graph-branch features as the key and value to extract the correspondence between two types of information.

[0076] After merging, perform residual normalization: ; in, Presentation layer normalization. This step avoids information imbalance caused by simple splicing.

[0077] The project also includes a time-attention enhancement module:

[0078] This expression represents a further attention redistribution within the fused features to highlight the representations most sensitive to predictions in the next H steps.

[0079] In this embodiment, the specific steps of step S5 are as follows: A time series folding model is constructed by training the model based on fused feature vectors. Wind speed probability prediction is performed based on a time series model, generating multiple prediction results; The wind speed prediction result is obtained by performing second-order moment fusion on multiple prediction results.

[0080] In this embodiment, a time series K-fold cross-validation method is used to construct a "time series fold model". For example, K=5 is set, dividing the continuous time series into 5 non-overlapping sub-intervals. Each fold uses the first K-1 intervals as the training set and the last interval as the validation set, and training is performed in chronological order (to avoid future information leakage). The model structure can use a neural network with a Bayesian inference mechanism, such as introducing probability distribution parameters (mean μ and variance σ²) into the output layer, or using Monte Carlo Dropout (e.g., dropout=0.2) to simulate Bayesian uncertainty. During training, the loss function can be a combination of negative log likelihood (NLL) or mean squared error (MSE) and uncertainty regularization term, such as Loss = MSE + λ·σ² constraint term, where λ is 0.01 to balance accuracy and uncertainty. The optimizer can be Adam, with a learning rate of 0.001, a batch size of 32, and 50–100 training epochs. An early stopping strategy (patience=10) is used to prevent overfitting. Each fold training yields an independent model, resulting in a total of K sub-models (e.g., 5), which constitute the time series fold model set.

[0081] For the test set or new input data, it is sequentially fed into each folded model to obtain K sets of predicted outputs. For example, for each sample, each model outputs a predicted mean μ_k and the corresponding uncertainty (such as variance σ_k²), thus forming K probability prediction results. In terms of implementation, if a Bayesian neural network or the MC Dropout method is used, multiple forward propagations can be performed within each model (such as T=20 random dropout samplings) to further obtain a richer prediction distribution, so that each model outputs a set of samples instead of a single value. For example, with K=5 and T=20, 100 predicted samples can be obtained to characterize the uncertainty distribution of wind speed. To ensure prediction consistency, the input data needs to be transformed using the normalizer parameters saved during the training phase, and inverse normalization is performed after the output to restore the true dimensions (such as m / s).

[0082] A second-moment ensemble method is employed, simultaneously considering the predicted mean (first moment) and the predicted variance (second moment). First, the predicted mean μ_k of the K model outputs is averaged to obtain the ensemble mean: μ = (1 / K)∑μ_k. Second, the total uncertainty is calculated, consisting of two parts: intra-model uncertainty (average variance) and inter-model inconsistency (mean variance). Specifically, it is calculated as: σ² = (1 / K)∑σ_k² + (1 / K)∑(μ_k − μ)², where the first term represents the uncertainty of each model itself, and the second term represents the uncertainty arising from the prediction differences between models. In practical implementation, this method effectively separates "data noise" and "model bias," thus providing a more reliable prediction interval. For example, a 95% confidence interval can be calculated as μ ± 1.96σ to assess prediction reliability. To prevent individual anomalous models from affecting the results, anomaly detection can be performed before fusion (e.g., removing predicted values ​​that deviate from the mean by more than 2σ).

[0083] In one specific embodiment, based on the fusion features, the present invention simultaneously outputs the mean and standard deviation: ; in, This is the future forecast mean. The standard deviation of future forecasts To prevent small constants from having a value of zero. In words, the mean answers "what is most likely to happen in the future," while the standard deviation answers "how uncertain this prediction is."

[0084] The ImprovedBNNOutput function in the project can also output MDN parameters:

[0085] From this, the conditional probability distribution can be obtained:

[0086] in, denoted as the number of Gaussian mixture components. This scheme is used to express complex uncertainty patterns that are difficult to cover with a single Gaussian.

[0087] In this embodiment, the model training calculation specifically includes: The mean prediction head and variance prediction head calculate and output the fused feature vector to obtain the predicted mean, predicted standard deviation, and optional mixture distribution parameters; The first stage of training involves optimizing the MSE loss for the predicted mean. The second phase of training involves calibrating the predicted standard deviation using Gaussian negative log-likelihood loss. The optimal model parameters are output based on two-stage training. Construct a time series fold model based on the optimal model parameters.

[0088] In this embodiment, the fused features are processed using linear mapping or a multilayer perceptron (MLP), for example, a two-layer fully connected structure (128 hidden layer dimensions, GELU activation function) to construct two independent branches. The mean prediction head outputs the predicted mean μ∈R^(B×N×H), representing the predicted value at the next H steps (e.g., H=2, corresponding to 20 min); the variance prediction head outputs the log-variance logσ² or standard deviation σ to ensure numerical stability (usually constrained by the Softplus function σ>0). The initial range of σ can be limited to [0.1, 5] to avoid extreme values ​​in the early stages of training. To enhance the model's expressive power, it can also be extended to a mixed distribution output (e.g., Gaussian mixture model GMM), that is, outputting the mean μ_k, variance σ_k², and mixture weight π_k (k=1~K, usually K=3) of multiple components to characterize the multimodal distribution characteristics. Only the mean prediction head is enabled to participate in backpropagation training, while the parameters of the variance prediction head are frozen or not involved in loss calculation. The loss function used is Mean Squared Error (MSE Loss), which minimizes the squared error between the predicted mean μ and the actual observed value y. In the experimental setup, the MSE loss is averaged over time steps and wind turbine nodes, i.e., the mean is calculated over all samples. During training, the Adam optimizer is used with a learning rate of 1e-3, a batch size of B=64, and approximately 30-50 training epochs. A learning rate decay strategy (e.g., decaying to 0.5 times every 10 epochs) is employed to improve convergence. To prevent overfitting, L2 regularization (weight decay of 1e-5) and Dropout (scale 0.2) are introduced. The focus of this stage is to allow the model to fully learn the mapping relationship between the fused features and the target variable, resulting in a high accuracy in the predicted mean. In practice, this stage typically reduces the prediction error (e.g., RMSE) to a relatively stable level (e.g., below 1.5 m / s).

[0089] The Gaussian Negative Log-Likelihood (NLL) loss is used as the optimization objective, with the form: Loss = (1 / N)∑[log(σ²) + (y_true − μ_pred)² / σ²]. This loss function considers both prediction error and prediction uncertainty: when the prediction error is large, the model tends to increase σ to reduce the penalty; when the prediction is accurate, it tends to decrease σ to increase confidence. In implementation, to avoid numerical instability caused by excessively small σ, a lower bound constraint is usually set (e.g., σ_min = 1e-3). For training strategies, a small learning rate (e.g., 0.0005) can be used for fine-tuning, and both the mean head and variance head are used in training, but the mean parameter is updated with a small weight (e.g., through different learning rates or weight decay). To prevent excessive variance inflation, a regularization term (e.g., λ·σ², where λ is 0.01) can be added for constraint.

[0090] Based on validation set performance, models across multiple training epochs or with different hyperparameter configurations are compared. Evaluation metrics include root mean square error (RMSE), mean absolute error (MAE), and negative log-likelihood (NLL). A weighted evaluation method is typically used, such as Score = RMSE + α·NLL, where α can range from 0.1 to 0.5 to balance accuracy and uncertainty quality. When selecting the optimal model, not only should the prediction error be low, but the uncertainty estimation should also be reasonable (e.g., confidence intervals should not be too wide or too narrow). Reliability diagrams or Prediction Interval Coverage Probability (PICP) metrics can also be used for auxiliary evaluation; for example, a 95% confidence interval coverage of 92%–98% is required. In implementation, the weights of the model with the best performance on the validation set during training can be saved (checkpointing mechanism), and the corresponding normalized parameters and model structure configuration can be recorded.

[0091] The complete dataset is divided into K consecutive sub-intervals (e.g., K=5) in chronological order, and training is performed using a rolling window approach: the first fold uses intervals 1–4 for training and interval 5 for validation; the second fold uses intervals 1–3 for training and interval 4 for validation; and so on. Each fold uses the optimal model parameters as initial weights (i.e., transfer learning), thereby accelerating convergence and ensuring model stability. During training, a two-stage training strategy is still used, but the number of training rounds can be appropriately reduced (e.g., 30 rounds in the first stage and 20 rounds in the second stage) to reduce computational costs. Finally, K sub-models are obtained, each capable of predicting the mean and uncertainty. These models will participate in inference together in subsequent prediction stages, forming a multi-model ensemble framework.

[0092] In a specific embodiment, the first stage focuses on mean accuracy, using mean squared error:

[0093] in, Indicates the number of samples. This represents the mean of the model output. This represents the true value. The goal of this stage is to first accurately learn the "predicted center location".

[0094] After loading the optimal weights from stage A, stage B freezes the mean output header and then calibrates the variance using Gaussian negative log-likelihood.

[0095] In this formula, if the mean deviation is large and the variance is too small, the loss will increase significantly; if the variance is too large, the logarithmic term will also be penalized. Therefore, this training method can constrain the probability output to be neither too confident nor too conservative.

[0096] S10 Multi-Fold Gaussian Integration Let the first The output of the individual fold model is The weight is ,satisfy The ensemble mean is:

[0097] The integrated variance is:

[0098] The first equation expresses the average trend, while the second equation considers both "intra-model variance" and "mean differences between different models." In the project, weights can be set uniformly or weighted according to the train_size recorded in folds_meta.json.

[0099] Further, the Gaussian quantile intervals are obtained:

[0100] This range is used to represent the possible range of future values ​​falling under different risk levels.

[0101] This step maps the wind speed prediction results to the power prediction results, and there are two parallel schemes in engineering.

[0102] Option A: MDN Power Mapping. This option takes wind speed or wind speed-derived features as input and outputs the weights, mean, and standard deviation of several mixed Gaussian components. Its conditional distribution is as follows:

[0103] Its training objective is:

[0104] This solution is suitable for scenarios where the relationship between wind speed and power has obvious multi-peak characteristics or where the same wind speed corresponds to multiple power states.

[0105] Option B: Quantile Regression Power Mapping. This option directly outputs multiple power quantile values, such as q10, q50, and q90. Its objective function is:

[0106] in, This is the set of quantiles. The advantage of this scheme is that the results are directly interpretable, and each curve in the output corresponds to a specific quantile level.

[0107] It should be noted that: Option A and Option B are two alternative second-phase solutions. Either one can be chosen for deployment. They are not required to be executed simultaneously, nor are they in any particular order in the process.

[0108] In this embodiment, the specific steps of step S6 are as follows: The latest raw SCADA observation data is collected based on the multi-source monitoring data; Based on the wind speed prediction results, the unit numbers are precisely aligned using the latest SCADA raw observation data, and the power prediction deviation is calculated to obtain error indicators and probability indicators. Based on the error index and probability index, a prediction deviation assessment is performed to obtain a prediction assessment report; Based on the prediction and evaluation report, the time series folding model is trained and corrected in a targeted manner, and an optimized prediction model is output to perform ultra-short-term power prediction tasks.

[0109] In this embodiment, the operating status parameters of each wind turbine are continuously acquired at high-frequency sampling periods (e.g., 1 s or 5 s) through the interface between the wind turbine control system and the Supervisory Control and Data Acquisition (SCADA) system. These parameters include key variables such as wind speed (m / s), active power (kW), rotational speed (rpm), pitch angle (°), and nacelle temperature (°C). To adapt to ultra-short-term prediction tasks, the high-frequency data needs to be time-aggregated, for example, by statistically analyzing it in 10-minute time windows to calculate the mean, maximum, and standard deviation, thus forming time-granularity data consistent with the prediction model input. During data transmission, a buffer queue (e.g., length 1000) and a breakpoint resume mechanism are used to ensure data integrity and real-time performance. Simultaneously, rapid quality checks are performed on newly acquired data, including missing rate detection (threshold 5%) and outlier filtering (based on physical constraints), to ensure the reliability of the input data. In the experimental setup, the most recent 24 hours of data are typically continuously collected as a real-time input buffer, and the most recent L time steps (e.g., L=6, corresponding to the past 60 minutes) are extracted as the input for the current prediction window.

[0110] Based on a unified wind turbine numbering system (e.g., WTG_01~WTG_50), the predicted results are matched one-to-one with the actual observed data to ensure comparisons are made at the same timestamp t and the same wind turbine node i. For wind speed prediction results (e.g., predicted mean μ_v and standard deviation σ_v), the corresponding predicted power value P_pred(i,t+h) can be converted using a power curve model (e.g., wind speed-power mapping function), and then compared with the actual SCADA-recorded power value P_real to calculate the prediction deviation ΔP = P_pred − P_real. Regarding error metrics, the absolute error (MAE), mean squared error (MSE), and relative error (e.g., MAPE) are calculated, and can be statistically analyzed at both the wind turbine level and the overall level. Regarding probability metrics, based on the predicted distribution (μ, σ²), the negative log-likelihood (NLL) and confidence interval coverage (e.g., the proportion of true values ​​covered within a 95% interval) are calculated to assess the quality of uncertainty estimation. A target coverage rate of 90%–95% is typically set; a lower actual coverage rate indicates insufficient variance estimation.

[0111] Statistical analysis of errors is performed over time, such as calculating the average MAE, RMSE, and maximum error over the past 24 hours, and identifying the time periods when error peaks occur (e.g., periods of sudden wind speed changes). Secondly, a comparative analysis of different wind turbines is conducted spatially to identify units with significant errors (e.g., areas with significant wake effects) and analyze their relationship with their location or prevailing wind direction. In terms of probability assessment, statistical confidence interval coverage (PICP) and interval width (MPIW) are used to evaluate the reasonableness of the model's uncertainty expression. A comprehensive scoring index (e.g., CRPS) can be introduced to quantify the overall performance of probability prediction. Performance thresholds can be set, such as MAE ≤ 150 kW and coverage ≥ 90%; failure to meet these thresholds is marked as performance degradation.

[0112] Targeted optimization of the time series folding model is performed to improve prediction accuracy and stability. First, the source of the problem is identified based on the error distribution results. For example, if the error is large in a specific time period, the sample weights for that time period can be increased (e.g., the weighting coefficient is increased to 1.5 times) for retraining. If certain wind turbines have consistently high errors, their adjacency relationships or graph structure weights can be adjusted to enhance spatial modeling capabilities. At the model training level, incremental learning or fine-tuning strategies can be adopted. That is, based on the original optimal parameters, training can continue for 10-20 rounds with a smaller learning rate (e.g., 1e-4) to adapt the model to the latest data distribution changes. Simultaneously, the variance prediction head is recalibrated to ensure that the uncertainty estimate is consistent with the latest error distribution. Within the time series folding model framework, poorly performing folding models can be replaced or retrained to maintain the stability of the overall model set.

[0113] The final technical effect achieved in this invention is as follows, based on the results from results / crps / crps_summary.csv: Model; CRPS; MAE; RMSE Improved Fusion; 0.3990; 0.5650; 0.7359 Fusion; 0.4993; 0.7095; 0.9077 HighRes Fusion; 0.5713; 0.8053; 1.0546 CNN Only; 0.6890; 0.9882; 1.2297 The improvement of the improved model compared to CNN Only is as follows: Δ_CRPS=42.09%, Δ_MAE=42.82%, Δ_RMSE=40.15%.

[0114] The above results demonstrate that, under the same verification criteria, the present invention can simultaneously improve the accuracy of mean prediction and the quality of probability prediction.

[0115] In this embodiment, a wind farm ultra-short-term power prediction system based on graph Bayesian fusion is provided to execute the wind farm ultra-short-term power prediction method based on graph Bayesian fusion as described above, including: The preprocessing module is used to collect multi-source monitoring data of wind turbines in wind farms, and to standardize the multi-source monitoring data to obtain a standard time series dataset. The segmentation module is used to perform continuous temporal segmentation on the standard time series dataset and construct a standardized sample tensor set. The graph creation module is used to identify the location coordinates of wind turbines in a wind farm; and to construct a wind turbine graph structure based on the location coordinates. The convolution module is used to inject the standardized sample tensor set into the wind turbine graph structure for graph convolution propagation to obtain the fused feature vector; The prediction module is used to perform wind speed probability prediction based on the fused feature vector to obtain the wind speed prediction result; The training correction module is used to calculate the power prediction deviation and perform targeted training correction on the wind speed prediction results in order to perform ultra-short-term power prediction operations.

[0116] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0117] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A wind farm ultra-short-term power prediction method based on Bayesian fusion, characterized in that, Includes the following steps: Step S1: Collect multi-source monitoring data of wind turbines in the wind farm, and standardize the multi-source monitoring data to obtain a standard time series dataset; Step S2: Perform continuous temporal segmentation on the standard time series dataset to construct a standardized sample tensor set; Step S3: Identify the location coordinates of the wind turbines in the wind farm; construct a wind turbine diagram structure based on the location coordinates; Step S4: Inject the standardized sample tensor set into the wind turbine graph structure and perform graph convolution propagation to obtain the fused feature vector; Step S5: Perform wind speed probability prediction based on the fused feature vector to obtain the wind speed prediction result; Step S6: Calculate the power prediction deviation and perform directional training correction on the wind speed prediction results to execute the ultra-short-term power prediction operation.

2. The wind farm ultra-short-term power prediction method based on Bayesian fusion according to claim 1, characterized in that, The specific steps of step S1 are as follows: Collect multi-source monitoring data from wind turbines in wind farms; the multi-source monitoring data includes raw SCADA observation data, reanalysis meteorological data, and topographic raster data; Extract the fan number and time coordinates from the multi-source monitoring data; Based on the wind turbine number and time coordinates, a standardized data table is obtained by uniform standardization. Standard preprocessing is performed on the standardized data table to obtain a standard time series dataset.

3. The wind farm ultra-short-term power prediction method based on Bayesian fusion according to claim 2, characterized in that, The standard preprocessing specifically includes; Identify wind speed and power fields based on standardized data tables; Perform physical constraint range checks on the wind speed and power fields, mark outliers and remove them; Detect the temporal continuity of the wind speed and power fields; Based on the time continuity, abnormal fields are distinguished to obtain short-term missing measurement segments and long missing measurement segments that disrupt continuity. Linear interpolation is used to repair short-term missing measurement segments; the long-term missing measurement segments are truncated. Physical constraint range checks and time continuity verifications are performed on wind speed and power fields. Outliers are eliminated and short-term missing data segments are repaired using linear interpolation. Long missing data segments that disrupt continuity are directly truncated.

4. The wind farm ultra-short-term power prediction method based on Bayesian fusion according to claim 2, characterized in that, The specific steps of step S2 are as follows: Set both the fixed input duration and the prediction duration as sliding windows; The standard time-series dataset is continuously segmented into time sequences according to the sliding window to obtain supervised learning samples. Robust normalization is performed on the dynamic input channel and target power channel of the supervised learning samples respectively, and the normalizer parameters are saved; Time series are partitioned based on normalizer parameters to construct a standardized sample tensor set; the standardized sample tensor set includes a training set, a validation set, a test set, and a time series K-fold set.

5. The wind farm ultra-short-term power prediction method based on Bayesian fusion according to claim 4, characterized in that, Step S3 is as follows: Identify the location coordinates of wind turbines in a wind farm; Based on the location coordinates, perform wind turbine layout analysis and construct a node set; Based on the location coordinates and the distance between computer groups, the angle between the wind turbine and the main wind direction is identified; The side set is determined based on the distance between the units and the angle between the fan and the main wind direction; Construct the wind turbine graph structure based on the set of nodes and the set of edges.

6. The wind farm ultra-short-term power prediction method based on Bayesian fusion according to claim 5, characterized in that, The specific steps of step S4 are as follows: One-dimensional convolution, dilated convolution, and residual stacking are performed on the standardized sample tensor set to extract local change trends and multi-scale temporal features. By coupling local change trends and multi-scale temporal features, time-coded features are obtained. The time-encoded features are injected into the wind turbine graph structure and graph convolution propagation is performed to obtain the wind turbine graph propagation features; Multi-head attention cross-modal fusion is performed on the propagation features and time-coded features of the wind turbine diagram, followed by attention enhancement and normalization to obtain the fused feature vector.

7. The Bayesian fusion-based wind farm ultra-short-term power prediction method according to claim 6, characterized in that, The specific steps of step S5 are as follows: A time series folding model is constructed by training the model based on fused feature vectors. Wind speed probability prediction is performed based on a time series model, generating multiple prediction results; The wind speed prediction result is obtained by performing second-order moment fusion on multiple prediction results.

8. The wind farm ultra-short-term power prediction method based on graph Bayesian fusion according to claim 7, characterized in that, The model training calculation is specifically as follows: The mean prediction head and variance prediction head calculate and output the fused feature vector to obtain the predicted mean, predicted standard deviation, and optional mixture distribution parameters; The first stage of training involves optimizing the MSE loss for the predicted mean. The second phase of training involves calibrating the predicted standard deviation using Gaussian negative log-likelihood loss. The optimal model parameters are output based on two-stage training. Construct a time series fold model based on the optimal model parameters.

9. The wind farm ultra-short-term power prediction method based on graph Bayesian fusion according to claim 8, characterized in that, The specific steps of step S6 are as follows: The latest raw SCADA observation data is collected based on the multi-source monitoring data; Based on the wind speed prediction results, the unit numbers are precisely aligned using the latest SCADA raw observation data, and the power prediction deviation is calculated to obtain error indicators and probability indicators. Based on the error index and probability index, a prediction deviation assessment is performed to obtain a prediction assessment report; Based on the prediction and evaluation report, the time series folding model is trained and corrected in a targeted manner, and an optimized prediction model is output to perform ultra-short-term power prediction tasks.

10. A wind farm ultra-short-term power prediction system based on graph Bayesian fusion, characterized in that, The method for performing ultra-short-term power prediction of wind farms based on graph Bayesian fusion as described in claim 1 includes: The preprocessing module is used to collect multi-source monitoring data of wind turbines in wind farms, and to standardize the multi-source monitoring data to obtain a standard time series dataset. The segmentation module is used to perform continuous temporal segmentation on the standard time series dataset and construct a standardized sample tensor set. The graph creation module is used to identify the location coordinates of wind turbines in a wind farm; and to construct a wind turbine graph structure based on the location coordinates. The convolution module is used to inject the standardized sample tensor set into the wind turbine graph structure for graph convolution propagation to obtain the fused feature vector; The prediction module is used to perform wind speed probability prediction based on the fused feature vector to obtain the wind speed prediction result; The training correction module is used to calculate the power prediction deviation and perform targeted training correction on the wind speed prediction results in order to perform ultra-short-term power prediction operations.