Fusion algorithm-based wind speed refined short-term and imminent forecasting method under complex terrain

By constructing a high spatiotemporal resolution grid dataset through fusion algorithms and dynamic downscaling techniques, and combining it with a machine learning model based on an ensemble learning framework, the problem of large wind speed forecasting errors under complex terrain was solved, achieving high-precision wind speed forecasting and meeting the power prediction requirements of wind farms.

CN121454650APending Publication Date: 2026-02-03甘肃省气象服务中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511581681.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Traditional wind speed forecasting techniques suffer from large errors and insufficient local wind field capture in complex terrain, especially in the Hexi Corridor. Existing numerical models have low spatiotemporal resolution, and physical parameterization schemes are difficult to adapt to complex terrain and local circulation, resulting in large deviations in wind power prediction. Furthermore, the generalization ability of single statistical methods is weak, which cannot meet the needs of refined wind speed forecasting.

Method used

A fusion algorithm-based approach is adopted to acquire large-scale wind field background data, ground-based wind field data, and high-precision topographic data. A high spatiotemporal resolution grid dataset is constructed using dynamic downscaling technology. Wind field characteristics are analyzed by combining high-resolution grid data and ground-based wind field data, and wind speed-related weather and climate influencing factors are extracted. A machine learning fusion forecast model with an integrated learning stacking framework is constructed to improve forecast accuracy and stability.

Benefits of technology

It enables refined short-term wind speed forecasts under complex terrain, improves spatiotemporal resolution and forecast accuracy, meets the needs of wind power dispatch, reduces power prediction bias, and supports the goal of "carbon peaking and carbon neutrality".

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121454650A_ABST
    Figure CN121454650A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of weather forecasting, and particularly discloses a complex terrain wind speed refined short-term and imminent forecasting method based on a fusion algorithm, which comprises the following steps: S1, acquiring large-scale wind field background data, ground live wind field data and high-precision terrain data of a target complex terrain; s2, performing fine processing on the large-scale wind field background data through a dynamic downscaling technology; s3, analyzing space-time distribution and evolution characteristics of the complex terrain wind field, and extracting weather climate influence factors related to the wind speed; s4, adopting an integrated learning stacking framework to construct a machine learning fusion forecasting model; according to the method, large-scale wind field background data and complex terrain data are combined through a dynamic downscaling technology, a high-temporal-spatial-resolution lattice point data set is constructed, meanwhile, an integrated fusion frame with LightGBM, RF and LSTM heterogeneous algorithms as a base learning device and ridge regression as a meta learning device is adopted, and the prediction precision of a local wind field under the complex terrain is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of meteorological forecasting, specifically involving a method for refined short-term wind speed forecasting under complex terrain based on a fusion algorithm. Background Technology

[0002] Accurate wind speed forecasting, especially short-term forecasting under complex terrain conditions, is crucial for wind farm power prediction, power grid safety dispatching, and early warning of severe wind disasters. Complex terrain areas, such as the Hexi Corridor region, formed by the Qilian Mountains and the Alashan Plateau, are characterized by abundant wind resources but frequent strong winds, and are key large-scale clean energy bases for the development of new energy industries in Gansu Province.

[0003] Refined short-term wind speed forecasts are a core prerequisite for ensuring high-precision wind power prediction and stable power system dispatch.

[0004] However, traditional wind speed forecasting techniques have significant limitations: single numerical models (such as the global-scale ECMWF and GFS models) have low spatiotemporal resolution (e.g., the initial spatial resolution of ECMWF is only 0.125°×0.125°) and physical parameterization schemes are difficult to adapt to complex terrain and local circulation (e.g., the funneling effect in the Hexi region), resulting in large deviations between surface wind speed forecasts and actual observations.

[0005] Although some studies have used a single statistical method (such as traditional regression or a single artificial intelligence algorithm) for correction, such methods rely on station data and have spatial limitations, weak generalization ability, poor forecast stability under different weather systems, and difficulty in capturing all the characteristics of nonlinear wind speed changes, thus failing to meet the actual needs of refined wind speed forecasts under complex terrain. Summary of the Invention

[0006] The purpose of this invention is to provide a refined short-term wind speed forecasting method for complex terrain based on a fusion algorithm, so as to solve the problems of large wind speed forecasting error and insufficient local wind field capture in the traditional forecasting methods mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A refined short-term wind speed forecasting method based on a fusion algorithm for complex terrain includes:

[0009] S1. Acquire large-scale wind field background data, ground real-time wind field data, and high-precision terrain data of the target complex terrain;

[0010] Large-scale wind field background data provides the initial background of the regional wind field;

[0011] Ground-based real-time wind field data ensures data authenticity verification;

[0012] High-precision terrain data characterizes the impact of terrain on wind fields, laying a data foundation for subsequent processing.

[0013] S2. Based on the high-precision terrain data, the large-scale wind field background data is refined using dynamic downscaling technology to construct a high spatiotemporal resolution grid dataset under complex terrain.

[0014] Large-scale wind field background data has low spatiotemporal resolution and cannot be adapted to local wind fields in complex terrain. Therefore, based on high-precision terrain data, dynamic downscaling technology (using physical laws such as atmospheric motion equations) is used to reduce the initial wind field resolution of large-scale wind field background data to a scale that matches the terrain data, thus constructing a high spatiotemporal resolution gridded dataset.

[0015] S3. Based on the high spatiotemporal resolution grid dataset and the ground real-time wind field data, analyze the spatiotemporal distribution and evolution characteristics of wind fields in complex terrain, extract weather and climate influencing factors related to wind speed, and establish a wind speed forecast influencing factor database.

[0016] By combining high-resolution gridded datasets (reflecting refined wind field distribution) with ground-based real-time wind field data (ensuring analytical accuracy), we can uncover the spatiotemporal patterns of wind fields in complex terrains. At the same time, we can extract key weather and climate factors that affect wind speed (such as circulation and jet stream) and establish a factor library—which not only clarifies the physical mechanisms of wind field changes but also provides "physically meaningful input features" for subsequent models.

[0017] S4. Using the high spatiotemporal resolution grid dataset and the wind speed forecast influencing factor library as input, construct a machine learning fusion forecast model using an ensemble learning stacking framework.

[0018] An integrated learning stacking framework is adopted—by leveraging the complementary advantages of multiple algorithms, high-resolution gridded data and an impact factor library are used as inputs to construct a fusion model, thereby improving forecast accuracy and stability.

[0019] S5. Use the trained machine learning fusion forecast model to perform refined grid-based forecasting of wind speed for future periods in the complex terrain.

[0020] Using the trained fusion model, wind speed forecasts for complex terrains are output. "Gridding" ensures that the forecast covers the entire area, "refinement" corresponds to high-resolution data support, and short-term (such as 0-6h short-term effect) meets the immediate needs of wind power dispatch and energy supply, ultimately achieving the technical goals.

[0021] Preferably, the large-scale wind field background data includes ECMWF fine-grid numerical forecast data and ERA5 reanalysis data;

[0022] ECMWF fine-grid data (reporting time 20:00, resolution 0.125°×0.125°) provides near real-time large-scale initial wind fields, supporting short-term forecasts;

[0023] ERA5 reanalysis data (resolution 0.25°×0.25°, time resolution 1h) covers nearly 30 years of historical data and is used for long-term wind field evolution characteristic analysis. The combination of the two achieves dual support of "real-time forecast + historical analysis".

[0024] The high-precision terrain data includes SRTM3 terrain data;

[0025] SRTM3's 90m spatial resolution can accurately depict the local topographic undulations of complex terrain, providing "topographic constraints" for dynamic downscaling technology. This ensures that the downscaled grid data can match the impact of terrain on the wind field (such as local strong winds caused by the funnel effect), avoiding forecast bias caused by coarse terrain data.

[0026] The ground-based real-time wind field data includes CLDAS ground-based wind speed analysis data and observation data from the complex terrain meteorological station.

[0027] CLDAS data (resolution 0.0625°×0.0625°, 1-hour lead time) provides regional real-time wind fields with wide coverage; meteorological station observation data (data from 20 stations in complex terrain over the past 30 years) provides fixed-point high-precision real-time data. The two together form a "regional + fixed-point" real-time verification system.

[0028] Preferably, the dynamic downscaling technique is used to reduce the resolution scale of the initial wind field in the large-scale wind field background data to a high-resolution grid that matches the high-precision terrain data.

[0029] The dynamic downscaling technique takes the initial wind field in the large-scale wind field background data as the processing object (the initial wind field is the basic input for forecasting), and combines it with SRTM3 high-precision terrain data (which provides constraints such as terrain height and slope). By solving the atmospheric dynamic equations, the resolution of the initial wind field is reduced from a large scale (e.g., 14km) to a high resolution (e.g., 90m) that matches the SRTM3 terrain, and finally forms a high spatiotemporal resolution grid dataset.

[0030] Large-scale meteorological signals are decomposed into fine signals adapted to local terrain, solving the problem of data-terrain compatibility and providing a data foundation that reflects local wind field characteristics for subsequent wind field analysis and model building.

[0031] Preferably, the method for analyzing the spatiotemporal distribution and evolution characteristics of wind fields in complex terrain includes empirical orthogonal function analysis, Mann-Kendall trend test, and sliding T-test, which are used to analyze the spatial distribution, temporal evolution, and abrupt changes of surface wind fields (including the number of windy days, maximum wind speed, average wind speed, etc. in daily, monthly, seasonal, annual, and interannual terms) under complex terrain conditions, and to analyze the distribution pattern of surface wind fields in complex terrain within the year using wind concentration degree and concentration period index.

[0032] Among them, Empirical Orthogonal Function Analysis (EOF) is used to "decompose the spatial distribution modes and temporal evolution trends of wind fields". For example, EOF can extract the main spatial features of complex wind fields (such as strong winds in the western section and weak winds in the eastern section) and their corresponding temporal changes (such as the wind field strengthening trend in spring), thereby achieving dimensionality reduction and focusing of the spatiotemporal features of wind fields.

[0033] The Mann-Kendall (MK) trend test combined with the sliding T-test is used to "detect the climate change characteristics of wind fields". For example, the MK test can be used to determine whether the number of windy days in complex terrain has increased significantly after a certain year, and the sliding T-test can be used to locate the time point of the change and identify the turning point of long-term wind field changes.

[0034] The concentration of strong winds and the concentration period index are used to "analyze the annual distribution pattern of the number of days with strong winds". For example, it is calculated that strong winds in complex terrain are mainly concentrated in March to May (concentration period) and the degree of concentration is high (large concentration index), which provides a basis for the "seasonal key attention period" in short-term forecasts.

[0035] The method for extracting weather and climate influencing factors includes principal component analysis, which is used to analyze the relationship between surface wind field and influencing factors in complex terrain, and to screen key influencing factors that play a dominant role in wind speed changes. The influencing factors include circulation index, upper and lower level jet stream index, and vertical velocity.

[0036] Principal component analysis (PCA) calculates the correlation between surface wind field and circulation indices (such as the East Asian circulation index), upper and lower level jet indices (such as jet intensity and location), and vertical velocity (such as uplift / downlift motion), eliminates redundant factors (such as elements with low correlation), and selects key factors that play a dominant role in wind speed changes (such as upper level jet intensity and vertical uplift velocity), and finally establishes a wind speed forecast influencing factor library.

[0037] By extracting effective features from numerous meteorological elements, we can provide "physically meaningful input variables" for subsequent integrated models, thus avoiding forecast errors caused by redundant model input features.

[0038] Preferably, the integrated learning stacking framework includes a first layer of base learners and a second layer of meta-learners;

[0039] The first base learner consists of three heterogeneous algorithms: LightGBM, Random Forest, and Long Short-Term Memory Neural Network.

[0040] Among them, LightGBM and Random Forest (RF) are good at capturing nonlinear features and the interaction between features. For example, they can effectively explore the nonlinear relationship between wind speed and terrain slope and circulation index (such as the nonlinear change of wind speed increase when the slope exceeds 15°).

[0041] Long Short-Term Memory (LSTM) neural networks are good at capturing time-series dependencies. Wind speed has obvious time correlation (e.g., wind speed in the previous hour affects wind speed in the next hour). LSTM can remember historical wind speed information through gating mechanisms, avoiding the loss of time information in traditional models.

[0042] The three are "heterogeneous algorithms," which extract information from two dimensions: "nonlinear features" and "time series features," respectively, avoiding the limitations of a single algorithm and improving the model's ability to capture complex wind speed changes.

[0043] The second-layer meta-learner uses the ridge regression algorithm to fuse the outputs of the base learners.

[0044] Multi-base learner outputs may exhibit collinearity (e.g., the prediction results of LightGBM and RF may be highly correlated). Ridge regression algorithm can effectively handle the collinearity problem by introducing L2 regularization, while weighting and fusing the prediction results of the base learners (instead of simple averaging) – for example, giving higher weight to the superior results of LSTM in time series forecasting, and optimizing the weight of the superior results of LightGBM in nonlinear fitting, ultimately outputting more stable and accurate fused forecast results.

[0045] Preferably, the step of constructing a machine learning fusion prediction model using an ensemble learning stacking framework includes:

[0046] Wind speed data and key factors from the wind speed forecast influencing factor library are extracted from the high spatiotemporal resolution grid dataset. An M×N dimension sample feature matrix is ​​constructed by combining the ground real-time wind field data, and the training dataset and test dataset are divided, where M is the number of samples and N is the number of features.

[0047] Core wind speed data (such as historical hourly wind speed) is extracted from high spatiotemporal resolution gridded datasets, and key factors (such as circulation index and jet stream parameters) are extracted from the wind speed forecast influencing factor library. Combined with ground-based real-time wind field data (such as CLDAS real-time wind speed as labels), an M×N dimension sample feature matrix is ​​constructed (M is the number of samples, such as hourly samples of the past 5 years; N is the number of features, such as historical wind speed values ​​+ 5 key factors) – this process “transforms data into features that the model can recognize”, providing input for training; the training / test datasets are divided according to a preset ratio (such as 7:3 or 8:2) to ensure the model’s adaptability to new data.

[0048] Processing the training dataset:

[0049] Vertical stacking: The “M×N feature matrix” and “M×1 real-time label” of the training set are matched and integrated in the vertical dimension to form a “feature-label” training data pair that can be directly used by each base learner. This ensures that the features of each sample are accurately matched with the real wind speed label and avoids training bias caused by feature and label mismatch.

[0050] Horizontal stacking: The complete training data that has undergone "vertical stacking" (feature-label matching) is input into three heterogeneous base learners (LGBM, RF, LSTM) in parallel for independent training. The three learners do not interfere with each other and run synchronously during training.

[0051] Processing the test dataset:

[0052] Averaging: For the "independent prediction results of the test set" output by the three base learners, calculate the arithmetic mean of each sample. By averaging the results of multiple algorithms, the random error of a single base learner can be offset (e.g., LGBM predicts a strong wind event more strongly, while RF predicts a weak wind event more strongly; averaging can reduce the bias) and form a more stable "intermediate prediction result".

[0053] Horizontal overlay: The "independent prediction results of the three base learners", the "average prediction result of the base learners", and the "final prediction result of the meta-learner" (the output of the ridge regression meta-learner after inputting the test set) are horizontally integrated along the sample dimension to form a "multi-result comparison matrix". All horizontally integrated prediction results can be compared with the "CLDAS real-time wind speed data corresponding to the test set" to quantitatively evaluate the errors of different prediction results (such as mean absolute error and correlation coefficient), and finally determine whether the meta-learner model meets the actual needs of refined short-term wind speed forecasting under the complex terrain of the Hexi region.

[0054] Base learner training: The three base learners, LightGBM, Random Forest and Long Short-Term Memory Neural Network, are trained independently using training samples. After training, the training prediction results of each base learner are output.

[0055] The training dataset is fed into three base learners—LightGBM, Random Forest, and LSTM—for independent training. LightGBM uses gradient boosting trees to iteratively optimize nonlinear fitting, Random Forest uses multiple decision trees to vote and reduce overfitting, and LSTM uses recurrent neural networks to learn time series dependencies. Independent training of the three can fully leverage the advantages of each algorithm, avoid mutual interference, and ultimately output their respective training prediction results (such as the wind speed prediction value of each base learner for the training set), providing "multi-dimensional prediction basis" for subsequent fusion.

[0056] The training prediction results are used as new features and input into the ridge regression learner for training to obtain the final stacked fusion model.

[0057] Meta-learner training: The training prediction results of the three base learners are used as "new features" to input the ridge regression meta-learner. At this time, the learning goal of the meta-learner is to "fit the ground real wind speed label through the new features (base learner prediction values). In essence, it is to learn "how to perform optimal weighted fusion of the base learner results" (such as learning that the weight of LSTM prediction values ​​should be higher in the early morning and the weight of LightGBM should be optimized in the afternoon), and finally forming a complete stacked fusion model.

[0058] To achieve the upgrade from single-model learning to multi-model fusion learning, ensuring that the model integrates the advantages of various algorithms.

[0059] Preferably, during the construction of the machine learning fusion prediction model, cross-validation and grid search methods are used to adjust the hyperparameters of the base learner and the meta learner, and the test dataset is input into the optimized machine learning fusion prediction model to obtain the model prediction results;

[0060] The hyperparameters of machine learning models (such as the learning rate of LightGBM, the number of trees in random forest, the number of hidden nodes in LSTM, and the regularization parameter of ridge regression) directly affect the model performance.

[0061] The "cross-validation + grid search" method is adopted: cross-validation (such as 5-fold cross-validation) divides the training set into multiple sub-samples and trains multiple times to avoid the randomness of a single training; grid search traverses preset hyperparameter combinations (such as learning rate 0.01-0.1) to select the hyperparameter combination that minimizes the model error (such as minimizing root mean square error), optimizes the model, reduces generalization error, and improves the model's adaptability to complex terrain and wind speed.

[0062] =Compare the model forecast results with the actual ground wind field data in the test dataset to quantitatively evaluate the model forecast capability.

[0063] The optimized fusion model is input into the test dataset to obtain the model's forecast results. Simultaneously, the corresponding ground-based real-time wind field data (such as CLDAS real-time wind speed and meteorological station observed wind speed) in the test dataset is called, and the forecast results are compared with the real-time data through quantitative indicators (such as root mean square error and mean absolute error). For example, if the root mean square error decreases from 2.5 m / s before optimization to 1.6 m / s, it indicates that the model accuracy has improved. This process is the "final verification of model performance" to ensure that the model can maintain high accuracy on unseen test data, avoid the overfitting problem of the model being "only effective on the training set", and ultimately ensure that the output of refined short-term wind speed forecast results is reliable and can serve practical needs such as wind power prediction and energy dispatch.

[0064] Compared with the prior art, the beneficial effects of the present invention are:

[0065] This invention combines large-scale wind field background data with complex terrain data through dynamic downscaling technology to construct a high spatiotemporal resolution grid dataset. At the same time, it adopts an integrated fusion framework with heterogeneous algorithms such as LightGBM, RF, and LSTM as base learners and ridge regression as meta learners, which significantly improves the forecast accuracy of local wind fields under complex terrain. Attached Figure Description

[0066] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0067] Figure 1 This is a flowchart of the method steps of the present invention;

[0068] Figure 2 This is a flowchart illustrating the technical process of the present invention. Detailed Implementation

[0069] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0070] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0071] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0072] As attached Figure 1 To be continued Figure 2 As shown:

[0073] Example 1:

[0074] Background: This project addresses the complex terrain of the Hexi Corridor region in Gansu Province (geographical area: Jiuquan City, Zhangye City, Wuwei City, Jinchang City, and Jiayuguan City, 92°-104°E, 37°-42°N). This region is surrounded by the Qilian Mountains (south side), the Alashan Plateau (north side), and the Beishan Mountains (east side), forming a narrow corridor terrain with significant local "tunneling effect" (e.g., the average annual wind speed in the Jiuquan-Yumen area reaches 3.5 m / s, with an average of 25-30 days of strong winds (≥10.2 m / s) per year). Furthermore, existing large-scale numerical weather prediction (such as ECMWF 0.125°×0.125° resolution) cannot capture local wind field changes, leading to significant deviations in wind power prediction (an extreme case occurred in March 2021 where the predicted generating capacity was 500 MW, while the actual output was only 50 MW).

[0075] This embodiment provides a method for refined short-term wind speed forecasting under complex terrain based on a fusion algorithm, achieving refined gridded wind speed forecasting for the aforementioned region from 0 to 6 hours, including:

[0076] S1. Obtain ECMWF fine-grid numerical forecast data for the complex terrain of the target: the forecast start time is 20:00 Beijing time, the spatial resolution is 0.125°×0.125° (about 14km), the temporal resolution is 3h, the forecast lead time is 0-48h, and the data period is 2021-2022.

[0077] ERA5 global reanalysis data: spatial resolution 0.25°×0.25° (approximately 28km), temporal resolution 1h, data period 1993-2022 (nearly 30 years).

[0078] CLDAS ground 10m wind speed real-time analysis data: spatial resolution 0.0625°×0.0625° (approximately 7km), temporal resolution 1h, data period 2021-2022;

[0079] Observational data from 20 meteorological stations in the Hexi Corridor, including Jiuquan, Zhangye, and Wuwei, with observational elements including 10m wind speed and direction, a time resolution of 1 hour, and a data period from 1993 to 2022.

[0080] SRTM3 terrain data: spatial resolution 90m, covering the entire Hexi region, including terrain height, slope, aspect and other elements.

[0081] All data were obtained through legitimate channels. ECMWF data came from the European Centre for Medium-Range Weather Forecasts (ECMWF) website, ERA5 data came from the Copernicus Climate Change Service (C3S), CLDAS data came from the China Meteorological Administration Information Center, SRTM3 data came from the U.S. Geological Survey (USGS), and weather station data came from the Gansu Provincial Meteorological Bureau.

[0082] S2. The dynamic downscaling technique uses the WRF (Weather Research and Forecasting) mesoscale numerical model as the downscaling tool, and the model configuration is as follows:

[0083] Horizontal resolution: 3km×3km (adapted to 90m terrain data for SRTM3, ensuring the capture of the funnel effect in areas such as Yumen and Guazhou).

[0084] Number of vertical layers: 30 layers (from near ground to 50 hPa, with a focus on optimizing the boundary layer (1000-850 hPa) wind field simulation);

[0085] Physical parameterization schemes: WSM6 is used for microphysics, Kain-Fritsch is used for cumulus parameterization, YSU is used for boundary layer, and Noah is used for land surface processes (all of which are commonly used regional model schemes implicit in the appendix and are suitable for the arid and semi-arid climate of Hexi).

[0086] ECMWF fine-grid numerical forecast data (reported from 20:00 onwards) were used as the initial and boundary fields of the WRF model, and SRTM3 topographic data (90m) and land use data (MODIS300m) were input.

[0087] The model integration time is 48 hours, and it outputs 10m wind speed and wind direction grid data every 1 hour.

[0088] Quality control was performed on the downscaling results for 2021-2022 (outliers were removed, such as unreasonable values ​​with wind speeds > 30 m / s), and deviation correction was performed with the CLDAS real-time data (using linear correction method). Finally, a high spatiotemporal resolution grid dataset of 3 km × 3 km with a temporal resolution of 1 h was formed for the Hexi region (covering longitude 92°-104° E and latitude 37°-42° N, with a total of 120 × 150 grid points).

[0089] S3. Based on the high-resolution gridded dataset (2021-2022) constructed in step S2 and the ERA5 data from the past 30 years, the following methods were used for analysis:

[0090] Empirical Orthogonal Function (EOF) Analysis:

[0091] Results: The first mode (variance contribution 62%) showed an "east-west difference" pattern—high wind speeds (annual average 3.2-3.8 m / s) in the western section of the Hexi Corridor (Jiuquan, Jiayuguan) and low wind speeds (annual average 1.8-2.2 m / s) in the eastern section (Wuwei), reflecting the dominant influence of topography on the wind field; the second mode (variance contribution 18%) showed a "seasonal fluctuation" pattern, with wind speeds significantly higher in spring (March-May) than in other seasons (average 3.5 m / s in spring and 2.1 m / s in winter).

[0092] Mann-Kendall (MK) test and sliding T-test:

[0093] Results: The annual average wind speed in the Hexi region has shown a slight downward trend over the past 30 years (MK statistic Z=-1.8, p<0.1), and a sudden change occurred in 2005 (sliding T test t=2.5, p<0.05). After the sudden change, the annual average wind speed decreased by 0.3 m / s compared with the previous period (which may be related to the weakening of circulation caused by global climate change).

[0094] Strong wind concentration and concentration period index:

[0095] Calculation method: Concentration index ( (where n is the number of windy days, and n is the horn corresponding to a windy day, with the peak period being...) ;

[0096] Results: The concentration of strong winds in the Hexi region was C=0.65 (0<C<1, the closer to 1 the more concentrated). The concentration period D corresponds to late March (at an ecliptic angle of 50°), indicating that strong winds are mainly concentrated from March to May (accounting for 65% of the year), which is consistent with the frequent cold air activity and enhanced funneling effect in spring.

[0097] Principal component analysis (PCA) was used to screen key factors from 20 candidate factors (including circulation, dynamic, topographic, and thermal factors). The steps are as follows:

[0098] Candidate factors: East Asian trough intensity index, Siberian high pressure intensity index, 300hPa upper-level jet axis position / intensity, 500hPa vertical velocity, 850hPa temperature advection, topographic slope, surface temperature, relative humidity, etc.

[0099] PCA screening: Calculate the correlation coefficient between each factor and the wind speed in the high-resolution grid dataset, remove factors with correlation |r| < 0.3, perform principal component extraction on the remaining factors, and select the original factors corresponding to the principal components with eigenvalues ​​> 1;

[0100] The final influencing factor database contains five key factors: ① 300 hPa upper-level jet intensity (r = 0.68), ② 500 hPa vertical velocity (r = -0.52, downdrafts favor increased wind speed), ③ East Asian trough intensity index (r = 0.45), ④ topographic slope (r = 0.41), and ⑤ 850 hPa temperature advection (r = 0.38). All of these factors passed the significance test with α = 0.01.

[0101] S4. Using the high spatiotemporal resolution grid dataset and the wind speed forecast influencing factor library as input, construct a machine learning fusion forecast model using an ensemble learning stacking framework.

[0102] Specifically: Sample feature engineering

[0103] Sample period: January 1, 2021 to December 31, 2022, with a total of 43,800 hourly samples (M=43,800).

[0104] Feature selection (N=12):

[0105] Key feature: The first 6 hours of wind speed sequence (t-1 to t-6, reflecting temporal correlation) in the high-resolution grid dataset of Step 2;

[0106] Auxiliary features: Step 3: 5 key factors from the influencing factor library (jet intensity, vertical velocity, etc.) + terrain height (extracted by SRTM3);

[0107] Tag data: CLDAS ground 10m wind speed real-time data (time t, as the model prediction target);

[0108] Data partitioning: The data was divided into a training set (30,660 samples) and a test set (13,140 samples) in a 7:3 ratio, using stratified sampling (to ensure that the sample ratio is consistent across seasons).

[0109] Implementation of an integrated learning stacking framework

[0110] The framework is built using Python (Scikit-learn, TensorFlow libraries) and consists of two layers:

[0111] First layer: Base learner training

[0112] LightGBM (Light Booster): Hyperparameter settings: learning rate 0.05, number of decision trees 100, maximum depth 8, number of leaf nodes 20 (optimized via grid search); training objective: minimize root mean square error (RMSE).

[0113] Random Forest (RF): Hyperparameter settings - number of decision trees 200, maximum depth 10, feature sampling rate 0.8, minimum number of samples for node splitting 5;

[0114] Long Short-Term Memory Neural Network (LSTM): Network structure: Input layer (12-dimensional) → Hidden layer (64 nodes, ReLU activation function) → Dropout layer (0.2, to prevent overfitting) → Output layer (1-dimensional, linear activation); Training parameters: Batch size 32, number of iterations 50, optimizer Adam, learning rate 0.001;

[0115] Training output: For each sample in the training set, the three base learners each output a predicted value, forming three columns of "base learner predicted features" (30660×3 dimensions).

[0116] Second layer: Meta-learner training

[0117] Ridge regression (RR) was selected as the meta-learner (to address the collinearity of features predicted by the base learners; its advantages in handling collinearity are emphasized in the appendix).

[0118] Input: The three columns of "base learner predicted features" from the first layer output;

[0119] Output: Final wind speed prediction (matched with training set labels);

[0120] Hyperparameter optimization: Through 5-fold cross-validation, the regularization parameter α was determined to be 0.1 (at which point the RMSE is minimized).

[0121] LSTM performs best in capturing wind speed time series dependence (such as the trend of increasing wind speed in the afternoon) (training set RMSE=1.3m / s), LightGBM performs best in fitting nonlinear terrain-wind speed relationships (RMSE=1.4m / s), and RF has the strongest stability (RMSE=1.5m / s). The advantages of the three can be integrated by fusing them through meta-learners.

[0122] Model optimization

[0123] Hyperparameters were optimized using a combination of 5-fold cross-validation and grid search.

[0124] Cross-validation: Divide the training set into 5 groups, each group is used as the validation set in turn, and the rest are used as the training set. Calculate the average RMSE of the 5 validations.

[0125] Optimization results: The average RMSE of the fusion model training set is 1.1m / s, which is 15.4% and 21.4% lower than that of single LSTM (1.3m / s) and LightGBM (1.4m / s), respectively, and the generalization error is significantly reduced.

[0126] Model Validation

[0127] Using the test set (2022 data) as the object, the following indicators were used for verification (compared with CLDAS real-world data):

[0128]

[0129] Typical case: 0-6h wind speed forecast for Jiuquan-Yumen area (40°N, 97°E) on April 15, 2022 - The actual wind speed at 10:00 was 10.5 m / s, the fusion model predicted 10.2 m / s (error 0.3 m / s), and the single LSTM predicted 9.1 m / s (error 1.4 m / s), demonstrating the advantages of the fusion model.

[0130] S5. Use the trained machine learning fusion forecast model to perform refined grid-based forecasting of wind speed for future periods in the complex terrain.

[0131] Specifically: Forecast lead time: 0-6h (short-term lead time, adapted to wind power dispatching needs);

[0132] Output format: 3km×3km grid product (120×150 grid points), including hourly 10m wind speed and wind direction forecast values, in NetCDF format (easy for meteorological and wind power departments to read);

[0133] Visualization: Using Python matplotlib, wind speed forecast contour maps of the Hexi region are drawn, and strong wind areas (≥10.2m / s) are marked. For example, if the forecast shows that strong winds of 12-14m / s will occur in the Guazhou area of ​​Jiuquan (40.5°N, 95°E) on a certain day, month, and hour, a warning is issued to the local wind farm 6 hours in advance.

[0134] As can be seen from the above, this embodiment achieves refined short-term wind speed forecasting in the Hexi region through the technical route of "multi-source data downscaling → wind field pattern mining → integrated fusion modeling". The core effects are as follows:

[0135] The spatiotemporal resolution has been improved from a large scale of 14km (ECMWF) to 3km grid points, which can capture local topographic wind fields;

[0136] The test set RMSE is 1.2 m / s, which is 25%-40% lower than the current industry level, improving forecast accuracy and meeting the need to improve the utilization rate of wind energy resources;

[0137] This helps wind farms reduce power prediction bias, aligning with the goal of "carbon peaking and carbon neutrality".

[0138] It should be understood that numerous specific implementation decisions can be made during the development of any practical implementation, such as in any engineering or design project. Such development efforts may be complex and time-consuming, but for those skilled in the art who benefit from this disclosure, the development effort will be a routine work of design, manufacturing, and production without requiring much experimentation.

[0139] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for refined short-term wind speed forecasting under complex terrain based on a fusion algorithm, characterized in that, include: S1. Acquire large-scale wind field background data, ground real-time wind field data, and high-precision terrain data of the target complex terrain; S2. Based on the high-precision terrain data, the large-scale wind field background data is refined using dynamic downscaling technology to construct a high spatiotemporal resolution grid dataset under complex terrain. S3. Based on the high spatiotemporal resolution grid dataset and the ground real-time wind field data, analyze the spatiotemporal distribution and evolution characteristics of wind fields in complex terrain, extract weather and climate influencing factors related to wind speed, and establish a wind speed forecast influencing factor database. S4. Using the high spatiotemporal resolution grid dataset and the wind speed forecast influencing factor library as input, construct a machine learning fusion forecast model using an ensemble learning stacking framework. S5. Use the trained machine learning fusion forecast model to perform refined grid-based forecasting of wind speed for future periods in the complex terrain.

2. The method for refined short-term wind speed forecasting under complex terrain based on a fusion algorithm as described in claim 1, characterized in that, The large-scale wind field background data includes ECMWF fine-grid numerical forecast data and ERA5 reanalysis data; The high-precision terrain data includes SRTM3 terrain data; The ground-based real-time wind field data includes CLDAS ground-based wind speed analysis data and observation data from the complex terrain meteorological station.

3. The method for refined short-term wind speed forecasting under complex terrain based on a fusion algorithm as described in claim 1, characterized in that, The dynamic downscaling technique is used to reduce the resolution scale of the initial wind field in the large-scale wind field background data to a high-resolution grid that matches the high-precision terrain data.

4. The method for refined short-term wind speed forecasting under complex terrain based on a fusion algorithm according to claim 1, characterized in that, The methods for analyzing the spatiotemporal distribution and evolution characteristics of wind fields in complex terrain include empirical orthogonal function analysis, Mann-Kendall trend test and sliding T test, which are used to analyze the spatial distribution, temporal evolution and abrupt changes of surface wind fields under complex terrain conditions. The wind concentration degree and concentration period index are used to analyze the distribution pattern of surface wind fields in complex terrain within the year. The method for extracting weather and climate influencing factors includes principal component analysis, which is used to analyze the relationship between surface wind field and influencing factors in complex terrain, and to screen key influencing factors that play a dominant role in wind speed changes. The influencing factors include circulation index, upper and lower level jet stream index, and vertical velocity.

5. A method for refined short-term wind speed forecasting under complex terrain based on a fusion algorithm as described in claim 1, characterized in that, The integrated learning stacking framework includes a first layer of base learners and a second layer of meta-learners. The first base learner consists of three heterogeneous algorithms: LightGBM, Random Forest, and Long Short-Term Memory Neural Network. The second-layer meta-learner uses the ridge regression algorithm to fuse the outputs of the base learners.

6. A method for refined short-term wind speed forecasting under complex terrain based on a fusion algorithm as described in claim 5, characterized in that, The method of constructing a machine learning fusion prediction model using an ensemble learning stacking framework includes: Wind speed data and key factors from the wind speed forecast influencing factor library are extracted from the high spatiotemporal resolution grid dataset. An M×N dimension sample feature matrix is ​​constructed by combining the ground real-time wind field data, and the training dataset and test dataset are divided, where M is the number of samples and N is the number of features. The three base learners, LightGBM, Random Forest and Long Short-Term Memory Neural Network, are trained independently using training samples. After training, the training prediction results of each base learner are output. The training prediction results are used as new features and input into the ridge regression learner for training to obtain the final stacked fusion model.

7. A method for refined short-term wind speed forecasting under complex terrain based on a fusion algorithm as described in claim 6, characterized in that, During the construction of the machine learning fusion prediction model, cross-validation and grid search methods are used to adjust the hyperparameters of the base learner and meta-learner, and the test dataset is input into the optimized machine learning fusion prediction model to obtain the model prediction results. The model's forecasting capability is quantitatively evaluated by comparing the model's forecasting results with the actual ground wind field data in the test dataset.