10-meter wind forecasting method and system
Through the Stacking model combined with multiple basic models and feature parameter extraction processing, the error and uncertainty problems of a single algorithm for predicting 10-meter wind are solved, and the accuracy and applicability of 10-meter wind forecast is improved.
Patent Information
- Application Number
- CN202510222508.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, a single algorithm is used to predict 10-meter wind with errors and uncertainties, which is difficult to meet the needs of meteorological forecasting services.
The Stacking model is used to combine multiple base models (such as lasso regression model, decision tree model, random forest model, support vector machine model) for 10-meter wind forecasting, and the prediction accuracy is improved through feature parameter extraction and Yeo-Johnson transformation processing.
It improves the accuracy of 10-meter wind forecast, reduces errors and uncertainties, and can better meet the needs of different application scenarios.
Smart Images

Figure CN120143307A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of meteorology, and more particularly, to a 10-meter wind forecasting method and system. Background Art
[0002] In meteorological observations, it is stipulated that an anemometer be installed at a height of 10 meters above the ground to measure wind speed, which is a globally accepted standard height. Such a unified standard height ensures the comparability of wind speed data in different regions and at different times. The wind at a height of 10 meters has a very direct impact on human daily life and some production activities. In urban planning, it is necessary to consider the wind environment at a height of 10 meters to ensure good ventilation between buildings and avoid local strong winds or eddies that may affect the living comfort and safety of residents. In agricultural production, the wind at a height of 10 meters affects the growth of crops. For example, strong winds may cause crops to lodge, affecting yields. In the field of transportation, wind speed changes affect the driving speed and safety of vehicles and trains. Strong winds may lead to accidents such as vehicle out of control and train derailment. 10m wind forecasting can help traffic management departments issue early warnings and take measures such as speed limits and service suspensions to ensure transportation safety. In maritime transportation, the sea wind conditions are complex, and 10m wind forecasting is crucial for ship navigation. Ships can adjust their routes and speeds according to the forecast to avoid sailing in bad sea conditions, reducing the probability of maritime accidents and ensuring the safety of crew members and cargo.
[0003] For the industrial and energy sectors, the wind speed at a height of 10 meters is an important reference index for evaluating wind energy resources. By long-term observation and analysis of the wind speed at a height of 10 meters, the characteristics of the local wind conditions, such as the average wind speed and the range of wind speed changes, can be roughly understood, so as to determine whether the area is suitable for building a wind farm. In industries such as aviation and petrochemicals, 10m wind forecasting also affects production operation safety. For example, in the aviation field, the ground wind speed needs to be considered during aircraft takeoff and landing, and 10m wind forecasting can assist airports in reasonably arranging flight takeoffs and landings; outdoor operations in industries such as petrochemicals also need to take safety protection measures according to wind condition forecasts to avoid safety accidents caused by strong winds.
[0004] Studying the wind at a height of 10 meters has important scientific value for understanding the dynamic processes of the atmospheric boundary layer, the propagation laws of wind, etc. 10m wind forecasting plays an important role in meteorological observations, human activities, wind speed highly sensitive related fields, and atmospheric science research. In-depth research and accurate understanding of it are of great significance.
[0005] The forecast of 10m wind can be made based on the observed data of meteorological elements such as long-term sequence of wind speed, wind direction, air pressure, temperature, relative humidity, etc. at meteorological observation stations. By using statistical methods such as time series analysis and regression analysis, and according to the variation laws and correlations among these elements, a statistical equation is established to predict the future 10m wind. However, this method is only based on the statistical relationship of historical data and has the characteristics of locality. Especially for some special weather changes or sudden meteorological events that are not fully reflected in historical data, it is difficult to accurately predict, and its adaptability is relatively weak. In addition, with the development of meteorological science, various meteorological large models have emerged one after another. These models have their own advantages and characteristics in the forecast of meteorological elements. However, when using a single model to forecast the 10m wind in a specific area such as an economic development zone, a key ecological area, an offshore engineering project, etc., there are often certain errors and uncertainties, which are difficult to meet the requirements of meteorological forecast services. Summary of the Invention
[0006] The embodiments of the present application provide a 10m wind forecasting method and system to at least solve the problem of errors and uncertainties in predicting 10m wind using a single algorithm in the related art.
[0007] According to one aspect of the present application, a 10m wind forecasting method is provided, including: collecting meteorological data, and extracting characteristic parameters from the collected meteorological data, where the characteristic parameters are characteristic parameters related to the 10m wind forecast; establishing a correspondence between the characteristic parameters and the 10m wind, and using the data with the established correspondence as the first training data set, where the first training data set includes multiple groups of training data; dividing the first training data set into multiple training data sets, and inputting one training data set for each base model in the first layer of the Stacking model; where at least two base models are included in the first layer of the Stacking model, and the types of the base models are at least one of the following: lasso regression model, decision tree model, random forest model, support vector machine model; obtaining the prediction results output by each base model, and using the prediction results as the second training data set to be input into the meta-model in the second layer of the Stacking model for training to obtain a trained Stacking model, where the meta-model in the second layer is a lasso regression model, and the trained Stacking model is used to output the predicted value of the 10m wind.
[0008] Further, inputting the prediction result as the second training data set into the meta-model of the second layer in the Stacking model for training includes: integrating the prediction results of each base model, performing Yeo-Johnson transformation processing on the integrated data to obtain the second training data; and inputting the second training data into the meta-model of the second layer in the Stacking model for training.
[0009] Further, extracting feature parameters from the collected meteorological data includes: performing CatBoost processing on all the feature parameters extracted from the meteorological data to obtain the feature parameters with a correlation degree higher than a threshold with respect to the 10m wind among all the feature parameters; and using the feature parameters higher than the threshold as the feature parameters for generating training data.
[0010] Further, the feature parameters for generating training data include at least one of the following: air pressure, temperature, terrain, humidity, vorticity, divergence, vertical velocity, temperature advection.
[0011] Further, the first layer in the Stacking model includes four base models, and the four base models are respectively: a lasso regression model, a decision tree model, a random forest model, and a support vector machine model.
[0012] According to another aspect of the present application, there is also provided a 10m wind forecasting system, including: a collection module, configured to collect meteorological data and extract feature parameters from the collected meteorological data, where the feature parameters are feature parameters related to 10m wind forecasting; a generation module, configured to establish a correspondence between the feature parameters and the 10m wind, and use the data with the established correspondence as the first training data set, where the first training data set includes multiple groups of training data; a first processing module, configured to divide the first training data set into multiple training data sets, and input one training data set for each base model in the first layer of the Stacking model; where the first layer in the Stacking model includes at least two base models, and the types of the base models are at least one of the following: a lasso regression model, a decision tree model, a random forest model, and a support vector machine model; a second processing module, configured to obtain the prediction results output by each base model, input the prediction results as the second training data set into the meta-model of the second layer in the Stacking model for training to obtain a trained Stacking model, where the meta-model of the second layer is a lasso regression model, and the trained Stacking model is used to output the predicted value of the 10m wind.
[0013] Further, the second processing module is configured to: integrate the prediction results of each base model, perform Yeo-Johnson transformation processing on the integrated data to obtain second training data; and input the second training data into the meta-model of the second layer in the Stacking model for training.
[0014] Further, the acquisition module is configured to: perform CatBoost processing on all feature parameters extracted from the meteorological data to obtain feature parameters with a relevance to the 10m wind higher than a threshold among all feature parameters; and use the feature parameters higher than the threshold as feature parameters for generating training data.
[0015] Further, the feature parameters for generating training data include at least one of the following: air pressure, temperature, terrain, humidity, vorticity, divergence, vertical velocity, temperature advection.
[0016] Further, the first layer of the Stacking model includes four base models, which are respectively: a lasso regression model, a decision tree model, a random forest model, and a support vector machine model.
[0017] In the embodiment of the present application, meteorological data is collected, feature parameters are extracted from the collected meteorological data, where the feature parameters are feature parameters related to the 10m wind forecast; a correspondence between the feature parameters and the 10m wind is established, and the data with the established correspondence is used as a first training data set, where the first training data set includes multiple groups of training data; the first training data set is divided into multiple training data sets, and each training data set is input into each base model in the first layer of the Stacking model; where the first layer of the Stacking model includes at least two base models, and the types of the base models are at least one of the following: a lasso regression model, a decision tree model, a random forest model, and a support vector machine model; prediction results output by each base model are obtained, and the prediction results are used as a second training data set and input into the meta-model of the second layer in the Stacking model for training to obtain a trained Stacking model, where the meta-model of the second layer is a lasso regression model, and the trained Stacking model is used to output a predicted value of the 10m wind. By means of the present application, the problems of errors and uncertainties in predicting the 10m wind using a single algorithm in the related art are solved, and the accuracy of the 10m wind forecast is improved to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0019] Figure 1 It is a schematic structural diagram of an improved Stacking integration model according to an embodiment of the present application;
[0020] Figure 2 It is a schematic flowchart of an improved Stacking integration method according to an embodiment of the present application;
[0021] Figure 3 It is a 1h wind speed prediction curve graph according to an embodiment of the present application;
[0022] Figure 4 It is a 3h wind speed prediction curve graph according to an embodiment of the present application;
[0023] Figure 5 It is a schematic structural diagram of the principle of the Stacking model according to an embodiment of the present application; and,
[0024] Figure 6 It is a flowchart of a 10m wind prediction method according to an embodiment of the present application. Detailed implementation manners
[0025] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0026] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0027] In the following embodiments, a solution for realizing the 10m wind prediction in the target area is provided by using the meteorological large model prediction data, combining multi-source meteorological observation data, etc., through an improved Stacking integration and result correction. The Stacking model will be described first below.
[0028] In the Stacking model, feature extraction determines the upper limit of the model, and the choice of the model is to continuously approach this upper limit. When the features have been extracted and the machine learning model algorithm has also been selected, to further improve the performance of the model, the Stacking method is an algorithm that can help the model improve further. Stacking is a hierarchical fusion model. Ensemble learning can integrate the excellent features of different models to improve the final prediction effect. Generally speaking, the requirements for each base learner model in ensemble learning are to be good, but also different. That is to say, the performance of each base learner model cannot be too poor, and different base learner models should also have different learning characteristics. Its main idea is to learn and generalize the prediction results of the base learners through training the model. To make the Stacking fusion effect obvious, the following conditions need to be met: ① The dataset is sufficient; ② The performance of the base learner models is excellent; ③ The number of base models is not too small.
[0029] Figure 5 is a schematic diagram of the principle structure of the Stacking model according to an embodiment of the present application, as Figure 5 shown, there are two layers of learners in the Stacking model, including the first-layer learner and the second-layer learner. In the first-layer learner, there are base learner 1, base learner 2 up to base learner n, and the second-layer learner is the meta-learner. First, the original data is divided into several sub-datasets, and these sub-datasets are input into each base learner of the first-layer prediction model. Each base learner outputs its own prediction result, and the prediction results are integrated into a result matrix. Then the output prediction result is used as the input of the meta-learner of the second-layer model and is trained, and the meta-learner outputs the final prediction result. Combining with neural networks (NNs), it can be understood that because the first-layer model extracts deeper and more relevant features of the result, the output of the second layer will be better. Stacking ensemble learning generalizes and learns the results of multiple models, which is equivalent to the deep feature extraction in NNs, thereby improving the overall prediction accuracy. Precautions for Stacking are: ① The models are strong and weak. Strong models should be selected for the first layer, and weak models (such as linear model LR, etc.) should be selected for the second layer to further avoid overfitting; ② Different types of models should be selected as much as possible for the models in the first layer, good but different, to achieve feature extraction in multiple aspects.
[0030] The Stacking algorithm is an effective ensemble method. It uses the predictions generated by different classifiers as the input of the next-layer learning algorithm. The Stacking algorithm is generated by different classification algorithms, can integrate the learning mechanisms of different learning algorithms, and can produce a higher accuracy than any of the base classification algorithms that make it up. By conducting empirical analysis on the Stacking algorithm and other general ensemble methods on simulated data and real data, it is found that the Stacking method performs better than other general ensemble methods.
[0031] In the following embodiments, a 10-meter wind forecasting method is provided. Figure 6 It is a flowchart of the 10-meter wind forecasting method according to the embodiments of the present application. The steps involved in the method in Figure 6 will be described below.
[0032] Step S602: Collect meteorological data, and extract characteristic parameters from the collected meteorological data. Among them, the characteristic parameters are characteristic parameters related to 10-meter wind forecasting.
[0033] In this step, as a relatively preferred embodiment, all the characteristic parameters extracted from the meteorological data can be processed by CatBoost to obtain the characteristic parameters whose correlation with the 10m wind among all the characteristic parameters is higher than the threshold; the characteristic parameters higher than the threshold are used as the characteristic parameters for generating training data.
[0034] Step S604: Establish the corresponding relationship between the characteristic parameters and the 10-meter wind, and use the data with the established corresponding relationship as the first training data set. Among them, the first training data set includes multiple groups of training data.
[0035] Step S606: Divide the first training data set into multiple training data sets, and input one training data set for each base model in the first layer of the Stacking model; among them, at least two base models are included in the first layer of the Stacking model, and the types of the base models are at least one of the following: lasso regression model, decision tree model, random forest model, support vector machine model. For example: four base models are included in the first layer of the Stacking model, and the four base models are respectively: lasso regression model, decision tree model, random forest model, support vector machine model.
[0036] Step S608: Obtain the prediction results output by each base model, and use the prediction results as the second training data set to be input into the meta-model in the second layer of the Stacking model for training to obtain a trained Stacking model. Among them, the meta-model in the second layer is a lasso regression model, and the trained Stacking model is used to output the predicted value of the 10-meter wind.
[0037] In this step, as a relatively preferred embodiment, the prediction results of each base model can be integrated, and the integrated data is processed by Yeo-Johnson transformation to obtain the second training data; the second training data is input into the meta-model in the second layer of the Stacking model for training.
[0038] By the above steps, the problems of errors and uncertainties in predicting 10-meter wind in the related art using a single algorithm are solved, and the accuracy of 10-meter wind forecasting is improved to a certain extent.
[0039] The following uses examples to illustrate the above Figure 6 steps. In the following examples, an improved Stacking method is used to integrate and correct the forecasting results of multiple large models based on multiple mainstream meteorological large models, aiming to solve the problems of errors in the 10m wind forecasting of existing single meteorological models and the difficulty in fully meeting the requirements of different application scenarios for 10m wind forecasting. Through integrating multi-source meteorological observation data and multiple meteorological large models, a 10m wind forecasting method based on improved Stacking ensemble learning and result correction is provided. In this embodiment, a 10m wind forecasting method based on improved Stacking ensemble learning and result correction using the forecasting data of meteorological large models is provided. The following describes the steps included in this method.
[0040] Step S1: Multi-source data collection and processing. In this step, meteorological observation data, historical operation data of the wind farm, and 10m wind forecasting data from meteorological large models in the target area are obtained, and the data is cleaned and standardized. It should be noted that data cleaning here may include deleting abnormal data or incorrect data, etc., and standardization processing may include unifying multi-source data into a predetermined format.
[0041] In this step S1, 10m wind forecasting data of multiple meteorological large models for the next three days in the target area can be collected to ensure that the data formats output by each model's forecast meet the requirements of subsequent processing, and the data covers key information such as corresponding time series and spatial coordinates. At the same time, multi-source meteorological observation data is collected. The multi-source meteorological observation data here includes but is not limited to: measured 10m wind data of ground meteorological observation stations, Doppler radar wind profile data, satellite remote sensing inversion 10m wind data, etc. Quality control is performed on these observation data. The quality control here includes data cleaning, such as removing outliers and incorrect data, and unified spatio-temporal registration is performed to make it match the meteorological large model forecasting data in terms of spatio-temporal scale. It should be noted that the spatio-temporal registration here may include unifying the data according to the time series and spatial coordinates.
[0042] Step S2: Feature parameter optimization. In this step, according to the data obtained in step S1, the feature parameters affecting 10m wind forecasting are optimized. For example, the CatBoost method can be used to optimize the feature parameters affecting 10m wind forecasting, and according to the correlation procedure, the ranking of the correlations of each feature is obtained.
[0043] CatBoost (categorical boosting) is a gradient boosting algorithm library that can handle categorical features well. It is a machine learning algorithm open-sourced by Yandex. It can be easily integrated with deep learning frameworks. It can handle various data types. CatBoost is a GBDT framework with fewer parameters, support for categorical variables, and high accuracy, implemented based on symmetric decision trees (oblivious trees) as the base learner. It mainly solves the problem of efficiently and reasonably handling categorical features, which can be seen from its name, as CatBoost is composed of Categorical and Boosting. In addition, CatBoost also solves the problems of Gradient Bias and Prediction Shift, thereby reducing the occurrence of overfitting and improving the accuracy and generalization ability of the algorithm.
[0044] In step S2, CatBoost is used to perform feature selection on the data affecting the 10m wind forecast, rank the importance of each highly correlated parameter, and eliminate redundant parameters. This method is an advanced gradient boosting algorithm with powerful categorical feature processing capabilities, high accuracy, and generalization ability. It makes up for the inefficiency of the parameter tuning process in XGBoost when there are many parameters, and also overcomes the problems of easy overfitting and excessive number of iterations in GBDT. By using the CatBoost algorithm, the air pressure, temperature, terrain, humidity, vorticity, divergence, vertical velocity, temperature advection, etc. predicted by each major model are selected as the preferred input models to reduce the impact of too many parameters on the model training efficiency.
[0045] Step S3: Selection of the base model (i.e., the base learner) and the meta-model (i.e., the original learner) in the Stacking model. In this step, the lasso regression model, decision tree model, random forest model, and support vector machine are used as the base models for the initial layer (i.e., the first layer), the lasso regression model is used as the meta-model for the second layer, and the optimized feature parameters in step S2 are used as the training data and input into the initial layer (i.e., the first layer) for training.
[0046] In this step, the Stacking ensemble learning algorithm is used to select a base model suitable for the 10m wind forecast, construct a two-layer learning architecture, use the S2 feature parameters as the input of the training data for the initial layer, and after training, input it into the second-layer meta-model for training and testing. The purpose is to enhance the normal distribution of the training data of the second-layer meta-model and improve the prediction accuracy.
[0047] Step S4: Improvement of the Stacking model. In S3, the training data of the initial layer base model is processed by the Yeo-Johnson transformation and then input into the second-layer model for continued training and testing, and finally the forecast result is obtained.
[0048] The Yeo-Johnson transformation is mainly applied to the data preprocessing stage of data mining and machine learning. When the normality of a random variable is poor, using the Yeo-Johnson transformation for preprocessing is beneficial for performing statistical analysis on the random variable based on the normal assumption. In the context of machine learning, the optimization method of the Yeo-Johnson transformation usually involves determining the transformation coefficient. This transformation coefficient can be determined by maximum likelihood estimation. In practical applications, the transformation coefficient can be estimated from the learning samples and then substituted into the test samples for calculation.
[0049] Step S5: Forecast result correction.
[0050] In this step, the forecast result obtained in S4 can be used with R 2 to test the goodness of fit of the model, and E RMSE , E MAE to test the prediction accuracy, and achieve the 10m wind forecast in the target area.
[0051] In the above steps, one of the ways to improve the accuracy of the 10m wind forecast is to improve the data normality and feature correlation in the training and testing of the initial layer-based model. An effective way is to perform the Yeo-Johnson transformation on the training set of the second layer input to the initial layer. This effectively improves the normality, additivity, homoscedasticity, stability, and regularity of the training data, thereby improving the fitting and forecasting ability of the data in the second-layer meta-model. At the same time, it improves the internal relationship between each feature parameter and the 10m wind data, helps the model better capture the complex relationship between the feature parameter and the 10m wind, and thus more accurately forecasts the 10m wind. When dealing with the correlation between the feature parameter and the 10m wind in the above steps, the CatBoost method is used to optimize the feature parameter. This method can automatically handle the order problem of categorical features, so that the model will not have large fluctuations in results due to changes in the order of categorical features, ensuring the robustness and prediction accuracy of the model. Then it is input into the initial layer-based model as training data; when the training set of the second layer is input to the initial layer, the Yeo-Johnson transformation is used to improve the normal distribution and forecasting accuracy of the training data; finally, the forecast result is tested to obtain the 10m wind forecast in the target area.
[0052] In the above steps, from the utilization of meteorological large model data, to the optimization of characteristic parameters, then to the conversion and processing of training data, and finally to the inspection and evaluation of the forecast results, the 10m wind forecast becomes more accurate. The above steps also comprehensively utilize multiple meteorological large models, integrate the advantages of different models, overcome the limitations of single-model forecasting, enable the 10m wind forecast to consider comprehensively from multiple perspectives and multiple information sources, and improve the accuracy and comprehensiveness of the forecast. For the algorithms in the above steps, an improved ensemble learning algorithm is used for model integration, which can perform reasonable weighting or optimized integration according to the performance characteristics of different models, further explore the potential effective information of each model, and enhance the accuracy and stability of the 10m wind forecast results. In summary, by combining multi-source meteorological observation data for bias correction, the actual observed real meteorological conditions are fully utilized, the systematic error between the forecast data and the actual 10m wind conditions is effectively reduced, and the final output of the 10m wind forecast data for the target area is more reliable, which can better serve many fields with high requirements for the 10m wind forecast accuracy, such as wind energy development, aviation and navigation, and urban environmental meteorological support.
[0053] Figure 1 is a schematic structural diagram of an improved Stacking ensemble model according to an embodiment of the present application, as Figure 1 shown, first, a data set including original characteristic parameters is obtained. The data set includes meteorological observation data, forecast data in a meteorological large model, terrain data, etc. Then, the relevant characteristic parameters are screened according to the Pearson correlation coefficient, and the optimal characteristic parameters are extracted through CatBoost. It should be noted that the extraction of the optimal characteristic parameters here can be to sort various parameters according to the correlation degree of the 10-meter wind prediction, and then select a predetermined number of characteristic parameters that have the greatest impact on the 10-meter wind prediction. These characteristic parameters are used to generate training data and input into the initial layer base model learning library. The initial layer base model learning library can include: a lasso regression model, a decision tree model, a random forest model, and a support vector machine model. After the initial layer training, an initial layer base learner training set is obtained. In the second-layer model, the training set of the initial layer base learner is processed through transformation and then input into the lasso regression model. This model outputs the 10m wind prediction result, and finally, the 10m wind forecast error value is calculated.
[0054] Figure 2 is a schematic flow diagram of an improved Stacking ensemble method according to an embodiment of the present application, as Figure 2As shown, after obtaining the optimal feature parameter data set, it can be divided into a test set and a training set, wherein the training set is divided into four parts and respectively input into the lasso regression model, decision tree model, random forest model and support vector machine model, and the data input into these four models are further divided into test data 1 to 4 and training data 1 to 4, and these four models respectively predict the results 1, 2, 3 and 4, and these four prediction results are transformed as the training set again as the training set of the second-layer meta-model, and the test set split from the optimal feature parameter data set is transformed as the test set of the second-layer meta-model, and then the second-layer meta-model outputs the prediction value.
[0055] Figure 1 and Figure 2 The model shown in the figure is based on the short-term wind forecasts of multiple large meteorological models, based on the improved Stacking method and result correction, to establish a wind prediction model to improve the forecast accuracy. Figure 1 and Figure 2 The method steps involved are described in detail.
[0056] (1) Obtain large-scale model forecast data, meteorological observation data, and terrain data.
[0057] Collect 10m wind forecast data for the next three days for the target area from multiple large meteorological models to ensure that the data format of each model forecast output meets the requirements of subsequent processing and that the data covers key information such as the corresponding time series and spatial coordinates. At the same time, collect multi-source meteorological observation data, including but not limited to the measured 10m wind data from ground meteorological observation stations, Doppler radar wind profile data, satellite remote sensing inversion 10m wind data, etc., clean these observation data, remove outliers and erroneous data, and uniformly perform spatiotemporal registration to match the forecast data of the large meteorological model in terms of spatiotemporal scale.
[0058] (2) Improvement of forecast model parameter characteristics.
[0059] In terms of parameter feature optimization, the CatBoost method can be used to optimize the data affecting the 10m wind forecast. This method is an advanced gradient boosting algorithm with powerful classification feature processing capabilities and high accuracy and generalization capabilities. It makes up for the shortcomings of XGBoost in the case of more parameters, the parameter adjustment process is not efficient, and it also overcomes the problem of GBDT's easy overfitting and too many iterations. The CatBoost algorithm selects single or composite forecast parameters of each model, mainly 3-hour pressure change, air pressure, temperature, temperature advection, humidity, vorticity, divergence, vertical velocity and terrain as the preferred feature parameters input model to reduce the impact of too many parameters on model training efficiency.
[0060] In this embodiment, the CatBoost algorithm consists of two parts: the objective function and gradient boosting iteration:
[0061] ① The Obj(θ) is the objective function, which consists of a loss function and a regularization term:
[0062]
[0063] represents the loss function, which is used to measure the difference between the predicted wind speed value and the true value y. The loss function adopts logarithmic loss (Log Loss):
[0064]
[0065] Ω(θ) is the regularization term, which is used to control the complexity of the CatBoost model, prevent overfitting, and impose regularization constraints related to the structure of the CatBoost tree, such as the depth of the tree and the number of leaf nodes.
[0066] ② The gradient boosting iteration process.
[0067] CatBoost uses gradient boosting to iteratively train the model and gradually constructs weak learners (decision trees) to minimize the objective function. At the m-th iteration, given the training dataset x i represents the input feature vector, y i represents the corresponding target value, and n is the number of samples. First, calculate the negative gradient of the loss function with respect to the previous round's predicted value as the pseudo-residuals for this round of training. For regression tasks:
[0068]
[0069] r im represents the pseudo-residual of the i-th sample at the m-th iteration. The pseudo-residual is a value calculated in the gradient boosting algorithm to enable the newly constructed weak learner (decision tree) to fit the deficiencies of the previous round's model. It reflects the difference direction and degree between the current model's predicted value and the true value, and is an important basis for subsequent training. is the predicted value of the model for the i-th sample at the (m - 1)-th iteration. is the (m - 1)-th iteration, and the model's prediction for the i-th sample. is the partial derivative, which is used to measure the rate and direction of change of the loss function
[0070] is the partial derivative, which is used to measure the rate and direction of change of the loss function as changes. L represents the loss function, which is used to measure the difference between the model's predicted value and the true value yi Degree of difference Indicates the difference between the true value of the nth sample and the predicted value of the previous round
[0071] For classification tasks (taking the logistic regression loss of binary classification):
[0072]
[0073] Then, based on these pseudo-residuals, train a new decision tree T m (x i ), which maps the input feature vector x to a real value (for regression tasks, it is directly the predicted value, and for classification tasks, usually subsequent logical transformation and other steps are required). Then update the predicted value ŷ, and its update formula is as follows:
[0074] For regression tasks:
[0075]
[0076] For classification tasks:
[0077]
[0078] where v is the learning rate, which controls the influence degree of the newly added tree on the final prediction result in each iteration. The value range is usually between (0, 1). A smaller learning rate means a slower learning speed, but it may make the model training more stable and avoid overfitting
[0079] For categorical feature x ij (indicating the jth categorical feature of the ith sample), its value range is {c 1 , c 2 , …, c n} (n is the number of categories). First, arrange the training data in a certain order (such as random order or order based on time, etc.), and then calculate statistics such as the mean of the target values in the samples before the current sample for each category c (taking the mean as an example) in the following way:
[0080]
[0081] where, [x sj = c] is the indicator function, which takes the value of 1 if x sj is equal to category c, otherwise 0; represents the mean statistic of the target value corresponding to category c at the tth sample
[0082] The correlation index of the optimal feature parameters selected by the CatBoost algorithm:
[0083] Pearson correlation coefficient
[0084]
[0085] This is an index that measures the strength of the correlation between wind speed and other data. In the formula, x represents the wind speed, y represents a single parameter or composite prediction parameter predicted by each large model, n is the number of parameters, and cov is the covariance between the wind speed and the feature parameters. The range of this coefficient is (-1, 1). The larger the value and the closer it is to 1, the stronger the positive correlation between the wind speed and the eigenvalue; the larger the value and the closer it is to -1, the stronger the negative correlation between the wind speed and the eigenvalue; the closer the value is to 0, the weaker the correlation between the wind speed and the eigenvalue.
[0086] The selection of the Stacking base model and the meta-model will be described below.
[0087] The 10m wind prediction data of various large models collected and the selected optimal feature parameters are respectively input into each base model for training. According to learning the features and laws in the wind speed prediction data for each base model, the main base models selected in the improved Stacking method are as follows:
[0088] 1) Lasso Regression model: It is a method similar to ridge regression. It can reduce the degree of variation and improve the accuracy of the linear regression model, and also penalize the absolute value of the regression coefficient. It adds a regularization term, that is, adds a constraint term, reduces unnecessary independent variables, and is convenient for model interpretation. Its objective function is the sum of the least squares fitting measure and the penalty term:
[0089]
[0090] where λ > 0 is the lasso penalty parameter, y is the result variable, x contains n potential covariates, β i is the i-th eigenvariable of β, ω j is the parameter-level weight, which is the penalty load, and n is the sample size.
[0091] 2) Decision Tree Model: It mainly deals with non - linear relationships, can automatically select important feature variables, has no strict assumptions about the distribution and features of input data, and has strong model interpretability. By learning a large amount of historical 10m wind data and related characteristic parameter data, a decision tree structure is constructed to predict the 10m wind, which is suitable for situations where data features are relatively complex. The decision tree algorithm uses the Gini index as an indicator for field selection. The smaller the index, the better the feature.
[0092] Gini index formula:
[0093]
[0094] R in the formula n is the probability that the nth value may occur. This probability is expressed by empirical probability and can also be simplified as:
[0095]
[0096] |F| represents all sample points in the forecast, and |C n | represents the number of times the nth value in the forecast appears. So the probability value R n is the frequency represented by.
[0097] 3) Random Forest Model: An ensemble learning model based on decision trees. By randomly sampling and constructing multiple decision trees and integrating their prediction results, the forecast accuracy is improved. It can effectively reduce the variance of the model, reduce the risk of overfitting, has good performance in dealing with high - dimensional data and complex non - linear relationships, and can capture more interactions and variation laws among meteorological elements in the 10m wind forecast.
[0098] Define the training data set X for wind speed prediction i→ Y i ,X i represents the feature vector value established from the meteorological element values of the ith sample, and Y i is the true value in the prediction model, mapped to the actual wind speed value of the ith sample in the data.
[0099] By searching for split variables and split values through the feature vector X and its corresponding true value Y in the training samples, the regression decision tree divides the entire vector space into j partitions {F 1 , F 2 , …, F j}. For any of these partitions, it can be mapped to the model R n . By the value of a certain feature, the vector space is divided into two parts, and the expression is:
[0100] F 1 (m, n) = {G|G m ≤n} (13);
[0101] F 2 (m, n) = {G|G m >n} (14);
[0102] Where m is a feature description and n represents the value at the time of splitting. The objective function for searching the vector empty table splitting variable and splitting value is:
[0103]
[0104] In the formula, w is the minimum variance of the measured wind speed value, Y i represents the actual observed value of the i-th sample, X i is the characteristic value corresponding to predicting the i-th sample, R 1 is the mean value of the first part of the actual wind speed value, R 2 is the mean value of the second part of the actual wind speed value.
[0105] 4) Support Vector Machine: A supervised learning algorithm that performs binary classification on data by finding an optimal hyperplane, which has great advantages in solving small-sample, high-dimensional, and non-linear regression problems. In the 10m wind forecast, the characteristic parameters can be used as input features and the 10m wind as the output target to construct a support vector regression model, which can better fit the complex relationship between the 10m wind and meteorological elements and improve the forecast accuracy.
[0106] The support vector machine searches for the optimal hyperplane in the feature space within the linear range based on the given data to achieve the optimal effect. The hyperplane equation formula is:
[0107] y(x) = ωx + b (16);
[0108] where ω is the weight vector and b is the bias term. For each data point x i , ideally, it is y i (ωx i +b) ≥ 1 to ensure that the data point x i is correctly classified and the distance from the hyperplane is at least 1,
[0109] Set it as the constraint condition and calculate the straight-line distance d from the hyperplane to the nearest point:
[0110]
[0111] ||ω|| is the parameter of the hyperplane.
[0112] Then, the Lagrange function is used to transform the established objective function into a dual problem, and the sequential minimal optimization algorithm is adopted to solve it, and ω and b are deduced:
[0113]
[0114] In the formula, the value of y can be 1 or -1. When the sample point is on the opposite side of the plane, it is -1, and vice versa, so as to ensure that the distance value is positive.
[0115] Stacking ensemble learning constructs a multi-layer model structure. The initial layer consists of 4 different base models that independently learn and train the given samples, trying to mine different patterns and rules from the data, and then output the corresponding prediction results. The meta-model set in the second layer (in this embodiment, the lasso regression model is selected) aims to use the output of the base models in the first layer as input, further integrate and refine the information, so as to generate the final prediction result.
[0116] Next, the Stacking ensemble model based on Yeo-Johnson transformation processing will be described.
[0117] (1) Yeo-Johnson transformation is a generalized power transformation method, which is applicable to data containing zero values and negative values, but does not require the data to be positive. It is useful for features with non-constant variance or non-normal distribution in standardization, and is especially suitable for data containing negative values and zero values and data for which Box-Cox is not applicable.
[0118]
[0119] (2) In this embodiment, innovatively, on the basis of the initial layer base model inputting the training set x=(x 1 ,x 2 ……x n ) to the second layer meta-model, a suitable optimal parameter λ is found after Yeo-Johnson transformation, and the data distribution is changed through the variation function X=(x,λ), aiming to enhance the normal distribution of the training data of the second layer meta-model and improve the prediction accuracy.
[0120] Taking python as an example
[0121] import numpy as np
[0122] from scipy.stats import YeoJohnson
[0123] base_model_outputs = np.array([1, 2, -3, 4, 0])
[0124] # Example of training set data output by the base model
[0125] transformed_data, λ = YeoJohnson(base_model_outputs)
[0126] # `transformed_data` is the data after Yeo - Johnson transformation, and `λ` is the optimal transformation parameter \(\lambda\) found.
[0127] **Apply the transformation to the training set data**
[0128] Substitute the training set data output by the base model into the above formula one by one (using the determined \(\lambda\)) to obtain the transformed training set data. These transformed data will be used as the input of the second - layer meta - model (lasso regression model).
[0129] **Precautions in meta - model training and evaluation**
[0130] When training the meta - model using the transformed data, note down the transformation parameter \(\lambda\), because when processing new data (such as test set data or data during actual prediction), the same parameter \(\lambda\) needs to be used for Yeo - Johnson transformation.
[0131] (3) The training set processed by Yeo - Johnson transformation is input into the second - layer meta - model. After the training and testing of the meta - model, and then through weighted averaging, the 10m wind forecast result is obtained.
[0132] (4) Specific steps of the improved Stacking ensemble model based on Yeo - Johnson:
[0133] Input: Training set \(L=\{(x_1,y_1),(x_2,y_2),...,(x_N,y_N)\}\) (obtain the initial - layer base model through training data \(D\));
[0134] for \(t = 1,2,...,T\)
[0135] for \(i = 1,2,...,k\)
[0136] The \(t\) - th homogeneous base learner in the initial layer is trained \(k\) times
[0137] Construct a new data sample \(D_{new}=\{(Z_{11},A_{11}),...,(Z_{TT},A_{TT})\}\)
[0138] After performing Yeo - Johnson transformation on the new data sample, input it into the first layer \(D_{transformed}=Yeo - Johnson\_Transform(D_{new})\)
[0139] The first - layer prediction model D_new is trained to obtain the model, D_new = {(Z_11,A_11),...,(Z_1T,A_1T)}
[0140] endfor
[0141] Calculate the training of the second - layer meta - model
[0142] {(y1,y2,...,yN),...,(y1,y2,...,yN)}-{(Z_11,Z_21,...,Z_N1),...,(Z_1T,Z_2T,...,Z_NT)}={(e_11,e_21,...,e_N1),...,(e_1T,e_2T,...,e_NT)}
[0143] Calculate the weighted average P_11=(e_21 +...+e_N1) / (e_11+e_21 +...+e_N1) ρ_1k=(e_1k +...+e_Nk) / (e_11+e_21 +...+e_N1)
[0144] Output P(x)=p(p_1k,...,p_Nk)
[0145] End
[0146] The following is an explanation of the evaluation of the 10m wind forecast result.
[0147] The coefficient of determination R2 is used to test the goodness of fit of the model, and the root - mean - square error ERMSE and the mean absolute error EMAE are used to test the 10m wind forecast accuracy, to solve the error and uncertainty problems in wind speed prediction. By combining information such as actual observed data, the powerful non - linear fitting ability of deep learning is used to optimize the wind speed prediction effect.
[0148] Root - mean - square error:
[0149]
[0150] Mean absolute error:
[0151]
[0152] Coefficient of determination:
[0153]
[0154] In the formula, is the predicted wind speed at the i - th moment of the day; is the measured wind speed at the - th moment of the day; n is the total number of 10m wind speed samples.
[0155] Figure 3 is the 1h wind speed prediction curve graph according to the embodiments of the present application, Figure 4 is the 3h wind speed prediction curve graph according to the embodiments of the present application, Figure 3 and Figure 4 forecast the 10m wind speed for 1h and 3h, and R 2 respectively reach 0.94 and 0.98, and the prediction results are excellent.
[0156] In Figure 3 the data is as follows in the table:
[0157] EMAE MSE <![CDATA[R 2 > 0.58 0.69 0.94
[0158] In Figure 4 the data is as follows in the table:
[0159] EMAE MSE <![CDATA[R 2 > 0.38 0.24 0.98
[0160] In this embodiment, an electronic device is provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method in the above embodiments.
[0161] The above program can run in the processor, or can also be stored in the memory (or referred to as a computer-readable medium). The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0162] These computer programs can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate computer-implemented processing. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. Corresponding to different steps, different modules can be implemented.
[0163] In this embodiment, such a device or system is provided. The system is called a 10-meter wind prediction system, including: a collection module for collecting meteorological data and extracting characteristic parameters from the collected meteorological data, where the characteristic parameters are characteristic parameters related to 10-meter wind prediction; a generation module for establishing a correspondence between the characteristic parameters and the 10-meter wind, and using the data with the established correspondence as a first training data set, where the first training data set includes multiple groups of training data; a first processing module for dividing the first training data set into multiple training data sets, and inputting one training data set for each base model in the first layer of the Stacking model; where at least two base models are included in the first layer of the Stacking model, and the types of the base models are at least one of the following: lasso regression model, decision tree model, random forest model, support vector machine model; a second processing module for obtaining the prediction results output by each base model, and using the prediction results as a second training data set to be input into the meta-model in the second layer of the Stacking model for training to obtain a trained Stacking model, where the meta-model in the second layer is a lasso regression model, and the trained Stacking model is used to output the predicted value of the 10-meter wind.
[0164] This system or device is used to implement the functions of the method in the above embodiment. Each module in this system or device corresponds to each step in the method, and those that have been described in the method will not be repeated here.
[0165] Optionally, the second processing module is used to: integrate the prediction results of each base model, perform Yeo-Johnson transformation processing on the integrated data to obtain second training data; and input the second training data into the meta-model in the second layer of the Stacking model for training.
[0166] Optionally, the collection module is used to: perform CatBoost processing on all the characteristic parameters extracted from the meteorological data to obtain the characteristic parameters with a relevance to the 10m wind higher than a threshold among all the characteristic parameters; and use the characteristic parameters higher than the threshold as the characteristic parameters for generating training data.
[0167] By the above embodiments, the problems of errors and uncertainties in predicting the 10-meter wind using a single algorithm in the related art are solved, and the accuracy of 10-meter wind prediction is improved to a certain extent.
[0168] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A 10-meter wind forecast method, characterized in that: include: Collecting meteorological data, and extracting characteristic parameters from the collected meteorological data, wherein the characteristic parameters are characteristic parameters related to 10-meter wind forecast; Establishing a corresponding relationship between the characteristic parameter and the 10-meter wind, and using the data for which the corresponding relationship is established as a first training data set, wherein the first training data set includes multiple sets of training data; Divide the first training data set into a plurality of training data sets, and input one training data set into each base model in the first layer of the Stacking model; wherein the first layer of the Stacking model includes at least two base models, and the type of the base model is at least one of the following: a lasso regression model, a decision book model, a random forest model, and a support vector machine model; The prediction results output by each base model are obtained, and the prediction results are input as the second training data set into the meta-model of the second layer in the Stacking model for training to obtain a trained Stacking model, wherein the meta-model of the second layer is a lasso regression model, and the trained Stacking model is used to output the prediction value of 10-meter wind.
2. The method according to claim 1, characterized in that Inputting the prediction result as a second training data set into the meta-model of the second layer in the Stacking model for training includes: The prediction results of each base model are integrated, and the integrated data are processed by Yeo-Johnson transformation to obtain the second training data; The second training data is input into the meta-model of the second layer in the Stacking model for training.
3. The method according to claim 1, characterized in that Extracting characteristic parameters from the collected meteorological data includes: Perform CatBoost processing on all characteristic parameters extracted from the meteorological data to obtain characteristic parameters whose correlation with the 10m wind is higher than a threshold value among all characteristic parameters; The feature parameters that are higher than the threshold are used as the feature parameters for generating training data.
4. The method according to claim 3, characterized in that The characteristic parameters used to generate training data include at least one of the following: air pressure, temperature, terrain, humidity, vorticity, divergence, vertical velocity, and temperature advection.
5. The method according to any one of claims 1 to 4, characterized in that The first layer of the Stacking model includes four base models, which are: lasso regression model, decision book model, random forest model, and support vector machine model.
6. A 10-meter wind forecast system, characterized in that: include: A collection module, used for collecting meteorological data, and extracting characteristic parameters from the collected meteorological data, wherein the characteristic parameters are characteristic parameters related to the 10-meter wind forecast; A generating module, used for establishing a corresponding relationship between the characteristic parameter and the 10-meter wind, and using the data for which the corresponding relationship is established as a first training data set, wherein the first training data set includes a plurality of sets of training data; A first processing module is used to divide the first training data set into a plurality of training data sets, and input a training data set for each base model in the first layer of the Stacking model; wherein the first layer of the Stacking model includes at least two base models, and the type of the base model is at least one of the following: a lasso regression model, a decision book model, a random forest model, and a support vector machine model; The second processing module is used to obtain the prediction results output by each base model, and input the prediction results as the second training data set into the meta-model of the second layer in the Stacking model for training to obtain a trained Stacking model, wherein the meta-model of the second layer is a lasso regression model, and the trained Stacking model is used to output the prediction value of 10-meter wind.
7. The system according to claim 6, characterized in that The second processing module is used for: The prediction results of each base model are integrated, and the integrated data are processed by Yeo-Johnson transformation to obtain the second training data; The second training data is input into the meta-model of the second layer in the Stacking model for training.
8. The system according to claim 6, characterized in that The acquisition module is used for: Perform CatBoost processing on all characteristic parameters extracted from the meteorological data to obtain characteristic parameters whose correlation with the 10m wind is higher than a threshold value among all characteristic parameters; The feature parameters that are higher than the threshold are used as the feature parameters for generating training data.
9. The system according to claim 8, characterized in that The characteristic parameters used to generate training data include at least one of the following: air pressure, temperature, terrain, humidity, vorticity, divergence, vertical velocity, and temperature advection.
10. The system according to any one of claims 6 to 9, characterized in that The first layer of the Stacking model includes four base models, which are: lasso regression model, decision book model, random forest model, and support vector machine model.