Wind power interval prediction method in combination with multi-source feature data
By combining the wind power power interval prediction method with multi-source feature data, using stacking integrated learning model and genetic algorithm optimization, the problem of conflict between single feature data and interval prediction indicators in wind power prediction is solved, and more accurate wind power power interval prediction is achieved.
Patent Information
- Application Number
- CN202510276906.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-08
AI Technical Summary
The existing wind power power prediction methods have single feature data, making it difficult to effectively learn the high volatility and complexity of wind power. There is a contradiction between PICP and AW indicators in interval prediction, and it is difficult to reduce AW while maintaining PICP.
The wind power power interval prediction method combined with multi-source feature data is adopted, stacking integrated learning model and genetic algorithm optimization is used, and the prediction interval is corrected through a conformal correction method. Four basic learners with different characteristics are selected, and the weight is optimized by pinball loss function and genetic algorithm to perform weighting and interval correction.
It effectively alleviates the contradiction between PICP and AW, improves the reliability and accuracy of the prediction interval, ensures that the probability of the prediction interval coverage is close to the set confidence, and at the same time reduces the average width and improves the prediction effect.
Smart Images

Figure CN120277604A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power prediction, and particularly relates to a wind power interval prediction method combining multi-source feature data. Background Art
[0002] Wind power, as a high-quality green energy, has received great attention and investment. The number of wind farms and wind turbines under construction has been increasing year by year, resulting in an increasing amount of wind power being incorporated into the power grid. However, when the generated power of a wind farm is incorporated into the power grid, there are some problems. Wind power generation is characterized by high volatility and complexity, which directly leads to unstable output power, making it difficult to predict and control. When a large amount of wind power is injected into the power grid, it may cause fluctuations in the grid voltage and instability in the frequency, posing a threat to the safe and stable operation of the power system. Predicting the wind power can help solve this problem and provide strong support for the dispatching and control of the power system. The core of wind power prediction lies in using advanced data analysis techniques and models, combined with meteorological conditions, historical operation data of wind farms, and physical characteristics of wind turbine equipment, etc., to accurately estimate the future output power of the wind farm. Such prediction can help grid dispatchers understand the changing trend of wind power in advance, so as to formulate a more reasonable dispatching plan and ensure the stable operation of the power grid when wind power is injected.
[0003] The methods of wind power prediction are generally classified according to the prediction principle into physical methods, statistical analysis methods, learning-based methods, and combined methods. Physical methods use numerical weather prediction models to predict wind speed based on information such as the terrain and meteorological conditions around the wind farm, and then use the results for wind power prediction. Statistical analysis methods are based on the mapping relationship between historical statistical data and the output power of the wind farm for prediction, including linear regression algorithms, time series analysis methods, etc. Learning-based methods use machine learning or deep learning algorithms to predict wind power, such as neural networks, support vector machines, etc. The combined method further improves the prediction effect by combining multiple processing means, such as time series decomposition, optimization algorithms, etc. Wind power prediction methods are constantly innovating. Currently, the mainstream prediction methods are learning-based methods and combined methods.
[0004] Although many research results have been achieved in the field of wind power prediction, there are still some problems in the existing prediction methods. One of the problems is the single feature data source. Many methods only use historical data features for prediction. Due to the high volatility, complexity, and intermittency of wind power, it is difficult for current learning models to learn the laws of historical time series data. Introducing numerical weather prediction data helps to improve the prediction effect. In addition, most current research focuses on point power prediction, while interval prediction of power is relatively less. Therefore, interval prediction is still a field worthy of research. In the problem of interval prediction, there is a contradiction between two indicators: the prediction interval coverage probability PICP and the average width AW. For the prediction interval, when the PICP is the same, the smaller the AW, the better, because the narrower the prediction interval, the more valuable it is as a reference. Generally, in general prediction methods, when the PICP increases, the AW also increases. Therefore, the proposed interval prediction method should make the AW as small as possible while keeping the PICP near the set interval confidence level. Summary of the Invention
[0005] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and provide a wind power interval prediction method combining multi-source feature data, which predicts the wind farm power generation power interval for the next several hours, provides reference information for grid dispatching, and ensures the stable operation of the grid when wind power is injected. To alleviate the problem of a single data source, this method introduces numerical weather prediction data, not only using historical data, but more importantly, using feature data representing the future. And to improve the prediction effect, the present invention adopts a combined prediction method based on machine learning algorithms. Specifically, four base models with different characteristics are selected for the stacking ensemble learning model, and the optimization process of the genetic algorithm is used to replace the meta-learner of the stacking ensemble learning model. Further, in order to ensure that the PICP index of the prediction interval can be maintained near the set interval confidence level, the conformal correction method is used to correct the prediction interval. The prediction results show that the proposed scheme alleviates the contradiction between the two indicators of the prediction interval PICP and AW, and at the same time effectively improves or guarantees the PICP index, obtaining a better prediction effect.
[0006] To achieve the above object, the present invention adopts the following technical solutions: A wind power interval prediction method combining multi-source feature data, the wind power interval prediction method is divided into a training stage, a testing stage, and a using stage, and includes the following steps: S1. Collect weather and actual power generation data at the wind farm, obtain numerical weather prediction data from a professional meteorological agency, save the data to a database, retrieve the data from the database, perform data cleaning, feature screening, and sample generation, and the generated samples are divided into a training set, a conformal correction set, and a testing set; S2. In the model training stage, first use the training set to train each base learner in the stacking ensemble learning model respectively; S3. Use the trained base learners to predict the training set, then optimize the weighted weights of the prediction results of each base learner through the genetic algorithm to obtain a set of weights, and combine the prediction results of each base learner with this set of weights for weighted average to obtain a preliminary prediction result; S4. Input the conformal correction set into the trained stacking ensemble learning model to obtain the weighted average prediction results of each base learner on the conformal correction set, calculate the conformal error according to this prediction result, and obtain a set of conformal error scores; S5. In the usage stage, first obtain the preliminary weighted average prediction result as the preliminary prediction interval through the stacking ensemble learning model, then use the set of conformal error scores to correct the preliminary prediction interval to obtain the final prediction interval, and finally display the wind power prediction interval for the reference and use of the wind farm staff.
[0007] Furthermore, the wind power interval prediction method further includes a verification step after step S4: Input the test set into the trained ensemble learning model to obtain the weighted average prediction interval; then, perform interval correction on the weighted average prediction interval according to the principle of conformal correction; finally, use three indicators, namely the prediction interval coverage probability PICP, the average width AW, and the Winkler score WS, to verify the effectiveness of the proposed scheme. Among the indicators used, PICP is the core indicator for measuring the probability of the prediction interval covering the true value in interval prediction, reflecting the reliability of the prediction interval. The closer PICP is to the nominal coverage probability PINC of the prediction interval, the more reliable the prediction interval is; AW is a measure of the average width of the prediction interval. When PICP can reach the set PINC indicator, the smaller AW is, the better. The smaller AW is, the stronger the physical meaning of the prediction interval; the Winkler score WS indicator comprehensively evaluates the reliability and accuracy of the prediction interval and can more comprehensively evaluate the performance of the prediction interval. A lower WS value means that the prediction interval can effectively cover the true value and will not be too wide, thus achieving a good balance between the coverage rate and accuracy. Therefore, the three selected indicators can comprehensively evaluate the prediction interval.
[0008] Furthermore, the process of step S1 is as follows: S101. Collect weather and actual power generation data at the wind farm, obtain numerical weather prediction data from a professional meteorological agency, and upload the data to the database. Obtain feature data including historical weather, historical actual power generation, and numerical weather prediction at the wind farm from the database, interpolate and fill in the missing values in the feature data, and use the recursive feature elimination method to screen out the features that have the greatest impact on power prediction from various data, and remove the unimportant features to reduce the calculation amount and improve the prediction effect. S102. Then generate training samples. A sample pair includes feature data and a label value. Regarding the problem of how long historical data should be used for historical features, the autocorrelation function and partial autocorrelation function are used to determine the time series correlation, and according to this method, the historical data of the previous several hours is selected as features. The total features also need to add the numerical weather prediction data, which is obtained from a professional meteorological agency. The label value of the sample is the wind power generation power for the corresponding time period. The numerical weather prediction data obtained from the meteorological agency is usually updated and released several times a day, and each release includes weather predictions for the next several hours, so it is suitable for multi-step interval prediction. S103. Divide the generated samples into a training set, a conformal calibration set, and a test set according to the ratio of 4:1:1.
[0009] Furthermore, in step S2, the advantage of the stacking ensemble learning model lies in integrating the advantages of different base learners. Therefore, selecting base learners with different properties helps to maximize the advantages of the stacking ensemble learning model. This method selects four base learners with their own characteristics, including the fully connected neural network FCNN, the random forest RF, the light gradient boosting machine LGBM, and the linear quantile regression model QR. Among these four base learners, QR belongs to the linear regression model, FCNN belongs to the deep learning model, and RF and LGBM belong to the bagging and boosting ensemble learning methods respectively. Therefore, the selected base learners are diverse in structure and principle, enabling the stacking ensemble learning model to comprehensively utilize their advantages and make up for their deficiencies at the same time. The interval prediction of the present invention forms the upper and lower bounds of the interval by predicting a pair of quantiles. For a prediction interval with a confidence level of ( ), the quantile level of the lower bound quantile is , and the quantile level of the upper bound quantile is . Since the RF, LGBM, and QR in the selected base learners are all single-output models, but the prediction interval consists of two components, the upper and lower quantiles. Therefore, for each of these three base learners, two models need to be trained separately, one to predict the lower quantile and the other to predict the upper quantile. Since the FCNN can output multiple values, a single model is used to output the two quantiles. Since it is a quantile prediction, the loss function used during model training is the pinball loss function, which is defined as follows:
[0010] In the above formula, is the label value of the i-th predicted sample, is the quantile of the i-th predicted sample, represents the error between the label value and the predicted quantile, is the value of the pinball loss function.
[0011] Furthermore, the process of step S3 is as follows: S301. Obtain the prediction results of each group of trained base learners on the training set, and use each pair of upper and lower quantiles to form a prediction interval; S302. Use the genetic algorithm to replace the learning process of the meta-learner in the stacking ensemble learning model, and use the genetic algorithm to find a set of weights for the 4 base learners, and perform a weighted average of the prediction results of the 4 base learners; the optimization objective and constraints of the genetic algorithm are shown in the following formulas:
[0012] In the above formula, PINC represents the nominal coverage probability of the prediction interval, which is equivalent to the set interval confidence level in the discrete statistical sense , for example, if PINC is set to 90%, then it is required that the coverage probability PICP of the prediction interval can reach 90%, , , , respectively represent the weights of the 4 base learners, abs( ) represents taking the absolute value, PICP represents the actual coverage probability of the prediction interval, and its mathematical definition is as follows:
[0013] In the above formula, is the number of samples, is the decision function, and are the lower and upper bounds of the i-th predicted sample interval respectively. If the label value falls within the prediction interval, then has a value of 1, otherwise 0; S303. Use the prediction intervals of each group of base learners for the training set as the training data of the genetic algorithm. After multiple rounds of iterative optimization of the genetic algorithm, obtain the optimal weighted weights. ; Compared with the prediction results using equal division and averaging, using the optimal weights to perform weighted averaging on the prediction results of the four base learners of the stacking ensemble learning model can make the PICP index closer to the set PINC value, and at the same time, it will also further reduce the AW index and WS index.
[0014] Further, the process of step S4 is as follows: S401. Use the trained stacking ensemble learning model to predict the conformal correction set to obtain the weighted average prediction result; S402. Calculate the conformal error scores for the prediction results of the conformal correction set to obtain the conformal error score set , where there are two implementation methods for the calculation formula of the conformal error score. The first implementation method is as follows:
[0015] The second implementation method is as follows:
[0016] Among them, represents the th prediction sample, is the conformal error score of the i-th prediction sample, is the feature data of the i-th prediction sample, is the label value of the i-th prediction sample, is the quantile of the lower bound of the interval, is the quantile of the upper bound of the interval. When , is the error that causes not to be in the interval; when , is the error that causes not to be in the interval; when neither of the above two situations occurs, that is, is in the interval, in the first implementation method is and the larger one of them, and in the second implementation method is 0.
[0017] Further, the process of the verification step is as follows: Sa. Use the trained stacking ensemble learning model to predict the test set to obtain the weighted average prediction interval; Sb, obtain the conformal error score set of the quantile , which is the confidence level of the prediction interval. Use this quantile to correct the preliminary prediction interval. The mathematical formula for predicting interval correction is as follows:
[0018] where, represents the interval after conformal correction, is the quantile of the lower bound of the interval, is the quantile of the upper bound of the interval; is the conformal error score set of the empirical quantile. That is to say, in , there are data that are less than or equal to . Theoretically, satisfies .
[0019] In the first implementation of the conformal error score calculation formula, when is not in the prediction interval, is a positive number; while when is in the prediction interval, is a negative number. Therefore, when more than of the falls within the prediction interval, is negative. According to the prediction interval correction formula, the upper bound of the prediction interval will decrease, the lower bound will increase, and the width of the prediction interval will become smaller, which will cause the falling within the prediction interval to decrease a little. And when less than of the falls within the prediction interval, is positive. According to the prediction interval correction formula, the upper bound of the prediction interval will increase, the lower bound will decrease, and the width of the prediction interval will become larger, so that more falls within the prediction interval. In this way, the conformal correction method of the prediction interval takes into account both the over-coverage and under-coverage of PICP. When there is under-coverage, it increases the width of the prediction interval, and when there is over-coverage, it decreases the width of the prediction interval; In the second implementation of the conformal error score calculation formula, when is not in the prediction interval, the calculation of is the same as the first implementation; while when is in the prediction interval, is 0; therefore, when more than of the When it falls within the prediction interval, is 0. According to the prediction interval correction formula, the prediction interval will not be corrected. of When it falls within the prediction interval, the correction of the prediction interval is the same as that shown in the first implementation. The reason why the second implementation of the conformal error score calculation formula is provided is that considering the complexity of the prediction data, the first implementation may lead to over-correction of the prediction interval, that is, the prediction interval coverage probability PICP is less than the set nominal coverage probability PINC. Therefore, the second implementation is suitable for scenarios with strict requirements for guaranteeing the PICP indicator; Sc, use the three indicators of prediction interval coverage probability PICP, average width AW, and Winkler score WS to verify the effectiveness and superiority of the proposed scheme, where the mathematical definition of AW is as follows:
[0020] In the above formula, Indicates prediction samples, is the sample size; and are the lower and upper bounds of the prediction interval, respectively. Indicates the distance between the upper and lower bounds of the prediction interval, that is, the width of the interval; The Winkler score WS is a comprehensive indicator that comprehensively evaluates the reliability and narrowness of the prediction interval. The mathematical definition of WS is as follows:
[0021] In the above formula, Indicates The true value of the predicted sample, and the definitions of other symbols are consistent with those described above.
[0022] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. The wind power interval prediction method proposed by the present invention combines characteristic data from multiple sources, which is a combined prediction method based on machine learning. The stacking ensemble learning model is used to fully combine the advantages of different types of base learners. From the experimental results, the prediction interval obtained by using the stacking ensemble learning model performs better in terms of PICP, AW, and WS indicators than the method that does not use ensemble learning. In general, the stacking ensemble learning model can make the PICP of the prediction interval meet the set PINC as much as possible, while its AW and WS indicators are also small.
[0023] 2. For the selection of base learners in the stacking ensemble learning model, four models are adopted: Random Forest (RF), Fully Connected Neural Network (FCNN), Light Gradient Boosting Machine (LGBM), and Linear Quantile Regression Model (QR). Each of these four models has its own characteristics. The fully connected neural network belongs to the deep learning model and is suitable for processing complex non-linear relationships. The random forest integrates multiple decision trees through the bagging mechanism and has good stability and noise resistance. The light gradient boosting machine is integrated through the boosting mechanism, and its serial integration method makes it have the characteristic of high prediction accuracy. The linear regression model is simple and easy to use, can well capture the linear relationship in the data, and has good interpretability. Each of the four models has its own advantages, and through stacking integration, their strengths can be fully utilized to improve the prediction ability and generalization performance of the overall model.
[0024] 3. The optimization process of the genetic algorithm is used as the meta-learner in the stacking ensemble learning model to optimize the weighted average weights of the prediction results of the base learners. The optimization goal set by the genetic algorithm can make the PICP of the prediction interval as close as possible to the set PINC. At the same time, it can also further narrow the AW, thus effectively alleviating the trade-off contradiction between PICP and AW. In terms of the Winkler Score (WS) index, compared with using equal-weight averaging for the prediction results of the base learners or comparing with the WS index of other model prediction results, this operation can make the WS index reach the minimum, which also indicates that the stacking ensemble learning + genetic algorithm as the meta-learner scheme has the best comprehensive effect.
[0025] 4. Due to the complexity of prediction, it is still difficult to ensure that the interval coverage probability PICP is greater than or equal to the set interval confidence level PINC only by using the above methods to predict the wind power interval. Therefore, the present invention corrects the prediction interval by using the conformal calibration method and provides two implementation methods of the conformal error score calculation formula. In the first implementation method, when PICP is less than PINC, the PICP index can be improved by slightly increasing the interval width; when PICP is greater than PINC, the PICP index can be reduced by slightly decreasing the interval width, making PICP closer to PINC and at the same time reducing the AW index. The second implementation method takes into account the complexity of the prediction data. When PICP is less than PINC, the PICP index can be improved by slightly increasing the interval width, thereby ensuring the PICP index; when PICP is greater than PINC, no prediction interval correction is performed to avoid overcorrection, which is suitable for scenarios with relatively strict requirements for ensuring the PICP index. Description of the Drawings
[0026] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0027] Figure 1 It is the basic flowchart of the wind power interval prediction method disclosed in the present invention in the training, testing, and usage stages; Figure 2 It is the prediction interval segment display diagram of the En-A and En-GA schemes for the test set of dataset #1 under the condition that the wind power interval prediction method disclosed in the present invention is in dataset #1 and PINC is set to 85%; Figure 3 It is the prediction interval segment display diagram of the En-A and En-GA schemes for the test set of dataset #2 under the condition that the wind power interval prediction method disclosed in the present invention is in dataset #2 and PINC is set to 90%. Detailed implementation manners
[0028] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope protected by the present application.
[0029] Referring to "embodiments" in the present application means that the specific features, structures, or characteristics described in combination with the embodiments can be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.
[0030] Embodiment 1 A wind power interval prediction method combining multi-source feature data provided by an embodiment of the present invention is based on the prediction and usage process as shown in Figure 1 and includes a training stage, a testing stage, and a usage stage.
[0031] The training stage of a wind power interval prediction method combining multi-source feature data provided by an embodiment of the present invention includes the following steps: S1. Obtain historical weather, power generation, and numerical weather prediction data from the database, perform data cleaning, feature screening, and sample generation. The generated samples are divided into a training set, a conformal calibration set, and a test set; In this embodiment, after feature screening, historical wind speed, wind direction, and wind speed and wind direction in numerical weather prediction data are left as features in terms of weather factors; and in this embodiment, the generated samples are divided into a training set, a conformal calibration set, and a test set in a ratio of 4:1:1 in chronological order; S2. Use the training set to train the four base learners of stacking ensemble learning, namely the fully connected neural network FCNN, the random forest RF, the light gradient boosting machine LGBM, and the linear quantile regression model QR; In this embodiment, the construction of the prediction interval is completed by predicting a pair of quantiles. For a prediction interval with a confidence level of , the quantile level of the quantile forming the lower bound of the interval is , and the quantile level of the quantile forming the upper bound of the interval is . Therefore, during model training, quantile regression learning is performed, and the pinball loss function is used for model training, which is defined as follows:
[0032] In the above formula, is the label value of the i-th prediction sample, is the quantile of the i-th prediction sample, represents the error between the label value and the predicted quantile, is the value of the pinball loss function; S3. Use the trained base learners to predict the training set, and optimize the weighted weights of the prediction results of each base learner through a genetic algorithm to obtain a set of optimal weights. This set of weights is used to perform a weighted average of the prediction results of each base learner, thereby obtaining a preliminary prediction interval; In this embodiment, the optimization parameter of the genetic algorithm is the weight of the base learner, and the optimization objective and constraints are shown in the following formula:
[0033] In the above formula, PINC represents the nominal coverage probability of the prediction interval, which is equivalent to the set interval confidence level , , , , respectively represent the weights of the 4 base learners, abs( ) represents taking the absolute value, and PICP represents the actual prediction interval coverage probability, and its mathematical definition is as follows:
[0034] In the above formula, is the number of samples, is the decision function, and are respectively the lower bound and the upper bound of the prediction interval of the i-th sample; if the label value falls within the prediction interval, then has a value of 1, otherwise 0; S4. Input the conformal calibration set into the trained stacking ensemble learning model to obtain the weighted average prediction results of each base learner on the conformal calibration set, and calculate the conformal error according to the prediction results to obtain the conformal error score set ; In this embodiment, the calculation formula of the conformal error score adopts the second implementation method:
[0035] where represents the th prediction sample, is the conformal error score of the i-th prediction sample, is the feature data of the i-th prediction sample, is the label value of the i-th prediction sample, is the quantile of the lower bound of the interval, is the quantile of the upper bound of the interval.
[0036] The test and verification stage of a wind power interval prediction method combining multi-source feature data provided by an embodiment of the present invention includes the following steps: Sa. Use the trained stacking ensemble learning model to predict the test set to obtain the weighted average prediction interval; Sb. In this embodiment, obtain the th quantile of the conformal error score set ,
[0037] where represents the interval after conformal correction, is the quantile of the lower bound of the interval, is the quantile of the upper bound of the interval; Sc. In this embodiment, use three indicators, namely the prediction interval coverage probability PICP, the average width AW, and the Winkler score WS, to verify the effectiveness and superiority of the proposed scheme. Among them, the mathematical definition of AW is as follows:
[0038] In the above formula, represents the th prediction sample, and is the number of samples; and are the lower and upper bounds of the prediction interval respectively, represents the distance between the upper and lower bounds of the prediction interval, that is, the interval width; .
[0039] Example 2 Based on the wind power interval prediction method combining multi-source feature data disclosed in Example 1, in this example, a dataset (#1) of a wind farm is used for simulation verification, and the nominal coverage probability index PINC of the prediction interval is set to 85%.
[0040] First, in the training stage, feature screening, sample generation, etc. are performed on dataset #1. Then, a stacking ensemble learning model is trained and the optimal weighted average weights are obtained by optimizing through the genetic algorithm. Finally, a conformal error score set is obtained according to the second implementation method of the conformal error calculation formula. Then, it enters the test and verification stage, the final prediction interval of the test set is obtained, and the PICP, AW, and WS indicators are used to evaluate the experimental results.
[0041] Table 1 shows the values of the PICP, AW, and WS indicators of the prediction interval of the test set of dataset #1 when using dataset #1 and setting PICN to 85%. In Table 1, 'En' in 'En-A' and 'En-GA' is the abbreviation of EnsembleLearning, representing the use of the stacking ensemble learning model; the meaning of 'A' is 'Average', indicating that the prediction results of the four base learners are summed using an equal-division average weighting method; 'GA' is the abbreviation of GeneticAlgorithm, indicating that the prediction results of the four base learners are weighted and averaged using the optimal weights obtained by optimizing through the genetic algorithm. The meaning of 'C' is 'Conformal', indicating the use of conformal correction.
[0042] According to Table 1, first, it can be clearly seen that before and after using conformal correction, the indicators of the experimental scheme did not change because the second implementation method of the conformal error score calculation formula was used, and when PICN is greater than PINC, the prediction interval is not corrected.
[0043] In the comparison between the En-A and En-GA schemes, the PICP of En-GA is closer to the set PINC, which is determined by the optimization objective of the genetic algorithm GA. Through the action of GA, not only is the PICP closer to the PINC, but also the average width of the interval AW is reduced, and the best result is also achieved in the comprehensive index WS.
[0044] Figure 2 Shows the prediction intervals for the test set of dataset #1, which is an intuitive manifestation of the reduction in the width of the prediction intervals under the action of GA. In Figure 2 'up' represents the upper boundary of the prediction interval, and 'low' represents the lower boundary of the prediction interval.
[0045] Table 1. Evaluation scores of the prediction results of different experimental schemes under the condition of #1 - 85%
[0046] Example 3 Based on a wind power interval prediction method combining multi-source feature data disclosed in Example 1, this example uses the dataset (#2) of a certain wind farm for simulation verification, and sets the nominal coverage probability index PINC of the prediction interval to 90%.
[0047] First, in the training stage, feature screening, sample generation, etc. are performed on dataset #2. Then, the stacking ensemble learning model is trained and the optimal weighted average weights are obtained through genetic algorithm optimization. Finally, the conformal error score set is obtained according to the second implementation method of the conformal error calculation formula. Then, it enters the test and verification stage, the final prediction intervals of the test set are obtained, and the experimental results are evaluated using the PICP, AW, and WS indicators.
[0048] Table 2 gives the values of the PICP, AW, and WS indicators of the prediction intervals for the test set of dataset #2 when using dataset #2 and setting PICN to 90%. In Table 1, the meanings of 'En-A', 'En-GA', 'En-A-C', and 'En-GA-C' are the same as in Table 1.
[0049] According to Table 2, it can be seen that before conformal correction, the PICPs of the En-A and En-GA schemes are both less than the PINC, which is determined by the complexity of the prediction data. At the same time, due to the action of GA, the PICP of En-GA is even lower than the PINC, and the AW also becomes smaller. After conformal correction, the PICPs of the En-A and En-GA schemes are both improved, and the AW becomes larger. However, after conformal correction, the PICP of the En-GA-C scheme is closer to the PINC than that of the En-A-C scheme. At the same time, the AW and WS indicators of En-GA-C are both smaller than those of En-A-C.
[0050] Based on the above discussion, the following conclusions can be drawn. First, when the PICP is less than the PINC, conformal calibration effectively improves the PICP, making it closer to the PINC. Moreover, the lower the PICP is than the PINC, the greater the improvement. Second, the role of GA in reducing the prediction interval width remains even after conformal calibration, and GA minimizes the comprehensive index WS.
[0051] Figure 3 shows the prediction intervals for the test set of dataset #2, which is an intuitive manifestation of the reduction in the prediction interval width under the action of GA. In Figure 3 , the meanings of 'up' and 'low' are the same as in Figure 2 .
[0052] Table 2. Evaluation scores of prediction results of different experimental schemes under the condition of #2 - 90%
[0053] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.
[0054] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0055] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A wind power interval prediction method combining multi-source feature data, characterized in that The wind power interval prediction method includes the following steps: S1. Collect weather and actual power generation data at the wind farm, obtain numerical weather prediction data from a professional meteorological structure, save the data to a database, retrieve the data from the database, perform data cleaning, feature screening, and sample generation. The generated samples are divided into a training set, a conformal calibration set, and a test set; S2. In the model training stage, first use the training set to train each base learner in the stacking ensemble learning model; S3. Use the trained base learners to predict the training set, and then optimize the weighted weights of the prediction results of each base learner through a genetic algorithm to obtain a set of weights, which are used to perform weighted averaging on the prediction results of each base learner combined with this set of weights to obtain a preliminary prediction result; S4. Input the conformal calibration set into the trained stacking ensemble learning model to obtain the weighted average prediction results of each base learner on the conformal calibration set, calculate the conformal error based on this prediction result to obtain a conformal error score set; S5. In the usage stage, first obtain a preliminary weighted average prediction result through the stacking ensemble learning model as a preliminary prediction interval, then use the conformal error score set to correct the preliminary prediction interval to obtain a final prediction interval, and finally display the wind power prediction interval.
2. The wind power interval prediction method combining multi-source feature data according to claim 1, wherein The wind power interval prediction method further includes a verification step after step S4: Input the test set into the trained ensemble learning model to obtain a weighted average prediction interval; then, perform interval correction on the weighted average prediction interval according to the principle of conformal calibration; finally, use three indicators, namely the prediction interval coverage probability PICP, the average width AW, and the Winkler score WS, to verify the effectiveness of the proposed scheme.
3. According to a wind power interval prediction method combining multi-source data as described in claim 1, the process of step S1 is as follows: S101. Collect weather and actual power generation data at the wind farm, obtain numerical weather prediction data from a professional meteorological agency, and upload the data to the database; retrieve feature data including historical weather, historical actual power generation, and numerical weather prediction at the wind farm from the database, interpolate and fill in the missing values in the feature data, and use the recursive feature elimination method to screen out the features that have the greatest impact on power prediction from various data; S102. Generate training samples. One training sample pair includes feature data and a label value. Use the autocorrelation function and the partial autocorrelation function to determine the temporal correlation of historical features to determine the sampling window size of historical feature data, and then combine it with the numerical weather prediction data to form the total features. The label value of the training sample is the wind power generation power corresponding to the corresponding time period; S103. Divide the generated training samples into a training set, a conformal calibration set, and a test set according to a ratio.
4. A wind power interval prediction method combining multi-source data according to claim 1, wherein in the step S2, the stacking ensemble learning model adopts 4 base learners, namely fully connected neural network FCNN, random forest RF, light gradient boosting machine LGBM, and linear quantile regression model QR; the interval prediction forms the upper and lower bounds of the interval by predicting a pair of quantiles. For a prediction interval with a confidence level of the lower bound quantile has a quantile level of , and the upper bound quantile has a quantile level of ; ; The loss function used during model training is the pinball loss function, which is defined as follows: In the above formula, is the label value of the i-th predicted sample, is the quantile of the i-th predicted sample, represents the error between the label value and the predicted quantile, is the value of the pinball loss function.
5. According to a wind power interval prediction method combining multi-source data as described in claim 1, the process of step S3 is as follows: S301. Obtain the prediction results of each group of trained base learners for the training set, and use each pair of upper and lower quantiles to form a prediction interval; S302. Replace the learning process of the meta-learner of the stacking ensemble learning model with a genetic algorithm, and use the genetic algorithm to find a set of weights for the 4 base learners , and perform weighted averaging on the prediction results of the 4 base learners; the optimization objective and constraint conditions of the genetic algorithm are shown in the following formula: In the above formula, PINC represents the nominal coverage probability of the prediction interval, which is equivalent to the set interval confidence level in the discrete statistical sense , , , , respectively represent the weights of the four base learners, abs( ) represents taking the absolute value, and PICP represents the actual prediction interval coverage probability. The mathematical definition is as follows: In the above formula, is the number of samples, is the decision function, and are the lower and upper bounds of the prediction interval of the i-th predicted sample respectively. If the label value falls within the prediction interval, then has a value of 1, otherwise 0; S303. Use the prediction intervals of each group of base learners for the training set as the training data of the genetic algorithm, and obtain the optimal weighted weights after multiple rounds of iterative optimization of the genetic algorithm. .
6. For a wind power interval prediction method combining multi-source data according to claim 1, the process of step S4 is as follows: S401. Use the trained stacking ensemble learning model to predict the conformal calibration set to obtain the prediction result after weighted averaging; S402. Calculate the conformal error score for the prediction results of the conformal correction set to obtain a conformal error score set , where There are two implementation methods for the calculation formula of the conformal error score. The first implementation method is as follows: The second implementation method is as follows: In the above formula, represents the th prediction sample, is the conformal error score of the i-th prediction sample, is the feature data of the i-th prediction sample, is the label value of the i-th prediction sample, is the quantile of the lower bound of the interval, is the quantile of the upper bound of the interval.
7. For a wind power interval prediction method combining multi-source data according to claim 2, the process of the verification step is as follows: Sa. Use the trained stacking ensemble learning model to predict the test set to obtain the prediction interval after weighted averaging; Sb, Obtain the conformal error score set the quantile , where is the confidence level of the prediction interval, and the preliminary prediction interval is corrected using this quantile. The mathematical formula for this correction process is as follows: Among them, Indicates the interval after conformal correction, is the quantile of the lower bound of the interval, is the quantile of the upper bound of the interval; Sc. Use three indicators, namely the prediction interval coverage probability PICP, the average width AW, and the Winkler score WS, to verify the effectiveness and superiority of the proposed scheme. Among them, the mathematical definition of AW is as follows: In the above formula, represents the th prediction sample, is the number of samples; and are the lower and upper bounds of the prediction interval respectively, represents the distance between the upper and lower bounds of the prediction interval, that is, the interval width; The Winkler score WS comprehensively evaluates the reliability and narrowness of the prediction interval. The mathematical definition of WS is as follows: 。
Citation Information
Cited By
Wind power probability load prediction method, system and equipment based on meta-learning and medium
CN122051950A