Method and device for evaluating influence factors of prediction errors, electronic equipment and medium
The wind power prediction data is preprocessed and feature dimension reduction is performed through the random forest regression model, an initial random forest regression model is constructed, and the contribution of meteorological characteristics to the prediction error is calculated. This solves the problem of inaccurate wind power prediction error assessment and achieves accurate assessment of wind power prediction error.
Patent Information
- Application Number
- CN202510684299.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-05
AI Technical Summary
Existing wind power forecasting research cannot accurately analyze the nonlinear relationship between wind power forecast error and meteorological characteristics. Traditional methods find it difficult to capture the coupling effect between meteorological characteristics, resulting in inaccurate forecast error assessment and inability to meet the needs of wind farm operation and maintenance.
The random forest regression model is used to preprocess the original data and reduce the feature dimension. The initial random forest regression model is constructed. The contribution of each meteorological feature to the prediction error is calculated through contribution analysis, and then the influencing factors of wind power prediction error are evaluated.
It realizes the accurate evaluation of wind power prediction error, can directly mine the nonlinear relationship between complex meteorological characteristics and prediction error, and provides an accurate evaluation method for wind power prediction error.
Smart Images

Figure CN120596867A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of wind power prediction, and in particular to a method, device, electronic device, and medium for evaluating factors affecting prediction errors. Background Art
[0002] With the growing demand for renewable energy in the global economy, wind power generation has become a vital component of power systems. However, because wind power is affected by a variety of complex meteorological factors, errors in short-term wind power forecast models remain significant. This error not only impacts wind farm operations and management but also poses significant challenges to grid scheduling and stability. Therefore, in-depth analysis of the meteorological factors influencing short-term wind power forecast errors is crucial for improving forecast accuracy and optimizing wind power resource utilization.
[0003] Existing research on wind power forecasting primarily focuses on methods such as deep learning neural networks, which are unable to produce accurate analytical expressions with practical physical meaning. This often leads to complex nonlinear relationships between forecast errors and meteorological characteristics. On the one hand, traditional error analysis methods mostly focus on the confidence intervals and uncertainties of the forecast results, directly making assumptions about the distribution of the errors. This makes them inapplicable to actual wind power data with unknown error distributions. On the other hand, traditional regression analysis methods either ignore the coupling effects between meteorological characteristics, making it difficult to capture nonlinear relationships, or lose some information during the dimensionality reduction process, resulting in a weak explanatory power for the specific impact of meteorological characteristics. This makes it difficult to meet the needs of wind farm operations and maintenance for detailed error influencing factors, and it is impossible to accurately evaluate the factors affecting wind power forecasting errors. Summary of the Invention
[0004] In order to solve the above technical problems, the present disclosure provides a method, device, electronic device and medium for evaluating factors affecting prediction error.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for evaluating factors affecting prediction error, the method comprising:
[0006] Preprocessing the original data to obtain preprocessed data, and constructing an input data set based on the preprocessed data;
[0007] Based on the input data set, construct an initial random forest regression model;
[0008] Performing feature dimensionality reduction processing on the input data set to obtain a target input data set, and training the initial random forest regression model based on the target input data set to obtain a target random forest regression model;
[0009] Performing a contribution analysis on the target random forest regression model to calculate the contribution of each meteorological feature in the target input data set to the prediction error;
[0010] Based on the contribution of each meteorological feature to the prediction error, the factors affecting the wind power prediction error are evaluated.
[0011] In some embodiments, preprocessing the raw data to obtain preprocessed data, and constructing the input data set based on the preprocessed data, includes:
[0012] Arrange the original data in chronological order to obtain an original data set, wherein the original data set includes measured data and predicted data; wherein the measured data includes meteorological measured data and wind power measured data, and the predicted data includes meteorological predicted data and wind power predicted data;
[0013] Processing duplicate data in the measured data to obtain first measured data, and processing missing data in the first measured data to obtain second measured data;
[0014] Processing abnormal data in the second measured data to obtain third measured data, and obtaining preprocessed data based on the third measured data and the predicted data;
[0015] Calculating the prediction error of the meteorological data and the prediction error of the wind power based on the preprocessed data;
[0016] An input data set is constructed based on the wind power prediction error, the absolute value of the meteorological data prediction error, and the preprocessed data.
[0017] In some embodiments, constructing an initial random forest regression model based on the input data set includes:
[0018] Divide the input dataset into training and test sets based on a preset ratio;
[0019] Sampling the training set a preset number of times based on a sampling with replacement method to obtain a first sub-training set, using the first sub-training set as the training set for a first decision tree, performing node splitting on the first decision tree based on an optimal splitting point determined by impurity to obtain a first target decision tree, and repeating the current step until a preset number of target decision trees are obtained;
[0020] Generate an initial random forest regression model based on the preset number of target decision trees;
[0021] The test set is input into the initial random forest regression model to obtain an error prediction value of wind power for each test sample in the test set.
[0022] In some embodiments, performing feature dimensionality reduction processing on the input data set to obtain a target input data set includes:
[0023] Calculating the importance score of each meteorological feature in the input data set, and normalizing the importance score of each meteorological feature to obtain a normalized importance score of each meteorological feature;
[0024] Calculate the Pearson correlation coefficient between any two meteorological characteristics and calculate the variance inflation factor of each meteorological characteristic;
[0025] Based on the normalized importance score of each meteorological feature, the Pearson correlation coefficient between any two meteorological features, and the variance inflation factor of each meteorological feature, meteorological features that meet the conditions are removed from each meteorological feature in the input data set to obtain a target input data set.
[0026] In some embodiments, the training of the initial random forest regression model based on the target input dataset to obtain a target random forest regression model includes:
[0027] Dividing the target input data set into a target training set and a target test set, and training the initial random forest regression model based on the target training set to obtain a pre-trained random forest regression model;
[0028] Determining whether the prediction accuracy of the pre-trained random forest regression model meets a preset accuracy threshold, and if so, determining the pre-trained random forest regression model as the target random forest regression model;
[0029] If not, cross-validation and hyperparameter tuning are performed on the pre-trained random forest regression model to obtain a target random forest regression model.
[0030] In some embodiments, performing contribution analysis on the target random forest regression model to calculate the contribution of each meteorological feature in the target input data set to the prediction error includes:
[0031] For each test sample in the target test set, sampling different combinations of meteorological features of each test sample, and calculating the change in the output of the target random forest regression model under the different combinations of meteorological features;
[0032] Based on the change in the output of the target random forest regression model under different combinations of meteorological characteristics, the marginal contribution of each meteorological characteristic of each test sample is obtained;
[0033] Based on the marginal contribution of each meteorological feature of each test sample, the error prediction result of each test sample is calculated.
[0034] In some embodiments, the evaluating the factors affecting the wind power prediction error based on the contribution of each meteorological feature to the prediction error includes:
[0035] Based on the marginal contribution of each meteorological characteristic of each test sample, locally evaluate the factors affecting the wind power forecast error to obtain a local evaluation result;
[0036] For any meteorological feature, calculate the average of the absolute values of the marginal contributions of the meteorological feature under all test samples;
[0037] Based on the average of the absolute values of the marginal contributions of the meteorological characteristics under all the test samples, a global evaluation is performed on the factors affecting the wind power prediction error to obtain a global evaluation result.
[0038] In a second aspect, an embodiment of the present disclosure provides a device for evaluating factors affecting prediction errors, the device comprising:
[0039] A construction module is used to preprocess the original data to obtain preprocessed data, and construct an input data set based on the preprocessed data;
[0040] A construction module, configured to construct an initial random forest regression model based on the input data set;
[0041] An obtaining module is used to perform feature dimensionality reduction processing on the input data set to obtain a target input data set, and train the initial random forest regression model based on the target input data set to obtain a target random forest regression model;
[0042] a calculation module, configured to perform a contribution analysis on the target random forest regression model and calculate a contribution result of each meteorological feature in the target input data set to the prediction error;
[0043] The evaluation module is used to evaluate the factors affecting the wind power prediction error based on the contribution result of each meteorological feature to the prediction error.
[0044] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:
[0045] Memory;
[0046] processor; and
[0047] computer programs;
[0048] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect.
[0049] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method as described in the first aspect.
[0050] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, implements the method described in the first aspect.
[0051] The method, device, electronic device and medium for evaluating factors affecting prediction error provided by the embodiments of the present disclosure preprocess the original data to obtain preprocessed data, construct an input data set based on the preprocessed data, and construct an initial random forest regression model based on the input data set. Further, feature dimensionality reduction is performed on the input data set to obtain a target input data set, the initial random forest regression model is trained based on the target input data set to obtain a target random forest regression model, a contribution analysis is performed on the target random forest regression model, and the contribution of each meteorological feature in the target input data set to the prediction error is calculated. Then, based on the contribution of each meteorological feature to the prediction error, the factors affecting the wind power prediction error are evaluated. Through the method of this embodiment, the nonlinear relationship between complex meteorological features and prediction errors can be directly mined, and the factors affecting the wind power prediction error can be accurately evaluated. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0053] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0054] Figure 1 A flow chart of a method for evaluating factors influencing prediction error provided by an embodiment of the present disclosure;
[0055] Figure 2 A schematic diagram of the overall process of the method for evaluating factors affecting prediction error provided by an embodiment of the present disclosure;
[0056] Figure 3A flow chart of a method for evaluating factors affecting prediction error provided by another embodiment of the present disclosure;
[0057] Figure 4 A schematic diagram of the structure of a device for evaluating factors affecting prediction error provided by an embodiment of the present disclosure;
[0058] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0059] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.
[0060] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0061] With the growing demand for renewable energy in the global economy, wind power generation has become a vital component of power systems. However, because wind power is affected by a variety of complex meteorological factors, errors in short-term wind power forecast models remain significant. This error not only impacts wind farm operations and management but also poses significant challenges to grid scheduling and stability. Therefore, in-depth analysis of the meteorological factors influencing short-term wind power forecast errors is crucial for improving forecast accuracy and optimizing wind power resource utilization.
[0062] Existing research on wind power forecasting primarily focuses on methods such as deep learning neural networks, which are unable to produce accurate analytical expressions with practical physical meaning. This often leads to complex nonlinear relationships between forecast errors and meteorological characteristics. On the one hand, traditional error analysis methods mostly focus on the confidence intervals and uncertainties of the forecast results, directly making assumptions about the distribution of the errors. This makes them inapplicable to actual wind power data with unknown error distributions. On the other hand, traditional regression analysis methods either ignore the coupling effects between meteorological characteristics, making it difficult to capture nonlinear relationships, or lose some information during the dimensionality reduction process, resulting in a weak explanatory power for the specific impact of meteorological characteristics. This makes it difficult to meet the needs of wind farm operations and maintenance for detailed error influencing factors, and it is impossible to accurately evaluate the factors affecting wind power forecasting errors.
[0063] To address this issue, an embodiment of the present disclosure provides a method for evaluating factors affecting prediction errors, which is described below in conjunction with specific embodiments.
[0064] Figure 1This is a flow chart of a method for evaluating factors influencing prediction error provided by an embodiment of the present disclosure. This method can be applied to electronic devices, including portable mobile devices such as tablets and laptops, as well as stationary devices such as personal computers and servers. The server can be a single server or a server cluster, which can be a distributed or centralized cluster. This method can be applied to scenarios where factors influencing wind power prediction error are evaluated. It can directly exploit the nonlinear relationship between complex meteorological characteristics and prediction error, thereby accurately evaluating the factors influencing wind power prediction error.
[0065] It is understandable that the method for evaluating factors affecting prediction errors provided by the embodiments of the present disclosure may also be applied in other scenarios.
[0066] Below Figure 1 The method for evaluating the factors affecting the prediction error shown in FIG. 1 is introduced, and the method includes the following steps:
[0067] S101 , preprocessing original data to obtain preprocessed data, and constructing an input data set based on the preprocessed data.
[0068] In this step, the electronic device acquires raw data, preprocesses the raw data to obtain preprocessed data, and then constructs an input dataset based on the preprocessed data. The raw data includes at least measured meteorological data, measured wind power data, meteorological forecast data, and wind power forecast data. This is optional. Preprocessing may include duplicate data processing, missing data processing, abnormal data processing, and other processing methods, which are not limited here.
[0069] In some embodiments, S101 may include but is not limited to S1011, S1012, S1013, S1014, and S1015:
[0070] S1011. Arrange the original data in chronological order to obtain an original data set, where the original data set includes measured data and predicted data;
[0071] In this step, the original data of the wind farm are collected and sorted by time to generate the original data set, such as Figure 2 As shown, the original data set includes measured data and predicted data. The measured data includes measured meteorological data and measured wind power data, and the predicted data includes predicted meteorological data and predicted wind power data. Meteorological features in the measured meteorological data and predicted meteorological data include, but are not limited to, wind speed, wind direction, temperature, air pressure, and relative humidity.
[0072] S1012: Process duplicate data in the measured data to obtain first measured data, and process missing data in the first measured data to obtain second measured data;
[0073] In this step, if Figure 2 As shown, the electronic device processes duplicate and missing data. Specifically, the electronic device traverses each row of the measured data, removes duplicate data with the same time stamp, and removes long periods of continuous duplicate data due to communication failures, thereby obtaining the first measured data. Furthermore, the electronic device traverses each column of the first measured data, locates missing data, and interpolates the data using a polynomial constructed using the Lagrange interpolation method to obtain the second measured data.
[0074] S1013: Process abnormal data in the second measured data to obtain third measured data, and obtain preprocessed data based on the third measured data and the predicted data;
[0075] In this step, if Figure 2 As shown, the electronic device processes the abnormal data.
[0076] Specifically, the second measured data is recorded as dataset S. Each column of measured data in S, except for wind speed and wind power, is traversed. The quartiles and interquartile ranges (IQR) of each meteorological characteristic dataset are calculated. Data outside the range (Q1 - 1.5*IQR, Q3 + 1.5*IQR) are recorded as abnormal data O1. Q1 and Q3 are the first and third quartiles of the dataset, sorted from smallest to largest.
[0077] The measured wind speed and wind power data set is denoted as D = {(v1, p1),…, (v i ,p i ),…,(v n ,p n )}, where n is the total number of samples of D, v i is the wind speed of the i-th sample, p i For the wind power of the i-th sample, calculate the local outlier factor (LOF) of each point, and output the k points with the highest degree of outliers, which are recorded as abnormal data O2.
[0078] The dataset D is segmented according to the measured wind speed, and the interquartile range method is applied to the wind power in each wind speed interval to locate abnormal data, which is recorded as O3.
[0079] Take the indexes of O1, O2, and O3 in the data set S to obtain the index union OI of the outliers. Set all the data corresponding to OI to null values based on the original data set S, and then process the missing data to complete the processing of the outliers and obtain the third measured data.
[0080] Furthermore, preprocessed data is obtained based on the third measured data and the predicted data.
[0081] S1014. Calculating a prediction error of meteorological data and a prediction error of wind power based on the preprocessed data;
[0082] In this step, the forecast error of meteorological data and wind power is calculated based on the preprocessed data. The formula is as follows:
[0083]
[0084] Among them, E f is the forecast error of meteorological data, E p is the wind power prediction error, f real and f pred are the measured and predicted values of meteorological data, respectively, real and p pred are the measured and predicted values of wind power data respectively.
[0085] S1015: Construct an input data set based on the wind power prediction error, the absolute value of the meteorological data prediction error, and the preprocessed data.
[0086] In this step, we need to use the predicted value of the error to correct p pred , E p Take the absolute error. f It is highly correlated with the actual meteorological data and cannot be directly used as the input feature of the prediction model, but its absolute value can be used as a new feature. p 、|E f |Together with the preprocessed data, it constitutes the input dataset.
[0087] S102: Based on the input data set, construct an initial random forest regression model.
[0088] In this step, if Figure 2 As shown in FIG, after constructing the input data set, an initial random forest regression model, specifically a random forest nonlinear regression model, can be constructed based on the input data set.
[0089] In some embodiments, S102 may include S1021, S1022, S1023, and S1024:
[0090] S1021. Divide the input data set into a training set and a test set based on a preset ratio;
[0091] In this step, divide the input data set T = {(x1, y1), …, (x i , y i ), …, (x n , y n )} into a training set and a test set, where n is the number of samples in the data set T, x i is a multi-dimensional vector composed of all meteorological features in the i-th sample, and y i is the wind power prediction error of the i-th sample.
[0092] S1022. Sample the training set a preset number of times in a sampling-with-replacement manner to obtain a first sub-training set, use the first sub-training set as the training set of the first decision tree, and perform node splitting on the first decision tree based on the best split point determined by impurity, to obtain the first target decision tree. Repeat the current step until a preset number of target decision trees are obtained;
[0093] In this step, perform Bootstrap sampling with replacement on the training set, and sample N times to form a sub-training set A j , which is used as the training set of the j-th decision tree. During the splitting process of each node of the decision tree, randomly select m features (m < M, where M is the total number of meteorological features), and use the feature that maximizes the decrease in impurity as the best split point. Repeat the splitting step until the maximum depth or the minimum number of samples in the leaf node is reached, so as to obtain the j-th target decision tree. Repeat the current step until a preset number, that is, k, of target decision trees are obtained.
[0094] S1023. Generate an initial random forest regression model based on the preset number of target decision trees;
[0095] In this step, an initial random forest regression model can be generated according to the preset number k of target decision trees.
[0096] S1024. Input the test set into the initial random forest regression model to obtain the error prediction values of the wind power of each test sample in the test set.
[0097] In this step, input the test set into the initial random forest regression model to obtain the error prediction values of the wind power of each test sample in the test set. For each test sample, its error prediction value of the wind power is:
[0098]
[0099] Among them, is the wind power error predicted by the jth decision tree.
[0100] The present disclosure embodiment uses the determination coefficient R 2 The prediction accuracy of the initial random forest regression model is evaluated by the root mean square error (RMSE), and the calculation formula is as follows:
[0101]
[0102] Among them, n is the number of samples in the test set, E av,i is the error prediction value of wind power of the i-th test sample, E p,i is the wind power prediction error of the i-th test sample.
[0103] S103 , performing feature dimensionality reduction processing on the input data set to obtain a target input data set, and training the initial random forest regression model based on the target input data set to obtain a target random forest regression model.
[0104] In this step, if Figure 2 As shown, the electronic device performs feature dimensionality reduction processing on the input data set, that is, processes redundant meteorological features, trains and optimizes the random forest regression model after dimensionality reduction, and obtains the target random forest regression model.
[0105] S104: Perform contribution analysis on the target random forest regression model to calculate the contribution of each meteorological feature in the target input data set to the prediction error.
[0106] In this embodiment, Figure 2 As shown, the electronic device performs a contribution analysis (ie, SHAP analysis) on the target random forest regression model to calculate the contribution of each meteorological feature to the prediction error.
[0107] S105 . Evaluate factors affecting wind power prediction error based on the contribution of each meteorological feature to the prediction error.
[0108] In this step, if Figure 2 As shown, after obtaining the contribution result of each meteorological feature to the prediction error, the influencing factors of the wind power prediction error can be evaluated based on the contribution result of each meteorological feature to the prediction error.
[0109] The disclosed embodiment preprocesses the original data to obtain preprocessed data, constructs an input data set based on the preprocessed data, and constructs an initial random forest regression model based on the input data set. Further, feature dimensionality reduction is performed on the input data set to obtain a target input data set, the initial random forest regression model is trained based on the target input data set to obtain a target random forest regression model, and a contribution analysis is performed on the target random forest regression model to calculate the contribution of each meteorological feature in the target input data set to the prediction error. Then, based on the contribution of each meteorological feature to the prediction error, the influencing factors of the wind power prediction error are evaluated. Through the method of this embodiment, the nonlinear relationship between complex meteorological features and prediction errors can be directly explored, and the influencing factors of the wind power prediction error can be accurately evaluated.
[0110] On the basis of the above embodiment, in S103, feature dimensionality reduction processing is performed on the input data set to obtain a target input data set, including but not limited to S1031, S1032, and S1033:
[0111] S1031. Calculate the importance score of each meteorological feature in the input data set, and normalize the importance score of each meteorological feature to obtain a normalized importance score of each meteorological feature.
[0112] In this step, the electronic device calculates the importance score of each meteorological feature based on the Gini index, and normalizes the importance score of each meteorological feature to obtain a normalized importance score of each meteorological feature.
[0113] Specifically, 1) First, calculate the Gini index of the i-th tree node q:
[0114]
[0115] Where C is the number of categories in the q node, is the proportion of category c in the i-th tree node q.
[0116] 2) Then, calculate the importance score of the jth meteorological feature based on the Gini index
[0117]
[0118] in, and They represent the Gini index of the new nodes l and r after the i-th tree branches from the node q, Q is the set of nodes where the j-th meteorological feature appears in the i-th tree, and K is the number of decision trees.
[0119] 3) Finally, normalize the importance scores of all meteorological features:
[0120]
[0121] Among them, M is the total number of meteorological characteristics, Score j is the importance score of the jth meteorological feature after normalization.
[0122] Sort the normalized importance scores of each meteorological feature from large to small, and calculate the cumulative score. The feature with a cumulative score greater than 85% is marked as X d1 .
[0123] S1032. Calculate the Pearson correlation coefficient between any two meteorological characteristics, and calculate the variance inflation factor of each meteorological characteristic.
[0124] In this step, the electronic device calculates the Pearson correlation coefficient and the variance inflation factor.
[0125] Specifically, the Pearson correlation coefficient between any two meteorological features is calculated, that is,
[0126]
[0127] Among them, X i and X j are the i-th and j-th meteorological characteristics, cov(X i ,X j ) is the covariance of the two, and are the standard deviations of the two respectively.
[0128] Calculate the variance inflation factor of the jth meteorological feature, that is
[0129]
[0130] in, It is the square of the complex correlation coefficient of the jth meteorological feature to the other meteorological features.
[0131] If the Pearson correlation coefficient and variance inflation factor of each meteorological characteristic are too large, they will affect the prediction accuracy and interpretability of the model respectively. According to the empirical rule, The two features of VIF are strongly correlated. j The features with ≥10 have multicollinearity, so the meteorological features that meet the above inequality conditions are recorded as X d2 .
[0132] S1033. Based on the normalized importance score of each meteorological feature, the Pearson correlation coefficient between any two meteorological features, and the variance inflation factor of each meteorological feature, the meteorological features that meet the conditions are removed from each meteorological feature in the input data set to obtain a target input data set.
[0133] In this step, according to the normalized importance score, Pearson correlation coefficient, and variance inflation factor, the meteorological features that meet the conditions are removed and X d =X d1 ∪X d2 , delete redundant features X from all meteorological features d , generate the target input dataset.
[0134] In some embodiments, in S103, training the initial random forest regression model based on the target input data set to obtain a target random forest regression model may include S1034, S1035, and S1036:
[0135] S1034, dividing the target input data set into a target training set and a target test set, and training the initial random forest regression model based on the target training set to obtain a pre-trained random forest regression model;
[0136] Further, after obtaining the target input data set, the target input data set is divided into a target training set and a target test set, that is, the target training set and test set Among them, n is the total number of samples, n train is the number of samples in the training set, x' i is the multidimensional vector consisting of the remaining features after removing redundant meteorological features in the i-th sample, y i The initial random forest regression model is trained with T1 to obtain a pre-trained random forest regression model, and the pre-trained random forest regression model is tested with T2.
[0137] S1035. Determine whether the prediction accuracy of the pre-trained random forest regression model meets a preset accuracy threshold. If so, determine the pre-trained random forest regression model as the target random forest regression model.
[0138] In this step, if Figure 2 As shown, it is determined whether the prediction accuracy of the pre-trained random forest regression model meets the accuracy requirement, that is, meets the preset accuracy threshold. If the prediction accuracy meets the preset accuracy threshold, the pre-trained random forest regression model is determined as the target random forest regression model.
[0139] S1036. If not, cross-validate and optimize hyperparameters of the pre-trained random forest regression model to obtain a target random forest regression model.
[0140] If the prediction accuracy does not meet the preset accuracy threshold, cross-validation and hyperparameter tuning are performed on the pre-trained random forest regression model to obtain a target random forest regression model.
[0141] As a rule of thumb, the range of adjustable hyperparameters for random forests is usually as follows:
[0142]
[0143] Among them, n estimators is the number of decision trees, depth max is the maximum depth of the decision tree, is the minimum number of split samples, is the minimum number of leaf node samples, m is the number of optional features for splitting, and M is the total number of features.
[0144] First, we randomly sample hyperparameter combinations for large-scale optimization to narrow the range of optimal hyperparameters. Then, we use grid search to traverse the hyperparameter combinations one by one to obtain the optimal solution. Finally, we cross-validate the model trained with this combination, also using R 2 and RMSE to evaluate the generalization ability of the model and avoid overfitting or underfitting.
[0145] After cross-validation and hyperparameter tuning of the pre-trained random forest regression model, the target random forest regression model is obtained.
[0146] This embodiment removes qualified meteorological features from the various meteorological features in the input data set based on the normalized importance score of each meteorological feature, the Pearson correlation coefficient between any two meteorological features, and the variance inflation factor of each meteorological feature, removes redundant meteorological features, and trains a target random forest regression model. This can further explore the nonlinear relationship between complex meteorological features and prediction errors, and thus accurately evaluate the factors affecting wind power prediction errors.
[0147] Figure 3 A flow chart of a method for evaluating factors affecting prediction error provided by another embodiment of the present disclosure is shown in FIG. Figure 3 As shown, the method includes the following steps:
[0148] S301 : Preprocessing original data to obtain preprocessed data, and constructing an input data set based on the preprocessed data.
[0149] Specifically, the implementation process and principle of S301 and S101 are the same and will not be described in detail here.
[0150] S302: Based on the input data set, construct an initial random forest regression model.
[0151] Specifically, the implementation process and principle of S302 and S102 are the same and will not be described in detail here.
[0152] S303 , performing feature dimensionality reduction processing on the input data set to obtain a target input data set, and training the initial random forest regression model based on the target input data set to obtain a target random forest regression model.
[0153] Specifically, the implementation process and principle of S303 and S103 are the same, and will not be repeated here.
[0154] S304: For each test sample in the target test set, sample the combination of different meteorological features of each test sample, and calculate the change in the output of the target random forest regression model under the combination of different meteorological features.
[0155] In this step, the electronic device samples different combinations of meteorological features for each test sample of the target test set T2, and calculates the change in the output of the target random forest regression model under the different combinations of meteorological features.
[0156] S305. Based on the variation of the output of the target random forest regression model under the combination of different meteorological features, obtain the marginal contribution of each meteorological feature of each test sample.
[0157] In this step, the electronic device will obtain the marginal contribution of each meteorological feature of each test sample based on the change in the output of the target random forest regression model under different combinations of meteorological features. The marginal contribution of each meteorological feature is:
[0158]
[0159] in, is the SHAP value of the jth feature in the i-th sample, S is the exclusion All feature subsets of , M is the total number of features, val(S) is the contribution function of the feature, which indicates the degree of common influence of the features in the set S on the error prediction results, and is calculated as follows:
[0160]
[0161] Among them, x i is all the features of the i-th sample, that is, n is the number of samples in the test set, It means that the features that do not belong to the set S in the i-th sample are integrated under the premise of determining the eigenvalues in the set S. Represents the error prediction value of the mean of all features of n samples.
[0162] S306: Calculate the error prediction result of each test sample based on the marginal contribution of each meteorological feature of each test sample.
[0163] In this step, the electronic device can calculate the error prediction result of each test sample based on the marginal contribution of each meteorological feature of each test sample. The error prediction result of each test sample can be expressed as the prediction mean and the SHAP value of each feature:
[0164]
[0165] Among them, f(x i ) is the error prediction value of the i-th sample, φ0 is the predicted mean, that is, the sample composed of the average value of all features The corresponding error prediction value.
[0166] S307 : Based on the marginal contribution of each meteorological feature of each test sample, locally evaluate the factors affecting the wind power prediction error to obtain a local evaluation result.
[0167] In this step, for a single test sample, first, the marginal contribution of each feature of each test sample, namely the SHAP value, is calculated. The positive or negative and the size of the SHAP value represent the specific impact direction (positive or negative push) and impact size of the meteorological feature on the prediction error at the current moment; then, a waterfall chart is drawn to intuitively explain the cause of the error at this time; finally, multiple samples are grouped according to the time series, and the SHAP value of each meteorological feature is statistically analyzed to complete the local evaluation of the factors affecting the prediction error.
[0168] S308. For any meteorological feature, calculate the average of the absolute values of the marginal contributions of the meteorological feature under all test samples.
[0169] In this step, for any meteorological feature, the marginal contribution of the meteorological feature under all test samples of the target test set, i.e., the SHAP value, is taken as the absolute value and then averaged, i.e.
[0170]
[0171] Among them, α j is the mean absolute SHAP value, which can characterize the global importance of the j-th feature to the power prediction error. j The larger the value of , the more significant the global impact of this feature on the prediction error.
[0172] S309: Based on the average of the absolute values of the marginal contributions of the meteorological characteristics under all the test samples, a global evaluation is performed on the factors affecting the wind power prediction error to obtain a global evaluation result.
[0173] After calculating the average absolute value of the marginal contribution of each meteorological feature (i.e. α j ) can be used to calculate the α of each meteorological feature. j The values are sorted to intuitively show the global impact of each meteorological feature on the power forecast error, complete the global evaluation, and obtain the global evaluation results.
[0174] In this embodiment, based on the contribution of each meteorological feature to the prediction error, the global importance of each meteorological feature to the prediction error can be ranked, and combined with the local evaluation of the factors affecting the prediction error, a feasible optimization strategy can be provided for the wind power prediction model, such as weighted adjustment of the input features of the prediction model according to the degree of influence of different meteorological characteristics, or using the error prediction value to correct the predicted power, so that the short-term predicted wind power is closer to the actual situation.
[0175] The embodiment of the present disclosure preprocesses the original data to obtain preprocessed data, constructs an input data set based on the preprocessed data, and constructs an initial random forest regression model based on the input data set. Further, feature dimensionality reduction processing is performed on the input data set to obtain a target input data set, and the initial random forest regression model is trained based on the target input data set to obtain a target random forest regression model. Then, for each test sample of the target test set, a combination of different meteorological features of each test sample is sampled, and the change in the output of the target random forest regression model under the combination of different meteorological features is calculated. Based on the change in the output of the target random forest regression model under the combination of different meteorological features, the marginal contribution of each meteorological feature of each test sample is obtained, and based on the marginal contribution of each meteorological feature of each test sample, the error prediction result of each test sample is calculated. Then, based on the marginal contribution of each meteorological characteristic of each test sample, a local evaluation is performed on the factors affecting the wind power prediction error to obtain a local evaluation result. For any meteorological characteristic, the average value of the absolute value of the marginal contribution of the meteorological characteristic under all test samples is calculated. Based on the average value of the absolute value of the marginal contribution of the meteorological characteristic under all test samples, a global evaluation is performed on the factors affecting the wind power prediction error to obtain a global evaluation result. Through this method, the embodiment of the present disclosure can evaluate the global impact of each meteorological characteristic on the wind power prediction error, and evaluate the local impact of each meteorological characteristic on the prediction error, thereby providing a feasible optimization strategy for the wind power prediction model and making the predicted wind power more accurate.
[0176] Figure 4 Schematic diagram of the structure of the prediction error influencing factor evaluation device provided by the embodiment of the present disclosure. The prediction error influencing factor evaluation device can be an electronic device as described in the above embodiment, or the prediction error influencing factor evaluation device can be a component or assembly in the electronic device. The prediction error influencing factor evaluation device provided by the embodiment of the present disclosure can execute the processing flow provided by the prediction error influencing factor evaluation method embodiment, such as Figure 4 As shown, the prediction error influencing factor evaluation device 40 includes: a construction module 41, a construction module 42, an acquisition module 43, a calculation module 44, and an evaluation module 45; wherein the construction module 41 is used to preprocess the original data to obtain preprocessed data, and construct an input data set based on the preprocessed data; the construction module 42 is used to construct an initial random forest regression model based on the input data set; the acquisition module 43 is used to perform feature dimensionality reduction processing on the input data set to obtain a target input data set, and train the initial random forest regression model based on the target input data set to obtain a target random forest regression model; the calculation module 44 is used to perform contribution analysis on the target random forest regression model, and calculate the contribution result of each meteorological feature in the target input data set to the prediction error; the evaluation module 45 is used to evaluate the influencing factors of the wind power prediction error based on the contribution result of each meteorological feature to the prediction error.
[0177] Optionally, the construction module 41 preprocesses the original data to obtain preprocessed data, and when constructing the input data set based on the preprocessed data, it is specifically used to: arrange the original data in chronological order to obtain the original data set, and the original data set includes measured data and predicted data; wherein the measured data includes meteorological measured data and wind power measured data, and the predicted data includes meteorological predicted data and wind power predicted data; process the repeated data in the measured data to obtain first measured data, process the missing data in the first measured data to obtain second measured data; process the abnormal data in the second measured data to obtain third measured data, and obtain preprocessed data based on the third measured data and the predicted data; calculate the prediction error of meteorological data and the prediction error of wind power based on the preprocessed data; construct the input data set based on the prediction error of wind power, the absolute value of the prediction error of meteorological data and the preprocessed data.
[0178] Optionally, when the construction module 42 constructs the initial random forest regression model based on the input data set, it is specifically used to: divide the input data set into a training set and a test set based on a preset ratio; sample the training set a preset number of times based on a sampling method with replacement to obtain a first sub-training set, use the first sub-training set as the training set of the first decision tree, split the nodes of the first decision tree based on the optimal splitting point determined by the impurity to obtain the first target decision tree, repeat the current step until a preset number of target decision trees are obtained; generate an initial random forest regression model based on the preset number of target decision trees; input the test set into the initial random forest regression model to obtain the error prediction value of the wind power of each test sample in the test set.
[0179] Optionally, the obtaining module 43 performs feature dimensionality reduction processing on the input data set to obtain the target input data set, which is specifically used to: calculate the importance score of each meteorological feature in the input data set, normalize the importance score of each meteorological feature, and obtain the normalized importance score of each meteorological feature; calculate the Pearson correlation coefficient between any two meteorological features, and calculate the variance inflation factor of each meteorological feature; based on the normalized importance score of each meteorological feature, the Pearson correlation coefficient between any two meteorological features, and the variance inflation factor of each meteorological feature, remove the meteorological features that meet the conditions from the meteorological features in the input data set to obtain the target input data set.
[0180] Optionally, when the obtaining module 43 trains the initial random forest regression model based on the target input data set to obtain the target random forest regression model, it is specifically used to: divide the target input data set into a target training set and a target test set, and train the initial random forest regression model based on the target training set to obtain a pre-trained random forest regression model; determine whether the prediction accuracy of the pre-trained random forest regression model meets a preset accuracy threshold; if so, determine the pre-trained random forest regression model as the target random forest regression model; if not, cross-validate and hyperparameter tune the pre-trained random forest regression model to obtain a target random forest regression model.
[0181] Optionally, when the calculation module 44 performs a contribution analysis on the target random forest regression model and calculates the contribution of each meteorological feature in the target input data set to the prediction error, it is specifically used to: for each test sample of the target test set, sample the combination of different meteorological features of each test sample, and calculate the change in the output of the target random forest regression model under the combination of different meteorological features; based on the change in the output of the target random forest regression model under the combination of different meteorological features, obtain the marginal contribution of each meteorological feature of each test sample; based on the marginal contribution of each meteorological feature of each test sample, calculate the error prediction result of each test sample.
[0182] Optionally, when the evaluation module 45 evaluates the influencing factors of the wind power prediction error based on the contribution result of each meteorological feature to the prediction error, it is specifically used to: perform a local evaluation on the influencing factors of the wind power prediction error based on the marginal contribution of each meteorological feature of each test sample to obtain a local evaluation result; for any meteorological feature, calculate the average value of the absolute value of the marginal contribution of the meteorological feature under all test samples; perform a global evaluation on the influencing factors of the wind power prediction error based on the average value of the absolute value of the marginal contribution of the meteorological feature under all test samples to obtain a global evaluation result.
[0183] Figure 4 The prediction error influencing factor evaluation device of the illustrated embodiment can be used to implement the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effects are similar and will not be repeated here.
[0184] Figure 5 The electronic device provided by the embodiment of the present disclosure can execute the processing flow provided by the embodiment of the method for evaluating the factors affecting the prediction error, such as Figure 5 As shown, the electronic device 80 includes: a memory 81, a processor 82, a computer program and a communication interface 83; wherein the computer program is stored in the memory 81 and is configured so that the processor 82 executes the above-mentioned method for evaluating factors affecting the prediction error.
[0185] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the method for evaluating factors affecting prediction errors described in the above embodiment.
[0186] In addition, an embodiment of the present disclosure further provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, implements the above-mentioned method for evaluating factors affecting the prediction error.
[0187] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0188] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an adhoc peer-to-peer network), as well as any currently known or future developed network.
[0189] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0190] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0191] Preprocessing the original data to obtain preprocessed data, and constructing an input data set based on the preprocessed data;
[0192] Based on the input data set, construct an initial random forest regression model;
[0193] Performing feature dimensionality reduction processing on the input data set to obtain a target input data set, and training the initial random forest regression model based on the target input data set to obtain a target random forest regression model;
[0194] Performing a contribution analysis on the target random forest regression model to calculate the contribution of each meteorological feature in the target input data set to the prediction error;
[0195] Based on the contribution of each meteorological feature to the prediction error, the factors affecting the wind power prediction error are evaluated.
[0196] In addition, the electronic device may also execute other steps in the above-mentioned method for evaluating factors affecting prediction error.
[0197] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0198] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0199] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0200] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0201] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0202] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0203] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for evaluating factors affecting prediction error, characterized in that: The method comprises: Preprocessing the original data to obtain preprocessed data, and constructing an input data set based on the preprocessed data; Based on the input data set, construct an initial random forest regression model; Performing feature dimensionality reduction processing on the input data set to obtain a target input data set, and training the initial random forest regression model based on the target input data set to obtain a target random forest regression model; Performing a contribution analysis on the target random forest regression model to calculate the contribution of each meteorological feature in the target input data set to the prediction error; Based on the contribution of each meteorological feature to the prediction error, the factors affecting the wind power prediction error are evaluated.
2. The method according to claim 1, characterized in that The preprocessing of the original data to obtain preprocessed data, and constructing an input data set based on the preprocessed data, includes: Arrange the original data in chronological order to obtain an original data set, wherein the original data set includes measured data and predicted data; wherein the measured data includes meteorological measured data and wind power measured data, and the predicted data includes meteorological predicted data and wind power predicted data; Processing duplicate data in the measured data to obtain first measured data, and processing missing data in the first measured data to obtain second measured data; Processing abnormal data in the second measured data to obtain third measured data, and obtaining preprocessed data based on the third measured data and the predicted data; Calculating the prediction error of the meteorological data and the prediction error of the wind power based on the preprocessed data; An input data set is constructed based on the wind power prediction error, the absolute value of the meteorological data prediction error, and the preprocessed data.
3. The method according to claim 1, characterized in that The initial random forest regression model is constructed based on the input data set, comprising: Divide the input dataset into training and test sets based on a preset ratio; Sampling the training set a preset number of times based on a sampling with replacement method to obtain a first sub-training set, using the first sub-training set as the training set for a first decision tree, performing node splitting on the first decision tree based on an optimal splitting point determined by impurity to obtain a first target decision tree, and repeating the current step until a preset number of target decision trees are obtained; Generate an initial random forest regression model based on the preset number of target decision trees; The test set is input into the initial random forest regression model to obtain an error prediction value of wind power for each test sample in the test set.
4. The method according to claim 1, wherein The step of performing feature dimensionality reduction processing on the input data set to obtain a target input data set includes: Calculating the importance score of each meteorological feature in the input data set, and normalizing the importance score of each meteorological feature to obtain a normalized importance score of each meteorological feature; Calculate the Pearson correlation coefficient between any two meteorological characteristics and calculate the variance inflation factor of each meteorological characteristic; Based on the normalized importance score of each meteorological feature, the Pearson correlation coefficient between any two meteorological features, and the variance inflation factor of each meteorological feature, meteorological features that meet the conditions are removed from each meteorological feature in the input data set to obtain a target input data set.
5. The method according to claim 1, wherein The training of the initial random forest regression model based on the target input data set to obtain a target random forest regression model includes: Dividing the target input data set into a target training set and a target test set, and training the initial random forest regression model based on the target training set to obtain a pre-trained random forest regression model; Determining whether the prediction accuracy of the pre-trained random forest regression model meets a preset accuracy threshold, and if so, determining the pre-trained random forest regression model as the target random forest regression model; If not, cross-validation and hyperparameter tuning are performed on the pre-trained random forest regression model to obtain a target random forest regression model.
6. The method according to claim 1, wherein The performing contribution analysis on the target random forest regression model to calculate the contribution of each meteorological feature in the target input data set to the prediction error includes: For each test sample in the target test set, sampling different combinations of meteorological features of each test sample, and calculating the change in the output of the target random forest regression model under the different combinations of meteorological features; Based on the change in the output of the target random forest regression model under different combinations of meteorological characteristics, the marginal contribution of each meteorological characteristic of each test sample is obtained; Based on the marginal contribution of each meteorological feature of each test sample, the error prediction result of each test sample is calculated.
7. The method according to claim 6, characterized in that The evaluating of factors affecting the wind power prediction error based on the contribution of each meteorological feature to the prediction error includes: Based on the marginal contribution of each meteorological characteristic of each test sample, locally evaluate the factors affecting the wind power forecast error to obtain a local evaluation result; For any meteorological feature, calculate the average of the absolute values of the marginal contributions of the meteorological feature under all test samples; Based on the average of the absolute values of the marginal contributions of the meteorological characteristics under all the test samples, a global evaluation is performed on the factors affecting the wind power prediction error to obtain a global evaluation result.
8. A device for evaluating factors affecting prediction error, characterized in that: The device comprises: A construction module is used to preprocess the original data to obtain preprocessed data, and construct an input data set based on the preprocessed data; A construction module, configured to construct an initial random forest regression model based on the input data set; An obtaining module is used to perform feature dimensionality reduction processing on the input data set to obtain a target input data set, and train the initial random forest regression model based on the target input data set to obtain a target random forest regression model; a calculation module, configured to perform a contribution analysis on the target random forest regression model and calculate a contribution result of each meteorological feature in the target input data set to the prediction error; The evaluation module is used to evaluate the factors affecting the wind power prediction error based on the contribution result of each meteorological feature to the prediction error.
9. An electronic device, characterized in that: include: Memory; processor; as well as computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.