Wind power density prediction method and device based on mean shift clustering
Through the combination of mean drift clustering and long-term memory neural network, the problems of large data volume and large prediction errors in offshore wind power power prediction are solved, and high-precision wind power power density prediction is achieved.
Patent Information
- Application Number
- CN202510457142.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-13
- Publication Date
- 2025-07-11
AI Technical Summary
The existing technology has the problem of huge data volume and difficult to deal with in the prediction of offshore wind power. Traditional methods such as physical models and statistical models are difficult to adapt to the changing marine climate conditions, resulting in large prediction errors, and different prediction models need to be built under different sea areas and meteorological conditions.
The wind power parameter data is clustered and analyzed by means of mean drift clustering. Combined with Spearman correlation coefficient and long-term memory neural network, the clustering center is screened through the mean drift clustering algorithm to establish a wind power power density prediction model.
The data processing volume is simplified, the accuracy and accuracy of prediction results are improved, and the wind power prediction is adapted to different sea areas and meteorological conditions.
Smart Images

Figure CN120296445A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of wind power prediction, and particularly to a wind power density prediction and device based on mean shift clustering. Background Art
[0002] With the rapid development of the global low-carbon economy, it is necessary to further optimize the proportion of the energy structure, and it is more important to increase the resource allocation in the green and low-carbon fields and the development and utilization degree of renewable energy. With the continuous expansion of the current offshore wind power installation scale year by year, remarkable achievements have been made in global offshore wind power. However, the current development of wind power technology and the power scheduling plan are still crucial issues faced in the wind power field. Among them, the prediction of offshore wind power is the basic premise for doing a good job in power dispatching and an important guarantee for ensuring the safe, stable and economic operation of the power system.
[0003] The prediction of traditional wind power generally adopts the physical model method and the statistical model method. However, for the variable climate conditions of offshore wind power, when using the physical model for prediction, not only the influence of various meteorological quantities such as wind speed and wind direction on the instantaneous wind power density needs to be analyzed, but also complex physical effects such as the thermal effect of the sea current and the wake effect need to be considered. The existence of these effects makes it difficult to perform predictions using the physical model method. The statistical model method is to find the potential relationship between electrical quantities and meteorological information through neural networks and other methods to achieve prediction. However, for power prediction of wind farms under different sea areas and meteorological conditions, different prediction models need to be constructed and different prediction data need to be selected to reduce the prediction error. At present, there are still problems such as a huge amount of data and difficulty in processing for the statistical model method for offshore wind power. Summary of the Invention
[0004] To overcome the deficiencies of the above-mentioned prior art, the present application provides a wind power density prediction and device based on mean shift clustering, and specifically adopts the following technical solutions:
[0005] A wind power density prediction method based on mean shift clustering, the method includes the following steps:
[0006] Obtain the parameter types affecting wind power, and collect the historical parameter data of the corresponding parameters and the wind power data; the parameter types at least include wind speed, wind direction, temperature, humidity and atmospheric pressure;
[0007] Perform normalization processing on the historical parameter data and the wind power data of different parameter types respectively;
[0008] Use the Spearman correlation coefficient method to perform correlation analysis on the parameter types after normalization processing and the power of the wind turbine generator respectively;
[0009] Use the mean shift clustering algorithm to perform clustering analysis on the historical parameter data of the selected parameter types to obtain different clustering centers;
[0010] Calculate the Euclidean distance between the parameter data in the time period to be predicted and different clustering centers, and select the best clustering center;
[0011] Use the parameter data included in the best clustering center to train the long short-term memory neural network to obtain a prediction model;
[0012] Input the parameter data in the time period to be predicted into the trained prediction model to obtain the prediction result of the wind power density.
[0013] Optionally: The step of normalizing the correlation parameters includes:
[0014] Obtain the maximum and minimum values of the historical parameter data in each parameter type, and calculate the extreme values of different parameter types;
[0015] Calculate the difference between each parameter data in different parameter types and the minimum value of the corresponding parameter type respectively, and calculate the ratio of the difference to the extreme value of the corresponding parameter type;
[0016] Amplify the obtained ratio by a multiple to obtain the normalized value of the historical parameter data of the corresponding parameter type.
[0017] Optionally: The step of using the Spearman correlation coefficient to perform correlation analysis on the normalized parameter types and the wind turbine power respectively includes:
[0018] Sort the parameter data of each normalized parameter type and the wind turbine power in ascending order to obtain the parameter data set of the corresponding parameter type and the wind turbine power data set;
[0019] Subtract the elements of the parameter data set of each parameter type from the wind turbine power data set respectively to obtain the ranking difference set of the corresponding parameter type;
[0020] Calculate the correlation coefficient between the corresponding parameter type and the wind turbine power according to the ranking difference sets of different parameter types.
[0021] Optionally: The step of calculating the correlation coefficient between the corresponding parameter type and the wind turbine power according to the ranking difference sets of different parameter types includes:
[0022]
[0023] Where ρ s is the correlation coefficient between the corresponding parameter type and the wind turbine power; d iis the difference value between the i-th parameter data in the corresponding parameter type and the power of the wind turbine; n is the number of parameters in the corresponding parameter type.
[0024] Optionally: The step of performing clustering analysis on the historical parameter data of the selected parameter type by using the mean shift clustering algorithm and screening out different cluster centers includes:
[0025] Determine any parameter data as the initial cluster center, and set the search radius and convergence threshold;
[0026] Screen out the parameter data smaller than the search radius, and set it as the data classification. At the same time, mark the screened parameter data;
[0027] Calculate the drift vector based on the currently screened parameter data, and determine whether the calculated drift vector is less than the convergence threshold;
[0028] When it is judged that the drift vector is less than the convergence threshold, fix the current cluster center point, and classify the screened parameter data into the current data classification; when it is judged that the drift vector is greater than or equal to the convergence threshold, update the cluster center and perform the next screening iteration;
[0029] Calculate the distance between different cluster centers, and determine whether the distance between different cluster centers is greater than the convergence threshold: when the distance between two cluster centers is greater than the convergence threshold, a new data classification is added; when the distance between two cluster centers is less than or equal to the convergence threshold, the two cluster centers are merged into the same data classification, and the marking times are superimposed;
[0030] When it is judged that all parameter data are marked, assign each parameter data to the data classification with the most marking times; when it is judged that there are unmarked parameter data, determine any unmarked parameter data as the cluster center for iteration.
[0031] Optionally: The method for calculating the drift vector based on the currently screened parameter data is:
[0032]
[0033] where M r is the drift vector; x is the cluster center point; S is the range of the current search radius r; x i is the i-th sample point within S; K is the number of sample points within S; y represents the sample points within the range with x as the center and radius r; X represents the sample space, which is the set of all sample points.
[0034] Optionally: When it is judged that the drift vector is greater than or equal to the convergence threshold, the method for updating the cluster center is:
[0035] x t+1 = x t + Mr ;
[0036] where M r is the drift vector; x t is the cluster center point at the t-th iteration; x t+1 is the cluster center point at the (t + 1)-th iteration.
[0037] Optionally; the steps of calculating the Euclidean distances between the parameter data of the to-be-predicted time period and different cluster centers and screening the best cluster center include:
[0038] Normalize the parameter data of the to-be-predicted time period;
[0039] Calculate the Euclidean distances between the normalized parameter data of the to-be-predicted time period and different cluster centers respectively:
[0040]
[0041] where O j is the Euclidean distance between the parameter data of the to-be-predicted time period and the corresponding cluster center; x(k) is the cluster center point; x j (k) is the j-th parameter data in the k-th prediction time period, and m is the number of parameter data of the to-be-predicted time period;
[0042] Select the cluster center with the minimum Euclidean distance as the best cluster center.
[0043] Optionally: the steps of training the long short-term memory neural network with the parameter data included in the best cluster center to obtain a prediction model include:
[0044] Perform time sorting based on the parameter data included in the best cluster center and the corresponding wind power data;
[0045] Construct an input sequence for the sorted parameter data and the corresponding wind power data in a sliding window manner;
[0046] Input the input sequence into the long short-term memory neural network for training and verification to obtain a prediction model.
[0047] In addition, the present application also discloses a wind power density prediction device based on mean shift clustering, and the device includes:
[0048] A parameter acquisition module, configured to acquire the parameter types affecting wind power, and collect the historical parameter data and wind power data of the corresponding parameters; the parameter types at least include wind speed, wind direction, temperature, humidity, and atmospheric pressure;
[0049] A parameter preprocessing module, configured to normalize the historical parameter data and wind power data of different parameter types respectively;
[0050] A correlation analysis module, which is used to perform a correlation analysis on the parameter type after normalization and the wind turbine power respectively by using the Spearman correlation coefficient method;
[0051] A clustering analysis module, which is used to perform a clustering analysis on the historical parameter data of the filtered parameter type by using the mean shift clustering algorithm to obtain different clustering centers;
[0052] A clustering and screening module, which calculates the Euclidean distance between the parameter data in the period to be predicted and different clustering centers, and screens the best clustering center;
[0053] A model training module, which is used to train a long short-term memory neural network with the parameter data included in the best clustering center to obtain a prediction model;
[0054] A model prediction module, which is used to input the parameter data in the period to be predicted into the trained prediction model to obtain the prediction result of the wind power density.
[0055] Beneficial effects
[0056] The technical solution of this application has obtained the following beneficial effects:
[0057] The wind power density prediction method of this application performs a clustering analysis on the historical parameter data through the Mean shift clustering method to form different clustering centers. Subsequently, by comparing the Euclidean distance between the parameter data in the period to be predicted and each clustering center, the best clustering result that conforms to the current parameter data to be predicted is screened out, and this clustering result is used as the training sample of the LSTM neural network algorithm. Furthermore, a prediction model for predicting the wind power density is established, so as to predict the wind power density. This method greatly simplifies the data processing volume of model training, and at the same time, the prediction result is relatively more accurate and the result is more accurate. Description of the drawings
[0058] Figure 1 It is a flowchart of the wind power density prediction method based on mean shift clustering in the embodiment of this application.
[0059] Figure 2 It is a schematic diagram of the output power characteristic curve of the wind turbine in the embodiment of this application.
[0060] Figure 3 It is a matrix diagram of the parameter data obtained based on the Spearman correlation coefficient and the influence of WPD in the embodiment of this application.
[0061] Figure 4 It is a flowchart of adopting Mean shift clustering in the embodiment of this application.
[0062] Figure 5It is the internal structure diagram of the LSTM network in the embodiments of this application.
[0063] Figure 6 It is the coordinate diagram of 12 wind turbines in the Alpha ventus wind farm in the embodiments of this application.
[0064] Figure 7 It is the 3D diagram of the wind speed change at each wind turbine within 24 hours on a certain day in the Alpha ventus wind farm in the embodiments of this application.
[0065] Figure 8 It is the relationship diagram between the sum of Euclidean distances and the clustering situation after adjusting the search radius in the embodiments of this application.
[0066] Figure 9 It is the wind speed curve diagram of a certain period in the Alpha ventus wind farm and the corresponding clustering center in the embodiments of this application.
[0067] Figure 10 It is the schematic diagram of the Euclidean distances from the data of the Alpha ventus wind farm on February 1 to each clustering center point in the embodiments of this application.
[0068] Figure 11 It is the WPD prediction schematic diagram of the Alpha ventus wind farm on February 1 in the embodiments of this application.
[0069] Figure 12 It is the structure diagram of the wind power density prediction device based on mean shift clustering in the embodiments of this application.
[0070] Figure 13 It is the composition diagram of an electronic device in the embodiments of this application. Detailed implementation manners
[0071] The following further describes this application with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of this application and cannot be used to limit the protection scope of this application. It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations for this application.
[0072] Currently, for power prediction of wind farms under different sea areas and meteorological conditions, different prediction models need to be constructed and different prediction methods need to be selected to reduce the prediction error and provide a basis for power dispatching planning. Currently, for the problems of huge amount of data and difficult processing existing in offshore wind power prediction, the traditional method is to classify and simplify a large amount of offshore wind power data using statistical methods. For example, by using the method of K-means clustering analysis to find similar samples, the climate characteristics of different offshore wind farms are classified driven by data, and the method of typical day analysis is adopted. Although it reduces the amount of data and makes the prediction more targeted. However, K-means clustering requires specifying the number of clusters and the initial cluster centers in advance. Once these two values are not selected well, it is possible to obtain ineffective clustering results, and even lead to an infinite loop.
[0073] To solve the above problems, this application proposes a prediction method for offshore wind power density based on mean shift clustering. Since the wind speed distribution characteristics and wind power density WPD are key indicators for measuring the wind resource assessment of wind farms and their operation reliability and economy, which are directly related to the efficient utilization of wind energy. Therefore, this application mainly uses the data obtained from NWP (numerical weather prediction) for Mean shift clustering to reduce the amount of data, and then uses the LSTM neural network algorithm to establish a prediction model for wind power density, so as to predict the wind power density.
[0074] Combined with Figure 1 As shown, this embodiment specifically discloses a wind power density prediction method based on mean shift clustering. The method includes the following steps:
[0075] Parameter determination and acquisition:
[0076] Obtain the parameter types affecting wind power, and collect the historical parameter data and wind power data of the corresponding parameters; the parameter types at least include wind speed, wind direction, temperature, humidity and atmospheric pressure; usually, offshore wind power is affected by various factors. In addition to the limitations of various parameters of the unit itself, the output power of wind turbines is also affected by many natural conditions, such as wind speed, wind direction, temperature, humidity, atmospheric pressure, etc., and sometimes the natural conditions have a more significant impact on the output power of wind turbines. Generally, a piecewise function as shown in Figure 2 is used to describe the relationship between wind turbines and wind speed. Among them, in Figure 2 , P r is the rated power, v ci , v r , v co are the cut-in wind speed, rated wind speed and cut-out wind speed respectively. From Figure 2It can be obtained that when the wind speed is less than the cut-in wind speed, the grid connection condition cannot be achieved, and the wind power output is 0; when the wind speed is greater than the cut-out wind speed, in order to ensure the safety of the unit, it must be cut out, and the wind power output is also 0. Therefore, in the embodiments of the present application, natural factors such as wind speed, wind direction, temperature, humidity, and atmospheric pressure are generally selected to predict wind power.
[0077] Data preprocessing process:
[0078] In order to eliminate the influence of the dimension between each index in this embodiment, it is necessary to perform normalization processing on the historical parameter data and wind power data of different parameter types respectively. The steps of performing normalization processing on the historical parameter data in this embodiment include:
[0079] Obtain the maximum value and minimum value of the historical parameter data in each parameter type, and calculate the extreme values of different parameter types;
[0080] Calculate the difference between each parameter data and the minimum value of the corresponding parameter type in different parameter types respectively, and calculate the ratio of the difference to the extreme value of the corresponding parameter type;
[0081] Magnify the obtained ratio by a multiple to obtain the normalized value of the historical parameter data of the corresponding parameter type.
[0082] For example, in this embodiment, the method of range normalization is adopted to convert the parameter data into a number between 0 and 1, and then multiply it by the multiple 10, as shown in the following formula:
[0083]
[0084] where y i is the normalized parameter data; x i is the i-th parameter data in a certain parameter type; x min is the minimum value of the parameter data in a certain parameter type; x max is the maximum value of the parameter data in a certain parameter type. Through the above formula, all parameter data can be converted into numbers between 0 and 10, removing the influence of dimension.
[0085] Correlation analysis process:
[0086] In order to quantitatively describe the correlation degree between input features such as wind speed, wind direction, and air temperature and the wind power output in this embodiment, the Spearman correlation coefficient is used for visual processing, that is, the Spearman correlation coefficient method is used to perform correlation analysis on the parameter type after normalization processing and the wind turbine power respectively.
[0087] Specifically, in this embodiment, the process of using the Spearman correlation coefficient to analyze the correlation between the parameter types after normalization and the wind turbine power is as follows:
[0088] Sort the parameter data of each parameter type after normalization and the wind turbine power in ascending order to obtain the parameter data set of the corresponding parameter type and the wind turbine power data set; for example, when analyzing the correlation between any parameter type and the wind turbine power in this embodiment, the corresponding parameter type can be regarded as the variable set X, and the wind turbine power can be regarded as the variable set Y. The number of their elements is N, and the i-th (1 <= i <= N) value in the two variable sets is represented by X i and Y i respectively. Subsequently, sort the elements in the variable set X and the variable set Y, for example, both in ascending or descending order. In this embodiment, ascending order is preferably used, and two element permutation sets x and y are obtained.
[0089] Subsequently, subtract the elements of the parameter data set of the corresponding parameter type from the elements of the wind turbine power data set, that is, subtract the corresponding elements in the variable sets x and y to obtain the rank difference value d i of the corresponding parameter type; its specific calculation formula is:
[0090] d i = rg(X i ) - rg(Y i );
[0091] Finally, calculate the correlation coefficient between the corresponding parameter type and the wind turbine power according to the rank difference sets of different parameter types:
[0092]
[0093] where ρ s is the correlation coefficient between the corresponding parameter type and the wind turbine power; d i is the difference value between the i-th parameter data in the corresponding parameter type and the wind turbine power; n is the number of parameters in the corresponding parameter type.
[0094] Parameter screening process:
[0095] Based on the correlation analysis results, screen out the highly correlated parameter types and the lowly correlated parameter types; for example, in this embodiment, a correlation coefficient matrix can be obtained through environmental factors such as NWP wind speed, NWP wind direction, NWP air temperature, NWP relative humidity, NWP pressure, etc. of offshore wind power and WPD, such as Figure 3As shown, it can be seen that many influencing factors of offshore wind power are interrelated and complex and diverse. Among them, the correlation between wind speed and WPD is the highest, reaching 0.97. The correlations of other factors such as wind direction, pressure, relative humidity, etc. with WPD are relatively low, and the correlation between air temperature and WPD is the lowest. In subsequent analysis, the influence of each factor on WPD can be analyzed according to the corresponding weights.
[0096] Mean shift clustering analysis process:
[0097] In this embodiment, the mean shift clustering algorithm is used to perform clustering analysis on the historical parameter data of the selected parameter types to obtain different clustering centers. Among them, the Mean shift clustering algorithm is a hill-climbing algorithm based on kernel density estimation and is currently mainly applied to clustering, image segmentation, data tracking, etc. Compared with the K-means clustering algorithm, the Mean shift clustering algorithm does not require pre-setting the number of clusters and does not need to specify the position of the sample center point, both of which can be obtained through automatic iteration of the computer, and this algorithm has more stable clustering results.
[0098] Specifically, in combination with Figure 4 As shown, the steps of using the mean shift clustering algorithm to perform clustering analysis on the historical parameter data of the selected parameter types in this embodiment include:
[0099] Determine any parameter data as the initial clustering center, and set the search radius and convergence threshold;
[0100] Select the parameter data smaller than the search radius and set it as the data classification, and at the same time mark the selected parameter data;
[0101] Calculate the drift vector based on the currently selected parameter data, and determine whether the calculated drift vector is less than the convergence threshold; the method of calculating the drift vector based on the currently selected parameter data is:
[0102]
[0103] Where M r is the drift vector; x is the clustering center point; S is the range of the current search radius r; x i is the i-th sample point within the range of S; K is the number of sample points within the range of S; y represents the sample points within the range with x as the center and radius r; X represents the sample space, which is the set of all sample points.
[0104] When it is determined that the drift vector is less than the convergence threshold, fix the current clustering center point and classify the selected parameter data as the current data classification; when it is determined that the drift vector is greater than or equal to the convergence threshold, update the clustering center and perform the next screening iteration; the method of updating the clustering center is:
[0105] x t+1 = x t + M r ;
[0106] where M r is the drift vector; x t is the cluster center at the t-th iteration; x t+1 is the cluster center at the (t + 1)-th iteration.
[0107] Calculate the distances between different cluster centers and determine whether the distances between different cluster centers are greater than the convergence threshold: when the distance between two cluster centers is greater than the convergence threshold, new data classification is added; when the distance between two cluster centers is less than or equal to the convergence threshold, the two cluster centers are merged into the same data classification, and the marking times are superimposed;
[0108] When it is determined that all parameter data are marked, each parameter data is respectively assigned to the data classification with the most marking times; when it is determined that there is unmarked parameter data, any unmarked parameter data is determined as the cluster center for iteration.
[0109] It should be noted that there are two parameters in the process of clustering analysis in this embodiment, namely the search radius r and the threshold. Among them, the clustering result is not very sensitive to the threshold, so the threshold can be taken as a constant. By adjusting the value of the search radius r, different clustering results can be obtained. However, if the value of r is too small, there will be too many clustering results, and the number of samples in each category will be too small, resulting in an increase in the error of neural network prediction; if the value of r is too large, the clustering result will be inaccurate, and the sample points are far from the cluster center, making the representative ability of the cluster center poor.
[0110] Cluster center screening process:
[0111] Calculate the Euclidean distances between the parameter data in the time period to be predicted and different cluster centers, and screen the best cluster center.
[0112] Specifically, the steps of calculating the Euclidean distances between the parameter data in the time period to be predicted and different cluster centers in this embodiment include:
[0113] Normalize the parameter data in the time period to be predicted;
[0114] Calculate the Euclidean distances between the normalized parameter data in the time period to be predicted and different cluster centers respectively:
[0115]
[0116] where O j is the Euclidean distance between the parameter data in the time period to be predicted and the corresponding cluster center; x(k) is the cluster center; x j(k) is the j-th parameter data in the prediction time period k, and m is the number of parameter data in the time period to be predicted;
[0117] Select the clustering center with the smallest Euclidean distance as the best clustering center.
[0118] Model training process:
[0119] Train the long short-term memory neural network with the parameter data included in the best clustering center to obtain a prediction model. In this embodiment, in order to accurately predict the WPD (Wind Power Density, that is, the instantaneous wind power density) in each time period of the wind farm, a prediction model based on LSTM (long short-term memory neural network) is adopted. Compared with the traditional RNN model, LSTM introduces the concept of cell state. Its internal structure diagram is as follows Figure 5 shown. LSTM has a long-term memory function and solves the problems of gradient disappearance and gradient explosion existing in the long sequence training process. Therefore, the LSTM network is more suitable for learning classification, processing and predicting time series from experience when there is a long time of uncertainty between important events.
[0120] Specifically, the steps of training the long short-term memory neural network in this embodiment include:
[0121] Sort the parameter data and the corresponding wind power data based on the parameter data included in the best clustering center;
[0122] Use the sliding window method to construct an input sequence for the sorted parameter data and the corresponding wind power data;
[0123] Input the input sequence into the long short-term memory neural network for training and verification to obtain a prediction model.
[0124] It should be noted that the LSTM unit in this embodiment has 3 inputs at time t, which are: the input value x of the network at the current moment t , the output value h of the LSTM hidden layer at the previous moment t-1 , and the cell state c at the previous moment t-1 ; the LSTM unit has 2 outputs at time t: the output value h of the hidden layer at the current moment t and the cell state c t . LSTM allows information to be selectively remembered or forgotten through the cooperation of the forget gate, input gate, and output gate, affecting the cell state.
[0125] Among them, the function of the forget gate of the LSTM unit is to selectively forget the information in the cell state, as shown in the following formula:
[0126] f t =σ(W f·[h t-1 ,x t +b f );
[0127] where f t is the proportionality coefficient for controlling past information at the current moment, with a range of [0, 1]; σ is the sigmoid activation function; W f is the weight matrix of the forget gate; b f is the bias term of the forget gate.
[0128] The input gate of the LSTM cell is used to selectively record new information into the cell state, as shown in the following equation:
[0129] i t = σ(W i ·[h t-1 ,x t +b i );
[0130]
[0131] where, i t is the proportionality coefficient for controlling input information at the current moment, with a range of [0, 1]; W i is the weight matrix of the input gate; b i is the bias term of the input gate; is the candidate cell state at the current moment, with a range of [-1, 1]; tanh is the activation function; W c is the weight matrix of the candidate cell state; b c is the bias term of the candidate cell state.
[0132] Subsequently, the cell state is updated: the cell state at the current moment is obtained by operating on the data obtained in the previous two steps, as shown in the following equation:
[0133]
[0134] Finally, the output gate of the LSTM cell is used to calculate the output value h t of the hidden layer at the current moment, as shown in the following equation:
[0135] o t = σ(W o ·[h t-1 ,x t +b o );
[0136] h t = o t ·tanh(c t );
[0137] where, ot is the proportionality coefficient for controlling the output information at the current moment, with a range of [0, 1]; W o is the weight matrix of the output gate; b o is the bias term of the output gate; h t is the output at the current moment, with a range of [0, 1].
[0138] Power prediction process:
[0139] The parameter data of the period to be predicted is input into the trained prediction model to obtain the prediction result of the wind power density. This method converts complex meteorological data into structured input suitable for LSTM processing, and uses clustering technology to improve data homogeneity, ultimately achieving high-precision WPD prediction.
[0140] Furthermore, in this embodiment, to verify the accuracy of the above method, the data from January 1st to 31st and February 1st, 2021 of the German Alpha ventus offshore wind farm is selected to verify the effectiveness of the method. Alpha ventus is the first offshore wind farm in Germany, with 12 wind turbines and was put into production in April 2010. Figure 6 is the coordinate map of the 12 wind turbines of Alpha ventus. Figure 7 is the three-dimensional diagram of the wind speed change at each wind turbine within 24 hours of a certain day. From Figure 6 and the collected data, it can be seen that within the same wind farm, due to the close distance of different turbines, most of the natural condition data such as wind speed and wind direction are very similar. Therefore, the Alpha ventus offshore wind farm can be regarded as a point, and the data is analyzed and predicted using the average value.
[0141] Subsequently, the data from January 1st to 31st and February 1st, 2021 of the Alpha ventus offshore wind farm is analyzed for correlation using the Spearman correlation coefficient analysis method in this embodiment, as Figure 3 shown. The correlation between wind speed and WPD is the highest, reaching 0.97. The correlations of other factors such as wind direction, pressure, relative humidity, etc. with WPD are relatively low, and the correlation between air temperature and WPD is the lowest.
[0142] Referring to the results of the correlation coefficient analysis, considering the influence degree of each meteorological factor on WPD, the data from January 1st to 30th is subjected to Mean-shift clustering analysis. By adjusting different search radius r values, the relationship between the number of clusters and the sum of Euclidean distances is obtained, as Figure 8 shown.
[0143] According to Figure 8It can be seen that the more the number of classifications, the smaller the sum of Euclidean distances. When the number of classification exceeds 3, as the number of classifications continues to increase, the decrease in the sum of Euclidean distances is not obvious. In addition, the more the number of classifications, the fewer the number of data samples for neural network training, and the greater the prediction error. By weighing the Euclidean distance and the neural network training samples, this embodiment finally selects the number of clusters to be 3. The clustering results of the data of the Alpha ventus offshore wind farm in January 2021 are shown in Table 1. It can be seen that from January 1st to January 10th, 2021, a total of 10 days belong to the first category, from January 11th to January 19th, 2021, a total of 9 days belong to the second category, and from January 20th to January 30th, 2021, a total of 11 days belong to the third category.
[0144] Table 1
[0145] Date Category Date Category Date Category January 1st 1 January 11th 2 January 21st 3 January 2nd 1 January 12th 2 January 22nd 3 January 3rd 1 January 13th 2 January 23rd 3 January 4th 1 January 14th 2 January 24th 3 January 5th 1 January 15th 2 January 25th 3 January 6th 1 January 16th 2 January 26th 3 January 7th 1 January 17th 2 January 27th 3 January 8th 1 January 18th 2 January 28th 3 January 9th 1 January 19th 2 January 29th 3 January 10th 1 January 20th 3 January 30th 3
[0146] Among them, the data of the second category is used to test the clustering effect, and the clustering center point of the second category clustering and the wind speed curves of each time period from January 13th to 19th are drawn as follows Figure 9 shown. It can be Figure 9 seen that the error between the wind speed curves of each time period of the sample points and the wind speed curves of each time period of the clustering center is small, and the clustering result is good.
[0147] Subsequently, this embodiment uses the data on February 1st, 2021 to verify the predicted data. First, calculate the Euclidean distances from the data on February 1st to the 3 clustering centers, and the results are Figure 10 shown. It is found that the Euclidean distance from the data on this day to the clustering center point of the third category is the smallest, so February 1st belongs to the third category. Use the NWP wind direction, NWP wind speed, NWP temperature, NWP humidity, NWP atmospheric pressure from January 20th to January 30th and the WPD data of the same day to train the LSTM network, and then input the above data on February 1st into the LSTM network to predict the WPD on this day. The prediction results are as follows Figure 11 shown.
[0148] Since the value of WPD in this embodiment is relatively small, SMAPE (Symmetric Mean Absolute Percentage Error) is used to measure the accuracy of the measurement results. Compared with the traditional MAPE index (Mean Absolute Percentage Error), SMAPE can better avoid the problem that MAPE leads to too large calculation results due to too small true values. Its process is the same as that of MAPE, and the range of this value is also [0, +∞). Moreover, the smaller this value is, the higher the accuracy of the prediction model. When the predicted value and the true value are exactly the same, SMAPE is 0, that is, the model is a perfect model; when SMAPE is greater than 100%, the model is a poor model. The calculation of SMAPE is carried out according to the following formula:
[0149]
[0150] To further illustrate the effectiveness of the above method in this embodiment, quadratic regression analysis and cubic regression analysis methods are also used to predict WPD through wind speed respectively, and the results are shown in Figure Figure 11 as follows. According to the calculation, the SMAPE of the LSTM prediction of WPD on February 1 is 23.96%, the SMAPE of the quadratic regression prediction is 52.52%, and the SMAPE of the cubic regression prediction is 46.25%. It can be seen by comparison that the prediction method of this embodiment can predict WPD, and the prediction accuracy is good.
[0151] In addition, as shown in Figure 12 , this application also discloses a wind power density prediction device based on mean shift clustering. The device includes:
[0152] A parameter acquisition module, configured to acquire parameter types affecting wind power, and collect historical parameter data and wind power data of corresponding parameters; the parameter types at least include wind speed, wind direction, temperature, humidity, and atmospheric pressure;
[0153] A parameter preprocessing module, configured to perform normalization processing on historical parameter data and wind power data of different parameter types respectively;
[0154] A correlation analysis module, configured to perform correlation analysis on the normalized parameter types and the power of the wind turbine generator respectively by using the Spearman correlation coefficient method;
[0155] A clustering analysis module, configured to perform clustering analysis on the historical parameter data of the filtered parameter types by using the mean shift clustering algorithm to obtain different clustering centers;
[0156] The clustering and screening module calculates the Euclidean distances between the parameter data in the time period to be predicted and different clustering centers, and screens the optimal clustering center;
[0157] The model training module is used to train the long short-term memory neural network with the parameter data included in the optimal clustering center to obtain a prediction model;
[0158] The model prediction module is used to input the parameter data in the time period to be predicted into the trained prediction model to obtain the prediction result of the wind power density.
[0159] The device provided by the embodiment of the present application can implement Figure 1 each process implemented by the method embodiment. To avoid repetition, it will not be elaborated here.
[0160] As Figure 13 shown, the embodiment of the present application also provides an electronic device, including a processor and a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the method embodiment shown in Figure 1 and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0161] The embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above Figure 1 mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0162] The embodiment of the present application also provides a computer program product, including computer instructions. When the computer instructions are executed by the processor, they implement each process of the above Figure 1 mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0163] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiment of the present application. The sequence numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.
[0164] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the presence of additional identical elements in the process, method, article or device that includes such element.
[0165] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical functional division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another device, or some features can be ignored or not executed. Additionally, the couplings, direct couplings, or communication connections between the various components shown or discussed can be through some interfaces, and the indirect couplings or communication connections of devices or units can be electrical, mechanical, or in other forms.
[0166] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0167] In addition, in each embodiment of this application, the various functional units can all be integrated in one processing unit, or each unit can be a separate unit alone, or two or more units can be integrated in one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0168] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs that can store program codes.
[0169] Alternatively, if the above integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a device (which can be a terminal or a platform, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as removable storage devices, ROMs, magnetic disks, or optical discs.
[0170] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present application, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present application.
Claims
1. A wind power density prediction method based on mean shift clustering, characterized in that, The method includes the following steps: Obtain the parameter types that affect wind power, and collect the historical parameter data of the corresponding parameters and the wind power data; the parameter types at least include wind speed, wind direction, temperature, humidity, and atmospheric pressure; Perform normalization processing on the historical parameter data and wind power data of different parameter types respectively; Use the Spearman correlation coefficient method to perform correlation analysis on the parameter types after normalization processing and the wind turbine power respectively; Use the mean shift clustering algorithm to perform clustering analysis on the historical parameter data of the selected parameter types to obtain different cluster centers; Calculate the Euclidean distances between the parameter data in the period to be predicted and different cluster centers, and select the best cluster center; Train the long short-term memory neural network with the parameter data included in the best cluster center to obtain a prediction model; Input the parameter data in the period to be predicted into the trained prediction model to obtain the prediction result of the wind power density.
2. The wind power density prediction method according to claim 1, wherein The step of performing normalization processing on the associated parameters includes: Obtain the maximum and minimum values of the historical parameter data in each parameter type, and calculate the extreme values of different parameter types; Calculate the differences between each parameter data in different parameter types and the minimum value of the corresponding parameter type respectively, and calculate the ratios of the differences to the extreme values of the corresponding parameter types; Amplify the obtained ratios by a multiple to obtain the normalized values of the historical parameter data of the corresponding parameter types.
3. The wind power density prediction method according to claim 2, wherein The step of using the Spearman correlation coefficient to perform correlation analysis on the parameter types after normalization processing and the wind turbine power respectively includes: Sort the parameter data of each parameter type after normalization processing and the wind turbine power in ascending order to obtain the parameter data set of the corresponding parameter type and the wind turbine power data set; Subtract the elements of the parameter data set of each parameter type from the elements of the wind turbine power data set respectively to obtain the ranking difference set of the corresponding parameter type; Calculate the correlation coefficients between the corresponding parameter types and the wind turbine power according to the ranking difference sets of different parameter types.
4. The wind power density prediction method according to claim 3, characterized in that, The step of calculating the correlation coefficients between the corresponding parameter types and the wind turbine power according to the ranking difference sets of different parameter types includes: where ρ s is the correlation coefficient between the corresponding parameter type and the power of the wind turbine; d i is the difference value between the i-th parameter data in the corresponding parameter type and the power of the wind turbine; n is the number of parameters in the corresponding parameter type.
5. The wind power density prediction method according to claim 1, wherein The step of using the mean shift clustering algorithm to perform clustering analysis on the historical parameter data of the selected parameter types and screening out different cluster centers includes: Determine any parameter data as the initial cluster center, and set the search radius and convergence threshold; Screen out the parameter data smaller than the search radius, and set it as the data classification, and mark the screened parameter data at the same time; Calculate the drift vector based on the currently screened parameter data, and judge whether the calculated drift vector is less than the convergence threshold; When it is judged that the drift vector is less than the convergence threshold, fix the current cluster center point, and classify the screened parameter data into the current data classification; when it is judged that the drift vector is greater than or equal to the convergence threshold, update the cluster center and perform the next screening iteration; Calculate the distances between different clustering centers and determine whether the distances between different clustering centers are greater than the convergence threshold: When the distance between two clustering centers is greater than the convergence threshold, new data classification is added; when the distance between two clustering centers is less than or equal to the convergence threshold, the two clustering centers are merged into the same data classification, and the marking times are superimposed. When it is judged that all parameter data are marked, each parameter data is respectively assigned to the data classification with the most marking times; when it is judged that there is unmarked parameter data, any unmarked parameter data is determined as the clustering center for iteration.
6. The wind power density prediction method according to claim 5, wherein, The method for calculating the drift vector based on the currently screened parameter data is as follows: Among which M r is the drift vector; x is the clustering center point; S is the range of the current search radius r; x i is the i-th sample point within S; K is the number of sample points within S; y represents the sample points within the range with x as the center and radius r; X represents the sample space, which is the set of all sample points.
7. The wind power density prediction method according to claim 5, wherein When it is judged that the drift vector is greater than or equal to the convergence threshold, the method for updating the clustering center is as follows: x t+1 = x t + M r ; where M r is the drift vector; x t is the cluster center point at the t-th iteration; x t+1 is the cluster center point at the (t + 1)-th iteration.
8. The wind power density prediction method according to claim 1, wherein The steps for calculating the Euclidean distances between the parameter data in the to-be-predicted time period and different clustering centers and screening the optimal clustering center include: Normalize the parameter data in the to-be-predicted time period. Calculate the Euclidean distances between the normalized parameter data in the to-be-predicted time period and different clustering centers respectively: where O j is the Euclidean distance between the parameter data of the to-be-predicted time period and the corresponding cluster center; x(k) is the cluster center point; x j (k) is the j-th parameter data in the prediction time period k, and m is the number of parameter data of the to-be-predicted time period; Select the clustering center with the smallest Euclidean distance as the optimal clustering center.
9. The wind power density prediction method according to claim 8, characterized in that, The steps for training the long short-term memory neural network with the parameter data included in the optimal clustering center to obtain a prediction model include: Perform time sorting on the parameter data and the corresponding wind power data included in the optimal clustering center. Construct an input sequence for the sorted parameter data and the corresponding wind power data in a sliding window manner. Input the input sequence into the long short-term memory neural network for training and verification to obtain a prediction model.
10. A wind power density prediction device based on mean shift clustering, characterized in that, The device includes: A parameter acquisition module, configured to acquire parameter types affecting wind power, and collect historical parameter data and wind power data of corresponding parameters; the parameter types at least include wind speed, wind direction, temperature, humidity, and atmospheric pressure. A parameter preprocessing module, configured to perform normalization processing on the historical parameter data and wind power data of different parameter types respectively. A correlation analysis module, configured to perform correlation analysis between the normalized parameter types and the wind turbine power by using the Spearman correlation coefficient method respectively. A clustering analysis module, configured to perform clustering analysis on the historical parameter data of the screened parameter types by using the mean shift clustering algorithm to obtain different clustering centers. A clustering screening module, which calculates the Euclidean distances between the parameter data in the to-be-predicted time period and different clustering centers and screens the optimal clustering center. A model training module, configured to train the long short-term memory neural network with the parameter data included in the optimal clustering center to obtain a prediction model. A model prediction module, configured to input the parameter data in the to-be-predicted time period into the trained prediction model to obtain a prediction result of the wind power density.