An Outlier Monitoring and Correction Method for Wind Farm Wind Speed-Power Data
Through discretized wind speed-power data, abandoning low-density points, fitting growth curves and correlation analysis methods, the monitoring and correction of wind speed-power data outliers in wind farms is solved, and the accuracy and fitting accuracy of the data are improved.
Patent Information
- Application Number
- CN202211050136.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-08-31
AI Technical Summary
The prior art is difficult to accurately monitor and correct outliers in wind speed-power data in wind farms, affecting the accuracy of wind farm operating conditions assessment.
By discrete the historical wind speed data and power data into multiple intervals, count the frequency distribution, discard the low-density data, use the segmented cubic Hermite interpolation polynomial and growth curve function to fit the wind speed-power curve, and set up margin screening and Pearson correlation analysis method to correct the abnormal data.
It improves the accuracy and fitting accuracy of wind speed-power data, effectively identify and correct outliers, and enhances the accuracy of wind farm operation evaluation.
Smart Images

Figure CN115408860B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wind power generation, and particularly to a method for monitoring and correcting outliers in wind speed-power data of a wind farm. Background Art
[0002] The safe operation and dispatching optimization of a wind farm require accurate and effective wind speed-power data for support. However, due to factors such as environmental impact, sensor failure, curtailment of wind power, and unplanned shutdowns, abnormal data is difficult to avoid in the data obtained through the Supervisory Control and Data Acquisition (SCADA) system. These abnormal data will directly affect the accuracy of the prediction model and have a great impact on the assessment of the operation status of the wind farm.
[0003] Establishing a mathematical model of the wind power curve is an effective method for cleaning the wind speed-power operation data of wind turbines. The wind power curve is the basis for analyzing many wind turbines and describes the relationship between wind speed and the output power of the turbines. It is not only an important basis for designing the control system of wind turbines but also an important indicator for evaluating the power generation performance of wind turbines and the operation status of wind farms.
[0004] Currently, the parametric methods for wind power curve modeling include piecewise linear method, polynomial power curve method, maximum principle method, dynamic power curve method, probability model method, etc. These methods have their own advantages and disadvantages. For example, the maximum principle method has a simple model, but the value at the transition zone of its fitted curve to the rated power overestimates the value at this point, thus affecting the accuracy and the evaluation effect of the operation status of the wind farm. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the technical problem to be solved by the present invention is to provide a method for monitoring and correcting outliers in wind speed-power data of a wind farm, which can improve the accuracy of outlier monitoring and thus improve the accuracy of the data.
[0006] The present invention solves the above technical problem by adopting the following technical solutions:
[0007] A method for monitoring and correcting abnormal data of wind speed-power in a wind farm, the method includes the following:
[0008] The historical wind speed data and power data are discretized and divided into m1 wind speed intervals and m2 power intervals. One wind speed interval corresponds to m2 power intervals, forming a total of m1xm2 areas. The number of wind speed-power data in each area is counted respectively to obtain a wind speed-power joint frequency distribution histogram; the density of wind speed-power data in each area is obtained according to the frequency distribution in each area, and the wind speed-power data in the area with a wind speed-power data density below 0.00001 is discarded; for the retained wind speed-power data, the maximum power value in each area is obtained, and at the same time, the maximum wind speed value in each wind speed interval is obtained. One wind speed maximum value corresponds to the maximum power value of m2 areas. The power maximum values in the area are sorted in the order of the maximum wind speed values to obtain an array P_v of the maximum wind speed value-maximum power value, and an array interval is between adjacent wind speed maximum values;
[0009] The array interval of the array P_v is interpolated by piecewise cubic Hermite interpolation polynomial (PCHIP), the interpolation number is set, and the wind speed-power data is supplemented; the upper boundary of the wind power curve and its function expression are obtained by fitting the supplemented wind speed-power data using the growth curve function. The upper boundary of the wind power curve and its function expression are the wind speed-power curve model obtained according to the maximum principle method;
[0010] Initial screening: In the supplemented wind speed-power data, the data with a wind speed lower than the cut-in wind speed is marked as abnormal. The initial abnormal data is marked as 1 and directly eliminated.
[0011] Margin screening: The remaining wind speed-power data after removing the abnormal data from the initial screening are processed in two stages: the constant torque operation stage and the constant power operation stage. In each stage, the corresponding margin is set to make it conform to the distribution characteristics of the standard wind speed-power curve of the corresponding stage. Then, the function expression of the upper boundary of the wind power curve determined by the growth curve function is used to obtain the corresponding power data in each stage using the known wind speed, and the upper boundary of the wind speed-power value in each stage is obtained respectively. The upper boundary of the wind speed-power value is shifted to the right to determine the lower boundary of the wind speed-power value; finally, the corresponding margin is added to the upper and lower boundaries in each stage. The data within the margin range is normal data, and the data outside the margin range is margin abnormal data. The margin abnormal data is also marked as 1. After the initial margin screening, all the remaining wind speed-power data after removing the abnormal data from the initial screening are divided into normal wind power data and margin abnormal data, and the normal wind power data is marked as 0;
[0012] Abnormal data correction: Perform correlation analysis on the normal wind power data of all wind turbines in the wind farm using Pearson correlation analysis method to obtain the correlation coefficient matrix between the wind turbines; for the wind turbine with data to be supplemented, select the wind turbine with the largest correlation coefficient in the matrix, and then match the abnormal data corresponding to the selected wind turbine data with the wind turbine with data to be supplemented. If the selected wind turbine data is marked as 0, correct and supplement the margin abnormal data according to the normal wind speed - wind power data of the wind turbine with the largest correlation coefficient; if the selected wind turbine data is marked as 1, do not correct and supplement the corresponding data of the wind turbine with data to be supplemented.
[0013] A method for monitoring and correcting abnormal wind speed - power data in a wind farm, the method specifically includes the following steps:
[0014] Step 1: Generate an array of wind speed - power maximum values
[0015] 1 - 1 Arrange the historical wind speed data and power data from small to large, perform discretization processing, divide them into m1 wind speed intervals with an interval size of 1m / s and m2 power intervals with an interval size of 100kw; based on the intervals divided by wind speed, count the frequency of data points in each interval.
[0016] Count the number N = {N1, N2, N3 ·····} of wind speed - power data in each of the m1 x m2 regions respectively to generate a wind speed - power joint frequency distribution histogram.
[0017] 1 - 2 Obtain the density ρ of wind speed - power in the corresponding region according to formula (1).
[0018]
[0019] In formula (1), n is the total number of historical wind speed data, and N is the number of wind speed - power data in each region, that is, the regional frequency.
[0020] 1 - 3 Check the density of wind speed - power in each region. If the density ρ of wind speed - power is less than 0.00001, discard the region; otherwise, retain the region.
[0021] 1 - 4 Find the power maximum value and wind speed maximum value of each region respectively from the data that meet the density requirements. After combining the wind speed maximum values and power maximum values of all regions, obtain an array P_v of wind speed maximum value - power maximum value, and there is an array interval between adjacent wind speed maximum values.
[0022] Step 2: Fit the wind speed - power curve with a growth curve function
[0023] 2-1 Use piecewise cubic Hermite interpolation polynomials to perform interpolation in each array interval of the array P_v to supplement wind speed-power data;
[0024]
[0025]
[0026] Equation (2) is the expression of the cubic Hermite interpolation polynomial, where x0 and x1 are the positions of two adjacent points of the points to be interpolated, that is, the wind speed values at the endpoints of each array interval; y0 and y1 are the dependent variables corresponding to the independent variables x0 and x1, that is, the maximum power corresponding to the endpoint wind speeds; y'0 and y'1 are the corresponding derivatives, and x is the wind speed;
[0027] 2-2 Use the growth curve function to fit the supplemented wind speed-power data to obtain the upper boundary of the wind power curve and its function expression. The upper boundary of the wind power curve and its function expression are shown in formula (3);
[0028]
[0029] Among them, a, b, and K are the parameters of the Pearl model; y represents power, and x represents wind speed;
[0030] Step Three: Abnormal data marking
[0031] 3-1 Preliminary screening: The cut-in wind speed is 3 m / s. In the supplemented wind speed-power data, preliminarily mark the data less than the cut-in wind speed as abnormal. The preliminarily abnormal data is marked as 1 and directly removed;
[0032] 3-2 Margin screening: Divide the remaining wind speed-power data after removing the preliminarily screened abnormal data into two stages for processing: the constant torque operation stage and the constant power operation stage. Set corresponding margins in each stage to conform to the distribution characteristics of the standard wind speed-power curve in the corresponding stage. Then use the function expression of the upper boundary of the wind power curve determined by the growth curve function to obtain the corresponding power data using the known wind speed in each stage, and obtain the upper boundaries of the wind speed-power values in each stage respectively. Translate the upper boundaries of the wind speed-power values to the right to determine the lower boundaries of the wind speed-power values; Finally, add the corresponding margins to the upper and lower boundaries in each stage. The data within the margin range is normal data, and the data outside the margin range is margin abnormal data. Mark the margin abnormal data as 1. After margin preliminary screening, divide all the remaining wind speed-power data after removing the preliminarily screened abnormal data into wind power normal data and margin abnormal data. The wind power normal data is marked as 0;
[0033] Step Four: Abnormal data correction, supplement wind speed data
[0034] Using the normal wind power data, establish the wind speed correlation coefficient matrix of all wind turbines in the wind farm according to the Pearson correlation analysis method. For the wind turbines with missing data, select the wind turbine with the largest correlation coefficient in the matrix, and then match the abnormal data corresponding to the selected wind turbine with the data of the wind turbine with missing data. If the selected wind turbine data is marked as 0, correct and supplement the margin abnormal data according to the normal wind speed-wind power data of the wind turbine with the largest correlation coefficient; if the selected wind turbine data is marked as 1, do not correct and supplement the corresponding data of the wind turbine with missing data.
[0035] Compared with the prior art, the beneficial effects of the present invention are:
[0036] The method of the present invention divides the historical data into several intervals. Before applying the original maximum value principle method, based on the data point density, discard the far-outliers, that is, discard the points with small density, fit the remaining data points, and obtain the wind speed-power fitting curve, which improves the accuracy of the fitted wind speed-power curve and effectively avoids the problems caused by excessive errors due to small data volume.
[0037] The method of the present invention sets a certain margin in both the constant torque operation stage and the constant power operation stage, increasing the applicability, avoiding losing too many normal data points due to the fitting curve, and marking the abnormal data; realizing the identification and monitoring of abnormal values of wind speed-power data in the wind farm.
[0038] In the method of the present invention, the distribution characteristics of the standard wind speed-power curve are similar to those of the growth curve. The present invention creatively uses the growth curve function fitting method for wind speed-power fitting to describe the relationship between wind speed and power, improving the fitting accuracy. At the same time, the Pearson correlation analysis method is used to analyze the correlation of wind speed data of different wind turbines, and the data of abnormal wind turbines are corrected by the normal wind turbine data with high correlation, realizing high-precision data correction. Brief Description of the Drawings
[0039] Figure 1 For the discretization of historical wind speed data and power data, divide them into n wind speed intervals and m power intervals, respectively count the number of wind speed-power data in each area, and obtain the wind speed-power joint frequency distribution histogram.
[0040] Figure 2 After discarding the historical data points with a density lower than 0.00001, it is the upper and lower boundary diagram of the wind speed-power curve with all historical data obtained according to the maximum value principle.
[0041] Figure 3 It is the scatter plot of wind speed-power data obtained by using the cubic polynomial interpolation method in each area.
[0042] Figure 4 It is the upper boundary diagram of the normal wind speed-power curve obtained by fitting with the growth curve function.
[0043] Figure 5 It is the marked wind speed-power scatter plot obtained after setting a certain margin in the constant torque operation stage and the constant power operation stage. The black scatter points are the original data, and the gray scatter points are the normal wind speed-power data.
[0044] Figure 6 It is the wind speed line chart obtained after replacing the data of the abnormal fan at the same moment with the normal fan wind speed data with high correlation. The line is the wind speed data before supplementation, and the circle mark is the supplemented wind speed data. Specific implementation mode
[0045] The technical solution of the present invention will be further specifically described below through embodiments and the accompanying drawings, but this is not used as a limitation to the protection scope of this application.
[0046] The abnormal data in the wind farm wind speed-power abnormal data monitoring and correction method of the present invention can be divided into three categories. The first category is the data with very high wind speed and zero power caused by sensor failure or shutdown; the second category is the data caused by curtailment of wind or failure; the third category is the data with very low wind speed and very high power caused by sensor failure or communication error. Among them, the second type of abnormal data is the most. The third type of abnormal data can be excluded by the initial screening of the cut-in wind speed, and the first type and the second type of abnormal data are screened and marked by setting a margin.
[0047] This embodiment is a method for monitoring and correcting abnormal wind speed-power data in a wind farm, and this method includes the following content:
[0048] Step 1: Generate an array of wind speed-power maximum values
[0049] 1-1 Arrange the historical wind speed data and power data from small to large, perform discretization processing, and divide them into m1 wind speed intervals with an interval size of 1 m / s and m2 power intervals with an interval size of 100 kw; the two intervals are divided separately and have no association with each other; based on the intervals divided by wind speed, count the frequency of data points in each interval.
[0050] Respectively count the number N = {N1, N2, N3 ·····} of wind speed-power data in each of the m1 x m2 regions to generate a wind speed-power joint frequency distribution histogram.
[0051] 1-2 Further obtain the corresponding density distribution of wind speed-power.
[0052]
[0053] In formula (1), n is the total number of historical wind speed data, and N is the number of wind speed-power data in each region, that is, the regional frequency.
[0054] 1-3 Examine the density of wind speed-power in each region. If the density of wind speed-power scatter points is less than 0.00001, discard this point; otherwise, retain this point.
[0055] 1-4 In the data that meets the density requirements, find the maximum power and the maximum wind speed in each region respectively. After combining the maximum wind speeds and maximum powers of all regions, an array P_v of maximum wind speed-maximum power is obtained. The interval between adjacent maximum wind speeds is an array interval.
[0056] Step two: Fit the wind speed-power curve with a growth curve function
[0057] 2-1 Use the piecewise cubic Hermite interpolation polynomial (PCHIP) to perform interpolation in each array interval of the array P_v to supplement the wind speed-power data.
[0058]
[0059] Formula (2) is the expression of the cubic Hermite interpolation polynomial, where x0 and x1 are the positions of two adjacent points to be interpolated, that is, the endpoint wind speed values of each array interval; y0 and y1 are the dependent variables corresponding to the independent variables x0 and x1, that is, the maximum power corresponding to the endpoint wind speed; y'0 and y'1 are the corresponding derivatives, and x is the wind speed; the values of y0 and y1 are different in different regions within the same array interval.
[0060] 2-2 Use the growth curve function to fit the supplemented wind speed-power data to obtain the upper boundary of the wind power curve and its function expression. The upper boundary of the wind power curve and its function expression are shown in formula (3), which is the wind speed-power curve model obtained according to the maximum value principle method.
[0061]
[0062] Formula (3) is the Pearl growth curve model, where a, b, and K are the parameters of the Pearl model; y represents power, and x represents wind speed.
[0063] Step three: Mark abnormal data
[0064] 3-1 Preliminary screening: The cut-in wind speed is 3 m / s. The cut-in wind speeds of different wind farms may be different. Here, 3 m / s is taken as an example. In the supplemented wind speed-power data, preliminarily mark the data less than the cut-in wind speed as abnormal. The preliminarily abnormal data is marked as 1 and directly removed.
[0065] 3-2 Margin Screening: The remaining wind speed-power data after eliminating the abnormally screened data is processed in two stages: constant torque operation stage and constant power operation stage. Appropriate margins are set in each stage to conform to the distribution characteristics of the standard wind speed-power curve in the corresponding stage. Then, the function expression of the upper boundary of the wind power curve determined by the growth curve function is used to obtain the corresponding power data from the known wind speed in each stage, and the upper boundary of the wind speed-power value in each stage is obtained respectively. The upper boundary of the wind speed-power value is translated to the right to determine the lower boundary of the wind speed-power value. Finally, the corresponding margins are added to the upper and lower boundaries in each stage. The data within the margin range is normal data, and the data outside the margin range is margin abnormal data, which is also marked as 1. After margin pre-screening, all the remaining wind speed-power data after eliminating the abnormally screened data is divided into normal wind power data and margin abnormal data, and the normal wind power data is marked as 0;
[0066] The wind speed range in the constant torque operation stage is greater than the cut-in wind speed and less than the rated wind speed; the wind speed range in the constant power operation stage is greater than the rated wind speed; an initial margin a is set for the constant torque operation stage, and an initial margin b is set for the constant power operation stage. After determining the upper boundary of the wind power curve in step 2-2, the parameters of the growth curve function are known values at this time. The known wind speed data in the corresponding stage is directly substituted to obtain the corresponding power data, and the exact upper boundary of the wind speed-power value is obtained, and its lower boundary is translated. At this time, the left endpoint of the horizontal line of the lower boundary is near the sparse point. The corresponding margins are added to the upper and lower boundaries in each stage to obtain new boundaries, and the two new boundaries form the margin range. It is judged whether the power data is within the margin range. If it exceeds this margin range, the power data is further marked as margin abnormal, which is also marked as 1.
[0067] The setting of the margin value finally needs to conform to the distribution characteristics of the standard wind speed-power curve. In actual use, the initial margin value can be set and adjusted later to meet this standard. If the overall shape of the scatter plot of the obtained normal data does not conform to the distribution characteristics of the standard wind speed-power curve, the sizes of the margin values a and b need to be readjusted until they meet the standard, and then the final margin value is determined to obtain a scatter plot of normal wind speed-power data with a certain margin.
[0068] Step 4: Abnormal data correction and wind speed data supplementation
[0069] 4-1 Establish a wind speed correlation coefficient matrix for all the wind turbines in the wind farm according to the Pearson correlation analysis method. In this embodiment, taking a wind farm with 50 wind turbines as an example, the normal wind turbine data with a high correlation with the missing data is found. The Pearson correlation coefficient, also known as the Pearson product-moment correlation coefficient, is used to measure the linear correlation between two sets of data X and Y, and its value ranges from -1 to 1.
[0070] Its calculation formula is
[0071]
[0072] where: cov(X,Y) is the covariance of X and Y; σ X and σ Y are the standard deviations of X and Y respectively; X and Y are the wind speed values of two different wind turbines. The closer the value of ρ X,Y is to 1, the stronger the correlation between the two sets of data.
[0073] 4-2 Sort the correlation coefficients of the wind turbine to be supplemented with data and other wind turbines from large to small, and use the normal wind speed-power data with the highest correlation coefficient to replace the abnormal wind speed-power data of the wind turbine to be supplemented with data at this moment to obtain the corrected values v eq and P eq .
[0074] What is not described in this invention applies to the prior art.
Claims
1. A method for monitoring and correcting abnormal wind speed-power data in a wind farm, which includes the following steps: discretize historical wind speed data and power data, divide them into m1 wind speed intervals and m2 power intervals. One wind speed interval corresponds to m2 power intervals, forming a total of m1 x m2 regions. Count the number of wind speed-power data in each region respectively to obtain a wind speed-power joint frequency distribution histogram; obtain the density of wind speed-power data in each region according to the frequency distribution in each region, and discard the wind speed-power data in the regions where the density of wind speed-power data is below 0.00001; for the remaining wind speed-power data, find the maximum power within each region, and at the same time find the maximum wind speed within each wind speed interval. One maximum wind speed corresponds to the maximum power of m2 regions. Sort the maximum power within the regions according to the magnitude order of the maximum wind speeds to obtain an array P_v of maximum wind speed-maximum power. The interval between adjacent maximum wind speeds is an array interval. Interpolate the array intervals of the array P_v using piecewise cubic Hermite interpolation polynomials, set the number of interpolations, and supplement wind speed-power data; use the growth curve function to fit the supplemented wind speed-power data to obtain the upper boundary of the wind power curve and its function expression. Initial screening: In the supplemented wind speed-power data, preliminarily mark the data less than the cut-in wind speed as abnormal. The preliminarily abnormal data is marked as 1 and directly removed. Margin screening: Divide the remaining wind speed-power data after removing the initially screened abnormal data into two stages for processing: the constant torque operation stage and the constant power operation stage. Set corresponding margins in each stage to conform to the distribution characteristics of the standard wind speed-power curve in the corresponding stage. Then use the function expression of the upper boundary of the wind power curve determined by the growth curve function to obtain the corresponding power data using the known wind speed in each stage, and obtain the upper boundary of the wind speed-power value in each stage respectively. Translate the upper boundary of the wind speed-power value to the right to determine the lower boundary of the wind speed-power value; finally, add the corresponding margins to the upper and lower boundaries in each stage. The data within the margin range is normal data, and the data outside the margin range is margin abnormal data. Mark the margin abnormal data as 1. After margin initial screening, all the remaining wind speed-power data after removing the initially screened abnormal data are divided into wind power normal data and margin abnormal data. The wind power normal data is marked as 0. Abnormal data correction: Use the Pearson correlation analysis method to perform correlation analysis on the wind power normal data of all the wind turbines in the wind farm to obtain the correlation coefficient matrix between the wind turbines. For the wind turbine to be supplemented with data, select the wind turbine with the largest correlation coefficient in the matrix, and then match the corresponding abnormal data of the selected wind turbine with the wind turbine to be supplemented with data. If the selected wind turbine data is marked as 0, then correct and supplement the margin abnormal data according to the wind speed-wind power normal data of the wind turbine with the largest correlation coefficient; if the selected wind turbine data is marked as 1, then do not correct and supplement the corresponding data of the wind turbine to be supplemented with data.
2. The wind farm wind speed-power abnormal data monitoring and correction method according to claim 1, characterized in that The abnormal data is divided into three categories. The first category is the data with a very high wind speed and zero power caused by sensor failure or shutdown. The second category is the data caused by curtailment of wind power or failure. The third category is the data with a very low wind speed and high power caused by sensor failure or communication error.
3. The wind farm wind speed-power abnormal data monitoring and correction method according to claim 1, characterized in that, The wind speed range in the constant torque operation stage is greater than the cut-in wind speed and less than the rated wind speed. The wind speed range in the constant power operation stage is greater than the rated wind speed. Set the initial margin a in the constant torque operation stage and the initial margin b in the constant power operation stage. After determining the upper boundary of the wind power curve, the parameters of the growth curve function are known values at this time. Directly substitute the known wind speed data in the corresponding stage to obtain the corresponding power data, and obtain the upper boundary of the exact wind speed-power value. Translate to determine its lower boundary. Add the remaining margin to the upper and lower boundaries of each stage to obtain new boundaries, and the two new boundaries form the margin range. Judge whether the power data is within the margin range. If it exceeds this margin range, further mark the power data with margin abnormality, and also mark it as 1.
4. The wind farm wind speed-power abnormal data monitoring and correction method according to claim 3, characterized in that The setting of the margin value conforms to the distribution characteristics of the standard wind speed-power curve. In actual use, the initial margin value is set and adjusted later to meet this standard. If the overall shape of the obtained normal data scatter plot does not conform to the distribution characteristics of the standard wind speed-power curve, the sizes of the margin values a and b need to be readjusted until they meet the standard, and then the final margin value is determined to obtain a scatter plot of normal wind speed-power data with a certain margin.
Citation Information
Patent Citations
Combined wind power plant data cleaning method
CN111291032A
Power curve fitting data preprocessing method and device based on density distribution
CN111460360A