A grid-connected wind farm operation performance evaluation method and system based on correlation analysis
By constructing a main monitoring area model in a grid-connected wind farm, performing data preprocessing and correlation analysis, and establishing a regression model, the problem of insufficient power source assessment in existing technologies is solved, and a comprehensive quantitative assessment and accuracy improvement of wind farm operation performance are achieved.
Patent Information
- Application Number
- CN202411787951.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Existing methods for evaluating the operational performance of grid-connected wind farms mainly focus on grid-side indicators, lacking an assessment of the operational characteristics of wind turbine generators on the power supply side. Furthermore, existing methods have limitations in setting the weights of influencing factors, making it impossible to quantitatively describe the relevant relationships.
By establishing a main monitoring area model, historical sampling data of wind farms and wind turbines are obtained. After filtering, smoothing and standardization, correlation analysis is performed, regression models are constructed, quantitative relationships between influencing factors and performance indicators are explored, and multiple linear regression is used to solve the wind farm operation performance.
It enables a comprehensive quantitative assessment of wind farm operation performance, improves the assessment index system, provides accurate performance assessment basis, can identify abnormal operating conditions, and enhances the accuracy and comprehensiveness of the assessment.
Smart Images

Figure CN119651588B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power systems, specifically relating to a method and system for evaluating the operational performance of grid-connected wind farms based on correlation analysis. Background Technology
[0002] With the rapid development of wind power technology and the gradual increase in wind power penetration, the impact of grid-connected wind power on the stability of the power system is also increasing. The operational performance and production efficiency of wind turbine units have become key factors affecting the economic benefits of wind farms. However, due to the complex and variable operating environment of wind turbine units and the mutual influence of internal parameters, the volatility and randomness of wind power generation bring a series of problems that urgently need to be addressed after grid connection. Focusing on the operational performance evaluation of wind turbine units, fully exploring the potential data correlations between wind energy resource data, turbine power generation data, and grid-connected operation data, and conducting a more comprehensive analysis and evaluation of the operational performance of wind farms, establishing an indicator evaluation system covering the power generation characteristics of the entire wind power production process, is an effective way to efficiently utilize wind power resources and can also provide strong technical support for the grid to absorb wind power resources on a larger scale.
[0003] Reference 1, "Construction and Application of Evaluation Index System for Large-Scale Wind Power Grid Connection" (Power System Technology, Vol. 45, No. 3, 2021, p. 841), establishes an evaluation index system based on five rates—wind power installed capacity penetration rate, electricity penetration rate, power generation penetration rate, utilization rate, and curtailment rate—from the perspective of grid connection and consumption of large-scale wind turbine units. It evaluates the impact of different utilization rate indicators and constraints on wind power installed capacity penetration rate and illustrates the characteristics of wind power generation and the application method of the five-rate index system using actual grid operation data. This has certain guiding significance for wind power planning and improving the lean operation level of wind farms.
[0004] Reference 1 focuses its research on the grid-connected operation of wind turbines. The classification of indicators in the performance evaluation process is clearer. However, the formulation of the "five rates" indicators is only focused on the grid-side link of wind turbine grid-connected operation. It cannot evaluate the operating characteristics of the "power supply side" of wind turbines or the utilization of wind energy resources by the units. It can only supplement and improve the "grid-side" indicators of wind turbines participating in system operation.
[0005] Reference 2, "Evaluation Index System of Statistical Characteristics of Wind Power at Multiple Spatiotemporal Scales and Its Application" (Proceedings of the CSEE, Vol. 33, No. 13, 2013, p. 53), divides the performance evaluation indexes of wind power generation into two main categories: those reflecting the natural characteristics of wind power based on wind resources and the characteristics of wind power itself, and those reflecting the interaction characteristics of the power grid that require the integration of other data such as the power grid and load. This, to a certain extent, enriches the volume of performance evaluation indexes for wind farm operation and provides more statistical index references for research.
[0006] Reference 2, starting from the characteristics of wind resources / wind power itself and the physical interaction between wind power, the power grid, and loads, seeks expanded indicators for performance evaluation in the wind resource absorption and grid integration stages, thus improving the evaluation indicator resource library to some extent. However, this reference's evaluation stage focuses on the physical meaning of performance indicators, making it difficult to explore operational indicators whose correlations are currently uncertain and lack qualitative analysis or quantitative description, thus failing to conduct a more in-depth and extensive exploration of wind power.
[0007] Reference 3, "Application of Analytic Hierarchy Process in Economic Evaluation of Wind Farm Operation" (China Electric Power, Vol. 39, No. 9, p. 42, 2006), uses the analytic hierarchy process to classify the factors affecting the economic operation of wind farms. From the perspective of wind farm operation and management, it evaluates the economic aspects of wind farm operation, such as power generation, operating costs, and personnel efficiency, through pairwise comparisons. Then, combined with the work experience of the evaluators, it determines the overall ranking of the relative importance of each influencing factor.
[0008] Reference 3 attempts to evaluate the economic efficiency of wind farm operations from the perspective of assessment methodology by classifying and interpreting sampled data to comprehensively evaluate multiple indicators. However, when there are many uncertainties and ambiguities, directly assigning quantitative weight coefficients to each element in the judgment matrix based on the scorer's existing experience is not easy to verify in terms of accuracy, and the evaluation results lack rigorous persuasiveness.
[0009] In summary, existing methods for evaluating the operational performance of grid-connected wind farms have the following shortcomings:
[0010] At the level of evaluation indicators, existing methods for evaluating the operational performance of grid-connected wind farms mainly focus on calculating "grid-side" indicators related to the grid connection and participation of wind turbines in system operation, with less consideration given to the evaluation of the operational characteristics of the "power-side" wind turbines themselves. At the level of evaluation methods, the method of manually assigning weight coefficients to influencing factors has certain limitations, and there are currently no clear evaluation methods for indicators whose correlations cannot be determined or which lack quantitative descriptions. Summary of the Invention
[0011] Purpose of the invention: The purpose of this invention is to provide a method and system for evaluating the operational performance of grid-connected wind farms based on correlation analysis. By conducting data mining and comprehensive evaluation of the factors influencing the performance of wind turbines during operation, the quantitative relationship between performance evaluation indicators and various influencing factors can be found, thus overcoming the problem of limited indicator coverage in existing evaluation standards.
[0012] Technical solution: The present invention provides a method for evaluating the operational performance of grid-connected wind farms based on correlation analysis, comprising:
[0013] Data acquisition: A main monitoring area model is established in the wind farm operation performance evaluation system, and secondary models of wind farm, wind turbine, wind turbine rotor, transmission device and converter are derived from it; the information required for performance evaluation is obtained from the wind farm SCADA system, power metering system and meteorological station monitoring system to form historical sampling data based on time series.
[0014] Data preprocessing: The historical sampling data based on time series is filtered, smoothed and standardized to obtain the initial test dataset for wind farm grid connection performance evaluation;
[0015] Data Correlation Analysis: Correlation analysis is performed on the initial experimental dataset. Based on the performance evaluation requirements, corresponding feature columns are selected from the initial experimental dataset to form a new experimental dataset. This dataset forms the basis for constructing the input feature vector for wind farm operation performance evaluation. The correlation between the performance evaluation indicators to be analyzed and the new experimental dataset, as well as the input feature vectors for wind farm operation performance evaluation within the new experimental dataset, is explored. Based on the obtained correlation strength results, the sequence of influencing factors as independent variables and the performance evaluation indicators as dependent variables are selected to construct a regression model. Regression analysis is performed based on the regression model to obtain the quantitative descriptive relationship between the feature vectors, and to obtain the revised algorithm for the evaluation indicators or new quantitative evaluation indicators. Among them, the independent variable influencing factors are factors that affect the operation performance of grid-connected wind farms, specifically including wind farm ledger data, grid-connected operation data, and wind energy resource data.
[0016] Performance evaluation: Substitute the actual operating data of the wind farm into the wind farm operation performance evaluation model composed of the new quantitative evaluation index and the revised quantitative evaluation index, calculate the values of various performance indicators of the wind farm, and obtain the overall performance evaluation results of the wind farm.
[0017] Furthermore, correlation analysis is performed on the initial experimental dataset. Based on the performance evaluation requirements, corresponding feature columns are selected from the initial experimental dataset to form a new experimental dataset. This dataset forms the basis for constructing the input feature vector for wind farm operation performance evaluation. The correlation between the performance evaluation indicators to be analyzed and the new experimental dataset, as well as the input feature vectors for each wind farm operation performance evaluation within the new experimental dataset, is explored. Based on the obtained correlation strength results, the sequence of influencing factors as independent variables and the performance evaluation indicators as dependent variables are selected to construct a regression model. Regression analysis is performed based on the regression model to obtain the quantitative descriptive relationship between the feature vectors, leading to the acquisition of a revised algorithm for the evaluation indicators or new quantitative evaluation indicators, including:
[0018] Normality check: Perform normality check on the experimental dataset to verify the fit of the sampled data in the experimental dataset with a normal distribution. After normality check, the experimental dataset that meets the conditions for correlation analysis is obtained.
[0019] Correlation coefficient calculation: Based on the normality test results, determine the calculation method for the correlation coefficient between each influencing factor in the experimental dataset; whereby the correlation coefficient is used to determine the correlation strength between different feature columns in the experimental dataset;
[0020] Remove abnormal result sets: Use correlation coefficients to perform correlation analysis on all feature column data in the experimental dataset. Based on the correlation analysis results, select feature column data with correlation coefficients greater than a preset threshold from the experimental dataset to form a new experimental dataset. Construct the input feature vector for wind farm operation performance evaluation based on this dataset.
[0021] Performance evaluation modeling: Explore the correlation between the performance evaluation index to be analyzed and the new experimental dataset, as well as the input feature vector of the wind farm operation performance evaluation in the new experimental dataset. Based on the obtained correlation strength results, select the sequence of influencing factors as independent variables and the performance evaluation index as dependent variables, and construct a regression model.
[0022] Solving the regression equation: Considering the impact of wind turbine parameters, grid-connected operation data, and wind energy environmental factors on the operating performance of wind turbines, based on the regression model, multiple linear regression equations are established for each new experimental dataset; by solving the multiple linear regression equations, the correlation between various influencing factors is transformed into quantitative evaluation indicators for the operating performance of wind turbines and wind farms.
[0023] Overall verification: Perform significance verification on the calculation method of the correlation coefficient and the calculation results of the multiple linear regression equation.
[0024] Furthermore, the normality of the experimental dataset is verified to check the approximation of the normality distribution of the sampled data. The experimental dataset that meets the conditions for correlation analysis after normality verification includes:
[0025] Calculate the percentiles of the sampled data in the experimental dataset. The percentiles of the sampled data in the experimental dataset are combined with the percentiles of the normal distribution to form the percentile comparison space. Based on the distribution of the sampled data points in the experimental dataset, we can make a preliminary judgment on whether the sampled data follows a normal distribution.
[0026] Check the goodness-of-fit deviation between the sampled data in the experimental dataset and the standard distribution, perform normality verification on the historical SCADA sampled data in the experimental dataset, and use the KS hypothesis testing method to perform nonparametric tests on the sampled data in the experimental dataset to check whether the hypothesis about the population normal distribution is valid.
[0027] Furthermore, the goodness-of-fit deviation of the sampled data in the experimental dataset from the standard distribution was examined, and the normality of the historical SCADA sampled data in the experimental dataset was verified. The KS hypothesis testing method was used to perform nonparametric tests on the sampled data in the experimental dataset to check whether the hypothesis about the population normality holds, including:
[0028] After sorting the SCADA historical sampling data in the experimental dataset in ascending order, the percentile markers of the SCADA historical sampling data are determined. The mean and standard deviation of the percentile comparison space are calculated. Based on the deviation of the measurement points of the SCADA historical sampling data in the experimental dataset from the standard deviation of the experimental dataset, the distribution function of the analysis samples of the experimental dataset is constructed.
[0029]
[0030] In the formula, σ is the standard deviation of the historical sampling data; This represents the average value of the experimental dataset.
[0031] Compare the deviations of each percentile marker interval point from their corresponding theoretical normal distribution cumulative function values, and compare the maximum absolute value of each deviation with the KS critical value. If the maximum difference between the cumulative distribution function of the experimental dataset and the cumulative distribution function of the normal distribution is less than the critical value, the null hypothesis is considered to be true; otherwise, the null hypothesis is rejected.
[0032] The maximum difference between the cumulative distribution function of the experimental dataset and the cumulative distribution function of the theoretical normal distribution is the KS statistic.
[0033] A significance level is set to determine whether the maximum difference between the cumulative distribution function of the observed experimental dataset and the cumulative distribution function of the theoretical normal distribution belongs to random error;
[0034] The calculated KS statistic is compared with the KS critical value. If the statistic KS is less than the KS critical value, the null hypothesis is considered to be true, and the sampled data in the experimental dataset basically follows a normal distribution.
[0035] Furthermore, based on the normality test results, the method for calculating the correlation coefficients among the influencing factors in the experimental dataset is determined, including:
[0036] Experimental Dataset Normalization: A state data matrix is constructed based on the experimental data set of potential influencing factors of the selected grid-connected wind farm operation performance evaluation indicators within a given time window. When there are significant differences between the sampled data points in the state data matrix, a normalization method is used to process the sampled data in the state data matrix, specifically as follows:
[0037]
[0038] In the formula, x represents the standardized value in the correlation analysis data matrix; ij μ represents the stored value of the j-th potential influencing factor at the i-th sampling time. j σ is the sampled average of the j-th potential influencing factor within a given time window; j For x ij The standard deviation of the j-th potential influencing factor; the data in each column of the normalized state data matrix approximately follow a normal distribution with a mean of 0 and a variance of 1;
[0039] Given a selected experimental dataset that approximates a normal distribution, correlation analysis is performed on any two waveforms at a given time point. The degree of linear correlation is represented by the Pearson correlation coefficient, which is expressed in discretized form as follows:
[0040]
[0041] In the formula, r is the correlation coefficient between the two variable sequences in a single analysis; cov(X,Y) and σ X σ Y The covariances of the two experimental datasets are respectively; x i y i These are the values of the i-th data point in a single analysis sampling data; These are the average values of the two columns of sampled data, respectively; n is the total amount of data contained in the sampled data window;
[0042] When the Pearson correlation coefficient between two sets of feature columns in the experimental dataset is greater than the threshold value, it indicates that there is a significant correlation between the two sets of feature columns in the experimental dataset.
[0043] Furthermore, based on the normality test results, the method for calculating the correlation coefficients among the influencing factors in the experimental dataset is determined, including:
[0044] If the data to be analyzed does not follow a normal distribution, then rank division is needed to verify correlation. When a data point appears multiple times, its rank is the average of the two preceding and following ranks. The rank is calculated to restore the original value after data changes. The linear correlation between two waveforms is represented by the Spearman coefficient, whose discretization formula can be expressed as:
[0045]
[0046] In the formula, ρ is the Spearman correlation coefficient between the two sets of sampled datasets; r i s i Indicates the column rank of two columns of data; d is the average of the ranks of the two columns of data;i The rank difference of the i-th pair of data is represented by n; n is the total amount of data sampled in this analysis.
[0047] The larger the absolute value of the Spearman correlation coefficient, the stronger the correlation between the two types of data. The sign of the correlation coefficient determines whether the two types of data are positive or negative.
[0048] Furthermore, a significance check was performed on the method of calculating the correlation coefficient, including:
[0049] If the feature columns of the experimental dataset used to calculate the correlation coefficient follow a normal distribution, then a new statistic t is constructed:
[0050]
[0051] In the formula, r represents the correlation coefficient of the calculated feature data column; n is the number of samples selected in the experimental dataset; and the statistic t follows a t-distribution with (n-2) degrees of freedom.
[0052] When the probability p value corresponding to the statistic t is less than the selected significance level, it is considered that there is a significant linear correlation in the feature columns of the selected experimental dataset.
[0053] Furthermore, considering the impact of wind turbine parameters, wind turbine operating data, and wind energy environmental factors on the operating performance of wind turbines and wind farms, multiple linear regression equations are established for each of the new experimental datasets. By solving the multiple linear regression equations, the correlations between the various influencing factors are transformed into quantitative evaluation indicators for the operating performance of wind turbines and wind farms, including:
[0054] Establish a multiple linear regression equation:
[0055] Y = β0 + β1X1 + β2X2 + ... + β n X n +ε
[0056] In the formula, β0 is a constant term; β i (i = 1, 2, 3, ..., n) represents the other independent variables X. i The average change in dependent variable Y for every unit change in a specified independent variable X when X remains constant; ε represents the calculation of residuals.
[0057] In solving the multiple linear regression equation, its normalized equation system can be expressed as:
[0058] (X T X)B=X T Y
[0059] In the formula, Y is the historical sampling data sequence of the dependent variable; B represents the historical sampling data sequence of the independent variable; and X is the coefficient matrix of the normalized equation system in the regression model.
[0060] Furthermore, the calculation results of the multiple linear regression equation are subjected to significance verification, including:
[0061] When all elements in the independent variable data sequence B are 0, the original linear correlation hypothesis is rejected, and the quantitative relationship is re-determined using first-order linearity or other methods; the quantile plot is used to analyze the calculated residuals to check whether they approximately follow a normal distribution;
[0062] When not all elements in the independent variable data sequence B are zero, perform an overall significance check on the multiple linear regression equation:
[0063]
[0064] In the formula, SS R SS represents the sum of squared deviations of a multiple linear regression equation; E This represents the sum of squared residuals from the calculation results of the multiple linear regression equation; y represents the i-th component of the predicted value of the regression equation; i These are the actual values used for regression analysis; Represents the average of actual values;
[0065] If the p-value of the multiple linear regression equation is less than the significance level, then the multiple linear regression equation is significant; a significance test is performed on the regression coefficients; if the p-value is less than the significance level, then the regression coefficients are significant; where:
[0066]
[0067] df R df represents the degrees of freedom in the regression equation, which is the number of independent variables in the multiple linear regression equation. E The residual degrees of freedom are represented by the following formula:
[0068] df E =df T -df R
[0069] df T =n-1
[0070] Among them, df T df represents the total degrees of freedom of the multiple linear regression equation. R This represents the regression degrees of freedom of the regression equation; n is the number of dimensions of the feature column vector.
[0071] The value of P can be calculated using the following formula:
[0072] P = 1 - F X (x)
[0073] Among them, F X (x) represents the cumulative distribution function value, whose characteristics are determined by F and df. R df E A joint decision.
[0074] Based on the same inventive concept, the present invention provides a grid-connected wind farm operation performance evaluation system based on correlation analysis, comprising:
[0075] The data acquisition module is used to establish a main monitoring area model in the wind farm operation performance evaluation system, and derive secondary models of wind farm, wind turbine, wind turbine rotor, transmission device, and converter from this model; it also obtains the information required for performance evaluation from the wind farm SCADA system, power metering system, and meteorological station monitoring system to form historical sampling data based on time series.
[0076] The data preprocessing module is used to filter, smooth, and standardize the historical sampling data based on time series to obtain the initial test dataset for wind farm grid connection performance evaluation.
[0077] The data correlation analysis module is used to perform correlation analysis on the initial experimental dataset. Based on the performance evaluation requirements, corresponding feature columns are selected from the initial experimental dataset to form a new experimental dataset. This dataset forms the basis for constructing the input feature vector for wind farm operation performance evaluation. The module then mines the correlation between the performance evaluation indicators to be analyzed and the new experimental dataset, as well as the input feature vectors for each wind farm operation performance evaluation within the new experimental dataset. Based on the obtained correlation strength results, the module selects the sequence of influencing factors as independent variables and the performance evaluation indicators as dependent variables, constructing a regression model. Regression analysis is performed based on the regression model to obtain the quantitative descriptive relationship between the feature vectors, leading to a revised algorithm for the evaluation indicators or new quantitative evaluation indicators. The independent variable influencing factors are those factors that affect the operation performance of grid-connected wind farms, specifically including wind farm ledger data, grid-connected operation data, and wind energy resource data.
[0078] The performance evaluation module is used to input the actual operating data of the wind farm into the wind farm operation performance evaluation model composed of new quantitative evaluation indicators and revised quantitative evaluation indicators, calculate the values of various performance indicators of the wind farm, and obtain the overall performance evaluation results of the wind farm.
[0079] Beneficial effects: Compared with the prior art, the significant technical effects of the present invention are as follows:
[0080] Correlation analysis of the sampled data can identify potential individual factors affecting the evaluation indicators. By using multiple linear regression to solve the performance evaluation model, the quantitative relationship between influencing factors and evaluation indicators can be found, providing technical support for improving the wind farm operation performance evaluation indicator system.
[0081] Based on the actual needs of performance evaluation, this study selects factors with strong correlation and potential impact on evaluation indicators according to the results of correlation analysis to construct a performance evaluation model. This provides a basis for studying and explaining unknown areas of performance evaluation that are difficult to explain and cannot be quantified at present.
[0082] For performance evaluation indicators with multiple calculation methods, the accuracy of the calculation method can be verified by the calculation results of regression analysis. On this basis, the coefficients of influencing factors can also be corrected, providing a reference for more accurately describing the relationship between variables.
[0083] By utilizing the solved performance evaluation model and test dataset, a quantitative basis can be provided for determining whether a wind turbine is in an abnormal operating or faulty state, and a solution can be provided for the determination and classification of the health status of wind turbines.
[0084] The overall framework of the performance index evaluation system has been improved. The evaluation scope is no longer limited to the classification of indicators. Any physical quantity with observable values can be used as a performance index for evaluation. Reasonable index calculation models can be incorporated into the new index evaluation system. Attached Figure Description
[0085] Figure 1 This is a flowchart illustrating a method for evaluating the operational performance of a grid-connected wind farm based on correlation analysis, as disclosed in an embodiment of the present invention.
[0086] Figure 2 This is a matrix diagram illustrating the correlation analysis results of performance index influencing factors disclosed in an embodiment of the present invention;
[0087] Figure 3 This is a schematic diagram of the structure of a grid-connected wind farm operation performance evaluation system based on correlation analysis disclosed in an embodiment of the present invention; Detailed Implementation
[0088] The technical solution of the present invention will now be described in detail with reference to specific embodiments and accompanying drawings.
[0089] Example 1
[0090] like Figure 1 As shown, the present invention provides a method for evaluating the operational performance of a grid-connected wind farm based on correlation analysis, comprising the following steps:
[0091] S1. Data Acquisition: Establish a main monitoring area model in the wind farm operation performance evaluation system, and derive secondary models such as wind farm, wind turbine, wind turbine rotor, transmission device, and converter from this model; obtain the information required for performance evaluation from data sources such as wind farm SCADA system, power metering system, and meteorological station monitoring system to form historical sampling data based on time series.
[0092] In this embodiment, data such as longitude, latitude, humidity, temperature, air pressure, wind speed, wind direction, and air density are obtained from the meteorological station monitoring system;
[0093] Obtain operational data from the wind farm's SCADA system, including voltage, current, active power, reactive power, grid connection status (generation, maintenance, power curtailment, shutdown), real-time load, and planned power generation; as well as wind turbine operating parameters such as rated capacity, prototype turbine markings, and cut-in / cut-out wind speeds.
[0094] Data on the pitch system and converter, such as pitch angle, blade speed, cabinet temperature, and motor speed, are obtained from the wind farm's SCADA system.
[0095] Data on the transmission device, such as gearbox inlet oil temperature, outlet pressure, nacelle temperature, and cooling water temperature, are obtained from the wind farm's SCADA system.
[0096] Electricity data such as grid-connected electricity and on-site electricity consumption are obtained from the electricity metering system.
[0097] A main monitoring area and weather station model are established in the wind farm SCADA system to collect global data from the wind turbine side, the wind farm grid side, and wind energy resources. Based on this, secondary models such as the wind farm, turbine drive system, and electrical control system are established to improve the wind turbine ledger parameters and receive real-time data.
[0098] S2. Data preprocessing: Filtering, smoothing, and standardization are performed on the historical sampling data based on time series to obtain the initial test dataset for wind farm grid connection performance evaluation.
[0099] In this embodiment, a first-order low-pass filter is used to filter the high-frequency random components of the measurements acquired in the wind power SCADA system. These random components are typically noise generated by fluctuations in the industrial environment and the operating conditions of the acquisition system. The filtering algorithm can be specifically expressed as follows:
[0100] X FIL (k+1)=X FIL (k)+(X RAW (k+1)-X FIL (k))*DT F (1)
[0101] In the formula, k represents the counter for the number of times the filtering algorithm is executed; X FIL Indicates the filtered value; X RAW Represents the raw data to be filtered; DT F This represents the discrete filtering factor, with a value range between 0 and 1.
[0102] The filtering algorithm, based on SCADA data, can reduce interference from invalid data when combined with measurement point and measurement validity verification. Generally, the following attributes can be used for judgment: a) SCADA measurement has a defective quality indicator; b) Measurement exceeds the normal specified range; c) Measurement does not change within a specified time; d) Measurement data shows abrupt changes within two consecutive indicator evaluation periods.
[0103] Preprocessing operations such as filtering and smoothing are performed on historical sampling data based on time series to reduce the impact of noise and ensure the accuracy and comparability of the data.
[0104] Taking the operating conditions of wind turbine units as a reference, data groups were established for several types, including stable operation of wind turbine units, operation condition switching of wind turbine units, and operation of wind turbine units with faults. Performance evaluation test datasets were obtained as sample data for subsequent analysis.
[0105] S3. Data Correlation Analysis: Correlation analysis is performed on the initial experimental dataset. Based on the performance evaluation requirements, corresponding feature columns are selected from the initial experimental dataset to form a new experimental dataset. This dataset forms the basis for constructing the input feature vector for wind farm operation performance evaluation. The correlation between the performance evaluation index to be analyzed and the new experimental dataset, as well as the input feature vectors for wind farm operation performance evaluation within the new experimental dataset, is explored. Based on the obtained correlation strength results, the sequence of influencing factors as independent variables and the performance evaluation index as the dependent variable are selected to construct a regression model. Regression analysis is performed based on the regression model to obtain the quantitative descriptive relationship between the feature vectors, and to obtain the correction algorithm for the evaluation index or a new quantitative evaluation index. Among them, the independent variable influencing factors are factors that affect the operation performance of grid-connected wind farms, specifically including wind farm ledger data, grid-connected operation data, and wind energy resource data.
[0106] The preset performance evaluation indicators for grid-connected wind farms include basic indicators, resource indicators, power generation indicators, energy consumption indicators, operation indicators, maintenance indicators, and energy storage indicators.
[0107] In this context, the influencing factor is the independent variable x, the evaluation index is the dependent variable y, and the result of the correlation analysis is to form a functional expression y = f(x) for the independent and dependent variables by combining factors with strong correlation.
[0108] The specific process of step S3 is as follows:
[0109] S3.1. Data Normality Verification: The experimental dataset is normalized to verify the approximation of a normal distribution in the sampled data. After normality verification, an experimental dataset that meets the conditions for correlation analysis is obtained. This includes the following steps:
[0110] S3.1.1 Calculate the percentiles of the sampled data in the experimental dataset. The percentiles of the sampled data in the experimental dataset are combined with the percentiles of the normal distribution to form the percentile comparison space. Based on the distribution of the sampled data points in the experimental dataset, we can make a preliminary judgment on whether the sampled data follows a normal distribution.
[0111] For cases where the sample size of historical SCADA data in the experimental dataset is small, or where the wind farm is operating under steady-state conditions, the following three methods can be used to verify normality:
[0112] 1) Use the four quantiles and standard deviation for the first round of rough judgment. When the ratio of the two is greater than the given threshold, it is approximately considered that the experimental dataset satisfies a normal distribution.
[0113] 2) Calculate the cumulative probability and percentiles of the experimental data, and use histograms to compare the calculated values with the theoretical values of the standard distribution.
[0114] 3) Calculate the kurtosis and skewness values of the experimental data, and conduct a combined quantitative examination of the symmetry and deviation of the experimental data:
[0115]
[0116] In the formula, Z-Score skew Z-Score curtosis Z-scores representing skewness and kurtosis of the experimental data, respectively; V skew V curtosis These represent the skewness and kurtosis values of the experimental data, respectively; σ skew σ curtosis These represent the standard deviations of the skewness and kurtosis values in this data set, respectively.
[0117] S3.1.2 Check the goodness-of-fit deviation between the sampled data in the experimental dataset and the standard distribution, perform normality verification on the historical SCADA sampled data in the experimental dataset, and perform nonparametric tests on the sampled data in the experimental dataset using the KS hypothesis testing method to check whether the hypothesis about the population normal distribution is valid.
[0118] Because the sampling data covers a relatively long period of time, only a portion of it will be analyzed.
[0119] The specific process of step S3.1.2 is as follows:
[0120] S3.1.2.1 After arranging the SCADA historical sampling data in the experimental dataset in ascending order, the percentile markers of the SCADA historical sampling data are determined. The mean and standard deviation of the percentile comparison space are calculated. Based on the deviation between the measurement points of the SCADA historical sampling data in the experimental dataset and the standard deviation of the experimental dataset, the distribution function of the analysis samples of the experimental dataset is constructed.
[0121]
[0122] In the formula, σ is the standard deviation of the historical sampling data; The average value of the test dataset
[0123] S3.1.2.2 Compare the deviations of each percentile marker interval point from their corresponding theoretical normal distribution cumulative function values, and compare the maximum absolute value of each deviation with the KS critical value. If the maximum difference between the cumulative distribution function of the experimental dataset and the cumulative distribution of the normal distribution is less than the critical value, the null hypothesis is considered to be true; otherwise, the null hypothesis is rejected.
[0124] The cumulative distribution function of the corresponding theoretical normal distribution can be obtained from the mean and standard deviation of the characteristic data column.
[0125] It's important to note that the absolute value here refers to the absolute value of the deviation between each data point in the feature data column and the cumulative function value of its corresponding theoretical normal distribution. The number of deviation values and their absolute values corresponds to the number of data points in the feature data column.
[0126] S3.1.2.3 Calculate the maximum difference between the cumulative distribution function of the experimental dataset and the cumulative distribution function of the theoretical normal distribution, which is the KS statistic;
[0127] Each performance evaluation analysis involves taking a batch of data from historical sampling data to form an experimental dataset, which, in a longitudinal view, results in multiple experimental datasets.
[0128] S3.1.2.4. Set a significance level to determine whether the maximum value of the difference between the cumulative distribution function of the observed experimental dataset and the cumulative distribution function of the theoretical normal distribution belongs to random error;
[0129] S3.1.2.5. Compare the calculated KS statistic with the KS critical value. If the statistic KS is less than the KS critical value, the null hypothesis is considered to be true, and the sampled data in the experimental dataset basically follows a normal distribution.
[0130] To ensure data accuracy, the sampling attributes in the test dataset need to be compared and verified: 1) The sampling time interval of each evaluation quantity must be consistent. For individual missing data, interpolation algorithms or moving time windows can be used to obtain the data to ensure that each data point has a sampling record; 2) Check the difference between the sampled values and the mean in each test dataset. If the difference is too large, it is necessary to determine whether any wind turbine is in a state of operation switching or failure.
[0131] The Kolmogorov-Smithnov hypothesis is a hypothesis testing method used to check whether the distribution of sample data follows a normal distribution.
[0132] S3.2 Correlation Coefficient Calculation: Based on the normality test results, determine the calculation method for the correlation coefficient between various influencing factors in the experimental dataset; the correlation coefficient is used to determine the strength of the correlation between different feature columns in the experimental dataset. Specifically:
[0133] Experimental Dataset Normalization: A state data matrix is constructed based on the experimental data set of potential influencing factors of the selected grid-connected wind farm operation performance evaluation indicators within a given time window. When there are significant differences between the sampled data points in the state data matrix, a normalization method is used to process the sampled data in the state data matrix, specifically as follows:
[0134]
[0135] In the formula, x represents the standardized value in the correlation analysis data matrix; ij μ represents the stored value of the j-th potential influencing factor at the i-th sampling time. j σ is the sampled average of the j-th potential influencing factor within a given time window; j For x ij The standard deviation of the j-th potential influencing factor; the data in each column of the normalized state data matrix approximately follow a normal distribution with a mean of 0 and a variance of 1;
[0136] Given a selected experimental dataset that approximates a normal distribution, correlation analysis is performed on any two waveforms at a given time point. The degree of linear correlation is represented by the Pearson correlation coefficient, which is expressed in discretized form as follows:
[0137]
[0138] In the formula, r is the correlation coefficient between the two variable sequences in a single analysis; cov(X,Y) and σ X σ Y The covariances of the two experimental datasets are respectively; x i y iThese are the values of the i-th data point in a single analysis sampling data; These are the average values of the two columns of sampled data, respectively; n is the total amount of data contained in the sampled data window;
[0139] When the Pearson correlation coefficient between two sets of feature columns in the experimental dataset is greater than the threshold value, it indicates that there is a significant correlation between the two sets of feature columns in the experimental dataset.
[0140] The correlation strength between any two modules in the experimental dataset is determined based on the correlation calculation results. Generally, a Pearson correlation coefficient |r| ≤ 0.3 indicates a very low correlation between the two sets of data; 0.3 < |r| ≤ 0.5 indicates a low correlation between the feature columns of the two sets of experimental datasets; 0.5 < |r| ≤ 0.8 indicates a high correlation between the feature columns of the two sets of experimental datasets; and 0.8 < |r| ≤ 1 indicates a very high correlation between the feature columns of the two sets of experimental datasets. The value of r ranges from [-1, 1]. When r = 1, it indicates a perfect positive correlation between the two sets of experimental data; when r = -1, it indicates a perfect negative correlation between the two sets of experimental data.
[0141] If the data to be analyzed does not follow a normal distribution, correlation verification is required using rank division. When a data point appears multiple times, its rank is the average of the two preceding and following ranks. The rank is then calculated to restore its original value after data changes. The linear correlation between two waveforms is represented by the Spearman correlation coefficient, whose discretization formula can be expressed as:
[0142]
[0143]
[0144] In the formula, ρ is the Spearman correlation coefficient between the two sets of sampled datasets; r i s i Indicates the column rank of two columns of data; d is the average of the ranks of the two columns of data; i Let represent the rank difference of the i-th pair of data; n is the total amount of data sampled in this analysis.
[0145] If the Spearman correlation coefficient is chosen to assess the correlation between various influencing factors of the performance index, the larger the absolute value of the Spearman correlation coefficient, the higher the correlation of the feature columns of the selected experimental dataset. The sign of the correlation coefficient determines the positive or negative correlation of the analyzed data.
[0146] The correlation coefficients mentioned above are calculated based on sample data within a given time window, and are therefore sample correlation coefficients. Due to the randomness of sampling and limitations in sample size, sample correlation coefficients cannot usually be directly used to indicate whether there is a significant linear correlation between two feature columns.
[0147] By mining and calculating the correlation between the operating performance indicators of grid-connected wind farms and potential influencing factors, we can identify the variables that have a significant impact on the performance indicators.
[0148] By comparing the deviations of the normality verification results with the given threshold, the calculation method for the correlation coefficients of the influencing factors to be analyzed in the experimental dataset is determined.
[0149] Identify the data columns containing the performance evaluation indicators of wind turbines or wind farms and the potential influencing factors, calculate the corresponding correlation coefficients, and determine whether there is a strong correlation between the selected sampled data columns, which will serve as the basis for further analysis and judgment.
[0150] S3.3 Removing Abnormal Result Sets: Using correlation coefficients, perform correlation analysis on all feature column data in the experimental dataset. Based on the correlation analysis results, select feature column data with correlation coefficients greater than a preset threshold from the experimental dataset to form a new experimental dataset.
[0151] By removing outlier or temporarily unexplainable datasets from the correlation analysis results, an experimental dataset that approximately satisfies linear correlation is obtained. Based on this dataset, an input feature vector for wind farm operation performance evaluation is constructed.
[0152] S3.4 Performance evaluation modeling: Explore the correlation between the performance evaluation index to be analyzed and the new experimental dataset, as well as the input feature vector of the wind farm operation performance evaluation in the new experimental dataset. Based on the obtained correlation strength results, select the influencing factor sequence as the independent variable and the performance evaluation index as the dependent variable, and construct a regression model.
[0153] The regression model comprehensively considers the impact of factors such as unit parameters, operating data, and wind energy environment on the operating performance of the unit and wind farm.
[0154] S3.5 Solving the Regression Equations: Considering the impact of wind turbine parameters, grid-connected operation data, and wind energy environmental factors on the operating performance of wind turbines, based on the regression model, multiple linear regression equations are established for each of the new experimental datasets. By solving the multiple linear regression equations, the correlation between the various influencing factors is transformed into quantitative evaluation indicators for the operating performance of wind turbines and wind farms. Details are as follows:
[0155] Multiple linear regression analysis was used to determine the quantitative relationships of interdependence among various factors influencing wind turbine performance evaluation indicators. The multiple linear regression equation was established as follows:
[0156] Y = β0 + β1X1 + β2X2 + ... + β n X n +ε
[0157] In the formula, β0 is a constant term; β i (i = 1, 2, 3, ..., n) represents the other independent variables X. i When the independent variable X remains constant, ε represents the average change in the dependent variable Y for every unit change in the independent variable X; ε represents the residual, which is the difference between the actual value and the estimated value of the dependent variable, and is a random variable.
[0158] In solving the multiple linear regression equation, its normalized equation system can be expressed as:
[0159] (X T X)B=X T Y
[0160] In the formula, Y is the historical sampling data sequence of the dependent variable; B represents the historical sampling data sequence of the independent variable; and X is the coefficient matrix of the normalized equation system in the regression model.
[0161] Data sets that are determined to be nonlinear based on actual operational experience at wind farm sites are transformed into linear relationships through variable transformation. These linear relationships are then resubmitted into step S3.5 to calculate the quantitative relationships between the datasets under study.
[0162] Calculate the exact quantitative relationship between relevant wind farm performance indicators and significant influencing factors, and perform a comprehensive verification of the calculation results.
[0163] S3.6, Overall Verification: Perform significance verification on the calculation method of the correlation coefficient and the calculation results of the multiple linear regression equation.
[0164] The significance of the correlation coefficient calculation method is verified, including:
[0165] After the above process, if the feature columns of the experimental dataset used to calculate the correlation coefficient follow a normal distribution, and a new statistic t is constructed, then:
[0166]
[0167] In the formula, r represents the correlation coefficient of the calculated feature data column; n is the number of samples selected in the experimental dataset. The statistic t follows a t-distribution with (n-2) degrees of freedom.
[0168] When the probability p value corresponding to the statistic t is less than the selected significance level, it is considered that there is a significant linear correlation in the feature columns of the selected experimental dataset.
[0169] The significance of the results of the multiple linear regression equation is verified to identify variables in the influencing factor dataset that are related to the performance evaluation index, including:
[0170] (1) When all elements in the independent variable data sequence B are 0, the original linear correlation hypothesis is rejected, and the quantitative relationship is re-determined using first-order linear or other methods; the residuals are analyzed using quantile plots to check whether they approximately follow a normal distribution;
[0171] By analyzing the residuals of the regression model, we can test whether it meets the basic assumptions of linear regression. If the normal probability plot of the residuals does not show large drift, it indicates that the previous assumptions are not significantly wrong.
[0172] (2) When not all elements in the independent variable data sequence B are 0, perform an overall significance check on the multiple linear regression equation:
[0173]
[0174] In the formula, SS R SS represents the sum of squared deviations of a multiple linear regression equation; E This represents the sum of squared residuals from the calculation results of the multiple linear regression equation; y represents the i-th component of the predicted value of the regression equation; i This represents the actual value used for regression analysis; This represents the average of the actual values.
[0175] (3) If the p-value of the multiple linear regression equation is less than the significance level, then the multiple linear regression equation is significant; perform a significance test on the regression coefficients. If the p-value is less than the significance level, then the regression coefficients are significant; among them, we have:
[0176]
[0177] In the formula, df R df represents the degrees of freedom in the regression equation, which is the number of independent variables in the multiple linear regression equation. E The residual degrees of freedom are represented by the following formula:
[0178] df E =df T -df R
[0179] df T =n-1
[0180] Among them, df T df represents the total degrees of freedom of the multiple linear regression equation. R This represents the regression degrees of freedom of the regression equation; n is the number of dimensions of the feature column vector.
[0181] The value of P can be calculated using the following formula:
[0182] P = 1 - F X (x)
[0183] Among them, F X (x) represents the cumulative distribution function value, whose characteristics are determined by F and df. R df E A joint decision.
[0184] S4. Performance Evaluation: Substitute the actual operating data of the wind farm into the wind farm operation performance evaluation model composed of the new quantitative evaluation index and the revised quantitative evaluation index, calculate the values of various performance indicators of the wind farm, and obtain the overall performance evaluation results of the wind farm.
[0185] By solving multiple linear regression, we can obtain the specific functional relationships between evaluation indicators and influencing factors. The set of these relationships is the wind farm operation performance evaluation model.
[0186] Performance evaluation includes operational analysis, economic operation, and statistical analysis; among which, the statistical analysis submodule includes functions such as statistical indicator analysis, key indicator query, and benchmarking of key information;
[0187] The operations analysis submodule covers several aspects, including equipment efficiency, operation and maintenance rates, and wind energy resources.
[0188] The economic operation submodule includes functions such as theoretical power balance analysis, power loss analysis, economic operation overview, and station operation comparison.
[0189] The assessment method for the health status of wind turbines during grid-connected operation includes: when a wind turbine experiences a fault or abnormal operation, the correlation between its parameters will also change. The correlation coefficient between each parameter during normal operation can be compared with the correlation coefficient during abnormal operation. Based on the fault characteristic parameters with large deviations, the fault range that causes the wind turbine to be in an unhealthy state can be located.
[0190] The wind farm operation indicators are extracted and calculated by category to obtain the data source for evaluating the grid connection performance of wind farms.
[0191] Quantitatively described data relationships will be added to relevant categories to improve the existing wind turbine performance evaluation index system.
[0192] Improve the evaluation system for the operation performance indicators of grid-connected wind farms. Based on the specific quantitative representation method of the wind farm operation performance evaluation indicators, input the actual operation data of the wind farms to the evaluation model to calculate and evaluate the performance of various indicators of wind farm operation, and obtain the evaluation results of each evaluation indicator to provide data support for the economic and safe operation and daily maintenance of wind farms.
[0193] This invention comprehensively analyzes the coupling between wind turbine operating data and environmental resources and turbine parameters. It utilizes correlation coefficient and regression analysis to mine and comprehensively evaluate the factors influencing wind turbine performance during operation, seeking quantitative relationships between performance indicators and various factors from an evaluation methodological perspective. This method overcomes the limitation of limited indicator coverage in existing evaluation standards, enabling a more comprehensive and in-depth establishment of a wind turbine operating performance indicator evaluation system. It provides a scientific basis for efficient wind farm operation and maintenance management, wind turbine fault identification, and optimization, and also offers technical support for the subsequent improvement of relevant industry standards.
[0194] Furthermore, the method proposed in this invention addresses some shortcomings of the first literature from a mathematical perspective: correlation analysis is a mathematical method that does not have specific requirements regarding the physical meaning of eigenvectors. When measurement conditions permit, the eigenvector evaluation indicators can be either "grid-side" indicators related to wind turbine grid connection and system operation, or they can be selected based on the wind turbine's own operating characteristics, focusing on the turbine's operating parameters and indicators. These indicators can be evaluation quantities with existing calculation formulas, or they can be evaluation quantities that currently lack quantitative descriptions but are actively explored and analyzed. The analysis results can provide data support for improving the indicator system.
[0195] Example 2
[0196] like Figure 3 As shown, the present invention provides a grid-connected wind farm operation performance evaluation system based on correlation analysis, comprising:
[0197] The data acquisition module is used to establish a main monitoring area model in the wind farm operation performance evaluation system, and derive secondary models of wind farm, wind turbine, wind turbine rotor, transmission device, and converter from this model; it also obtains the information required for performance evaluation from the wind farm SCADA system, power metering system, and meteorological station monitoring system to form historical sampling data based on time series.
[0198] The data preprocessing module is used to filter, smooth, and standardize the historical sampling data based on time series to obtain the initial test dataset for wind farm grid connection performance evaluation.
[0199] The data correlation analysis module is used to perform correlation analysis on the initial experimental dataset. Based on the performance evaluation requirements, corresponding feature columns are selected from the initial experimental dataset to form a new experimental dataset. This dataset forms the basis for constructing the input feature vector for wind farm operation performance evaluation. The module then mines the correlation between the performance evaluation indicators to be analyzed and the new experimental dataset, as well as the input feature vectors for each wind farm operation performance evaluation within the new experimental dataset. Based on the obtained correlation strength results, the module selects the sequence of influencing factors as independent variables and the performance evaluation indicators as dependent variables, constructing a regression model. Regression analysis is performed based on the regression model to obtain the quantitative descriptive relationship between the feature vectors, leading to a revised algorithm for the evaluation indicators or new quantitative evaluation indicators. The independent variable influencing factors are those factors that affect the operation performance of grid-connected wind farms, specifically including wind farm ledger data, grid-connected operation data, and wind energy resource data.
[0200] The performance evaluation module is used to input the actual operating data of the wind farm into the wind farm operation performance evaluation model composed of new quantitative evaluation indicators and revised quantitative evaluation indicators, calculate the values of various performance indicators of the wind farm, and obtain the overall performance evaluation results of the wind farm.
[0201] In one optional implementation, the correlation analysis-based method for evaluating the operational performance of grid-connected wind farms includes: a) establishing a main monitoring area model, and deriving secondary models of the wind farm, wind turbine generators, turbine rotors, transmission devices, and converters; acquiring the information required for performance evaluation and forming historical sampling data based on time series; b) performing filtering, smoothing, and standardization processing on the historical sampling data based on time series to obtain an initial experimental dataset for evaluating the grid-connected performance of the wind farm; c) performing correlation analysis on the initial experimental dataset, using the correlation coefficient method and regression analysis to perform data mining and comprehensive evaluation of the factors influencing the performance of the wind turbine generators during operation, and finding the quantitative relationship between performance indicators and various factors from the perspective of evaluation methods; d) substituting the actual operating data of the wind farm into the wind farm operational performance evaluation model, calculating the values of various performance indicators of the wind farm according to the quantitative expression method of indicator evaluation obtained in the above process, and obtaining the overall performance indicator evaluation result of the wind farm.
Claims
1. A method for evaluating the operational performance of grid-connected wind farms based on correlation analysis, characterized in that, include: Data acquisition: A main monitoring area model is established in the wind farm operation performance evaluation system, and secondary models of wind farm, wind turbine, wind turbine rotor, transmission device and converter are derived from it; the information required for performance evaluation is obtained from the wind farm SCADA system, power metering system and meteorological station monitoring system to form historical sampling data based on time series. Data preprocessing: The historical sampling data based on time series is filtered, smoothed and standardized to obtain the initial test dataset for wind farm grid connection performance evaluation; Data Correlation Analysis: Correlation analysis is performed on the initial experimental dataset. Based on the performance evaluation requirements, corresponding feature columns are selected from the initial experimental dataset to form a new experimental dataset. This dataset forms the basis for constructing the input feature vector for wind farm operation performance evaluation. The correlation between the performance evaluation indicators to be analyzed and the new experimental dataset, as well as the input feature vectors for wind farm operation performance evaluation within the new experimental dataset, is explored. Based on the obtained correlation strength results, the sequence of influencing factors as independent variables and the performance evaluation indicators as dependent variables are selected to construct a regression model. Regression analysis is performed based on the regression model to obtain the quantitative descriptive relationship between the feature vectors, and to obtain the revised algorithm for the evaluation indicators or new quantitative evaluation indicators. Among them, the independent variable influencing factors are factors that affect the operation performance of grid-connected wind farms, specifically including wind farm ledger data, grid-connected operation data, and wind energy resource data. Performance evaluation: Substitute the actual operating data of the wind farm into the wind farm operation performance evaluation model composed of the new quantitative evaluation index and the revised quantitative evaluation index, calculate the values of various performance indicators of the wind farm, and obtain the overall performance evaluation results of the wind farm.
2. The method for evaluating the operational performance of grid-connected wind farms based on correlation analysis according to claim 1, characterized in that, Correlation analysis was performed on the initial experimental dataset. Based on the performance evaluation requirements, corresponding feature columns were selected from the initial experimental dataset to form a new experimental dataset. Based on this, the input feature vector for wind farm operation performance evaluation was constructed. The correlation between the performance evaluation index to be analyzed and the new experimental dataset, as well as the input feature vector for wind farm operation performance evaluation within the new experimental dataset, was explored. Based on the obtained correlation strength results, the sequence of influencing factors as independent variables and the performance evaluation index as dependent variables were selected to construct a regression model. Regression analysis based on regression models yields quantitative descriptive relationships between eigenvectors, leading to revised algorithms for evaluation indicators or new quantitative evaluation indicators, including: Normality check: Perform normality check on the experimental dataset to verify the fit of the sampled data in the experimental dataset with a normal distribution. After normality check, the experimental dataset that meets the conditions for correlation analysis is obtained. Correlation coefficient calculation: Based on the normality test results, determine the calculation method for the correlation coefficient between each influencing factor in the experimental dataset; whereby the correlation coefficient is used to determine the correlation strength between different feature columns in the experimental dataset; Remove abnormal result sets: Use correlation coefficients to perform correlation analysis on all feature column data in the experimental dataset. Based on the correlation analysis results, select feature column data with correlation coefficients greater than a preset threshold from the experimental dataset to form a new experimental dataset. Construct the input feature vector for wind farm operation performance evaluation based on this dataset. Performance evaluation modeling: Explore the correlation between the performance evaluation index to be analyzed and the new experimental dataset, as well as the input feature vector of the wind farm operation performance evaluation in the new experimental dataset. Based on the obtained correlation strength results, select the sequence of influencing factors as independent variables and the performance evaluation index as dependent variables, and construct a regression model. Solving the regression equation: Considering the impact of wind turbine parameters, grid-connected operation data, and wind energy environmental factors on the operating performance of wind turbines, based on the regression model obtained in the previous step, multiple linear regression equations are established for each of the new experimental datasets; by solving the multiple linear regression equations, the correlation between each influencing factor is transformed into a quantitative evaluation index for the operating performance of wind turbines and wind farms. Overall verification: Perform significance verification on the calculation method of the correlation coefficient and the calculation results of the multiple linear regression equation.
3. The method for evaluating the operational performance of grid-connected wind farms based on correlation analysis according to claim 2, characterized in that, The experimental dataset is subjected to normality verification to check the approximation of a normal distribution in the sampled data. The experimental dataset that meets the conditions for correlation analysis after normality verification includes: Calculate the percentiles of the sampled data in the experimental dataset. The percentiles of the sampled data in the experimental dataset are combined with the percentiles of the normal distribution to form the percentile comparison space. Based on the distribution of the sampled data points in the experimental dataset, we can make a preliminary judgment on whether the sampled data follows a normal distribution. Check the goodness-of-fit deviation between the sampled data in the experimental dataset and the standard distribution, perform normality verification on the historical SCADA sampled data in the experimental dataset, and use the KS hypothesis testing method to perform nonparametric tests on the sampled data in the experimental dataset to check whether the hypothesis about the population normal distribution is valid.
4. The method for evaluating the operational performance of grid-connected wind farms based on correlation analysis according to claim 3, characterized in that, The goodness-of-fit deviation of the sampled data in the experimental dataset from the standard distribution was examined. Normality was verified on the historical SCADA sampled data in the experimental dataset. Nonparametric tests were performed on the sampled data in the experimental dataset using the KS hypothesis testing method to check whether the hypothesis about the population's normal distribution holds, including: After sorting the SCADA historical sampling data in the experimental dataset in ascending order, the percentile markers of the SCADA historical sampling data are determined. The mean and standard deviation of the percentile comparison space are calculated. Based on the deviation of the measurement points of the SCADA historical sampling data in the experimental dataset from the standard deviation of the experimental dataset, the distribution function of the analysis samples of the experimental dataset is constructed. In the formula, σ is the standard deviation of the historical sampling data; This represents the average value of the experimental dataset; Compare the deviations of each percentile marker interval point from their corresponding theoretical normal distribution cumulative function values, and compare the maximum absolute value of each deviation with the KS critical value. If the maximum difference between the cumulative distribution function of the experimental dataset and the cumulative distribution function of the normal distribution is less than the critical value, the null hypothesis is considered to be true; otherwise, the null hypothesis is rejected. The maximum difference between the cumulative distribution function of the experimental dataset and the cumulative distribution function of the theoretical normal distribution is the KS statistic. A significance level is set to determine whether the maximum difference between the cumulative distribution function of the observed experimental dataset and the cumulative distribution function of the theoretical normal distribution belongs to random error; The calculated KS statistic is compared with the KS critical value. If the statistic KS is less than the KS critical value, the null hypothesis is considered to be true, and the sampled data in the experimental dataset basically follows a normal distribution.
5. The method for evaluating the operational performance of grid-connected wind farms based on correlation analysis according to claim 2, characterized in that, The method for calculating the correlation coefficients among various influencing factors in the experimental dataset is determined based on the results of the normality test, including: Experimental Dataset Normalization: A state data matrix is constructed based on the experimental data set of potential influencing factors of the selected grid-connected wind farm operation performance evaluation indicators within a given time window. When there are significant differences between the sampled data points in the state data matrix, a normalization method is used to process the sampled data in the state data matrix, specifically as follows: In the formula, x represents the standardized value in the correlation analysis data matrix; ij μ represents the stored value of the j-th potential influencing factor at the i-th sampling time. j σ is the sampled average of the j-th potential influencing factor within a given time window; j For x ij The standard deviation of the j-th potential influencing factor; the data in each column of the normalized state data matrix approximately follow a normal distribution with a mean of 0 and a variance of 1; Given a selected experimental dataset that approximates a normal distribution, correlation analysis is performed on any two waveforms at a given time point. The degree of linear correlation is represented by the Pearson correlation coefficient, which is expressed in discretized form as follows: In the formula, r is the correlation coefficient between the two variable sequences in a single analysis; cov(X,Y) and σ X σ Y The covariances of the two experimental datasets are respectively; x i y i These are the values of the i-th data point in a single analysis sampling data; These are the average values of the two columns of sampled data, respectively; n is the total amount of data contained in the sampled data window; When the Pearson correlation coefficient between two sets of feature columns in the experimental dataset is greater than the threshold value, it indicates that there is a significant correlation between the two sets of feature columns in the experimental dataset.
6. The method for evaluating the operational performance of grid-connected wind farms based on correlation analysis according to claim 2, characterized in that, The method for calculating the correlation coefficients among various influencing factors in the experimental dataset is determined based on the results of the normality test, including: If the data to be analyzed does not follow a normal distribution, correlation verification is required using rank division. When a data point appears multiple times, its rank is the average of the two preceding and following ranks. The rank is calculated to restore the original value after data changes. The linear correlation between two waveforms is represented by the Spearman coefficient, whose discretization formula is as follows: In the formula, ρ is the Spearman correlation coefficient between the two sets of sampled datasets; r i s i Indicates the column rank of two columns of data; d is the average of the ranks of the two columns of data; i The rank difference of the i-th pair of data is represented by n; n is the total amount of data sampled in this analysis. The larger the absolute value of the Spearman correlation coefficient, the stronger the correlation between the two types of data. The sign of the correlation coefficient determines whether the two types of data are positive or negative.
7. The method for evaluating the operational performance of grid-connected wind farms based on correlation analysis according to claim 2, characterized in that, The significance of the correlation coefficient calculation method is verified, including: If the feature columns of the experimental dataset used to calculate the correlation coefficient follow a normal distribution, then a new statistic t is constructed: In the formula, r represents the correlation coefficient of the calculated feature data column; n is the number of samples selected in the experimental dataset; the statistic t follows a t-distribution with (n-2) degrees of freedom; When the probability p value corresponding to the statistic t is less than the selected significance level, it is considered that there is a significant linear correlation in the feature columns of the selected experimental dataset.
8. The method for evaluating the operational performance of grid-connected wind farms based on correlation analysis according to claim 2, characterized in that, Considering the impact of wind turbine parameters, grid-connected operation data, and wind energy environmental factors on wind turbine operating performance, multiple linear regression equations are established for each of the new experimental datasets. By solving the multiple linear regression equations, the correlations between the influencing factors are transformed into quantitative evaluation indicators for wind turbine operating performance and wind farm operating performance, including: Establish a multiple linear regression equation: Y=β0+β1X1+β2X2+…+β n X n +e In the formula, β0 is a constant term; β i (i = 1, 2, 3, ..., n) represents the other independent variables X. i The average change in dependent variable Y for every unit change in a specified independent variable X when X remains constant; ε represents the calculation of residuals. In solving the multiple linear regression equation, the normalized equation system is expressed as: (X T X)B=X T Y In the formula, Y is the historical sampling data sequence of the dependent variable; B represents the historical sampling data sequence of the independent variable; and X is the coefficient matrix of the normalized equation system in the regression model.
9. The method for evaluating the operational performance of grid-connected wind farms based on correlation analysis according to claim 2, characterized in that, The significance of the calculation results of the multiple linear regression equation is verified, including: When all elements in the independent variable data sequence B are 0, the original linear correlation hypothesis is rejected, and the quantitative relationship is re-determined using first-order linearity or other methods; the quantile plot is used to analyze the calculated residuals to check whether they approximately follow a normal distribution; When not all elements in the independent variable data sequence B are zero, perform an overall significance check on the multiple linear regression equation: In the formula, SS R SS represents the sum of squared deviations of a multiple linear regression equation; E This represents the sum of squared residuals from the calculation results of the multiple linear regression equation; y represents the i-th component of the predicted value of the regression equation; i This represents the actual value used for regression analysis; This represents the average of the actual values; If the p-value of the multiple linear regression equation is less than the significance level, then the multiple linear regression equation is significant; a significance test is performed on the regression coefficients; if the p-value is less than the significance level, then the regression coefficients are significant; where: df R df represents the degrees of freedom in the regression equation, which is the number of independent variables in the multiple linear regression equation. E The residual degrees of freedom are represented by the following formula: df E =df T -df R df T =n-1 Among them, df T df represents the total degrees of freedom of the multiple linear regression equation. R This represents the regression degrees of freedom of the regression equation; n is the number of dimensions of the feature column vectors. The value of P is calculated using the following formula: P=1-F X (x) Among them, F X (x) represents the cumulative distribution function value, whose characteristics are determined by F and df. R df E A joint decision.
10. A grid-connected wind farm operation performance evaluation system based on correlation analysis, characterized in that, include: The data acquisition module is used to establish a main monitoring area model in the wind farm operation performance evaluation system, and derive secondary models of wind farm, wind turbine, wind turbine rotor, transmission device, and converter from this model; it also obtains the information required for performance evaluation from the wind farm SCADA system, power metering system, and meteorological station monitoring system to form historical sampling data based on time series. The data preprocessing module is used to filter, smooth, and standardize the historical sampling data based on time series to obtain the initial test dataset for wind farm grid connection performance evaluation. The data correlation analysis module is used to perform correlation analysis on the initial experimental dataset. Based on the performance evaluation requirements, corresponding feature columns are selected from the initial experimental dataset to form a new experimental dataset. This dataset forms the basis for constructing the input feature vector for wind farm operation performance evaluation. The module then mines the correlation between the performance evaluation indicators to be analyzed and the new experimental dataset, as well as the input feature vectors for each wind farm operation performance evaluation within the new experimental dataset. Based on the obtained correlation strength results, the module selects the sequence of influencing factors as independent variables and the performance evaluation indicators as dependent variables, constructing a regression model. Regression analysis is performed based on the regression model to obtain the quantitative descriptive relationship between the feature vectors, leading to a revised algorithm for the evaluation indicators or new quantitative evaluation indicators. The independent variable influencing factors are those factors that affect the operation performance of grid-connected wind farms, specifically including wind farm ledger data, grid-connected operation data, and wind energy resource data. The performance evaluation module is used to input the actual operating data of the wind farm into the wind farm operation performance evaluation model composed of new quantitative evaluation indicators and revised quantitative evaluation indicators, calculate the values of various performance indicators of the wind farm, and obtain the overall performance evaluation results of the wind farm.
Citation Information
Patent Citations
Method for evaluating real-time running state of wind turbine generator based on big data technology
CN111062508A
Energy efficiency state evaluation method and system of wind turbine generator and medium
CN114742363A