Nonlinear correlation pair screening method for photovoltaic power generation output power influence factors
By constructing a time window pair and calculating nonlinear correlation coefficients, the influencing factors of photovoltaic power output power were screened, which solved the problem of failure to fully consider the impact of time scale in the existing technology, and achieved more accurate significance in the screening and analysis results of influencing factors.
Patent Information
- Application Number
- CN202411096932.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-12
- Publication Date
- 2025-06-13
AI Technical Summary
When screening the factors affecting the output power of photovoltaic power generation, the prior art fails to fully consider the different regularity, periodicity and volatility of meteorological characteristics on the time scale, resulting in the scattered screening results and the significance is not strong.
A nonlinear correlation pair screening method for influencing factors of photovoltaic power output power is proposed. By obtaining the photovoltaic power output power data and meteorological factor data within a predetermined time period, a time window pair is constructed, a nonlinear correlation coefficient is calculated, and a candidate set is obtained, and the optimal nonlinear correlation pair is obtained through reduction processing.
This method can more accurately consider the influence of meteorological factors on the time scale, improve the reliability and significance of the analysis results, and obtain more accurate screening results for the influencing factors of photovoltaic power output power.
Smart Images

Figure CN120144993A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of photovoltaic power generation. Specifically, it relates to a screening method for non-linear correlation pairs of factors affecting the output power of photovoltaic power generation. Background Art
[0002] In order to improve the reliability, security, and stability of power grid operation, it is crucial to explore the influencing factors of photovoltaic power prediction. Photovoltaic power generation is a non-linear stochastic process with multi-variable coupling. The main meteorological influencing factors include: solar radiation, aerosol, sunshine duration, temperature, humidity, wind direction, wind speed, cloud cover, etc. With the gradual expansion of the scale of photovoltaic power generation connected to the grid, the inherent intermittency and uncertainty of photovoltaic output have brought huge challenges to the reliable operation of the power grid. Therefore, domestic and foreign scholars have also carried out research and analysis on the meteorological related influencing factors of photovoltaic power generation.
[0003] For example, some researchers used the Pearson correlation coefficient to analyze the influence mechanism of typical meteorological factors on photovoltaic power generation under different seasons and the coupling relationship between multiple variables. Some researchers used the Pearson correlation analysis method to analyze the correlation between the power generation power of a distributed photovoltaic power station in a certain area and seven meteorological influencing factors such as daily solar radiation, sunshine hours, daily average ambient temperature, daily maximum temperature, daily minimum temperature, diurnal temperature range, and daily average wind speed. Some researchers used the Pearson correlation coefficient to analyze the relationship between the photovoltaic power generation power curve and meteorological factors, screened out five important meteorological characteristics including diffuse radiation, ultraviolet irradiance, ultraviolet index, global radiation, and ambient temperature, and incorporated these screened weather characteristics into the photovoltaic power generation power modeling and prediction model, effectively improving the accuracy of power prediction and reducing the prediction error. Some researchers used a hybrid method combining filtering and packaging strategies to select meteorological influencing factors. According to the internal correlation between the power generation power and meteorological influencing factors, considering the criteria of feature filtering such as dependence, information, and distance consistency, the Pearson correlation coefficient was used to evaluate the influence of 17 feature subsets including solar radiation, total cloud cover, relative humidity, etc. on the full-field power. Some researchers used the K-Nearest Neighbor (KNN) algorithm to fully explore the key factors affecting photovoltaic output among many meteorological factors such as external temperature, humidity, and pressure, thereby reconstructing the multivariate data sequence, and inputting the excavated meteorological influencing factors into the BiLSTM network model to improve the prediction accuracy of the short-term power generation of photovoltaic power. Some researchers used the Spearman correlation coefficient to select an appropriate combination of meteorological characteristics, then used LSTM and Informer as basic models to extract the dependence relationship between variables and target values in short-time series and long-time series respectively to generate meta-features, and finally constructed multiple linear regression relationships as meta-models to fit the mapping relationship between meteorological influencing factors and target values, thereby outputting more accurate prediction results.
[0004] Although the above research has drawn some beneficial conclusions, it has not considered the different regularities, periodicity, and volatility of meteorological characteristics on the time scale, resulting in scattered screening of meteorological factors, and the conclusions drawn are inevitably too broad or not significant enough. Summary of the Invention
[0005] The technical problem solved by this application is: how to provide a non-linear correlation pair screening method for the influencing factors of photovoltaic power generation output power that considers the influence of the time scale and can improve the significance.
[0006] This application provides a non-linear correlation pair screening method for the influencing factors of photovoltaic power generation output power, and the non-linear correlation pair screening method includes:
[0007] Obtain the photovoltaic power generation output power data and the corresponding meteorological factor data within a predetermined time period;
[0008] Construct a number of time window pairs based on the photovoltaic power generation output power data and the meteorological factor data;
[0009] Calculate the non - linear correlation coefficient of the time window pairs, and obtain a candidate set of non - linear correlation pairs according to the calculation results;
[0010] Perform a reduction process on the candidate set of non - linear correlation pairs to obtain the optimal non - linear correlation pairs.
[0011] Optionally, the method for constructing a number of time window pairs based on the photovoltaic power generation output power data and the meteorological factor data includes:
[0012] Divide the photovoltaic power generation output power data and the meteorological factor data into power data time windows Y(s’, l) and meteorological data time windows X(s, l) respectively according to a preset time window length;
[0013] The constructed time window pair CP = <s, l, τ>, where s and s’ both represent the start time of the time window, l represents the time window length, and τ = s’ - s represents the delay between time windows.
[0014] Optionally, the method for calculating the non - linear correlation coefficient of the time window pairs and obtaining a candidate set of non - linear correlation pairs according to the calculation results is calculated based on a non - linear correlation analysis algorithm with time window reduction, including:
[0015] Split the time window pairs into a number of sub - window pairs according to the split time window;
[0016] Calculate the non - linear correlation coefficient of each sub - window pair, and add the sub - window pairs with non - linear correlation coefficients greater than the threshold to the candidate set of non - linear correlation pairs.
[0017] Optionally, the method for calculating based on the non - linear correlation analysis algorithm with time window reduction further includes:
[0018] Divide the split time window into a number of non - overlapping subset windows;
[0019] Reduce the sub - window pairs in the candidate set of non - linear correlation pairs according to the subset windows;
[0020] Calculate the non - linear correlation coefficient of each remaining sub - window pair after reduction, and retain the sub - window pairs with non - linear correlation coefficients greater than the threshold to form the final candidate set of non - linear correlation pairs.
[0021] Optionally, the method for reducing the candidate set of the non - linear correlation pairs to obtain the optimal non - linear correlation pairs is as follows: determine the time - window pair in the candidate set of the non - linear correlation pairs that contains the best start time and the best time - window length, and the non - linear correlation coefficient of the time - window pair is the largest.
[0022] The non - linear correlation pair screening method for influencing factors of photovoltaic power generation output power provided by this application has the following technical effects:
[0023] This method fully considers the different regularities, different periodicities, and volatility effects of the influencing factors of output power, that is, meteorological factors, on the time scale, and can obtain more reliable analysis results. Brief Description of the Drawings
[0024] Figure 1 It is a flowchart of the non - linear correlation pair screening method for influencing factors of photovoltaic power generation output power according to one or more embodiments;
[0025] Figure 2 It is a schematic diagram of the screening process based on time - window reduction according to one or more embodiments;
[0026] Figure 3 It is a schematic diagram of the reduction process based on time - window reduction according to one or more embodiments;
[0027] Figure 4 It is a schematic diagram of the screening process based on time - window expansion according to one or more embodiments;
[0028] Figure 5 It is a schematic diagram of the optimization process of the best starting point of the time window according to one or more embodiments;
[0029] Figure 6 It is a schematic diagram of the optimization process of the best time - window length according to one or more embodiments;
[0030] Figure 7 It is the correlation analysis result of the time - series pair based on the MI coefficient according to one or more embodiments;
[0031] Figure 8 It is the correlation analysis of the time - series pair based on the CCC coefficient according to one or more embodiments;
[0032] Figure 9 It is the analysis of the influencing factors of photovoltaic power generation power of Photovoltaic Power Station 1 according to one or more embodiments;
[0033] Figure 10 It is the analysis of the influencing factors of photovoltaic power generation power of Photovoltaic Power Station 2 according to one or more embodiments. Detailed Embodiments
[0034] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0035] Before describing the various embodiments of the present application in detail, the technical concept of the present application will be briefly described first: In the current non-linear correlation analysis method, the influence of different regularities, periodicities and fluctuations of meteorological characteristics on different time scales is not considered, resulting in scattered screening of meteorological factors, and the conclusions drawn are inevitably too broad or not significant enough. For this reason, the non-linear correlation pair screening method for influencing factors of photovoltaic power generation output power provided by the present application constructs a number of time window pairs based on the photovoltaic power generation output power data and meteorological factor data within a certain period of time, calculates the non-linear correlation coefficients of each time window pair, obtains a candidate set of non-linear correlation pairs, and after reduction processing, obtains the optimal non-linear correlation pair. This method fully considers the influence of different regularities, different periodicities and fluctuations of the influencing factors of the output power, that is, meteorological factors, on the time scale, and can obtain more reliable analysis results. The specific principle of the non-linear correlation pair screening method for influencing factors of photovoltaic power generation output power of the present application will be described below with more embodiments.
[0036] Specifically, as Figure 1 shown, the non-linear correlation pair screening method for influencing factors of photovoltaic power generation output power provided in the first embodiment of the present application includes the following steps:
[0037] Step S10: Obtain the photovoltaic power generation output power data and the corresponding meteorological factor data within a predetermined time period;
[0038] Step S20: Construct a number of time window pairs based on the photovoltaic power generation output power data and the meteorological factor data;
[0039] Step S30: Calculate the non-linear correlation coefficient of the time window pair, and obtain a candidate set of non-linear correlation pairs according to the calculation result;
[0040] Step S40: Perform reduction processing on the candidate set of non-linear correlation pairs to obtain the optimal non-linear correlation pair.
[0041] First, before describing each step, the relevant content of the study on non - linear correlation measurement is introduced. Currently, the classic methods of correlation measurement include Pearson, Spearman, and Kendall correlation coefficients. These correlation coefficients are very powerful for detecting linear and monotonic association relationships. However, even in the case of completely noiseless data, these three coefficients cannot effectively detect the relationships between non - linear factors. Therefore, numerous research works have proposed the maximal correlation coefficient, the correlation coefficient based on the joint cumulative distribution function and ranking, the correlation coefficient based on the kernel, the correlation coefficient based on information theory, the correlation coefficient based on copulas, and the correlation coefficient based on pairwise distances to overcome the defects of classical correlation coefficients in non - linear correlation detection. However, the current correlation coefficients are designed based on the independence of tests and cannot well evaluate the strength of the correlation between variables. Ideally, the correlation coefficient approaches the maximum value if and only if one variable increasingly resembles a noiseless function of another variable. Although the maximal correlation coefficient (MCC, maximal correlation ) especially the maximal information coefficient (MIC, maximal information ) have been widely applied to non - linear correlation analysis, MCC and MIC have been proven to draw incorrect conclusions about non - linear correlation. For example, even when the actual relationship between the variable and the dependent variable is very noisy, the calculation based on the MIC and MCC correlation coefficients can yield a noiseless conclusion. To solve the above problems, the Chatterjee correlation coefficient (CCC, Chatterjee correlation ) has been proposed in the industry. CCC is a simple and interpretable correlation coefficient measurement method that can always accurately estimate the degree of dependence between variables; the CCC measurement is 0 if and only if the variables are independent, and additionally, it is 1 if and only if one variable is a measurable function of another variable; meanwhile, under the independent hypothesis, the CCC coefficient has a simple asymptotic theory and is easy to calculate. Based on the above advantages of the CCC coefficient, the CCC coefficient has been widely applied to non - linear correlation detection work.
[0042] Next, let (X,Y) be a pair of random variables, where Y is not a constant. Let (X 1 ,Y 1 ),…,(X n ,Y n ) be independent and identically distributed, where n≥2, and X i and Y i have no equal values. Rearrange the data as (X (1) ,Y (1) )…,(X (n) ,Y (n) ), such that X(1) ≤ … ≤ X (n) Let r i be the rank of Y (i) , that is, j is the number of Y (j) ≤ Y (i) . Assume there is no autocorrelation between X i . Then the definition of the CCC correlation coefficient is as follows:
[0043]
[0044] CCC(X, Y) is a consistent estimator of a certain correlation measure between the random variables X and Y.
[0045] Assume there is autocorrelation between X i . An increasing sequence is selected by randomly and uniformly breaking the autocorrelation links. Let r i be the rank of Y (i) . Define l (i) to be the number of Y (j) ≥ Y (i) . Then the definition of the CCC correlation coefficient at this time is as follows:
[0046]
[0047] Since photovoltaic power generation is a multi-variable coupled non-linear stochastic process, if a point of the full-field power and meteorological related influencing factors is generated every 5 minutes, there are 105,120 (60÷5×24×365) long time series in a year. Two time series may be correlated at some time intervals, but not over the entire time period. Therefore, when analyzing time series data, an important task is to evaluate the correlation between time series. Based on the strategy of shrinking and expanding the time window to characterize the relationship between the full-field output power of photovoltaic power generation and related meteorological factors. In order to more accurately characterize the non-linear correlation between factors, the Chatterjee correlation coefficient is adopted in this embodiment.
[0048] Among them, the photovoltaic power generation output power data is represented by the time series Y, the meteorological factor data is represented by the time series X, the power data time window is represented as Y(s’, l), and the meteorological data time window is represented as X(s, l). The time window W = (s, l) is a segmentation on the entire continuous time period (1 ≤ s ≤ n - l + 1). The mapping of the time series X = {x 1 , x 2 , …, x n} on the time window W is denoted as X w = {x s , …, x s+l-1}, or denoted as X(s, l). The time series pair (X, Y) = ({x 1 , x2 ,…,x n},{y 1 ,y 2 ,…,y n}(X, Y) is a time series within the same observation time period. The length of the time pair is: n = |(X, Y)|. Assume that two time windows, X(s, l) and Y(s’, l), can form a time window pair (X(s, l), Y(s’, l)). The start times of the time windows are s and s’ respectively, and the time delay of the time relationship is defined as: τ = s’ - s. The time window pair can be represented as a triple <s, l, τ>. For any set of time window pairs, X(s, l) and Y(s’, l), if CCC(X(s, l), Y(s', l)) ≥ θ (θ is the minimum threshold of non - linear correlation), then X(s, l) and Y(s’, l) are considered non - linearly correlated. At the same time, define that the non - linear correlation pair is composed of triples CP = <s, l, τ>, where τ = s’ - s is the delay of the non - linear correlation relationship time. If there exists s i +l i ≤s j and s i +τ i +l i ≤s j +τ i , then the two correlation pairs CP i =<s i ,l i ,τ i >and CP j =<s j ,l j ,τ j >(assuming s i <s j ) are not connected. Otherwise, there will be an overlapping phenomenon for the correlation pairs. If there exists a relationship CCC(CP i )≥CCC(CP j ) for any CP i that has a connection intersection with CP j , the non - linear correlation pair CP i is significant.
[0049] In one or more embodiments, in step S30, a non - linear correlation analysis algorithm based on time window reduction is used for calculation: according to the split time window, the time window pair is split into several sub - window pairs, the non - linear correlation coefficients of each sub - window pair are calculated, and the sub - window pairs with non - linear correlation coefficients greater than the threshold are added to the candidate set of non - linear correlation pairs.
[0050] Exemplarily, the strategy based on time window shrinking is mainly for dealing with the non-linear correlation relationship with sparse distribution. The pseudo-code of Algorithm 1 is shown in Table 1: Non-Linear Correlation search Based on window shrinking Strategy and CCC non-linear correlation coefficient (CCC_NLC S ).
[0051] The purpose of Algorithm 1 is to find the relevant time window pairs in the time series interval [l min , l max . Considering the computational complexity, it is impossible to calculate the non-linear correlation coefficients between every pair of time series. Set a relatively large time region range w e , which is used to filter out the candidate sets that do not meet the requirements in the time series interval [l min , l max . w e is a time encapsulation window. Use the time window w e to split the time series X into several sub-series, denoted as: where X i = X(sp i , w e ), sp i = (i - 1)w e + 1.
[0052] Table 1. Non-Linear Correlation search algorithm based on time window shrinking strategy and CCC coefficient (CCC_NLCS)
[0053]
[0054] First of all, the algorithm assumes that there is at least one non-linearly correlated point pair in the time series pair (X, Y) over the entire time series. Algorithm 1 (step 2) initializes a time window pair <sp i , w e , 0> to be examined and initializes the non-linear coefficient CCC to 0. Then, Algorithm 1 (steps 3 - 6) gradually traverses each time window pair divided by the time window w e . Calculate the non-linear correlation coefficient CCC(X i , Y(sp i + τ, w e ) between the time window pair X i and Y(sp i + τ, w e), if all coefficient values within this time segment are less than the threshold θ, then there are no non - linear correlation pairs within this region, and the time series slides to the next candidate time series X i+1 and time window w e . Otherwise, if there exists CCC(X i , Y(sp i + τ, w e ) ≥ θ, then select the point with the maximum correlation coefficient as the candidate CP(sp i , w e , τ), which means that on the time series X i , there exists a time window pair such that X i and Y(sp i + τ, w e ) are non - linearly correlated under the condition of a time delay of τ.
[0055] This embodiment uses Figure 2 an example to illustrate the screening process of Algorithm 1. Assume that Algorithm 1 sets the minimum threshold θ to 0.6 and the time window w e = 1000. In the time series X 1 = (1001, 1000), since the calculated correlation coefficient is less than the minimum threshold θ, Y(1001, 1000) and Y(1051, 1000) will be excluded. Because Y(951, 1000) has the highest correlation coefficient, CP = <1001, 1000, - 50> will be added to the candidate subset. Similarly, in the time series X 2 , the correlation coefficients of Y(2001, 1000) and Y(2051, 1000) are both greater than the minimum threshold, so CP = <2001, 1000, 0> will be added to the candidate set.
[0056] Obviously, assuming that the set time window w e is too large and the threshold θ, then the generated candidate set CP = <s, l, τ> will inevitably have redundant items. Therefore, the method of calculating based on the non - linear correlation analysis algorithm with time window reduction also includes: splitting the time window into several non - overlapping subset windows; reducing the sub - window pairs in the candidate set of non - linear correlation pairs according to the subset windows; calculating the non - linear correlation coefficients of the remaining sub - window pairs after reduction, and retaining the sub - window pairs with non - linear correlation coefficients greater than the threshold to form the final candidate set of non - linear correlation pairs.
[0057] Exemplarily, a reduction link is designed in Algorithm 1 (step 8). Split the time window w e into several smaller non - overlapping subset windows, and perform reduction operations at both ends of the time window. Assume splitting the time window w eThe minimum scale window is miniL. The nonlinear correlation coefficient between the time series windows X(s, miniL) and Y(s+τ, miniL) is calculated. If it is less than the minimum threshold θ, this area is discarded and moved to the next miniL window. This process is continued until a nonlinear correlation coefficient greater than the threshold θ is found. Similarly, the same simplification strategy is used on the right. After simplifying the time window, the CCC value of the remaining time series is recalculated to form the final candidate item set CP 0 .
[0058] For example, see Figure 3 . Assume that the candidate time window pair is: <2001,1000,0>, and the minimum optimized window unit is miniL=100. On the left side of the window, remove the window pairs that do not meet the minimum threshold, and the first window pair that meets the requirements is: <2301,100>. At the same time, on the right side of the window, delete the window pairs that do not meet the conditions until <2701,100>. Therefore, the remaining part that needs to be considered is <2301,500>. Recalculate the value of CCC to obtain the nonlinear correlation coefficient CCC=0.7>θ, so add <2301,500,0> to the candidate item set CP 0 middle.
[0059] Furthermore, when the nonlinear correlation pairs are densely distributed, if the large window indentation method is still used, two different correlation pairs that meet the conditions may appear in one envelope window w e According to Algorithm 1, due to an envelope window w e Only one correlation pair is retained in the time window, so some correlation pairs will be ignored. The window expansion strategy is opposite to the window reduction strategy, and a small window is used to find candidate item sets. The pseudo code of the algorithm is shown in Algorithm 2: Non-Linear Correlation search Based on Window Extending Strategy and CCC nonlinear correlation coefficient (abbreviated as CCC_NLC E ).
[0060] Table 2. Nonlinear correlation search algorithm based on time window expansion strategy and CCC nonlinear correlation coefficient (CCC_NLCE)
[0061]
[0062] First, Algorithm 2 uses a minimum window of L min Expand the traversal of the time series X (step 2), and record the current time window as X(sp,L min ). Calculate the time series X and Y (sp+τ,Lmin ) The correlation coefficient between, if the current CCC(X, Y(sp + τ, L min )) < θ, then there is no current correlation pair, and move to the next time window X(sp + L min , L min )(Step 14 of the algorithm). Otherwise, record the time window pair <sp, L min , τ> when the CCC value is the largest.
[0063] Different from the time window shrinking strategy, the correlation pairs currently selected using the time window expansion strategy may be a segment of the entire time series. Therefore, at this time, in order to find the complete candidate set, the algorithm expands the time window from both ends. For the time series pair X(sp, L min ) and Y(sp + τ, L min ), the window expansion method is as follows:
[0064] CP 1 = <sp, L min + L max , τ>
[0065] CP 2 = <sp - L max , L min + L max , τ>
[0066] Then, in Steps 7 and 8, in order to obtain a more optimized candidate set, an optimization strategy is used to optimize CP 1 and CP 2 . Compare the values of the non - linear correlation coefficients generated by CP 1 and CP 2 respectively, and add the window pair that generates the larger value to the candidate set CP 0 (Steps 10 and 12).
[0067] Figure 4 shows the window expansion method. Assume that the minimum correlation threshold is set to θ = 0.6, the length constraint of the time window is [200, 200], and the minimum window length of 200 is used as the minimum scale for reduction. Calculate the correlation coefficients of three groups of time window pairs <X(2001, 200), Y(1951, 200)>, <X(2001, 200), Y(2001, 200)>, and <X(2001, 200), Y(2051, 200)>. It is found that the maximum value is obtained at <X(2001, 200), Y(2001, 200)>. Therefore, select the window pair <X(2001, 200), Y(2051, 200)> as the expansion object. Expand to the right to the time window CP 1=<2001, 400, 0>, calculate the correlation coefficient on the corresponding time series Y, and combine with the simplification strategy to obtain the candidate item set CP 1 =<2001, 350, 0>, CCC(CP 1 ) = 0.8. Similarly, expand CP to the left 2 =<1601, 400, 0>, after reducing the window, obtain at CP 2 =<1901, 400, 0>, obtain the maximum correlation coefficient CCC(CP 2 ) = 0.6. Obviously, CCC(CP 1 ) > CCC(CP 2 ), so add CP 2 to the candidate item set CP 0 .
[0068] Furthermore, the candidate item sets CP0 are all time windows that meet the condition threshold, and there are many redundant items. Therefore, it is necessary to reduce the candidate item sets. The purpose of reduction is to find a CP' = <s', l', τ'> with a smaller time boundary for any CP 0 =<s, l, τ>, where s' is the best starting point for the correlation of the two time series, and l' is the best time window length. The reduced CP' = <s', l', τ'> needs to meet the following two conditions: 1). 2). The non-linear correlation coefficient CCC(s', l', τ) reaches the maximum value.
[0069] This problem is transformed into enumerating all possible s and l and screening out the best values that meet the condition threshold. However, enumerating one by one is an NP-hard problem. To reduce the time complexity, based on the nested one-dimensional direct search strategy, the maximum value of each iteration is recorded, and at the same time, the maximum value of the non-linear correlation coefficient is considered in each iteration to determine the best s' and l'. The basic idea of the optimization strategy proposed in this embodiment is that first, in each iteration, the current time window is divided into 3 equal-sized rectangular blocks on the left, middle, and right. Then, gradually exclude the rectangular areas that do not meet the conditions by shrinking the window and record the S point with the maximum CCC value. Finally, after traversing the CCC values of all the rectangular blocks, retain the rectangular block with the maximum CCC. Specifically, initialize s = [lb, ub] = [s, s + l - L min , according to the search strategy, divide [lb, ub] into 3 equal-sized rectangular area blocks Find the best length l' that meets the following conditions:
[0070]
[0071] Then calculate S 1 , the CCC value of the central region. If there exists the following CCC(s 3 , l'(s 1 , l'(s 1 )), τ) < CCC(s 3 , l'(s 3 ), τ), then delete the relevant region with a smaller degree of relevance. Update the iteration and re-partition the relevant regions of S 1 and S 3 until [lb, ub] is less than the set region threshold.
[0072] Use Figure 5 as an example to illustrate how the algorithm determines the optimal starting point s of the time window. In Figure 5 , the CCC curve gives the values of the correlation coefficient between the time series at each time window point in the time interval [2001, 2900]. The purpose of the algorithm optimization is to find the optimal starting point s' with the maximum CCC coefficient in this interval. In the first iteration process, first split the time window into 3 equal-length intervals: S 1 = [2001, 2300], S 2 = [2301, 2600], S 2 = [2601, 2900]. Then, obtain the midpoint of each interval, s 1 , s 2 , s 3 , and calculate their corresponding CCC values. According to Figure 5 , the relationship of non-linear correlation can be obtained: CCC(s 3 , L(s 3 )) < CCC(s 1 , L(s 1 )) < CCC(s 2 , L(s 2 ). Record the highest point s 2 , and delete the lowest point s 3 , then set the examination interval [2001, 2600] for the next iteration. Conduct the second iteration. Similarly, split [2001, 2600] into 3 equal-length intervals: S 4 = [2001, 2200], S 5 = [2201, 2400], S 6 = [2401, 2600]. Calculate the non-linear correlation coefficient values of the corresponding midpoints (s 4 , s 5 , s 6 ) of each interval, and obtain the relationship: CCC(s 4 , L(s 4 )) < CCC(s6 , L(s 6 )) < CCC(s 5 , L(s 5 ))。Compare the highest point s 5 of the second iteration with the highest point s 2 of the first iteration, CCC(s 2 , L(s 2 )) < CCC(s 5 , L(s 5 )),so record the highest point as s 5 . At the same time, delete the lowest point s 4 to form the interval [2201, 2600] for the next iteration. Next, perform the third iteration, split the interval [2201, 2600] into 3 equal-length intervals (S 7 , S 8 , S 9 ), calculate the non-linear correlation coefficient values of the midpoints corresponding to the 3 intervals, and obtain the relationship CCC(s 9 , L(s 9 )) < CCC(s 7 , L(s 7 )) < CCC(s 8 , L(s 8 ))。Therefore, delete the lowest point s 9 and update the traversed area. Compare the highest point s 8 obtained at this time with the current highest point s 5 , CCC(s 8 , L(s 8 )) < CCC(s 5 , L(s 5 )),still record the highest point as s 5 . It can be seen that after three iterations, s 5 is already relatively close to the optimal starting point. The algorithm will continue to iterate until the examined area is less than the set threshold.
[0073] Further, in Figure 6 , this embodiment gives an optimization strategy for finding the optimal time window length l. Assume that we focus on finding the optimal time window length for the s 2 node in the first iteration. First, the time window length is [L min = 300, L max = 600], divide this time length into 3 equal-length windows, and select the midpoints of the 3 regions: l 1 = 550, l 2 = 450, l 3 = 350. Then calculate CCC(s 2 , 550), CCC(s2 , 450), CCC(s 2 , 350), delete the minimum value l 3 , record the maximum value l 2 . Then, iterate the interval [400, 600] again, using Figure 5 the same strategy, and iterate repeatedly until the L value when CCC reaches the maximum is recorded as the desired value.
[0074] Next, perform performance analysis and verification on the method of this embodiment.
[0075] First, use the synthetic non - linearly correlated dataset to verify the effectiveness of the algorithm. The synthetic correlation relationship consists of non - linear relationships such as logarithm, exponent, square, and quartic. The formation method of non - linearly correlated pairs is as follows: 1). Randomly select a time starting point and a custom non - linear relationship type; 2). The non - linearly correlated pair CP (set the parameter value range: l ∈ [400, 850] ∧ τ ∈ [-150, 150]) is generated according to the time point t. This dataset gives 40 pairs of non - linear correlation window pairs. The algorithm uses Precision, Recall, and F - score to verify the effectiveness of the algorithm. According to the results of the dataset, all time series pairs are divided into two types: relevant and irrelevant. For each pair of time series, there are four situations: true positive (the time series pair is calculated as relevant and marked as relevant in the non - linear relationship, denoted as TP), false positive (the time series pair is calculated as relevant and marked as irrelevant in the non - linear relationship, denoted as FP), false negative (the time series pair is calculated as irrelevant and marked as relevant in the non - linear relationship, denoted as FN), and true negative (the time series pair is calculated as irrelevant and marked as irrelevant in the non - linear relationship, denoted as TN).
[0076]
[0077] Table 3. Performance analysis of the improved time window shrinking and expanding algorithm
[0078] Precision Recall F-score Total running time MI_NLCS 0.9849 0.8444 0.9092 138.968 CCC_NLCS 0.9929 0.8532 0.9178 112.754 MI_NLCE 0.9362 0.7642 0.8415 102.283 CCC_NLCE 0.9994 0.8504 0.9189 86.652
[0079] For the analysis of this dataset, the experimental results are shown in Table 3. Analyzing from the algorithm performance, the performance of the improved time window shrinking and expanding algorithm using the CCC non - linear correlation coefficient is relatively high. It should be noted that CCC_NLC E achieves the highest performance in terms of Precision, F - score, and time performance, and CCC_NLC S achieves the best performance in terms of Recall.
[0080] To further analyze the correlation between time series pairs obtained by the improved time window shrinking and expanding algorithms, this study calculated the correlation coefficients between the preset (X, Y), the correlation coefficients obtained by the time window shrinking algorithm based on the MI coefficient (abbreviated as MI_NLC S ), the correlation coefficients obtained by the time window expanding algorithm based on the MI coefficient (abbreviated as MI_NLC E ), the correlation coefficients obtained by the time window shrinking algorithm based on the CCC coefficient (abbreviated as CCC_NLC S ), and the correlation coefficients obtained by the time window expanding algorithm based on the CCC coefficient (abbreviated as CCC_NLC E ), and normalized them. The results are shown in Figure 7 and Figure 8 . In Figure 7 and Figure 8 , only 36 points are shown because MI_NLC S only obtained 36 CP values. For the convenience of comparison, Figure 7 and Figure 8 show the first 36 points extracted by all algorithms. It can be seen from the graph that MI_NLC S , MI_NLC E , CCC_NLC S , CCC_NLC E are all relatively close to the given correlation coefficient values. Since the original time series pairs are labeled as relevant and irrelevant, the calculated correlation coefficients only need to be in the same direction as the given correlation to be detected. From the points in Figure 7 and Figure 8 , it can be seen that MI_NLC S , MI_NLC E , CCC_NLC S , CCC_NLC E and the given correlation coefficient are all in the same direction, which explains why the Precision, Recall, and F-score of all algorithms in Table 3 are all above 80%.
[0081] To verify the effectiveness of the proposed algorithms, the MI_NLC S , MI_NLC E , CCC_NLC S , CCC_NLC E algorithms were used to analyze the influencing factors related to the photovoltaic power of Photovoltaic Power Station 1 and Photovoltaic Power Station 2.
[0082] Among them, the data analyzed for Photovoltaic Power Station 1 is 242,784 pieces of data from January 1, 2021 to April 23, 2023 (one data point is collected every 5 minutes). This data is exported from the photovoltaic power generation system, and the influencing factors for export include: 'Station Irradiance Intensity', 'Station Temperature', 'Tilted Irradiance Intensity', 'Horizontal Irradiance Intensity', 'Ambient Temperature', 'Ambient Humidity', 'Wind Direction', 'Wind Speed', 'Horizontal Diffuse Radiation Intensity', and 'Daily Cumulative Reading of Tilted Irradiance'. The MI_NLC S and MI_NLC E and CCC_NLC S and CCC_NLC E are used to evaluate the non - linear correlation between each influencing factor and the photovoltaic power generation power. The results are shown in Figure 9 . Observing the conclusions obtained from MI_NLC S and MI_NLC E and CCC_NLC S and CCC_NLC E , it can be concluded that there is a strong correlation between 'Station Irradiance Intensity', 'Tilted Irradiance Intensity', and 'Horizontal Diffuse Radiation Intensity' and the overall photovoltaic power generation power. Different from this, CCC_NLC S and CCC_NLC E believe that there is a weak influence relationship for 'Station Temperature', 'Ambient Temperature', and 'Ambient Humidity' (the CCC value is between 0.4 and 0.6), and at the same time, 'Wind Direction' and 'Wind Speed' have the weakest influence (the CCC value is less than 0.3). However, MI_NLC S and MI_NLC E believe that 'Station Temperature', 'Ambient Temperature', and 'Ambient Humidity' are very weak (the MI values are all less than 0.2, and if the MI value is greater than 1, it is considered to have a strong correlation), and at the same time, it is considered that there is no relationship between 'Wind Direction' and 'Wind Speed' and the overall power (the MI is close to 0). In order to more accurately evaluate the correctness of the obtained conclusions, in this embodiment, the relationship between all influencing factors and the overall power is plotted using a scatter plot. It can be seen that there is indeed a certain correlation between 'Station Temperature', 'Ambient Temperature', and 'Ambient Humidity' and the overall power, while the relationship between 'Wind Direction' and 'Wind Speed' and the overall power is indeed weak. Therefore, the conclusions obtained by CCC_NLC S and CCC_NLC E are more credible.
[0083] Secondly, the data of Photovoltaic Power Station 2 are 105,120 pieces of data from January 1, 2021 to December 31, 2021 (one data point is collected every 5 minutes). The influencing factors in this dataset are: 'Station Irradiance', 'Station Temperature', 'Tilted Irradiance', 'Ambient Temperature', 'Ambient Humidity', 'Wind Direction', 'Wind Speed', and 'Horizontal Diffuse Irradiance'. MI_NLC S and MI_NLC E as well as CCC_NLC S and CCC_NLC E are used to evaluate the non - linear correlation between each influencing factor and the photovoltaic power generation. The results are shown in Figure 10 . Observing the conclusions obtained from MI_NLC S and MI_NLC E as well as CCC_NLC S and CCC_NLC E , it can be concluded that there is a strong correlation between 'Station Irradiance', 'Tilted Irradiance', and 'Horizontal Diffuse Irradiance' and the overall photovoltaic power generation. The difference is that CCC_NLC S and CCC_NLC E believe that there is a weak influence relationship for 'Station Temperature', 'Ambient Temperature', and 'Ambient Humidity' (the CCC value is between 0.4 and 0.65), and at the same time, 'Wind Direction' and 'Wind Speed' have the weakest influence (the CCC value is less than 0.4). However, MI_NLC S and MI_NLC E believe that 'Ambient Humidity' is weak (MI is between 0.6 and 0.8), and at the same time, it is considered that there is no relationship between 'Station Temperature', 'Ambient Temperature', 'Wind Direction', and 'Wind Speed' and the overall power (MI is close to 0). Similarly, in this embodiment, scatter plots of all relevant factors and the overall power are drawn, and it can be seen that the conclusions obtained by CCC_NLC S and CCC_NLC E are more in line with the actual situation.
[0084] The specific implementation manners of the present application have been described in detail above. Although some embodiments have been shown and described, those skilled in the art should understand that without departing from the principles and spirit of the present application defined by the claims and their equivalents, these embodiments can be modified and improved, and such modifications and improvements should also be within the protection scope of the present application.
Claims
1. A nonlinear correlation pair screening method for factors affecting photovoltaic power output power, characterized in that: The nonlinear correlation pair screening method comprises: Obtain photovoltaic power output data and corresponding meteorological factor data within a predetermined time period; Constructing a plurality of time window pairs according to the photovoltaic power generation output power data and the meteorological factor data; Calculating the nonlinear correlation coefficient of the time window pair, and obtaining a candidate set of nonlinear correlation pairs according to the calculation result; The candidate set of the non-linear correlation pairs is simplified to obtain the optimal non-linear correlation pair.
2. The nonlinear correlation pair screening method for factors affecting photovoltaic power generation output power according to claim 1, characterized in that: The method for constructing a plurality of time window pairs according to the photovoltaic power generation output power data and the meteorological factor data includes: According to the preset time window length, the photovoltaic power generation output power data and the meteorological factor data are divided into a power data time window Y(s',l) and a meteorological data time window X(s,l) respectively; The constructed time window pair CP =<s,l,τ> , where s and s' represent the start time of the time window, l represents the length of the time window, and τ = s'-s represents the delay between time windows.
3. The nonlinear correlation pair screening method for factors affecting photovoltaic power generation output power according to claim 2, characterized in that: The method of calculating the nonlinear correlation coefficient of the time window pair and obtaining the candidate set of the nonlinear correlation pair according to the calculation result is to calculate based on the nonlinear correlation analysis algorithm of time window reduction, including: Splitting the time window pair into a plurality of sub-window pairs according to the split time window; The nonlinear correlation coefficient of each sub-window pair is calculated, and the sub-window pairs whose nonlinear correlation coefficient is greater than a threshold are added to the candidate set of nonlinear correlation pairs.
4. The nonlinear correlation pair screening method for factors affecting photovoltaic power generation output power according to claim 3, characterized in that: The method for calculating based on the nonlinear correlation analysis algorithm with reduced time window also includes: Splitting the split time window into a plurality of mutually non-intersecting subset windows; Simplifying the sub-window pairs in the candidate set of non-linear correlation pairs according to the subset windows; The nonlinear correlation coefficients of the remaining sub-window pairs after simplification are calculated, and the sub-window pairs whose nonlinear correlation coefficients are greater than the threshold are retained to form the final candidate set of nonlinear correlation pairs.
5. The nonlinear correlation pair screening method for factors affecting photovoltaic power generation output power according to claim 2, characterized in that: The method of simplifying the candidate set of the nonlinear correlation pairs to obtain the optimal nonlinear correlation pair is: determining a time window pair in the candidate set of the nonlinear correlation pairs that contains the optimal start time and the optimal time window length, and the nonlinear correlation coefficient of the time window pair is the largest.
Citation Information
Cited By
Wind power prediction method and system based on wind speed inversion and Chatterje-GPT correction
CN120433203A