A photovoltaic anomaly data cleaning method based on bidirectional quartile and integrated anomaly detection

Through the bidirectional quartile method and integrated anomaly detection method, combined with local anomaly factors and nearest neighbor integrated isolation detectors, the dispersed and stacked anomalies in the photovoltaic power generation power data are cleaned, which solves the problem of inaccurate identification of abnormal data in the prior art and improves the data quality and prediction accuracy of photovoltaic power stations.

CN115935149BActive Publication Date: 2025-08-22HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211552221.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-05
Publication Date
2025-08-22
Estimated Expiration
2042-12-05

AI Technical Summary

Technical Problem

Existing photovoltaic power abnormal data cleaning algorithms are difficult to accurately identify local abnormal data and stacked abnormal data with similar distributions to normal data, resulting in the problems of misdeleting normal data and misdeleting abnormal data.

Method used

Using a method based on bidirectional quartile method and integrated anomaly detection, the dispersed anomaly data is cleaned through longitudinal and transverse quartile methods, and a basic anomaly detector pool composed of local anomaly factor detectors and nearest neighbor integrated isolation detectors are used to filter out effective anomaly detectors in combination with Pearson's correlation coefficient to clean the stacked anomaly data.

Benefits of technology

The dispersed and accumulated abnormal data in the photovoltaic power generation power data are effectively cleaned, the ability to identify local abnormal data is improved, the error deletion of normal data and the missed abnormal data is reduced, and the accuracy of power prediction of photovoltaic power stations is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115935149B_ABST
    Figure CN115935149B_ABST
Patent Text Reader

Abstract

The method for cleaning photovoltaic abnormal data based on bidirectional quartiles and integrated anomaly detection provided by the present invention first collects the actual operation history data of the photovoltaic power station, including the actual operation power generation data of the photovoltaic units and the corresponding meteorological data; then preprocesses the acquired photovoltaic power generation data; then uses the bidirectional quartile method to clean the dispersed abnormal data in the photovoltaic power generation data; and finally uses the integrated anomaly detection method to clean the accumulated abnormal data in the photovoltaic power generation data. The present invention combines the bidirectional quartile method and the integrated anomaly detection method to effectively clean dispersed abnormal data and accumulated abnormal data at the same time. Among them, the integrated anomaly detection method combines the advantages of the local anomaly factor and the nearest neighbor integrated isolation method, strengthens the cleaning effect of the local abnormal data, and can effectively identify abnormal operation data with similar spatial distribution characteristics to normal data and abnormal operation data parallel to the coordinate axis, with great application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of new energy power generation technology, and in particular to a method for cleaning photovoltaic abnormal data based on bidirectional quartile and integrated anomaly detection. Background Art

[0002] my country has significant potential for photovoltaic power generation. Furthermore, within rail transit infrastructure, concentrated spatial resources such as along-line rail lines, within station yards, and on rooftops offer significant potential for renewable energy development. Therefore, the geographical advantages of rail transit can be fully utilized to maximize renewable energy development. While solar energy offers advantages such as abundant resources, high efficiency, and cleanliness, photovoltaic power generation is intermittent, variable, and random, resulting in poor dispatchability and sustained power generation, which in turn impacts the operation and energy management of rail transit systems. Accurately and effectively predicting the output power of photovoltaic power plants is crucial for the safe and stable operation of rail transit systems and their energy management. The accuracy of photovoltaic power plant power forecasts is highly dependent on the quality of historical operating data. However, due to factors such as measurement errors, sensor failures, and curtailed solar power, a significant amount of anomalies can be found in the collected operating data. While operational status data on individual components within a photovoltaic power plant can be helpful in recovering anomaly data, such data is difficult to obtain, with most collected data consisting solely of site-level irradiance and power data. Accurately identifying anomalous photovoltaic power plant operating data in this context is crucial.

[0003] Currently, existing algorithms for cleaning abnormal photovoltaic power generation data fall into two categories: global probabilistic statistical methods and intelligent clustering methods. The first type of method cannot accurately identify datasets containing large amounts of accumulated abnormal data. Clustering methods, on the other hand, typically analyze the spatial distribution characteristics of data samples, but are unable to effectively clean abnormal operating data that has similar spatial distribution characteristics to normal data. Furthermore, focusing on global abnormal data can easily overlook local abnormal data, leading to the inadvertent deletion of normal data and the omission of abnormal data. Summary of the Invention

[0004] In order to solve the technical problem of difficulty in identifying local abnormal data and accumulated abnormal data with a distribution type similar to normal data and parallel to the coordinate axis when processing abnormal photovoltaic power generation data, the present invention provides a method for cleaning photovoltaic abnormal data based on bidirectional quartile and integrated anomaly detection.

[0005] In order to solve the above technical problems, the present invention adopts the following technical method: a method for cleaning photovoltaic abnormal data based on bidirectional quartile and integrated anomaly detection, comprising:

[0006] Step 1: Collect historical data on the actual operation of the photovoltaic power station, including the actual power generation data of the photovoltaic units and corresponding meteorological data, including solar irradiance;

[0007] Step 2: Preprocess the acquired photovoltaic power generation data;

[0008] Step 3: Use the bi-directional quartile method to clean the scattered abnormal data in the photovoltaic power generation data;

[0009] Step 4: Use the integrated anomaly detection method to clean the accumulated anomaly data in the photovoltaic power generation data.

[0010] Furthermore, in step 2, preprocessing the acquired photovoltaic power generation data includes:

[0011] Step 21: Eliminate the original data in step 1 that do not conform to the operating rules of the photovoltaic power station, including data where the photovoltaic power generation is not zero when the solar irradiance is zero; data where the photovoltaic power generation exceeds the rated output power of the photovoltaic cell when the solar irradiance exceeds the rated absorbed irradiance of the photovoltaic panel; and data where the photovoltaic power generation is less than or equal to zero when the solar irradiance is not zero.

[0012] Step 22: After elimination, the remaining photovoltaic power data in the original data are combined into a data set X T , X T =[x1…x i …x n ], according to the following formula (1) for the data set X T Each data point in is normalized;

[0013]

[0014] Where x i For dataset X T The i-th data point in , i∈[1,…,n], data point x i After normalization, we get x i * , μ is the dataset X T The mean value of the data set X T The mean square error of the data set X T After normalization, we get

[0015] Furthermore, in step 3, when using the bidirectional quartile method to clean the photovoltaic power generation power dispersion abnormal data, it includes:

[0016] Step 31: Use the longitudinal quartile method to clean the scattered abnormal data in each irradiance interval: First, the irradiance is 20W / m 2The interval is divided into several irradiance intervals, and then the interquartile range and upper and lower boundaries of abnormal data of photovoltaic power generation in each irradiance interval are calculated. Data outside the two boundaries are considered abnormal data. Among them, the calculation formula of the upper and lower boundaries of abnormal data of photovoltaic power generation in the i-th irradiance interval is as follows:

[0017]

[0018] Where, P li is the upper limit of abnormal data of photovoltaic power generation in the i-th irradiance interval; P ui is the lower boundary of abnormal data of photovoltaic power generation in the i-th irradiance interval; is the first quartile of photovoltaic power generation in the i-th irradiance interval; is the third quartile of photovoltaic power generation in the i-th irradiance interval; is the interquartile range of photovoltaic power generation in the i-th irradiance interval, and

[0019] Step 32: Use the horizontal quartile method to clean up the scattered abnormal data in each power interval: divide the photovoltaic power generation power into several power intervals at intervals of 2% of the rated installed capacity, and then calculate the interquartile range and the upper and lower boundaries of the abnormal data for each power interval. Data outside the two boundaries are considered abnormal data. Among them, the calculation formula for the upper and lower boundaries of the abnormal data of the irradiance in the i-th power interval is as follows:

[0020]

[0021] Where R li is the upper boundary of the abnormal data of irradiance in the i-th power interval; R ui is the lower boundary of the abnormal data of irradiance in the i-th power interval; is the first quartile of the irradiance in the i-th power interval; is the third quartile of the irradiance in the i-th power interval; is the interquartile range of irradiance in the ith power interval, and

[0022] Furthermore, in step 4, when using the integrated anomaly detection method to clean the photovoltaic power accumulation type abnormal data, it includes:

[0023] Step 41: Train t basic anomaly detectors: Assume is the set of real numbers, X trainFor a training set containing several data points, two basic anomaly detectors, local anomaly factor detector and nearest neighbor ensemble isolation detector with different hyperparameters, are used to form a basic anomaly detector pool C = {C1,...,C t}, t is the number of basic anomaly detectors, the training set X train Input into the basic anomaly detector pool to train all basic anomaly detectors and complete parameter debugging of each basic anomaly detector;

[0024] Step 42: Use K nearest neighbor method to obtain data set X T The local nearest neighbor region of all data points in the training set is: randomly select m groups of feature subspaces from d / 2 dimensions to d dimensions, and for each group of feature subspaces selected, find the nearest neighbor region in the feature subspace that is closest to the data point x in the training set. i The k nearest neighboring samples in the Euclidean distance are the samples that appear more than m / 2 times as the data point x i The local nearest neighbor region Ψ i ;

[0025]

[0026] Where x j is the sample contained in the local nearest neighbor area; represents the k nearest neighbor samples obtained by the K nearest neighbor method;

[0027] Step 43: Calculate the local anomaly score matrix of the local nearest neighbor region of each data point: i The local nearest neighbor region Ψ i The k neighboring samples in the t-threshold are detected by t basic anomaly detectors respectively, and t local anomaly score vectors are obtained, which are combined to form the local anomaly score matrix O(Ψ i );

[0028] O(Ψ i )=[C1(Ψ i ),...,C t (Ψ i )] (5)

[0029] Where C t (Ψ i ) represents the local anomaly score vector from the t-th basic anomaly detector;

[0030] Step 44: Generate local pseudo-anomaly labels for the local nearest neighbor region of each data point: Normalize each component vector of the local anomaly score matrix obtained in step 43:

[0031]

[0032] In the formula, the mean variance

[0033] Then calculate the corresponding local pseudo-anomaly label φ i (O(Ψ i ));

[0034]

[0035] Step 45: Detect the local ability of each basic anomaly detector on each data point by using the Pearson correlation coefficient: Calculate the local anomaly score matrix O(Ψ i ) and local pseudo-anomaly label φ i (O(Ψ i ))’s Pearson correlation coefficient, select s basic anomaly detectors with large correlation coefficients from t basic anomaly detectors;

[0036] Step 46: Combine the results of the selected s basic anomaly detectors to calculate the anomaly label score of the data point.

[0037] Furthermore, in step 43, the data point x is detected by a local anomaly factor detector. i The local nearest neighbor region Ψ i The process of abnormal data detection for k neighboring samples within is as follows:

[0038] 1) Calculate the data point x i k distance D k (x i ), assuming X N Represents the data point x i There are N sample points in the k-distance neighborhood of ;

[0039]

[0040] Where, Represents X N The t-th sample point in Represents the distance from the data point x i The k-th most distant data sample;

[0041] 2) Calculate samples To data point x i The reachable distance

[0042]

[0043] 3) Calculate the data point x i Locally reachable density LRD k (x i );

[0044]

[0045] 4) Calculate the data point x i The anomaly score LOF obtained after detection by the local anomaly factor detector k (x i ), as shown below.

[0046]

[0047] Furthermore, in step 43, the nearest neighbor integrated isolation detector is used to isolate the data point x i The local nearest neighbor region Ψ i The process of abnormal data detection for k neighboring samples within is as follows:

[0048] 1) Construct t sets of hyperspheres: from the dataset X T Randomly select data points from the set to form a subsample of size Ψ right Perform a nearest neighbor search for each data point in the equation, that is, find the point closest to itself among the remaining Ψ-1 sample points, and then draw Ψ hyperspheres with itself as the center and the distance to the nearest neighbor as the radius. The mathematical expression is as shown in formula (12). Repeat the above operation t times to obtain t sets of hyperspheres, as shown in formula (13);

[0049] {x:||xc||≤τ(c)} (12)

[0050]

[0051] In the formula, c, x is For any data point in c is the nearest neighbor of data point x, c is the center of the hypersphere B(c), τ(c)=||c-η c || is the radius of the hypersphere B(c), ||xc|| represents the Euclidean distance between x and c;

[0052] 2) Place each data point in the dataset into each set of hyperspheres and calculate the isolation score of all data points: If the data point x i is not contained by any hypersphere, then the data point x i The isolation score is 1; if the data point x i is included in the hypersphere B1 in a set of hyperspheres, and then find the hypersphere B2 closest to the hypersphere B1 in the set of hyperspheres, and record the radius τ(B1) of the hypersphere B1 and the radius τ(B2) of B2 respectively. Then the data point x i The isolation score is As shown below;

[0053]

[0054] 3) Calculate the sum of the isolation scores of each data point in the dataset when placed into different hypersphere sets, and then take the average to obtain the anomaly score of each data point after detection by the nearest neighbor ensemble isolation detector, as shown below;

[0055]

[0056] Where, I j (x i ) is the data point x i The isolation score obtained by placing it into the jth set of hyperspheres;

[0057] 4) The anomaly score of each data point is iteratively calculated and compared with the set threshold. If the anomaly score is greater than or equal to the threshold, the data point is judged as an anomaly point; if the anomaly score is less than the threshold, the data point is judged as a normal point.

[0058] Furthermore, the abnormality score of the data point is set to a threshold of -0.01.

[0059] Preferably, in step 46, the selected s basic anomaly detectors are used to calculate the data point x i If s is 1, the anomaly score obtained by the basic anomaly detector is the anomaly score of the data point x i If s is greater than 1, the maximum or average value of the anomaly scores obtained by s basic anomaly detectors is used as the anomaly label score of the data point x i The abnormal label score of .

[0060] The beneficial effects of the present invention are: combining the bidirectional quartile method and the integrated anomaly detection method, it can effectively clean up scattered abnormal data and accumulated abnormal data at the same time. Compared with the methods based on global probability statistics and distance clustering, the present invention adopts an integrated anomaly detection method that combines the advantages of the local anomaly factor method and the nearest neighbor integrated isolation method, enhances the cleaning effect of local abnormal data, and can effectively identify abnormal operation data with similar spatial distribution characteristics to normal data and abnormal operation data parallel to the coordinate axis, and has wide application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is a flow chart of a method for cleaning photovoltaic abnormal data based on bidirectional quartiles and integrated anomaly detection proposed in the present invention.

[0062] Figure 2 It is a flow chart of the integrated anomaly detection method in the present invention.

[0063] Figure 3 This is a graph showing the results of using the bidirectional quartile method to clean dispersed abnormal data in an embodiment of the present invention.

[0064] Figure 4 This is a graph showing the results of using the integrated anomaly detection method to clean accumulated anomaly data in an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and drawings. The contents mentioned in the embodiments are not intended to limit the present invention.

[0066] like Figure 1 As shown, a photovoltaic abnormal data cleaning method based on bidirectional quartile and integrated anomaly detection includes:

[0067] Step 1: Collect historical data on the actual operation of the photovoltaic power station, including the actual operating power generation data of the photovoltaic units and the corresponding meteorological data. The meteorological data includes solar irradiance. This implementation method uses the actual operating data of a photovoltaic power plant in my country in 2019, a total of 102,905 data items, with a sampling interval of 5 minutes. Table 1 shows a total of 48 operating data items from 8:00 to 11:55 on a certain day during its operation.

[0068] Table 1 Part of the operating data

[0069]

[0070]

[0071] Step 2: Preprocess the acquired photovoltaic power generation data.

[0072] Step 21: Eliminate the original data in step 1 that do not conform to the operating rules of the photovoltaic power station, including data where the photovoltaic power generation is not zero when the solar irradiance is zero; data where the photovoltaic power generation exceeds the rated output power of the photovoltaic cell when the solar irradiance exceeds the rated absorbed irradiance of the photovoltaic panel; and data where the photovoltaic power generation is less than or equal to zero when the solar irradiance is not zero.

[0073] Step 22: After elimination, the remaining photovoltaic power data in the original data are combined into a data set X T , X T =[x1…x i …x n ], according to the following formula (1) for the data set X T Each element in is normalized;

[0074]

[0075] Where x iFor dataset X T The i-th element in , i∈[1,…,n], element x i After normalization, we get x i * , μ is the dataset X T The mean value of the data set X T The mean square error of the data set X T After normalization, we get

[0076] Step 3: Use the bidirectional quartile method to clean the data set preprocessed in step 2 The scattered abnormal data in the cleaned data is as follows Figure 3 shown.

[0077] Step 31: Use the longitudinal quartile method to quantify the data set Clean the scattered abnormal data in each irradiance range: First, the irradiance is set at 20W / m 2 The interval is divided into several irradiance intervals, and then the interquartile range and upper and lower boundaries of abnormal data of photovoltaic power generation in each irradiance interval are calculated. Data outside the two boundaries are considered abnormal data. Among them, the calculation formula of the upper and lower boundaries of abnormal data of photovoltaic power generation in the i-th irradiance interval is as follows:

[0078]

[0079] Where, P li is the upper limit of abnormal data of photovoltaic power generation in the i-th irradiance interval; P ui is the lower boundary of abnormal data of photovoltaic power generation in the i-th irradiance interval; is the first quartile of photovoltaic power generation in the i-th irradiance interval; is the third quartile of photovoltaic power generation in the i-th irradiance interval; is the interquartile range of photovoltaic power generation in the i-th irradiance interval, and

[0080] Step 32: Use horizontal quartile method to analyze the data set Clean the scattered abnormal data in each power interval: divide the photovoltaic power generation power into several power intervals at intervals of 2% of the rated installed capacity, and then calculate the interquartile range and upper and lower boundaries of the abnormal data of the irradiance in each power interval. The data outside the two boundaries are considered abnormal data; the calculation formula for the upper and lower boundaries of the abnormal data of the irradiance in the i-th power interval is as follows:

[0081]

[0082] Where Rli is the upper boundary of the abnormal data of irradiance in the i-th power interval; R ui is the lower boundary of the abnormal data of irradiance in the i-th power interval; is the first quartile of the irradiance in the i-th power interval; is the third quartile of the irradiance in the i-th power interval; is the interquartile range of irradiance in the ith power interval, and

[0083] Step 4: Use the integrated anomaly detection method to clean the dataset X after the preprocessing in step 2 T The accumulated abnormal data in the cleaned data is as follows Figure 4 For detailed steps, see Figure 2 ,include:

[0084] Step 41: Train t basic anomaly detectors: Assume is the set of real numbers, X train For a training set containing several data points, two basic anomaly detectors, local anomaly factor detector and nearest neighbor ensemble isolation detector with different hyperparameters, are used to form a basic anomaly detector pool C = {C1,...,C t}, t is the number of basic anomaly detectors, the training set X train Input into the basic anomaly detector pool to train all basic anomaly detectors and complete the parameter debugging of each basic anomaly detector.

[0085] Here, it is worth mentioning that when each basic anomaly detector performs anomaly data detection on the same data set, the detection results obtained can be combined to obtain the anomaly score matrix O(X train ):

[0086] O(X train )=[C1(X train ),…,C t (X train )](16)

[0087] Where C t (·) represents the anomaly score vector from the t-th base anomaly detector.

[0088] Step 42: Use the K nearest neighbor method to obtain the local nearest neighbor regions of all data points in the data set: randomly select m groups of feature subspaces with dimensions d / 2 to d, and for each group of feature subspaces selected, find the nearest neighbor region in the training set that is closest to the data point x in the feature subspace. i The k nearest neighboring samples in terms of Euclidean distance, x i ∈X T, the samples that appear more than m / 2 times are taken as the data point x i The local nearest neighbor region Ψ i ;

[0089]

[0090] Where x j is the sample contained in the local nearest neighbor area; Represents the k nearest neighbor samples obtained by the K nearest neighbor method.

[0091] Step 43: Calculate the local anomaly score matrix of the local nearest neighbor region of each data point: j The local nearest neighbor region Ψ i The k neighboring samples in the t-threshold are detected by t basic anomaly detectors respectively, and t local anomaly score vectors are obtained, which are combined to form the local anomaly score matrix O(Ψ i );

[0092] O(Ψ i )=[C1(Ψ i ),...,C t (Ψ i )] (5)

[0093] Where C t (Ψ i ) represents the local anomaly score vector from the t-th basic anomaly detector.

[0094] Step 44: Generate local pseudo-anomaly labels for the local nearest neighbor region of each data point: Normalize each component vector of the local anomaly score matrix obtained in step 43:

[0095]

[0096] In the formula, the mean variance

[0097] Then, according to the pseudo anomaly label, the normalized local anomaly score matrix O(Ψ i ) is calculated based on the average or maximum value of the local pseudo-anomaly label φ i (O(Ψ i )), as shown in formula (7):

[0098]

[0099] Step 45: Detect the local ability of each basic anomaly detector on each data point by using the Pearson correlation coefficient: Calculate the local anomaly score matrix O(Ψ i ) and local pseudo-anomaly label φi (O(Ψ i )), and select s basic anomaly detectors with large correlation coefficients from t basic anomaly detectors.

[0100] Step 46: Combine the results of the selected s basic anomaly detectors to calculate the anomaly label score of the data point. Specifically, use the selected s basic anomaly detectors to calculate the anomaly label score of the data point x i If s is 1, the anomaly score obtained by the basic anomaly detector is the anomaly score of the data point x i If s is greater than 1, the maximum or average value of the anomaly scores obtained by s basic anomaly detectors is used as the anomaly label score of the data point x i The abnormal label score of .

[0101] In the above step 43, the data point x is detected by the local anomaly factor detector. i The local nearest neighbor region Ψ i The process of abnormal data detection for k neighboring samples within is as follows:

[0102] 1) Calculate the data point x i k distance D k (x i ), assuming X N Represents the data point x i There are N sample points in the k-distance neighborhood of ;

[0103]

[0104] Where, Represents X N The t-th sample point in Represents the distance from the data point x i The k-th most distant data sample;

[0105] 2) Calculate samples To data point x i The reachable distance

[0106]

[0107] 3) Calculate the data point x i Locally reachable density LRD k (x i );

[0108]

[0109] 4) Calculate the data point x i The anomaly score LOF obtained after detection by the local anomaly factor detector k(x i ), as shown below.

[0110]

[0111] Where LOF k (x i ) is close to 1, then the data point x i The more likely it is normal data, the greater its value is than 1. i The more likely it is an outlier.

[0112] In the aforementioned step 43, the nearest neighbor integrated isolation detector is used to isolate the data point x i The local nearest neighbor region Ψ i The process of abnormal data detection for k neighboring samples within is as follows:

[0113] 1) Construct t sets of hyperspheres: from the dataset X T Randomly select data points from the set to form a subsample of size Ψ right Perform a nearest neighbor search for each data point in the equation, that is, find the point closest to itself among the remaining Ψ-1 sample points, and then draw Ψ hyperspheres with itself as the center and the distance to the nearest neighbor as the radius. The mathematical expression is as shown in formula (12). Repeat the above operation t times to obtain t sets of hyperspheres, as shown in formula (13);

[0114] {x:||xc||≤τ(c)} (12)

[0115]

[0116] In the formula, c, x is For any data point in c is the nearest neighbor of data point x, c is the center of the hypersphere B(c), τ(c)=||c-η c || is the radius of the hypersphere B(c), ||xc|| represents the Euclidean distance between x and c;

[0117] 2) Place each data point in the dataset into each set of hyperspheres and calculate the isolation score of all data points: If the data point x i is not contained by any hypersphere, then the data point x i The isolation score is 1; if the data point x i is included in the hypersphere B1 in a set of hyperspheres, and then find the hypersphere B2 closest to the hypersphere B1 in the set of hyperspheres, and record the radius τ(B1) of the hypersphere B1 and the radius τ(B2) of B2 respectively. Then the data point x i The isolation score is As shown below;

[0118]

[0119] 3) Calculate the sum of the isolation scores of each data point in the dataset when placed into different hypersphere sets, and then take the average to obtain the anomaly score of each data point after detection by the nearest neighbor ensemble isolation detector, as shown below;

[0120]

[0121] Where, I j (x i ) is the data point x i The isolation score obtained by placing it into the jth set of hyperspheres;

[0122] 4) Iteratively calculate and compare the anomaly score of each data point with a set threshold. If the anomaly score is greater than or equal to the threshold, the data point is judged as an outlier; if the anomaly score is less than the threshold, the data point is judged as a normal point. Here, preferably, the anomaly score threshold of the data point is set to -0.01.

[0123] In summary, the abnormal data cleaning method based on bidirectional quartile and integrated anomaly detection provided by the present invention can effectively clean the scattered abnormal data and accumulated abnormal data in photovoltaic power generation data at the same time, enhance the cleaning effect of local abnormal data, and can effectively identify abnormal operation data with similar spatial distribution characteristics to normal data and abnormal operation data parallel to the coordinate axis. It has strong versatility and is suitable for most photovoltaic abnormal data processing occasions.

[0124] The above embodiments are preferred implementation schemes of the present invention. In addition, the present invention can also be implemented in other ways. Any obvious replacement without departing from the concept of the present technical solution is within the scope of protection of the present invention.

[0125] In order to make it easier for ordinary technicians in this field to understand the improvements of the present invention over the prior art, some drawings and descriptions of the present invention have been simplified, and for the sake of clarity, some other elements are omitted in this application document. Ordinary technicians in this field should realize that these omitted elements may also constitute the content of the present invention.

Claims

1. A method for cleaning photovoltaic abnormal data based on bidirectional quartile and integrated anomaly detection, characterized in that: include: Step 1: Collect historical data on the actual operation of the photovoltaic power station, including the actual power generation data of the photovoltaic units and corresponding meteorological data, including solar irradiance; Step 2: Preprocess the acquired photovoltaic power generation data; Step 3: Use the bi-directional quartile method to clean the scattered abnormal data in the photovoltaic power generation data; Step 31: Use the longitudinal quartile method to clean the scattered abnormal data in each irradiance interval: First, the irradiance is 20W / m 2 The interval is divided into several irradiance intervals, and then the interquartile range and upper and lower boundaries of abnormal data of photovoltaic power generation in each irradiance interval are calculated. Data outside the two boundaries are regarded as abnormal data. Step 32: Use the horizontal quartile method to clean up the scattered abnormal data in each power interval: divide the photovoltaic power generation power into several power intervals at intervals of 2% of the rated installed capacity, and then calculate the interquartile range of the irradiance in each power interval and the upper and lower boundaries of the abnormal data. Data outside the two boundaries are considered abnormal data; Step 4: Use the integrated anomaly detection method to clean the accumulated anomaly data in the photovoltaic power generation data; Step 41: Train t basic anomaly detectors; Step 42: Use K nearest neighbor method to obtain data set X T The local nearest neighbor region of all data points in ; Step 43: Calculate the local anomaly score matrix of the local nearest neighbor region of each data point: i The local nearest neighbor region Ψ i The k neighboring samples in the t-threshold are detected by t basic anomaly detectors respectively, and t local anomaly score vectors are obtained, which are combined to form the local anomaly score matrix O(Ψ i ); O(Ψ i )=[C1(Ψ i ),...,C t (P i )] (5) Where C t (Ψ i ) represents the local anomaly score vector from the t-th basic anomaly detector; Step 44: Generate a local pseudo-anomaly label for the local nearest neighbor region of each data point; then calculate the corresponding local pseudo-anomaly label φ i (O(Ψ i )); Step 45: Detect the local ability of each basic anomaly detector on each data point by using the Pearson correlation coefficient: Calculate the local anomaly score matrix O(Ψ i ) and local pseudo-anomaly label φ i (O(Ψ i ))’s Pearson correlation coefficient, select s basic anomaly detectors with large correlation coefficients from t basic anomaly detectors; Step 46: Combine the results of the selected s basic anomaly detectors to calculate the anomaly label score of the data point.

2. The photovoltaic abnormal data cleaning method based on bidirectional quartile and integrated anomaly detection according to claim 1 is characterized by: In step 2, preprocessing the acquired photovoltaic power generation data includes: Step 21: Eliminate the original data in step 1 that do not conform to the operating rules of the photovoltaic power station, including data where the photovoltaic power generation is not zero when the solar irradiance is zero; data where the photovoltaic power generation exceeds the rated output power of the photovoltaic cell when the solar irradiance exceeds the rated absorbed irradiance of the photovoltaic panel; and data where the photovoltaic power generation is less than or equal to zero when the solar irradiance is not zero. Step 22: After elimination, the remaining photovoltaic power data in the original data are combined into a data set X T , X T =[x1…x i …x n ], according to the following formula (1) for the data set X T Each data point in is normalized; Where x i For dataset X T The i-th data point in , i∈[1,…,n], data point x i After normalization, we get x i * , μ is the dataset X T The mean value of the data set X T The mean square error of the data set X T After normalization, we get 3. The photovoltaic abnormal data cleaning method based on bidirectional quartile and integrated anomaly detection according to claim 2 is characterized by: In step 31, the calculation formulas for the upper and lower bounds of abnormal data of photovoltaic power generation in the i-th irradiance interval are as follows: Where, P li is the upper limit of abnormal data of photovoltaic power generation in the i-th irradiance interval; P ui is the lower boundary of abnormal data of photovoltaic power generation in the i-th irradiance interval; is the first quartile of photovoltaic power generation in the i-th irradiance interval; is the third quartile of photovoltaic power generation in the i-th irradiance interval; is the interquartile range of photovoltaic power generation in the i-th irradiance interval, and In step 32, the calculation formulas for the upper and lower limits of abnormal irradiance data in the i-th power interval are as follows: Where R li is the upper boundary of the abnormal data of irradiance in the i-th power interval; R ui is the lower boundary of the abnormal data of irradiance in the i-th power interval; is the first quartile of the irradiance in the i-th power interval; is the third quartile of the irradiance in the i-th power interval; is the interquartile range of irradiance in the ith power interval, and 4. The photovoltaic abnormal data cleaning method based on bidirectional quartile and integrated anomaly detection according to claim 1, 2 or 3, characterized in that: In step 41, when training t basic anomaly detectors: assuming is the set of real numbers, X train For a training set containing several data points, two basic anomaly detectors, local anomaly factor detector and nearest neighbor ensemble isolation detector with different hyperparameters, are used to form a basic anomaly detector pool C = {C1,...,C t }, t is the number of basic anomaly detectors, the training set X train Input into the basic anomaly detector pool to train all basic anomaly detectors and complete parameter debugging of each basic anomaly detector; In step 42, the K nearest neighbor method is used to obtain the data set X T When the local nearest neighbor region of all data points in the training set is: randomly select m groups of feature subspaces from d / 2 dimensions to d dimensions, and for each group of feature subspaces selected, find the nearest neighbor region in the feature subspace that is closest to the data point x in the training set. i The k nearest neighboring samples in the Euclidean distance are the samples that appear more than m / 2 times as the data point x i The local nearest neighbor region Ψ i ; Where x j is the sample contained in the local nearest neighbor area; represents the k nearest neighbor samples obtained by the K nearest neighbor method; In step 44, when generating the local pseudo-anomaly labels for the local nearest neighbor region of each data point, each component vector of the local anomaly score matrix obtained in step 43 is normalized: In the formula, the mean variance 5. The photovoltaic abnormal data cleaning method based on bidirectional quartile and integrated anomaly detection according to claim 4 is characterized in that: In step 43, the data point x is detected by a local anomaly factor detector. i The local nearest neighbor region Ψ i The process of abnormal data detection for k neighboring samples within is as follows: 1) Calculate the data point x i k distance D k (x i ), assuming X N Represents the data point x i There are N sample points in the k-distance neighborhood of ; Where, Represents X N The t-th sample point in Represents the distance from the data point x i The k-th most distant data sample; 2) Calculate samples To data point x i The reachable distance 3) Calculate the data point x i Locally reachable density LRD k (x i ); 4) Calculate the data point x i The anomaly score LOF obtained after detection by the local anomaly factor detector k (x i ), as shown below.

6. The photovoltaic abnormal data cleaning method based on bidirectional quartile and integrated anomaly detection according to claim 5 is characterized by: In step 43, the nearest neighbor integrated isolation detector is used to isolate the data point x i The local nearest neighbor region Ψ i The process of abnormal data detection for k neighboring samples within is as follows: 1) Construct t sets of hyperspheres: from the dataset X T Randomly select data points from the set to form a subsample of size Ψ right Perform a nearest neighbor search for each data point in the equation, that is, find the point closest to itself among the remaining Ψ-1 sample points, and then draw Ψ hyperspheres with itself as the center and the distance to the nearest neighbor as the radius. The mathematical expression is as shown in formula (12). Repeat the above operation t times to obtain t sets of hyperspheres, as shown in formula (13); {x:||xc||≤τ(c)} (12) Where, x is For any data point in c is the nearest neighbor of data point x, c is the center of the hypersphere B(c), τ(c)=||c-η c || is the radius of the hypersphere B(c), ||xc|| represents the Euclidean distance between x and c; 2) Place each data point in the dataset into each set of hyperspheres and calculate the isolation score of all data points: If the data point x i is not contained by any hypersphere, then the data point x i The isolation score is 1; if the data point x i is included in the hypersphere B1 in a set of hyperspheres, and then find the hypersphere B2 closest to the hypersphere B1 in the set of hyperspheres, and record the radius τ(B1) of the hypersphere B1 and the radius τ(B2) of B2 respectively. Then the data point x i The isolation score is As shown below; 3) Calculate the sum of the isolation scores of each data point in the dataset when placed into different hypersphere sets, and then take the average to obtain the anomaly score of each data point after detection by the nearest neighbor ensemble isolation detector, as shown below; Where, I j (x i ) is the data point x i The isolation score obtained by placing it into the jth set of hyperspheres; 4) The anomaly score of each data point is iteratively calculated and compared with the set threshold. If the anomaly score is greater than or equal to the threshold, the data point is judged as an anomaly point; if the anomaly score is less than the threshold, the data point is judged as a normal point.

7. The photovoltaic abnormal data cleaning method based on bidirectional quartile and integrated anomaly detection according to claim 6 is characterized by: The anomaly score of the data point was set to a threshold of -0.

01.

8. The photovoltaic abnormal data cleaning method based on bidirectional quartile and integrated anomaly detection according to claim 7 is characterized in that: In step 46, the selected s basic anomaly detectors are used to calculate the data point x i If s is 1, the anomaly score obtained by the basic anomaly detector is the anomaly score of the data point x i If s is greater than 1, the maximum or average value of the anomaly scores obtained by s basic anomaly detectors is used as the anomaly label score of the data point x i The abnormal label score of .

Citation Information

Patent Citations

  • Method for cleaning abnormal data of photovoltaic power station

    CN114090559A

  • Wind power abnormal data detection method based on quartile and improved isolated nearest neighbor

    CN114841275A