Photovoltaic power station abandoned light data identification method, device and storage medium
By preprocessing the historical power generation and irradiation data of photovoltaic power stations and performing secondary clustering using the DBSCAN algorithm, the problems of automation and accuracy in identifying abandoned power data in photovoltaic power stations were solved, thereby improving the accuracy of power generation forecasts and reducing power dispatch costs.
Patent Information
- Application Number
- CN202210420774.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-04-21
AI Technical Summary
Existing technologies make it difficult to efficiently and automatically identify abandoned power data from photovoltaic power stations, resulting in low accuracy in photovoltaic power generation predictions and a high reliance on manual intervention.
By obtaining the historical power generation and irradiation data of photovoltaic power stations, the sample area is divided after preprocessing, the 3-sigma rule is combined to screen out outliers, and the secondary clustering method of the DBSCAN algorithm is used to identify the abandoned light data.
It has achieved the automated and accurate identification of abandoned light data in photovoltaic power stations, improved the accuracy of power generation predictions, and reduced power dispatch costs.
Smart Images

Figure CN114936590B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, device and storage medium for identifying abandoned light data in a photovoltaic power station, and belongs to the technical field of power systems. Background Art
[0002] In the context of large-scale renewable energy power generation connected to the grid, the problem of abandoned solar power caused by human factors (restrictions on photovoltaic grid connection), natural factors (dust accumulation and snow covering of photovoltaic panels, etc.) or device failure not only causes waste of clean energy, but also greatly damages the regularity of historical power generation data of photovoltaic power stations, thereby seriously affecting the analysis and prediction of subsequent photovoltaic power generation data. Therefore, analyzing and identifying abandoned solar power data is of great significance for improving the accuracy of photovoltaic power generation prediction, providing accurate boundary data for scheduling plans and spot markets, and reducing power scheduling costs.
[0003] Currently, there are few studies on the problem of identifying abandoned solar power data. Two methods are usually used: (1) using the 3-sigma criterion to analyze the fluctuation law of photovoltaic power generation itself, and treating outliers as abnormal data; (2) using copula theory to fit the boundary relationship curve of irradiance-photovoltaic power generation, and treating sample points outside the boundary as abnormal data.
[0004] Methods based on the 3-sigma criterion for identifying abnormal PV power generation data only consider the regularity of generated power itself, ignoring the impact of external factors on generated power. Furthermore, PV power generation patterns are affected by meteorological factors, and power generation only approximates a normal distribution during clear skies. Using only the 3-sigma criterion to identify PV power generation data can easily lead to inaccurate identification. Furthermore, the process of deriving the upper and lower quantile values corresponding to the conditional probability distribution of PV power based on copula theory relies heavily on the quality of the original sample. When a high proportion of abnormal data is present in the sample, it is necessary to manually filter out some "suspected" abnormal samples based on experience, otherwise this will significantly interfere with the copula function model fitting. In summary, existing methods for identifying curtailed PV power generation data rely heavily on empirical rules and even require manual intervention, making them difficult to meet the requirements of fully automated and efficient PV power generation data identification. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a method, device and storage medium for identifying abandoned photovoltaic power station data, which can automatically and efficiently identify abandoned photovoltaic power station data.
[0006] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0007] In a first aspect, the present invention provides a method for identifying abandoned light data in a photovoltaic power station, comprising:
[0008] Obtain historical power generation data and corresponding irradiation data of the photovoltaic power station and generate a sample point set;
[0009] Preprocess the sample point set;
[0010] Divide the preprocessed sample point set into multiple sample areas according to the irradiation data;
[0011] Screen out abnormal data in the sample area according to the 3-sigma rule;
[0012] The secondary clustering method based on the DBSCAN algorithm is used to perform cluster analysis on each sample area after the abnormal data are screened out to obtain the abandoned light data.
[0013] Optionally, obtaining historical power generation data and corresponding irradiation data of a photovoltaic power station and generating a sample point set includes:
[0014] Data acquisition:
[0015] Obtaining photovoltaic power station model data from the power grid model, the photovoltaic power station model data including the photovoltaic power station ID, installed capacity, and geographic information; obtaining historical power generation data of the photovoltaic power station at a preset quantity granularity based on the photovoltaic power station ID; and obtaining irradiation data corresponding to the historical power generation data based on the geographic information;
[0016] Generate sample point set :
[0017]
[0018] in, For the moment The sample points, , and Separate moments Irradiation data and power generation data, is the number of sample points.
[0019] Optionally, the preprocessing of the sample point set is:
[0020] For any sample point , power generation data Lower than the installed capacity of photovoltaic power plants or irradiation data Lower than , then the sample point From the sample point set Remove from .
[0021] Optionally, dividing the preprocessed sample point set into a plurality of sample areas according to the irradiation data includes:
[0022] Arrange the sample points in the preprocessed sample point set in ascending order according to the irradiation data;
[0023] Divide the irradiation data into multiple equal intervals according to the maximum and minimum values;
[0024] Generate a sample area based on the sample points in each interval.
[0025] Optionally, the filtering out abnormal data in the sample area according to the 3-sigma rule includes:
[0026] Calculate the average value of the sample points in the sample area and standard deviation , will satisfy Sample points Typical outliers were identified and screened out.
[0027] Optionally, performing cluster analysis on the sample area after abnormal data is screened out according to a secondary clustering method based on the DBSCAN algorithm to obtain abandoned light data includes:
[0028] The DBSCAN algorithm is used to cluster the sample points in the sample area after the abnormal data are screened out to obtain discrete samples and several sample clusters;
[0029] Calculate the cluster center of each sample cluster and record it as , is the number of sample clusters, For the The cluster centers of the sample clusters;
[0030] The sample cluster with the largest number of sample points is taken as the benchmark cluster, and the cluster center of the benchmark cluster is recorded as ;
[0031] Calculate cluster centers Cluster centers outside the cluster to cluster centers Distance: ;
[0032] The distance With preset threshold For comparison, if , then the cluster center The sample points in the corresponding sample cluster are identified as abandoned light data.
[0033] In a second aspect, the present invention provides a device for identifying abandoned light data in a photovoltaic power station, the device comprising:
[0034] The data acquisition module is used to obtain the historical power generation data and corresponding irradiation data of the photovoltaic power station and generate a sample point set;
[0035] A preprocessing module is used to preprocess the sample point set;
[0036] A data partitioning module is used to divide the pre-processed sample point set into multiple sample areas according to the irradiation data;
[0037] Data screening module, used to screen out abnormal data in the sample area according to the 3-sigma rule;
[0038] The data identification module is used to perform cluster analysis on each sample area after the abnormal data is screened out according to the secondary clustering method based on the DBSCAN algorithm to obtain the abandoned light data.
[0039] In a third aspect, the present invention provides a photovoltaic power station abandoned light data identification device, characterized in that it includes a processor and a storage medium;
[0040] The storage medium is used to store instructions;
[0041] The processor is configured to operate according to the instructions to execute the steps of the above method.
[0042] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the program implements the steps of the above method when executed by a processor.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] The present invention provides a method, device and storage medium for identifying abandoned power data in photovoltaic power stations. The method preprocesses the historical power generation data and corresponding irradiation data of the photovoltaic power station to remove dense zero values and zero drift values; typical outliers are screened out based on the 3-sigma rule; and finally, cluster analysis is performed based on the secondary clustering method of the DBSCAN algorithm to obtain abandoned power data. Compared with traditional methods, this method does not rely on manual labor, realizes automated identification, and is efficient and accurate, providing more reasonable historical data for photovoltaic power generation prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flow chart of a method for identifying abandoned power data in a photovoltaic power station provided in the first embodiment of the present invention;
[0046] Figure 2 This is an irradiance-power scatter plot after identification of abandoned light data provided in the first embodiment of the present invention;
[0047] Figure 3 This is a comparison chart of irradiance and power curves after identification of abandoned light data provided in the first embodiment of the present invention;
[0048] Figure 4This is a comparison chart of predicted data and measured data of a photovoltaic power station before and after identification of abandoned light data provided by the first embodiment of the present invention. DETAILED DESCRIPTION
[0049] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0050] Example 1:
[0051] like Figure 1 As shown, the present invention provides a method for identifying abandoned light data in a photovoltaic power station, comprising the following steps:
[0052] 1. Obtain historical power generation data and corresponding irradiation data of the photovoltaic power station and generate a sample point set;
[0053] 1.1 Data Acquisition
[0054] Photovoltaic power station model data is obtained from the power grid model. The photovoltaic power station model data includes the photovoltaic power station ID, installed capacity and geographic information. The historical power generation data of the photovoltaic power station is obtained at a preset quantity granularity based on the photovoltaic power station ID. The irradiation data corresponding to the historical power generation data is obtained based on the geographic information. In this embodiment, the data granularity is set to 15 minutes.
[0055] 1.2. Generate sample point set :
[0056]
[0057] in, For the moment The sample points, , and Separate moments Irradiation data and power generation data, is the number of sample points.
[0058] 2. Preprocess the sample point set;
[0059] For any sample point , power generation data Lower than the installed capacity of photovoltaic power plants or irradiation data Lower than , then the sample point From the sample point set Remove from .
[0060] Due to the particularity that photovoltaic power generation is determined by sunlight, there are dense zero values and zero drift values in the data, which affects the distribution pattern of the data. Therefore, preprocessing is required to remove zero values and zero drift values.
[0061] 3. Divide the preprocessed sample point set into multiple sample areas according to the irradiation data;
[0062] 3.1. Arrange the sample points in the preprocessed sample point set in ascending order according to the irradiation data;
[0063] 3.2. Divide the irradiation data into multiple equal intervals according to the maximum and minimum values;
[0064] 3.3. Generate a sample area based on the sample points in each interval.
[0065] 4. Screen out abnormal data in the sample area according to the 3-sigma rule;
[0066] Calculate the average value of the sample points in the sample area and standard deviation , will satisfy Sample points Typical outliers were identified and screened out.
[0067] 5. Perform cluster analysis on each sample area after abnormal data is screened out using the secondary clustering method based on the DBSCAN algorithm to obtain abandoned light data.
[0068] 5.1. Use the DBSCAN algorithm to cluster the sample points in the sample area after the abnormal data are screened out to obtain discrete samples and several sample clusters;
[0069] 5.2. Calculate the cluster center of each sample cluster and record it as , is the number of sample clusters, For the The cluster centers of the sample clusters;
[0070] 5.3. Take the sample cluster with the largest number of sample points as the benchmark cluster, and record the cluster center of the benchmark cluster as ;
[0071] 5.4. Calculating Cluster Centers Cluster centers outside the cluster to cluster centers Distance: ;
[0072] 5.5、Distance With preset threshold For comparison, if , then the cluster center The sample points in the corresponding sample cluster are identified as abandoned light data. In this embodiment, the preset threshold .
[0073] The application effect of this embodiment is as follows:
[0074] The photovoltaic power generation data of a photovoltaic power station from January to July 2018 was selected and the above method was used to identify abnormal data of abandoned light. The results are as follows: Figure 2 As shown, the horizontal axis is irradiance, the vertical axis is power generation, the dot sample points are normal data samples, and the plus sign sample points are abnormal data of abandoned light. Map the dot sample points to the power generation curve, as shown Figure 3 As shown, the solid line is the irradiance curve, and the dotted line is the power generation curve. It can be seen that the sample points marked as abnormal data of abandoned light by the method of the present invention are consistent with the actual abandoned light sample points.
[0075] The historical sample data processed by the method of the present invention and the unprocessed historical sample data are used to predict the power generation of the photovoltaic power station in August 2018. The results are as follows: Figure 4 The dotted line is the actual power generation curve, the solid dotted line is the power generation curve predicted based on unprocessed historical samples, and the solid line is the power generation curve predicted based on historical sample data processed by the method of the present invention. It can be seen that the solid line is significantly closer to the actual power generation curve than the solid dotted line.
[0076] The average prediction accuracy of the two prediction results is calculated as follows:
[0077]
[0078] in, is the number of prediction results, is the installed capacity, is the predicted data of the i-th point, is the measured data of the i-th point.
[0079] According to statistics, the average accuracy of the power generation predicted based on unprocessed historical samples is 93.26%, and the average accuracy of the power generation predicted based on historical samples processed by the method of the present invention is 95.11%, with the average accuracy increased by 1.85%. The method of the present invention has significant effects and broad application prospects in identifying abnormal data of abandoned light and improving the accuracy of photovoltaic power generation prediction.
[0080] Example 2:
[0081] An embodiment of the present invention provides a device for identifying abandoned light data in a photovoltaic power station, the device comprising:
[0082] The data acquisition module is used to obtain the historical power generation data and corresponding irradiation data of the photovoltaic power station and generate a sample point set;
[0083] A preprocessing module is used to preprocess the sample point set;
[0084] A data partitioning module is used to divide the pre-processed sample point set into multiple sample areas according to the irradiation data;
[0085] Data screening module, used to screen out abnormal data in the sample area according to the 3-sigma rule;
[0086] The data identification module is used to perform cluster analysis on each sample area after the abnormal data is screened out according to the secondary clustering method based on the DBSCAN algorithm to obtain the abandoned light data.
[0087] Example 3:
[0088] Based on the first embodiment, the present invention provides a photovoltaic power station abandoned light data identification device, characterized by including a processor and a storage medium;
[0089] The storage medium is used to store instructions;
[0090] The processor is configured to operate according to the instructions to execute the steps of the above method.
[0091] Example 4:
[0092] Based on the first embodiment, the embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program implements the steps of the above method when executed by a processor.
[0093] The purpose of the present invention is to effectively identify and clean the abandoned light data generated by photovoltaic power stations due to reasons such as data acquisition device failure, human factors, and natural factors, starting from the data distribution characteristics of the actual power curve of photovoltaic power generation and the meteorological data of the area where photovoltaic power generation is located. This method does not rely on the specific physical properties of photovoltaic components and the historical abandoned light information maintained by humans. It realizes the automatic identification of abnormal data of various types of photovoltaic power stations based on unsupervised algorithms, provides more reasonable data samples for the later photovoltaic power generation prediction and regional photovoltaic power generation prediction, improves the accuracy of photovoltaic power generation prediction, and reduces the cost of power dispatch. In the context of building a new power system with new energy as the main body, it has practical engineering significance.
[0094] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0095] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0096] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0098] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for identifying abandoned light data in a photovoltaic power station, characterized in that: include: Obtain historical power generation data and corresponding irradiation data of a photovoltaic power station and generate a sample point set, including: data acquisition: obtaining photovoltaic power station model data from a power grid model, the photovoltaic power station model data including the photovoltaic power station ID, installed capacity and geographic information; obtaining the historical power generation data of the photovoltaic power station at a preset quantity granularity according to the photovoltaic power station ID; obtaining the irradiation data corresponding to the historical power generation data according to the geographic information; generating a sample point set : ; in, For the moment The sample points, , and Separate moments Irradiation data and power generation data, is the number of sample points; Preprocess the sample point set, including for any sample point , power generation data Lower than the installed capacity of photovoltaic power plants or irradiation data Lower than , then the sample point From the sample point set Remove Arrange the sample points in the preprocessed sample point set in ascending order according to the irradiation data; Divide the irradiation data into multiple equal intervals according to the maximum and minimum values; Generate a sample area based on the sample points in each interval; filter out abnormal data in the sample area based on the 3-sigma rule; The secondary clustering method based on the DBSCAN algorithm is used to perform cluster analysis on each sample area after the abnormal data are screened out to obtain the abandoned light data.
2. The method for identifying abandoned photovoltaic power station data according to claim 1, characterized in that: The method of filtering out abnormal data in the sample area according to the 3-sigma rule includes: Calculate the average value of the sample points in the sample area and standard deviation , will satisfy Sample points Typical outliers were identified and screened out.
3. The method for identifying abandoned photovoltaic power station data according to claim 1, characterized in that: The method of performing cluster analysis on the sample area after the abnormal data is screened out according to the secondary clustering method based on the DBSCAN algorithm to obtain the abandoned light data includes: The DBSCAN algorithm is used to cluster the sample points in the sample area after the abnormal data are screened out to obtain discrete samples and several sample clusters; Calculate the cluster center of each sample cluster and record it as , is the number of sample clusters, For the The cluster centers of the sample clusters; The sample cluster with the largest number of sample points is taken as the benchmark cluster, and the cluster center of the benchmark cluster is recorded as ; Calculate cluster centers Cluster centers outside the cluster to cluster centers Distance: ; The distance With preset threshold For comparison, if , then the cluster center The sample points in the corresponding sample cluster are identified as abandoned light data.
4. A photovoltaic power station abandoned light data identification device, characterized in that: The device comprises: The data acquisition module is used to obtain the historical power generation data and corresponding irradiation data of the photovoltaic power station and generate a sample point set, including: data acquisition: obtaining photovoltaic power station model data from the power grid model, the photovoltaic power station model data including the photovoltaic power station ID, installed capacity and geographic information; obtaining the historical power generation data of the photovoltaic power station with a preset quantity granularity according to the photovoltaic power station ID; obtaining the irradiation data corresponding to the historical power generation data according to the geographic information; generating a sample point set : ; in, For the moment The sample points, , and Separate moments Irradiation data and power generation data, is the number of sample points; The preprocessing module is used to preprocess the sample point set, including any sample point , power generation data Lower than the installed capacity of photovoltaic power plants or irradiation data Lower than , then the sample point From the sample point set Remove The data partitioning module is used to arrange the sample points in the pre-processed sample point set in ascending order according to the irradiation data; divide the sample points into multiple equal intervals according to the maximum and minimum values of the irradiation data; and generate a sample area according to the sample points in each interval; Data screening module, used to screen out abnormal data in the sample area according to the 3-sigma rule; The data identification module is used to perform cluster analysis on each sample area after the abnormal data is screened out according to the secondary clustering method based on the DBSCAN algorithm to obtain the abandoned light data.
5. A photovoltaic power station abandoned light data identification device, characterized in that: including processors and storage media; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Phase identification method based on clustering algorithm and network search
CN111178679A