A photovoltaic power station power prediction method based on cluster division and related equipment
By using cluster partitioning and machine learning methods, combined with data on the geographical location of photovoltaic power plants and the characteristics of the power grid, benchmark power plants are selected for power prediction. This solves the problem of lack of historical data for newly built photovoltaic power plants, achieves accurate power prediction for photovoltaic power plants, and improves the efficiency of power grid dispatch and management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 华能(嘉峪关)新能源有限公司
- Filing Date
- 2024-11-29
- Publication Date
- 2026-06-02
AI Technical Summary
Existing statistical methods for predicting photovoltaic power plant power output lack historical data and therefore cannot effectively predict the accuracy of new photovoltaic power plant construction.
By collecting data on the geographical location, grid area, and geographical characteristics of photovoltaic power plants, clusters are divided and benchmark power plants are selected. Machine learning methods are used to combine historical and measured data of the benchmark power plants for prediction, and power prediction is performed using the nonlinear relationship between the clusters and the benchmark power plants.
It improves the accuracy and reliability of power prediction for newly built photovoltaic power plants, reduces prediction complexity, ensures that prediction results reflect the actual operating status of photovoltaic power plant clusters, and provides strong data support for grid dispatch and energy management.
Smart Images

Figure CN122136791A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of photovoltaic power plant power prediction technology, and particularly relates to a photovoltaic power plant power prediction method and related equipment based on cluster partitioning. Background Technology
[0002] Currently, with the large-scale investment and operation of photovoltaic power plants, accurate prediction of photovoltaic power output is crucial. This not only provides real-time and accurate power output forecasts for grid-connected operation, optimizing dispatch strategies, reducing curtailment, and improving energy efficiency, but also provides key input to the grid's energy management system (EMS), helping to achieve grid supply and demand balance and enhancing grid flexibility and reliability. In the context of smart grid development, this capability has immeasurable value for realizing large-scale grid connection of renewable energy and promoting energy structure transformation.
[0003] Existing photovoltaic (PV) power forecasting methods mainly employ statistical methods for accurate prediction of PV power plant output. These statistical methods are similar to those used for wind power forecasting, identifying the relationship between weather conditions and PV power plant output based on historical statistical data. Then, they combine measured data and numerical weather prediction data with machine learning models to predict the output power of PV power plants. However, these forecasting methods have certain limitations. Specifically, while these statistical methods can predict the power output of PV power plants that have been in operation for a certain period or have been in operation for a long time, they cannot predict the future power output of newly commissioned or newly built PV power plants due to the lack of large-scale historical data.
[0004] It is evident that existing statistical power prediction methods, due to the lack of historical data accumulation, cannot effectively and accurately predict photovoltaic power for newly built photovoltaic power plants. Summary of the Invention
[0005] The purpose of this invention is to provide a photovoltaic power plant power prediction method and related equipment based on cluster partitioning, so as to solve the technical problem that existing power prediction methods based on statistical methods cannot effectively achieve accurate prediction of photovoltaic power for newly built photovoltaic power plants due to the lack of historical data accumulation.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: Firstly, a photovoltaic power plant power prediction method based on cluster partitioning includes: Collect geographical location information, grid area, and geographical characteristics data of each photovoltaic power station; Based on the collected geographical location information, the power grid area to which they belong, and geographical characteristics, each photovoltaic power station is divided into multiple clusters; at the same time, a corresponding benchmark power station is selected in each cluster. Based on machine learning methods, and combined with historical and measured data of the benchmark power station, the power of the benchmark power station is predicted, and the power prediction result of the benchmark power station is obtained. Based on the nonlinear relationship between the cluster and the corresponding benchmark power station, and combined with the power prediction results of the benchmark power station, the power prediction value of the cluster is obtained, so as to realize the overall photovoltaic power station power prediction.
[0007] Furthermore, the step of dividing each photovoltaic power station into multiple clusters based on the collected geographical location information, the power grid area to which it belongs, and geographical characteristic data is as follows: Based on geographical location information, photovoltaic power stations within a preset radius will be initially screened; Based on the power grid region to which they belong, determine whether the photovoltaic power stations that have been initially screened belong to the same power grid region, and retain photovoltaic power stations that belong to the same power grid region; Based on geographical similarity, the remaining photovoltaic power stations were further screened to obtain multiple clusters.
[0008] Furthermore, the steps for selecting the corresponding benchmark power station in each cluster are as follows: Obtain the output data of each photovoltaic power station within the same cluster, as well as the output data of the benchmark power station; The correlation coefficients between multiple benchmark power plants and photovoltaic power plants within the cluster were calculated using a correlation machine learning method. The benchmark power plants corresponding to each cluster are obtained by screening based on the correlation coefficient.
[0009] Furthermore, the historical data of the benchmark power station includes historical operational data and historical meteorological data; the measured data of the benchmark power station includes measured operational data, measured meteorological data, and digital weather forecast data.
[0010] Furthermore, the specific steps for predicting the power output of a benchmark power station based on machine learning methods and combining historical and measured data include: Based on the historical data of the benchmark power station, the data is divided into training set, validation set, and test set. The initial benchmark power prediction model was trained using a training set, and the trained benchmark power prediction model was validated and tested using a validation set and a test set, respectively. The power prediction results of the benchmark power station are obtained by using the measured data of the benchmark power station with the benchmark power prediction model.
[0011] Furthermore, the initialization benchmark power prediction model is constructed based on one of the following algorithms: linear regression, random forest, gradient boosting tree, and support vector machine.
[0012] Furthermore, data preprocessing is performed on the collected geographical location information, power grid area, and geographical characteristics data of each photovoltaic power station; the data preprocessing includes at least: removing outliers and missing values, and normalizing and standardizing the data.
[0013] Secondly, the present invention also provides a photovoltaic power plant power prediction system based on cluster partitioning, comprising: The data acquisition module is used to collect geographical location information, power grid area, and geographical characteristics data of each photovoltaic power station; The cluster division module is used to divide each photovoltaic power station into multiple clusters based on the collected geographical location information, the power grid area to which they belong, and geographical characteristics; at the same time, it selects the corresponding benchmark power station in each cluster. The first power prediction module is used to predict the power of the benchmark power station based on machine learning methods and by combining historical data and measured data of the benchmark power station, so as to obtain the power prediction result of the benchmark power station. The second power prediction module is used to obtain the power prediction value of the cluster based on the nonlinear relationship between the cluster and the corresponding benchmark power station, combined with the power prediction result of the benchmark power station, so as to realize the overall photovoltaic power station power prediction.
[0014] Specifically used for: The steps for dividing each photovoltaic power station into multiple clusters based on the collected geographical location information, the power grid region to which it belongs, and geographical feature data are as follows: Based on geographical location information, photovoltaic power stations within a preset radius will be initially screened; Based on the power grid region to which they belong, determine whether the photovoltaic power stations that have been initially screened belong to the same power grid region, and retain photovoltaic power stations that belong to the same power grid region; Based on geographical similarity, the remaining photovoltaic power stations were further screened to obtain multiple clusters.
[0015] Furthermore, the steps for selecting the corresponding benchmark power station in each cluster are as follows: Obtain the output data of each photovoltaic power station within the same cluster, as well as the output data of the benchmark power station; The correlation coefficients between multiple benchmark power plants and photovoltaic power plants within the cluster were calculated using a correlation machine learning method. The benchmark power plants corresponding to each cluster are obtained by screening based on the correlation coefficient.
[0016] Furthermore, the historical data of the benchmark power station includes historical operational data and historical meteorological data; the measured data of the benchmark power station includes measured operational data, measured meteorological data, and digital weather forecast data.
[0017] Furthermore, the specific steps for predicting the power output of a benchmark power station based on machine learning methods and combining historical and measured data include: Based on the historical data of the benchmark power station, the data is divided into training set, validation set, and test set. The initial benchmark power prediction model was trained using a training set, and the trained benchmark power prediction model was validated and tested using a validation set and a test set, respectively. The power prediction results of the benchmark power station are obtained by using the measured data of the benchmark power station with the benchmark power prediction model.
[0018] Furthermore, the initialization benchmark power prediction model is constructed based on one of the following algorithms: linear regression, random forest, gradient boosting tree, and support vector machine.
[0019] Furthermore, data preprocessing is performed on the collected geographical location information, power grid area, and geographical characteristics data of each photovoltaic power station; the data preprocessing includes at least: removing outliers and missing values, and normalizing and standardizing the data.
[0020] Thirdly, a device is also provided, comprising: Memory, used to store computer programs; A processor is used to implement the steps of the above-described photovoltaic power plant power prediction method based on cluster partitioning when executing the computer program.
[0021] Fourthly, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, is used to implement the steps of the above-described photovoltaic power plant power prediction method based on cluster partitioning.
[0022] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a photovoltaic power plant power prediction method based on cluster partitioning. This method collects data on the geographical location of photovoltaic power plants, grid area, and geographical characteristics, and uses this data to partition clusters and select benchmark power plants, achieving efficient organization and centralized management of photovoltaic power plant power prediction. By utilizing machine learning methods and historical and measured data from benchmark power plants, the accuracy and reliability of power prediction are improved. Furthermore, the nonlinear relationship between clusters and benchmark power plants is used to derive the predicted power values, which not only reduces the overall complexity of the prediction but also ensures that the prediction results more comprehensively reflect the actual operating status of the photovoltaic power plant cluster, providing strong data support for grid dispatching, energy management, and operation and maintenance decisions. This method can effectively and accurately predict the photovoltaic power of newly built photovoltaic power plants. The method is simple in principle and easy to implement, possessing good potential for widespread application.
[0023] Preferably, in this invention, the cluster division process ensures the similarity and correlation of photovoltaic power stations within the cluster through gradual screening based on geographical location, power grid area, and geographical characteristics, providing a reliable foundation for subsequent cluster-based power prediction.
[0024] Preferably, in this invention, the correlation coefficient is calculated by the correlation machine learning method, and the benchmark power station with the highest correlation with the photovoltaic power station in each cluster is reasonably selected, thereby improving the pertinence and accuracy of power prediction.
[0025] Preferably, in this invention, historical operating data, historical meteorological data, measured operating data, measured meteorological data, and digital weather forecast data are comprehensively considered, providing a comprehensive and rich data source for power prediction and further improving the accuracy of the prediction.
[0026] Preferably, in this invention, the benchmark power prediction model is systematically trained and validated by dividing it into training, validation, and test sets, ensuring the model's stability and reliability. Simultaneously, advanced machine learning algorithms are employed to construct the prediction model, improving prediction accuracy and generalization ability.
[0027] Preferably, in this invention, a prediction model is constructed based on algorithms such as linear regression, random forest, gradient boosting tree, and support vector machine, providing multiple options for power prediction. The most suitable model can be selected for prediction according to the actual situation.
[0028] Preferably, in this invention, by preprocessing the collected data to remove outliers and missing values, and by normalizing and standardizing the data, the quality and consistency of the data are improved, providing a reliable data foundation for subsequent data analysis and prediction. Attached Figure Description
[0029] Figure 1 A flowchart illustrating a photovoltaic power plant power prediction method based on cluster partitioning provided by this invention; Figure 2 A schematic diagram of the structure of a photovoltaic power plant power prediction system based on cluster partitioning provided by the present invention; Figure 3 A flowchart illustrating a photovoltaic power plant power prediction method based on cluster partitioning, provided as an embodiment of the present invention. Detailed Implementation
[0030] Example 1 This embodiment provides a photovoltaic power plant power prediction method based on cluster partitioning, including the following steps: S1: Collect geographical location information, power grid area, and geographical characteristics data of each photovoltaic power station; perform data preprocessing on the collected geographical location information, power grid area, and geographical characteristics data of each photovoltaic power station; the data preprocessing includes at least: removing outliers and missing values, and normalizing and standardizing the data.
[0031] The specific steps are as follows: Based on geographical location information, photovoltaic power stations within a preset radius will be initially screened; Based on the power grid region to which they belong, determine whether the photovoltaic power stations that have been initially screened belong to the same power grid region, and retain photovoltaic power stations that belong to the same power grid region; Based on geographical similarity, the remaining photovoltaic power stations were further screened to obtain multiple clusters.
[0032] S2: Based on the collected geographical location information, the power grid area to which they belong, and geographical characteristics, each photovoltaic power station is divided into multiple clusters; at the same time, the corresponding benchmark power station is selected in each cluster. The specific steps are as follows: Obtain the output data of each photovoltaic power station within the same cluster, as well as the output data of the benchmark power station; The correlation coefficients between multiple benchmark power plants and photovoltaic power plants within the cluster were calculated using a correlation machine learning method. The benchmark power plants corresponding to each cluster are obtained by screening based on the correlation coefficient.
[0033] S3: Based on machine learning methods and combining historical and measured data of the benchmark power station, the power of the benchmark power station is predicted to obtain the power prediction result of the benchmark power station.
[0034] Specifically, the historical data of the benchmark power station includes historical operational data and historical meteorological data; the measured data of the benchmark power station includes measured operational data, measured meteorological data, and digital weather forecast data.
[0035] The specific steps for predicting the power output of a benchmark power station based on machine learning methods and combining historical and measured data include: Based on the historical data of the benchmark power station, the data is divided into training set, validation set, and test set. The initial benchmark power prediction model was trained using a training set, and the trained benchmark power prediction model was validated and tested using a validation set and a test set, respectively. The power prediction results of the benchmark power station are obtained by using the measured data of the benchmark power station with the benchmark power prediction model.
[0036] Here, the above-mentioned initialization benchmark power prediction model for power plants is constructed based on one of the following algorithms: linear regression, random forest, gradient boosting tree, and support vector machine.
[0037] S4: Based on the nonlinear relationship between the cluster and the corresponding benchmark power station, and combined with the power prediction results of the benchmark power station, the power prediction value of the cluster is obtained to realize the overall photovoltaic power station power prediction.
[0038] like Figure 2 As shown in the figure, this embodiment also provides a photovoltaic power plant power prediction system based on cluster partitioning, including: a data acquisition module for collecting geographical location information, grid area, and geographical characteristics data of each photovoltaic power plant; a cluster partitioning module for dividing each photovoltaic power plant into multiple clusters based on the collected geographical location information, grid area, and geographical characteristics; and simultaneously selecting a corresponding benchmark power plant in each cluster; a first power prediction module for predicting the power of the benchmark power plant based on machine learning methods and combining historical data and measured data of the benchmark power plant to obtain the benchmark power prediction result; and a second power prediction module for obtaining the cluster power prediction value based on the nonlinear relationship between the cluster and the corresponding benchmark power plant, combined with the benchmark power prediction result, so as to realize the overall photovoltaic power plant power prediction.
[0039] The present invention also provides an apparatus comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the photovoltaic power plant power prediction method based on cluster partitioning.
[0040] When the processor executes the computer program, it implements the above-mentioned steps for predicting the power of photovoltaic power plants based on cluster partitioning. For example: collecting geographical location information, power grid area, and geographical characteristics data of each photovoltaic power plant; dividing each photovoltaic power plant into multiple clusters based on the collected geographical location information, power grid area, and geographical characteristics; simultaneously selecting the corresponding benchmark power plant in each cluster; predicting the power of the benchmark power plant based on machine learning methods and combining historical data and measured data of the benchmark power plant to obtain the benchmark power prediction result; and obtaining the cluster power prediction value based on the nonlinear relationship between the cluster and the corresponding benchmark power plant, combined with the benchmark power prediction result, to achieve overall photovoltaic power plant power prediction.
[0041] Alternatively, when the processor executes the computer program, it implements the functions of each module in the above system, such as: a data acquisition module for collecting geographical location information, power grid area, and geographical characteristics data of each photovoltaic power station; a cluster division module for dividing each photovoltaic power station into multiple clusters based on the collected geographical location information, power grid area, and geographical characteristics; and simultaneously selecting the corresponding benchmark power station in each cluster; a first power prediction module for predicting the power of the benchmark power station based on machine learning methods and combining historical data and measured data of the benchmark power station to obtain the benchmark power prediction result; and a second power prediction module for obtaining the cluster power prediction value based on the nonlinear relationship between the cluster and the corresponding benchmark power station, combined with the benchmark power prediction result, so as to achieve overall photovoltaic power station power prediction.
[0042] Exemplarily, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing preset functions, the instruction segments describing the execution process of the computer program in the cluster-based photovoltaic power plant power prediction device. For example, the computer program can be divided into a data acquisition module, a cluster division module, a first power prediction module, and a second power prediction module; the specific functions of each module are as follows: the data acquisition module is used to collect geographical location information, the power grid area to which each photovoltaic power plant belongs, and geographical characteristic data; the cluster division module is used to divide each photovoltaic power plant into multiple clusters based on the collected geographical location information, the power grid area to which it belongs, and geographical characteristics; and simultaneously selects the corresponding benchmark power plant in each cluster; the first power prediction module is used to predict the power of the benchmark power plant based on machine learning methods and combined with historical data and measured data of the benchmark power plant to obtain the benchmark power prediction result; the second power prediction module is used to obtain the cluster power prediction value based on the nonlinear relationship between the cluster and the corresponding benchmark power plant, combined with the benchmark power prediction result, to achieve overall photovoltaic power plant power prediction.
[0043] The photovoltaic power plant power prediction device based on cluster partitioning can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The photovoltaic power plant power prediction device based on cluster partitioning may include, but is not limited to, processors and memory. Those skilled in the art will understand that the above are examples of photovoltaic power plant power prediction devices based on cluster partitioning and do not constitute a limitation on such devices. The device may include more components than described above, or combine certain components, or use different components. For example, the photovoltaic power plant power prediction device based on cluster partitioning may also include input / output devices, network access devices, buses, etc.
[0044] The processor referred to can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or any conventional processor. This processor is the control center of the cluster-based photovoltaic power plant power prediction system, connecting various parts of the system via various interfaces and lines.
[0045] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the cluster-based photovoltaic power plant power prediction device by running or executing the computer program and / or modules stored in the memory and calling the data stored in the memory.
[0046] The memory may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function (such as sound playback, image playback, etc.). The data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.
[0047] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the photovoltaic power plant power prediction method based on cluster partitioning.
[0048] If the modules / units integrated by the photovoltaic power plant power prediction system based on cluster partitioning are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0049] Based on this understanding, the present invention can implement all or part of the processes in the above-mentioned photovoltaic power plant power prediction method based on cluster partitioning, or it can be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-mentioned photovoltaic power plant power prediction method based on cluster partitioning. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or a preset intermediate form, etc.
[0050] The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0051] It should be noted that the content contained in the computer-readable storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0052] The present invention will be further described below with reference to embodiments and accompanying drawings: Example 2 As described in the background section, existing photovoltaic power prediction measures mainly employ statistical methods to accurately predict photovoltaic power plants. These statistical methods are similar to those used for wind power prediction, identifying the relationship between weather conditions and photovoltaic power plant output based on historical statistical data. Then, based on measured data and numerical weather forecast data, combined with machine learning models, the output power of the photovoltaic power plant is predicted. However, the above prediction methods have certain limitations. Specifically, while these statistical methods can predict the power output of photovoltaic power plants that have been in operation for a certain period or have been in operation for a long time, they cannot predict the future power output of newly commissioned or newly built photovoltaic power plants due to the lack of large-scale historical data.
[0053] To achieve the above objectives, this invention provides a photovoltaic power plant power prediction method based on cluster partitioning. This method is based on the prediction principle of physical model and achieves cluster power prediction by partitioning photovoltaic power plants into clusters, thereby realizing photovoltaic power prediction for the entire power grid area. Using this method, comprehensive prediction, management and evaluation functions related to power prediction can be provided for photovoltaic power plants in various regions.
[0054] The photovoltaic power prediction method based on cluster partitioning provided in this embodiment is based on the following idea: First, cluster modeling and prediction technology divides the cluster area into sub-regions to maximize the decomposition of the impact of different meteorological conditions and geographical features on the cluster output. Then, based on the cluster partitioning results, a benchmark power station is selected. Since the meteorological conditions and topography are similar within the cluster, the meteorological information of the benchmark power station can represent the meteorological information of the cluster. Machine learning of the correlation between the output of the benchmark power station and the output of the cluster is a key indicator for selecting the benchmark power station. Based on the results of selecting the benchmark power station, the cluster power curve, as well as the maximum and minimum output curves, are predicted. Finally, the correlation between the output of the benchmark power station and the output of the cluster is analyzed to achieve the prediction of the cluster power.
[0055] like Figure 3 As shown in the figure, this embodiment provides a photovoltaic power plant power prediction method based on cluster partitioning, specifically including: Step 1: Clustering of photovoltaic power plants: Step 1.1: Define the three main criteria for cluster division: similar geographical features, geographical location within a preset radius, and located in the same power grid area. In this embodiment, the preset radius of geographical location is 30 kilometers as a reference; the determination of similar geographical features is as follows: First, similar climatic conditions: The power generation of photovoltaic power plants is highly dependent on sunlight resources. Therefore, power plants with similar climate conditions often have similar sunshine duration, radiation intensity, and air quality. For example, power plants located in the same climate zone or adjacent climate regions, such as arid, semi-arid, or desert areas, are very suitable for building photovoltaic power plants due to their long sunshine duration and good air quality. Furthermore, the geographical characteristics of power plants in these regions are similar.
[0056] Second, the terrain and landforms are similar: Topography and landforms have a significant impact on the construction and operation of photovoltaic power plants. Similar topography (such as slope and aspect) and landforms (such as soil type, rock stability, and groundwater level) facilitate the unified planning, construction, and maintenance of photovoltaic power plants. For example, areas with moderate slopes (5°–30°) and good soil bearing capacity are suitable for large-scale installation of photovoltaic panels, and power plants in these areas share similar geographical characteristics.
[0057] Third, they have similar ecological environments: The construction and operation of photovoltaic power plants need to consider their impact on the ecological environment. Power plants located in similar ecological environments, such as non-arable land, barren hills, wastelands, or abandoned mines, can reduce the occupation of arable land resources and the damage to the ecological environment. Photovoltaic power plants in these areas face similar ecological and environmental problems during construction and operation, and therefore can be considered as power plants with similar geographical characteristics.
[0058] Fourth, the grid connection conditions are similar: Photovoltaic power plants need to be connected to the power grid to transmit and sell electricity. Power plants with similar grid connection conditions, such as those located in the same grid area, with similar grid structures and connection points, can be more easily clustered and their power output predicted. Power transmission and dispatch among these power plants are also more efficient and stable.
[0059] If all four points are similar, it indicates a high degree of similarity in geographical characteristics.
[0060] Step 1.2: Collect geographical location information, power grid area, and geographical features (such as terrain and climate) data for each photovoltaic power station.
[0061] Step 1.3: Perform preprocessing operations such as data cleaning, data normalization, and standardization on the data collected in Step 1.2.
[0062] Step 1.4: Based on the preprocessed data, power plants that meet the classification criteria are divided into the same cluster to ensure that power plants within the cluster have similar geographical and grid characteristics.
[0063] Step 2: Benchmark Power Plant Selection: Step 2.1: Calculation of correlation coefficient and correlation degree: Calculate the correlation coefficient and correlation degree of the output power between the power plants in the cluster to assess their similarity and mutual influence.
[0064] Step 2.2: Select a representative benchmark power station based on the correlation coefficient and degree of correlation. The benchmark power station is used to better reflect the overall power generation characteristics of the cluster.
[0065] Step 2.3: When selecting a benchmark power station, consider correcting the correlation coefficient to address situations where some power station data is lost or abnormal, ensuring accurate cluster power prediction even in these cases.
[0066] Among these methods, Pearson correlation can be used to correct the correlation coefficient, and the specific formula is as follows:
[0067] Where R represents the Pearson correlation coefficient; x represents the factor to be analyzed; and l represents the photovoltaic power time series.
[0068] Step 3: Baseline Power Plant Power Prediction In this embodiment, existing and mature power prediction methods, such as physical methods, statistical methods, machine learning or neural network methods, are selected for the benchmark power plant.
[0069] For example, the specific steps for using a neural network to predict the power of a benchmark power plant are as follows: The historical data of the benchmark power station is divided into training set, validation set and test set. The historical data includes historical operation data and historical meteorological data. The measured data of the benchmark power station includes measured operation data, measured meteorological data and digital weather forecast data. After acquiring the data, the above data are preprocessed to ensure the accuracy of model training.
[0070] The initial benchmark power prediction model was trained using a training set, and the trained benchmark power prediction model was validated and tested using a validation set and a test set, respectively. The power prediction results of the benchmark power station are obtained by using the measured data of the benchmark power station with the benchmark power prediction model.
[0071] Step 4: Cluster Power Prediction The nonlinear relationship between the benchmark power station and the cluster is obtained by using the neural network method. The power prediction value of the cluster is obtained by using the power prediction result of the benchmark power station. Finally, the prediction values of each cluster are added together to obtain the new energy power prediction value of the entire cluster, that is, the power prediction of the whole power grid area.
[0072] In this embodiment, the above method can also be used to construct a complete cluster prediction model. This model needs to collect digital weather forecast data at multiple spatiotemporal scales, as well as digital weather forecast data, historical operational data, and measured meteorological data reported by stations. Through big data analysis technologies such as physical methods, machine learning methods, statistical methods, and neural network methods, cluster division principles, benchmark power station screening models, cluster prediction models, power energy calculation models, and confidence intervals are established. By statistically analyzing uncertainties such as data completeness checks, power station-cluster output correlation, correlation degree, and prediction data errors, benchmark power stations are selected and cluster power prediction curves are calculated. The confidence interval curves are calculated using the power energy model to more accurately correct the cluster prediction data.
[0073] In summary, this invention provides a photovoltaic power plant power prediction method based on cluster partitioning, which has the following advantages compared with existing photovoltaic power plant power prediction methods based on cluster partitioning: First, it improves the accuracy of power prediction: By comprehensively considering the geographical location information of photovoltaic power plants, their respective power grid areas, and geographical characteristics, a refined cluster division is performed, and benchmark power plants for each cluster are selected for power prediction. This method can capture the similarities and differences between photovoltaic power plants, thereby more accurately predicting the power output of photovoltaic power plants and improving the accuracy and reliability of the prediction.
[0074] Second, it optimizes resource allocation and scheduling: the power forecasting method based on cluster partitioning enables the power grid dispatch center to more rationally arrange the power generation plan and dispatching strategy according to the power forecast results of different clusters. This helps to optimize the allocation of power resources, improve the operating efficiency and stability of the power grid, and reduce power shortages or surpluses caused by inaccurate forecasts.
[0075] Third, improve the operational efficiency of photovoltaic power plants: Through accurate power forecasting, photovoltaic power plants can develop more scientific operation and maintenance plans and rationally arrange maintenance work such as cleaning and overhauling, thereby improving the operating efficiency and power generation capacity of the power plant. At the same time, the forecast results can also provide a basis for decision-making on the expansion and renovation of the power plant, promoting the sustainable development of photovoltaic power plants.
[0076] Fourth, reduced operation and maintenance costs: Cluster partitioning and the selection of benchmark power plants reduce the number of photovoltaic power plants that need to be predicted individually, thereby reducing the costs of data collection, processing, and model training. Furthermore, accurate power prediction helps reduce power waste and losses caused by prediction errors, further lowering the operation and maintenance costs of photovoltaic power plants.
[0077] Fifth, enhancing system robustness: By introducing data preprocessing steps, such as removing outliers and missing values, and performing normalization and standardization operations on the data, the quality and consistency of the data are improved. This helps to enhance the robustness of the prediction model and reduce prediction errors and fluctuations caused by data problems.
[0078] Sixth, promoting intelligent management: This technical solution combines machine learning methods for power prediction, realizing intelligent management of photovoltaic power plants. By continuously learning and optimizing the prediction model, the system can automatically adapt to changes in the operating environment and conditions of the photovoltaic power plant, improving the adaptability and real-time performance of the prediction.
[0079] The above embodiments are merely one of the implementation methods for achieving the technical solution of the present invention. The scope of protection claimed by the present invention is not limited to this embodiment, but also includes any variations, substitutions and other implementation methods that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention.
Claims
1. A photovoltaic power plant power prediction method based on cluster partitioning, characterized in that, include: Collect geographical location information, grid area, and geographical characteristics data of each photovoltaic power station; Based on the collected geographical location information, the power grid area to which they belong, and geographical characteristics, each photovoltaic power station is divided into multiple clusters; at the same time, a corresponding benchmark power station is selected in each cluster. Based on machine learning methods, and combined with historical and measured data of the benchmark power station, the power of the benchmark power station is predicted, and the power prediction result of the benchmark power station is obtained. Based on the nonlinear relationship between the cluster and the corresponding benchmark power station, and combined with the power prediction results of the benchmark power station, the power prediction value of the cluster is obtained, so as to realize the overall photovoltaic power station power prediction.
2. The photovoltaic power plant power prediction method based on cluster partitioning according to claim 1, characterized in that, The steps for dividing each photovoltaic power station into multiple clusters based on the collected geographical location information, the power grid region to which it belongs, and geographical feature data are as follows: Based on geographical location information, photovoltaic power stations within a preset radius will be initially screened; Based on the power grid region to which they belong, determine whether the photovoltaic power stations that have been initially screened belong to the same power grid region, and retain photovoltaic power stations that belong to the same power grid region; Based on geographical similarity, the remaining photovoltaic power stations were further screened to obtain multiple clusters.
3. The photovoltaic power plant power prediction method based on cluster partitioning according to claim 1, characterized in that, The steps for selecting the corresponding benchmark power station in each cluster are as follows: Obtain the output data of each photovoltaic power station within the same cluster, as well as the output data of the benchmark power station; The correlation coefficients between multiple benchmark power plants and photovoltaic power plants within the cluster were calculated using a correlation machine learning method. The benchmark power plants corresponding to each cluster are obtained by screening based on the correlation coefficient.
4. The photovoltaic power plant power prediction method based on cluster partitioning according to claim 1, characterized in that, Historical data for benchmark power stations include historical operational data and historical meteorological data; measured data for benchmark power stations include measured operational data, measured meteorological data, and digital weather forecast data.
5. The photovoltaic power plant power prediction method based on cluster partitioning according to claim 1, characterized in that, The specific steps for predicting the power output of a benchmark power station based on machine learning methods and combining historical and measured data include: Based on the historical data of the benchmark power station, the data is divided into training set, validation set, and test set. The initial benchmark power prediction model was trained using a training set, and the trained benchmark power prediction model was validated and tested using a validation set and a test set, respectively. The power prediction results of the benchmark power station are obtained by using the measured data of the benchmark power station with the benchmark power prediction model.
6. The photovoltaic power plant power prediction method based on cluster partitioning according to claim 5, characterized in that, The initialization benchmark power prediction model for power plants is constructed based on one of the following algorithms: linear regression, random forest, gradient boosting tree, and support vector machine.
7. The photovoltaic power plant power prediction method based on cluster partitioning according to claim 1, characterized in that, Data preprocessing is performed on the collected geographical location information, power grid area, and geographical characteristics data of each photovoltaic power station. The data preprocessing includes at least the following: removing outliers and missing values, and normalizing and standardizing the data.
8. A photovoltaic power plant power prediction system based on cluster partitioning, characterized in that, include: The data acquisition module is used to collect geographical location information, power grid area, and geographical characteristics data of each photovoltaic power station; The cluster division module is used to divide each photovoltaic power station into multiple clusters based on the collected geographical location information, the power grid area to which they belong, and geographical characteristics; at the same time, it selects the corresponding benchmark power station in each cluster. The first power prediction module is used to predict the power of the benchmark power station based on machine learning methods and by combining historical data and measured data of the benchmark power station, so as to obtain the power prediction result of the benchmark power station. The second power prediction module is used to obtain the power prediction value of the cluster based on the nonlinear relationship between the cluster and the corresponding benchmark power station, combined with the power prediction result of the benchmark power station, so as to realize the overall photovoltaic power station power prediction.
9. A device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the photovoltaic power plant power prediction method based on cluster partitioning as described in any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it is used to implement the steps of the photovoltaic power plant power prediction method based on cluster partitioning as described in any one of claims 1-7.