Generating capacity evaluation method and device based on KMeans clustering
Through the power generation evaluation method based on KMeans clustering, a model is formed using historical data, and the characteristic variable with the highest similarity is selected to calculate the power generation efficiency, solving the accuracy and applicability of the power generation evaluation of photovoltaic power stations, achieving high accuracy evaluation under a small amount of data, and is suitable for a variety of power station types.
Patent Information
- Application Number
- CN202510435815.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-27
AI Technical Summary
The existing photovoltaic power generation evaluation methods have the problem that high evaluation accuracy and high applicability cannot be obtained. Traditional methods rely on simplified models or insufficient data to cause large deviations in prediction results, while large-scale model algorithms require a large amount of historical data and cannot be applied to power stations with missing historical data.
The power generation evaluation method based on KMeans clustering is adopted. By receiving the power generation data to be evaluated, it is converted into feature variables, and the pre-trained historical data KMeans clustering model is used to select feature variables, calculate the similarity, select the feature variable with the highest similarity, calculate the reference power generation efficiency, and determine the target power generation efficiency for evaluation.
It realizes high-accurate power generation assessment under a small amount of historical data, and is suitable for old power plants with missing historical data and new power plants with insufficient historical data, helping power plant managers make more reasonable operational decisions.
Smart Images

Figure CN120218752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power equipment evaluation, and particularly to a power generation evaluation method based on KMeans clustering, a device thereof, a training method for a historical data KMeans clustering model, and a device thereof. Background Art
[0002] With the increasing global demand for renewable energy, photovoltaic power stations, as an important source of clean energy, have developed rapidly. However, the power generation of photovoltaic power stations is affected by various factors, such as weather conditions, component performance, system efficiency, etc. The uncertainty of these factors poses challenges to the operation and management of the power stations. In order to optimize the performance of the power stations and improve the return on investment, it is necessary to accurately evaluate the daily power generation of photovoltaic power stations.
[0003] Traditional methods for evaluating the power generation of photovoltaic power stations rely on simplified models or insufficient data, resulting in a large deviation between the prediction results and the actual power generation. Moreover, the existing technologies fail to make full use of historical data and real-time monitoring data, and do not effectively integrate multi-source information to improve the accuracy of prediction. Although the accuracy of large model algorithms has been improved, the amount of data required for pre-training is too large. For some newly built power stations or the renovation of old power stations, there is simply not so much historical data left, resulting in too small an applicable range of large model evaluation and being unable to meet various needs in actual production.
[0004] Therefore, how to provide a method for evaluating the power generation of photovoltaic power stations with low data volume requirements, wide applicability, and high evaluation accuracy is an urgent problem to be solved in the existing technology. Summary of the Invention
[0005] The purpose of the present invention is to provide a power generation evaluation method based on KMeans clustering, a device thereof, a training method for a historical data KMeans clustering model, and a device thereof, so as to solve the problem in the existing technology that high evaluation accuracy and high applicability of the method for evaluating the power generation of photovoltaic power stations cannot be achieved simultaneously.
[0006] To solve the above technical problems, the present invention provides a power generation evaluation method based on KMeans clustering, including:
[0007] Receiving the power generation data of the day to be evaluated;
[0008] Converting the power generation data of the day to be evaluated into characteristic variables of the day to be evaluated;
[0009] Selecting a first number of time series characteristic variables in a pre-trained historical data KMeans clustering model; the selected time series characteristic variables belong to different clusters, and the historical data KMeans clustering model has the first number of clusters;
[0010] Calculate the similarity between each of the timing feature variables and the feature variables of the day to be evaluated, and select the second largest number of timing feature variables with the highest similarity as the template timing feature variables;
[0011] In the historical data KMeans clustering model, search for the third largest number of timing feature variables with the highest similarity for each of the template timing feature variables as the reference timing feature variables;
[0012] Calculate the reference power generation efficiency corresponding to each of the reference timing feature variables;
[0013] Determine the target power generation efficiency based on all the reference power generation efficiencies;
[0014] Evaluate the power generation data of the day to be evaluated according to the target power generation efficiency to obtain a power generation amount evaluation result.
[0015] Optionally, in the power generation amount evaluation method based on KMeans clustering, determining the target power generation efficiency based on all the reference power generation efficiencies includes:
[0016] Find the average value of all the reference power generation efficiencies as the target power generation efficiency.
[0017] Optionally, in the power generation amount evaluation method based on KMeans clustering, the range of the second largest number and / or the third largest number is from 4 to 8, including the end values.
[0018] Optionally, in the power generation amount evaluation method based on KMeans clustering, after receiving the power generation data of the day to be evaluated, it further includes:
[0019] Perform normalization processing on the power generation data of the day to be evaluated;
[0020] Correspondingly, the data corresponding to the timing feature variables in the historical data KMeans clustering model is data after normalization processing.
[0021] A power generation amount evaluation device based on KMeans clustering includes:
[0022] A first receiving module for receiving the power generation data of the day to be evaluated;
[0023] A transformation module for converting the power generation data of the day to be evaluated into feature variables of the day to be evaluated;
[0024] A variable selection module for selecting the first largest number of timing feature variables in a pre-trained historical data KMeans clustering model; the selected timing feature variables belong to different clusters, and the historical data KMeans clustering model has the first largest number of clusters;
[0025] A similarity module, configured to calculate the similarity between each of the time series feature variables and the feature variables of the day to be evaluated, and select the second largest number of time series feature variables with the highest similarity as the template time series feature variables;
[0026] A reference variable module, configured to search for the third largest number of time series feature variables with the highest similarity for each of the template time series feature variables in the historical data KMeans clustering model as the reference time series feature variables;
[0027] A reference power generation module, configured to calculate the reference power generation efficiency corresponding to each of the reference time series feature variables;
[0028] A target power generation module, configured to determine the target power generation efficiency according to all the reference power generation efficiencies;
[0029] An evaluation module, configured to evaluate the power generation data of the day to be evaluated according to the target power generation efficiency to obtain a power generation amount evaluation result.
[0030] A training method for a historical data KMeans clustering model, where the training method for the historical data KMeans clustering model is used for any one of the above power generation amount evaluation methods based on KMeans clustering, and includes:
[0031] Receiving historical power generation records; the historical power generation records include multiple historical daily power generation data;
[0032] Converting all the historical daily power generation data into time series feature variables corresponding to the dates;
[0033] Selecting the first largest number of time series feature variables from all the time series feature variables as the clustering centers, and performing clustering processing on the remaining time series feature variables so that each of the remaining time series feature variables has a corresponding clustering center, and the Euclidean distance from the remaining time series feature variables to the corresponding clustering center is not greater than the Euclidean distance from the remaining time series feature variables to the other clustering centers;
[0034] Classifying the clustering centers and the remaining time series feature variables corresponding to the clustering centers into the same cluster, and performing KMeans clustering iteration on the clustering centers within the same cluster until the clustering centers of all the clusters remain unchanged between two consecutive iterations, to obtain a historical data KMeans clustering model.
[0035] Optionally, in the training method for the historical data KMeans clustering model, before converting all the historical daily power generation data into time series feature variables corresponding to the dates, it further includes:
[0036] Determine that the entries in the historical power generation records lack power generation data; the entries lacking power generation data are historical daily power generation data with blank entry items;
[0037] Remove the entries lacking power generation data from the historical power generation records.
[0038] Optionally, in the training method of the historical data KMeans clustering model, the entry items of the historical daily power generation records include daily equivalent hours;
[0039] Before converting all the historical daily power generation data into time series feature variables corresponding to dates, it further includes:
[0040] Sequentially determine whether the daily equivalent hours of each piece of historical daily power generation data exceed a preset theoretical upper limit of sunshine;
[0041] When there is historical daily power generation data whose daily equivalent hours exceed the theoretical upper limit of sunshine, regard the historical daily power generation data whose daily equivalent hours exceed the theoretical upper limit of sunshine as data with incorrect power generation data;
[0042] Remove the data with incorrect power generation data from the historical power generation records.
[0043] Optionally, in the training method of the historical data KMeans clustering model, the determination method of the first quantity includes:
[0044] The first quantity is obtained by the following formula:
[0045] m = n / k;
[0046] Where m is the first quantity, n is the total number of historical daily power generation data included in the historical power generation records; k is a preset grouping coefficient;
[0047] The range of the grouping coefficient is from 5 to 20, including the endpoint values.
[0048] A training device for a historical data KMeans clustering model, the training device for the historical data KMeans clustering model is used for any one of the above-mentioned power generation amount evaluation methods based on KMeans clustering, and includes:
[0049] A second receiving module, configured to receive historical power generation records; the historical power generation records include multiple pieces of historical daily power generation data;
[0050] A date variable transformation module, configured to convert all the historical daily power generation data into time series feature variables corresponding to dates;
[0051] A preliminary clustering module, which is used to select the first number of time series feature variables from all the time series feature variables as clustering centers, and perform clustering processing on the remaining time series feature variables, so that each of the remaining time series feature variables has a corresponding clustering center, and the Euclidean distance from the remaining time series feature variable to the corresponding clustering center is not greater than the Euclidean distance from the remaining time series feature variable to the other clustering centers;
[0052] A clustering iteration module, which is used to classify the clustering centers and the remaining time series feature variables corresponding to the clustering centers into the same cluster, and perform KMeans clustering iteration on the clustering centers within the same cluster until the clustering centers of all clusters remain unchanged between two consecutive iterations, so as to obtain a historical data KMeans clustering model.
[0053] The power generation evaluation method based on KMeans clustering provided by the present invention includes receiving power generation data of the day to be evaluated; converting the power generation data of the day to be evaluated into feature variables of the day to be evaluated; selecting the first number of time series feature variables in a pre-trained historical data KMeans clustering model; the clusters to which the selected time series feature variables belong are different, and the historical data KMeans clustering model has the first number of clusters; calculating the similarity between each of the time series feature variables and the feature variables of the day to be evaluated, and selecting the second number of time series feature variables with the highest similarity as template time series feature variables; searching for the third number of time series feature variables with the highest similarity for each of the template time series feature variables in the historical data KMeans clustering model as reference time series feature variables; calculating the reference power generation efficiency corresponding to each of the reference time series feature variables; determining the target power generation efficiency according to all the reference power generation efficiencies; and evaluating the power generation data of the day to be evaluated according to the target power generation efficiency to obtain a power generation evaluation result.
[0054] The present invention utilizes historical power generation records to form a KMeans clustering model. When it is necessary to evaluate daily power generation data, the historical daily power generation data with the highest similarity (i.e., the corresponding template time series feature variables and reference time series feature variables) can be searched from the historical power generation records, and the power generation efficiency of the daily power generation data to be evaluated can be deduced based on the power generation efficiency of these historical daily power generation data (i.e., the reference power generation efficiency), thus completing the evaluation. Compared with other related technologies, the present invention can achieve high-accuracy power generation quantity evaluation by only using a small amount of historical power generation records, helping power station managers make more reasonable operation decisions; at the same time, the present invention also greatly reduces the requirement for the amount of historical power generation data, enabling the method of the present invention to be applicable to old power stations with missing historical data and newly built power stations with insufficient historical data, greatly broadening the applicable scope of the power generation quantity evaluation method. The present invention also provides a power generation quantity evaluation device based on KMeans clustering and a training method and device for a historical data KMeans clustering model with the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 It is a schematic flow chart of a specific implementation manner of the power generation quantity evaluation method based on KMeans clustering provided by the present invention;
[0057] Figure 2 It is a schematic structural diagram of a specific implementation manner of the power generation quantity evaluation device based on KMeans clustering provided by the present invention;
[0058] Figure 3 It is a schematic flow chart of a specific implementation manner of the training method for the historical data KMeans clustering model provided by the present invention;
[0059] Figure 4 It is a schematic structural diagram of a specific implementation manner of the training device for the historical data KMeans clustering model provided by the present invention.
[0060] Reference Signs:
[0061] 110 - First receiving module; 120 - Transformation module; 130 - Variable selection module; 140 - Similarity module; 150 - Reference variable module; 160 - Reference power generation module; 170 - Target power generation module; 180 - Evaluation module; 210 - Second receiving module; 220 - Date variable transformation module; 230 - Initial clustering module; 240 - Clustering iteration module. Detailed implementation manner
[0062] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0063] The core of the present invention is to provide a power generation evaluation method based on KMeans clustering. The schematic flow diagram of a specific implementation manner is as Figure 1 shown, which is called Specific Implementation Manner 1 and includes:
[0064] S101: Receive the power generation data of the day to be evaluated.
[0065] If it is necessary to evaluate the daily power generation data of a power station, it is at least necessary to know how much electricity the power station actually generated on this day (i.e., the actual daily power generation), and the data of each item that will affect the power generation on this day (i.e., the external factors affecting the power generation). The items include the daily equivalent hours of the power station, the daily equivalent hours of the inverter, the daily irradiance, the average daily temperature, the weather conditions, and the date (corresponding season). In other words, the power generation data of the day to be evaluated should record all the above items. Of course, in addition to the above items, the power generation data of the day to be evaluated should also include information such as the actual power generation of the day to be evaluated.
[0066] Further, after receiving the power generation data of the day to be evaluated, it further includes:
[0067] A1: Perform normalization processing on the power generation data of the day to be evaluated.
[0068] A2: Correspondingly, the data corresponding to the time series feature variables in the historical data KMeans clustering model is the data after normalization processing.
[0069] In this preferred implementation manner, the original data of all items is normalized by MinMaxScaler (i.e., Equation 1 below), so that it is scaled to between [0, 1], converted into dimensionless data, the influence of data fluctuations on the model performance is eliminated, and the model accuracy is improved.
[0070] The MinMaxScaler normalization calculation is as follows in Equation (1):
[0071] ; (1)
[0072] Where x* is the data after normalization processing, max is the maximum value of the original data corresponding to a single entry item, and min is the minimum value of the original data corresponding to a single entry item.
[0073] S102: Convert the power generation data of the day to be evaluated into characteristic variables of the day to be evaluated.
[0074] In this step, convert the power generation data of the day to be evaluated into characteristic variables of the day to be evaluated. Among them, each entry item in the previous text is each dimension of the characteristic variables of the day to be evaluated.
[0075] S103: Select the first number of time series characteristic variables in the pre-trained historical data KMeans clustering model; the clusters to which the selected time series characteristic variables belong are different, and the historical data KMeans clustering model has the first number of clusters.
[0076] In other words, in this step, select one of the time series characteristic variables in each cluster in the historical data KMeans clustering model. Of course, the time series characteristic variables selected from each cluster can be randomly selected or selected according to the actual situation, and the present invention does not make a limitation here. Of course, among each time series characteristic variable in the historical data KMeans clustering model, the included entry items should correspond to the entry items in the power generation data of the day to be evaluated.
[0077] It can be set that the set of all time series characteristic variables in the historical data KMeans clustering model is T. In this step, select the first number of time series characteristic variables, and these characteristic variables form the set K.
[0078] S104: Calculate the similarity between each of the time series characteristic variables and the characteristic variables of the day to be evaluated, and select the second number of time series characteristic variables with the highest similarity as the template time series characteristic variables.
[0079] Continuing from the previous text, in this step, sort the time series characteristic variables in the set K according to their similarity to the characteristic variables of the day to be evaluated, and take the second number of time series characteristic variables at the head after sorting as the template time series characteristic variables. The template time series characteristic variables form the set F. For example, after sorting according to similarity, the top 5 can be taken as the template time series characteristic variables, and the corresponding second number is 5 at this time.
[0080] S105: Search for the third largest number of time series feature variables with the highest similarity for each of the template time series feature variables in the historical data KMeans clustering model as the reference time series feature variables.
[0081] Continuing from the previous text, in this step, further search for the time series feature variables corresponding to the template time series feature variables in set T, and for each of the template time series feature variables, search for the third largest number of time series feature variables with the highest similarity. Assuming the third largest number is also 5, then 5×5 = 25 reference time series feature variables need to be found.
[0082] As a specific implementation, for the template time series feature variables in set F, calculate its similarity S with the remaining time series feature variables t (referred to as the to-be-matched time series feature variables) in set T according to formula (2). The similarity S is the cosine of the angle between the two time series feature variables, that is:
[0083] ; (2)
[0084] where is the cosine of the angle between the template time series feature variable and the to-be-matched time series feature variable, is the i-th element in the template time series feature variable and is the i-th element in the to-be-matched time series feature variable .
[0085] Calculate the similarity between the to-be-matched time series feature variable and a certain template time series feature variable and sort them. As described above, take the first few (determined according to the third largest number) as the reference time series feature variables corresponding to the template time series feature variable.
[0086] Furthermore, the range of the second largest number and / or the third largest number is from 4 to 8, including the endpoint values, such as any one of 4.0, 6.0, or 8.0. The values of the second largest number and the third largest number can be the same or different. The above range is the best range after a large number of theoretical calculations and practical tests. In the above range, both the accuracy of the result and the avoidance of excessive computing power consumption can be ensured. Of course, it can also be adjusted according to the actual situation.
[0087] S106: Calculate the reference power generation efficiency corresponding to each of the reference time series feature variables.
[0088] S107: Determine the target power generation efficiency based on all the reference power generation efficiencies.
[0089] As a specific implementation, this step includes:
[0090] Find the average value of all the reference power generation efficiencies as the target power generation efficiency.
[0091] In other words, in this step, the power generation efficiencies corresponding to all the reference timing feature variables (i.e., the reference power generation efficiencies) are averaged, and the average value is used as the target power generation efficiency. While improving the calculation accuracy, it also saves computing power. In actual production, other processes can also be performed on the reference power generation efficiencies to obtain the target power generation efficiency, such as methods like weighted average or taking the median value, etc. The present invention does not make any limitations here.
[0092] S108: Evaluate the power generation data of the day to be evaluated according to the target power generation efficiency to obtain a power generation amount evaluation result.
[0093] It should be noted that the finally obtained power generation amount evaluation result in this step can be the power generation efficiency expressed as a percentage, or can be expressed in other ways. Specifically, the power generation data of the day to be evaluated must include the actual power generation amount of the day to be evaluated. Divide the actual power generation amount of the day to be evaluated by the installed capacity to obtain the equivalent hourly number of the power station per day, and then divide it by the sunshine hours under peak sunshine conditions to obtain the power generation efficiency of the day to be evaluated. Compare it with the target power generation efficiency obtained in this step to conduct the evaluation. Or, directly compare the equivalent hourly number of the power station per day to also obtain the power generation amount evaluation result, which can be selected according to specific circumstances. The present invention will not elaborate here.
[0094] The power generation evaluation method based on KMeans clustering provided by the present invention includes receiving the power generation data of the day to be evaluated; converting the power generation data of the day to be evaluated into characteristic variables of the day to be evaluated; selecting a first number of time series characteristic variables in a pre-trained KMeans clustering model of historical data; the clusters to which the selected time series characteristic variables belong are different, and the KMeans clustering model of historical data has the first number of clusters; calculating the similarity between each time series characteristic variable and the characteristic variable of the day to be evaluated, and selecting a second number of time series characteristic variables with the highest similarity as template time series characteristic variables; searching for a third number of time series characteristic variables with the highest similarity for each template time series characteristic variable in the KMeans clustering model of historical data as reference time series characteristic variables; calculating the reference power generation efficiency corresponding to each reference time series characteristic variable; determining the target power generation efficiency according to all the reference power generation efficiencies; and evaluating the power generation data of the day to be evaluated according to the target power generation efficiency to obtain a power generation evaluation result. The present invention uses historical power generation records to form a KMeans clustering model. When it is necessary to evaluate the daily power generation data, the historical daily power generation data with the highest similarity can be searched from the historical power generation records, and the power generation efficiency of the power generation data of the day to be evaluated can be deduced according to the power generation efficiencies of these historical daily power generation data to complete the evaluation. Compared with other related technologies, the present invention can achieve high-accuracy power generation evaluation only by using a small amount of historical power generation records, helping power station managers make more reasonable operation decisions; at the same time, the present invention greatly reduces the requirement for the amount of historical power generation data, enabling the method of the present invention to be applicable to old power stations with missing historical data and newly built power stations with insufficient historical data, greatly expanding the applicable scope of the power generation evaluation method.
[0095] The power generation evaluation device based on KMeans clustering provided by the embodiments of the present invention will be introduced below. The power generation evaluation device based on KMeans clustering described below can be correspondingly referred to the power generation evaluation method based on KMeans clustering described above.
[0096] Figure 2 It is a structural block diagram of the power generation evaluation device based on KMeans clustering provided by the embodiments of the present invention, which is called the second specific implementation manner. Refer to Figure 2 The power generation evaluation device based on KMeans clustering may include:
[0097] A first receiving module 110, configured to receive the power generation data of the day to be evaluated;
[0098] A transformation module 120, configured to convert the power generation data of the day to be evaluated into characteristic variables of the day to be evaluated;
[0099] A variable selection module 130 is configured to select a first number of time series feature variables from a pre-trained historical data KMeans clustering model; the clusters to which the selected time series feature variables belong are different, and the historical data KMeans clustering model has the first number of clusters;
[0100] A similarity module 140 is configured to calculate the similarity between each of the time series feature variables and the feature variables of the day to be evaluated, and select a second number of time series feature variables with the highest similarity as the template time series feature variables;
[0101] A reference variable module 150 is configured to search for a third number of time series feature variables with the highest similarity for each of the template time series feature variables in the historical data KMeans clustering model as the reference time series feature variables;
[0102] A reference power generation module 160 is configured to calculate the reference power generation efficiency corresponding to each of the reference time series feature variables;
[0103] A target power generation module 170 is configured to determine a target power generation efficiency based on all the reference power generation efficiencies;
[0104] An evaluation module 180 is configured to evaluate the power generation data of the day to be evaluated based on the target power generation efficiency to obtain a power generation evaluation result.
[0105] As a preferred embodiment, the target power generation module 170 includes:
[0106] An average target power generation unit is configured to calculate the average value of all the reference power generation efficiencies as the target power generation efficiency.
[0107] As a preferred embodiment, the first receiving module 110 further includes:
[0108] A receiving normalization unit is configured to perform normalization processing on the power generation data of the day to be evaluated;
[0109] Correspondingly, the data corresponding to the time series feature variables in the historical data KMeans clustering model is data that has been normalized.
[0110] The power generation evaluation device based on KMeans clustering provided by the present invention includes a first receiving module 110 for receiving power generation data of the day to be evaluated; a transformation module 120 for converting the power generation data of the day to be evaluated into characteristic variables of the day to be evaluated; a variable selection module 130 for selecting a first number of time series characteristic variables in a pre-trained historical data KMeans clustering model, where the clusters to which the selected time series characteristic variables belong are different, and the historical data KMeans clustering model has the first number of clusters; a similarity module 140 for calculating the similarity between each of the time series characteristic variables and the characteristic variables of the day to be evaluated, and selecting a second number of time series characteristic variables with the highest similarity as template time series characteristic variables; a reference variable module 150 for searching for a third number of time series characteristic variables with the highest similarity for each of the template time series characteristic variables in the historical data KMeans clustering model as reference time series characteristic variables; a reference power generation module 160 for calculating the reference power generation efficiency corresponding to each of the reference time series characteristic variables; a target power generation module 170 for determining a target power generation efficiency based on all the reference power generation efficiencies; and an evaluation module 180 for evaluating the power generation data of the day to be evaluated according to the target power generation efficiency to obtain a power generation evaluation result. The present invention uses historical power generation records to form a KMeans clustering model. When it is necessary to evaluate the power generation data of a day, it can search for the historical day power generation data with the highest similarity from the historical power generation records, and deduce the power generation efficiency of the power generation data of the day to be evaluated based on the power generation efficiencies of these historical day power generation data to complete the evaluation. Compared with other related technologies, the present invention can achieve high-accuracy power generation evaluation only by using a small amount of historical power generation records, helping power station managers make more reasonable operation decisions. At the same time, the present invention also greatly reduces the requirement for the amount of historical power generation data, enabling the method of the present invention to be applicable to old power stations with missing historical data and newly built power stations with insufficient historical data, greatly broadening the applicable scope of the power generation evaluation method.
[0111] The power generation evaluation device based on KMeans clustering in this embodiment is used to implement the aforementioned power generation evaluation method based on KMeans clustering. Therefore, the specific implementation manners in the power generation evaluation device based on KMeans clustering can be seen in the embodiment part of the power generation evaluation method based on KMeans clustering in the foregoing text. For example, the first receiving module 110, the transformation module 120, the variable selection module 130, the similarity module 140, the reference variable module 150, the reference power generation module 160, the target power generation module 170, and the evaluation module 180 are respectively used to implement steps S101, S102, S103, S104, S105, S106, S107, and S108 in the aforementioned power generation evaluation method based on KMeans clustering. Therefore, the specific implementation manners can refer to the descriptions of the corresponding parts of the embodiments and will not be elaborated here.
[0112] The present invention also provides a training method for a historical data KMeans clustering model having the above beneficial effects. A schematic flowchart of a specific implementation manner thereof is as shown in Figure 3 Figure 3, which is called the third specific implementation manner. The training method for the historical data KMeans clustering model is used for any one of the above-mentioned power generation amount evaluation methods based on KMeans clustering, and includes:
[0113] S201: Receive historical power generation records; the historical power generation records include multiple pieces of historical daily power generation data.
[0114] The historical data KMeans clustering model in this specific implementation manner serves the power generation amount evaluation method based on KMeans clustering in the foregoing text. Therefore, the entry items in a single piece of the historical daily power generation data should correspond to the entry items of the to-be-evaluated daily power generation data in the foregoing text.
[0115] In other words, the historical power generation records in this step are a set of the historical daily power generation data, and the historical daily power generation data corresponds to a date, and one date corresponds to one piece of historical daily power generation data.
[0116] S202: Convert all the historical daily power generation data into time series feature variables corresponding to the dates.
[0117] The time series feature variables corresponding to the dates also correspond to the to-be-evaluated daily feature variables in the foregoing text. In other words, similarity comparison can be performed.
[0118] S203: Select the first number of time series feature variables from all the time series feature variables as clustering centers, and perform clustering processing on the remaining time series feature variables so that each of the remaining time series feature variables has a corresponding clustering center, and the Euclidean distance from the remaining time series feature variable to the corresponding clustering center is not greater than the Euclidean distance from the remaining time series feature variable to the other clustering centers.
[0119] S204: Classify the clustering centers and the remaining time series feature variables corresponding to the clustering centers into the same cluster, and perform KMeans clustering iteration on the clustering centers within the same cluster until the clustering centers of all the clusters remain unchanged between two consecutive iterations, so as to obtain a historical data KMeans clustering model.
[0120] Specifically, the following formula (3) can be used to update the clustering centers:
[0121] ; (3)
[0122] where is the center of the w-th iteration and the i-th cluster, is the number of time series feature variables in the i-th cluster at the (w - 1)-th iteration, indicates the j-th time series feature variable in the i-th cluster, is the i-th cluster at the (w - 1)-th iteration.
[0123] Preferably, before converting all historical daily power generation data into time series feature variables corresponding to dates, it further includes:
[0124] B1: Determine the historical power generation records with missing power generation data in the entries; the historical power generation records with missing power generation data are the historical daily power generation data with blank entry items.
[0125] Illustratively, assume that the entry items of the complete historical daily power generation data include 6 items: the daily equivalent hours of the power station filled with data, the daily equivalent hours of the inverter, the daily irradiance, the daily average temperature, the weather condition, and the date. Then the historical power generation records with missing power generation data are the historical daily power generation data with blank items among the above 6 items, such as the historical daily power generation data without recording the daily average temperature, or the historical daily power generation data without recording the daily equivalent hours of the power station in the historical daily power generation data set, etc.
[0126] B2: Exclude the historical power generation records with missing power generation data from the historical power generation records.
[0127] In this preferred embodiment, before converting the historical daily power generation data, the historical daily power generation data is first checked, and the historical daily power generation data with missing data in the entry items is excluded, which can further improve the working stability of the model, avoid data pollution, and improve the evaluation accuracy.
[0128] As another preferred embodiment, the entry items of the historical daily power generation records include the daily equivalent hours;
[0129] Before converting all historical daily power generation data into time series feature variables corresponding to dates, it further includes:
[0130] C1: Sequentially determine whether the daily equivalent hours of each piece of historical daily power generation data exceed the preset theoretical upper limit of sunshine.
[0131] Preferably, the range of the theoretical upper limit of sunshine is from 8 hours to 10 hours, including the endpoint values, such as any one of 8.0 hours, 9.3 hours, or 10.0 hours. The above range is a preferred range after a large number of theoretical calculations and actual tests. Of course, it can also be adjusted according to the actual situation, and the details are not elaborated in this invention.
[0132] C2: When there is historical daily power generation data with the daily equivalent hours exceeding the theoretical upper limit of sunshine, regard the historical daily power generation data with the daily equivalent hours exceeding the theoretical upper limit of sunshine as the power generation data with data errors.
[0133] C3: Eliminate the erroneous power generation data from the historical power generation record.
[0134] For power stations located in fixed geographical locations, there is an upper limit to the number of daily equivalent hours throughout the year, which is also known as the theoretical upper limit of sunshine. Therefore, if the number of daily equivalent hours exceeds the theoretical upper limit of sunshine, it indicates that there is a data error, and in this case, the data (i.e., the data-error power generation data) is discarded.
[0135] As a specific implementation, the method for determining the first quantity includes:
[0136] The first quantity is obtained by the following formula (4):
[0137] m=n / k; (4)
[0138] Wherein, m is the first number, n is the total number of historical daily power generation data included in the historical power generation record; k is a preset grouping coefficient;
[0139] The grouping coefficient ranges from 5 to 20, including endpoint values, such as any one of 5.0, 12.0 or 20.0. In other words, in this specific implementation, the number of clusters in the KMeans clustering model of the historical data can be adjusted by the grouping coefficient. In order to maintain the number of clusters within a reasonable range, the grouping coefficient can be adjusted according to the total number of historical daily power generation data included in the historical power generation record. In other words, the larger n is, the larger the value of k can be selected, so that m can maintain the historical power generation records in various situations within a certain reasonable range.
[0140] The training method of the historical data KMeans clustering model provided by the present invention includes receiving historical power generation records; the historical power generation records include multiple pieces of historical daily power generation data; converting all the historical daily power generation data into time series feature variables corresponding to the dates; selecting the first number of time series feature variables from all the time series feature variables as the clustering centers, and performing clustering processing on the remaining time series feature variables, so that each of the remaining time series feature variables has a corresponding clustering center, and the Euclidean distance from the remaining time series feature variable to the corresponding clustering center is not greater than the Euclidean distance from the remaining time series feature variable to the other clustering centers; classifying the clustering centers and the remaining time series feature variables corresponding to the clustering centers into the same cluster, and performing KMeans clustering iteration on the clustering centers within the same cluster until the clustering centers of all clusters remain unchanged between two consecutive iterations, thereby obtaining the historical data KMeans clustering model. The present invention uses historical power generation records to form a KMeans clustering model. When it is necessary to evaluate daily power generation data, the historical daily power generation data with the highest similarity can be searched from the historical power generation records, and the power generation efficiency of the daily power generation data to be evaluated can be deduced based on the power generation efficiency of these historical daily power generation data to complete the evaluation. Compared with other related technologies, the present invention can achieve high-accuracy power generation evaluation only by using a small amount of historical power generation records, helping power station managers make more reasonable operation decisions; at the same time, the present invention also greatly reduces the demand for historical power generation data volume, enabling the method of the present invention to be applicable to old power stations with missing historical data and newly built power stations with insufficient historical data, greatly broadening the applicable scope of the power generation evaluation method.
[0141] The training device of the historical data KMeans clustering model provided by the embodiments of the present invention will be introduced below. The training device of the historical data KMeans clustering model described below can be correspondingly referred to the training method of the historical data KMeans clustering model described above.
[0142] Figure 4 It is a structural block diagram of the training device of the historical data KMeans clustering model provided by the embodiments of the present invention, which is called the fourth specific implementation manner. Refer to Figure 4 The training device of the historical data KMeans clustering model may include:
[0143] A second receiving module 210, configured to receive historical power generation records; the historical power generation records include multiple pieces of historical daily power generation data;
[0144] A date variable transformation module 220, configured to convert all the historical daily power generation data into time series feature variables corresponding to the dates;
[0145] The preliminary clustering module 230 is used to select the first number of time series feature variables from all the time series feature variables as the clustering centers, and perform clustering processing on the remaining time series feature variables, so that each of the remaining time series feature variables has a corresponding clustering center, and the Euclidean distance from the remaining time series feature variable to the corresponding clustering center is not greater than the Euclidean distance from the remaining time series feature variable to the other clustering centers;
[0146] The clustering iteration module 240 is used to classify the clustering centers and the remaining time series feature variables corresponding to the clustering centers into the same cluster, and perform KMeans clustering iteration on the clustering centers within the same cluster until the clustering centers of all clusters remain unchanged between two consecutive iterations, obtaining a historical data KMeans clustering model.
[0147] As a preferred implementation manner, the date variable transformation module 220 further includes:
[0148] The missing determination unit is used to determine that there is missing power generation data in the entries of the historical power generation records; the entry missing power generation data is the historical daily power generation data with blank entry items;
[0149] The missing elimination unit is used to eliminate the entry missing power generation data from the historical power generation records.
[0150] As a preferred implementation manner, the entry items of the historical daily power generation records include daily equivalent hours;
[0151] The date variable transformation module 220 further includes:
[0152] The error judgment unit is used to sequentially judge whether the daily equivalent hours of each piece of historical daily power generation data exceed a preset theoretical upper limit of sunshine;
[0153] The error determination unit is used to, when there is historical daily power generation data whose daily equivalent hours exceed the theoretical upper limit of sunshine, use the historical daily power generation data whose daily equivalent hours exceed the theoretical upper limit of sunshine as data error power generation data;
[0154] The error elimination unit is used to eliminate the data error power generation data from the historical power generation records.
[0155] As a preferred implementation manner, it further includes:
[0156] The first number determination unit is used to obtain the first number through the following formula:
[0157] m = n / k;
[0158] Wherein, m is the first quantity, n is the total number of historical daily power generation data included in the historical power generation record; k is a preset grouping coefficient;
[0159] The range of the grouping coefficient is from 5 to 20, including the endpoint values.
[0160] The training device for the historical data KMeans clustering model provided by the present invention, through the second receiving module 210, is used to receive the historical power generation record; the historical power generation record includes multiple historical daily power generation data; the date variable transformation module 220 is used to convert all the historical daily power generation data into time series feature variables corresponding to the dates; the preliminary clustering module 230 is used to select the first quantity of time series feature variables from all the time series feature variables as the clustering centers, and perform clustering processing on the remaining time series feature variables, so that each remaining time series feature variable has a corresponding clustering center, and the Euclidean distance from the remaining time series feature variable to the corresponding clustering center is not greater than the Euclidean distance from the remaining time series feature variable to the other clustering centers; the clustering iteration module 240 is used to classify the clustering centers and the remaining time series feature variables corresponding to the clustering centers into the same cluster, and perform KMeans clustering iteration on the clustering centers within the same cluster until the clustering centers of all clusters remain unchanged between two consecutive iterations, so as to obtain the historical data KMeans clustering model. The present invention uses the historical power generation record to form a KMeans clustering model. When it is necessary to evaluate the daily power generation data, the historical daily power generation data with the highest similarity can be searched from the historical power generation record, and the power generation efficiency of the daily power generation data to be evaluated can be deduced according to the power generation efficiency of these historical daily power generation data to complete the evaluation. Compared with other related technologies, the present invention can achieve high-accuracy power generation evaluation only by using a small amount of historical power generation records, helping the power station manager to make more reasonable operation decisions; at the same time, the present invention also greatly reduces the demand for the amount of historical power generation data, enabling the method of the present invention to be applicable to old power stations with missing historical data and newly built power stations with insufficient historical data, greatly broadening the applicable range of the power generation evaluation method.
[0161] The training device for the historical data KMeans clustering model in this embodiment is used to implement the foregoing training method for the historical data KMeans clustering model. Therefore, the specific implementation manners in the training device for the historical data KMeans clustering model can be seen in the embodiment part of the training method for the historical data KMeans clustering model in the foregoing text. For example, the second receiving module 210, the date variable transformation module 220, the preliminary clustering module 230, and the clustering iteration module 240 are respectively used to implement steps S201, S202, S203, and S204 in the foregoing training method for the historical data KMeans clustering model. Therefore, the specific implementation manners can be referred to the descriptions of the corresponding various part embodiments and will not be elaborated herein.
[0162] The present invention also provides a power generation amount evaluation device based on KMeans clustering, including:
[0163] A memory for storing a computer program;
[0164] A processor for implementing the steps of any one of the above-mentioned power generation amount evaluation methods based on KMeans clustering when executing the computer program. The power generation amount evaluation method based on KMeans clustering provided by the present invention includes receiving power generation data of the day to be evaluated; converting the power generation data of the day to be evaluated into characteristic variables of the day to be evaluated; selecting a first number of time series characteristic variables in a pre-trained historical data KMeans clustering model, where the clusters to which the selected time series characteristic variables belong are different, and the historical data KMeans clustering model has the first number of clusters; calculating the similarity between each of the time series characteristic variables and the characteristic variables of the day to be evaluated, and selecting a second number of time series characteristic variables with the highest similarity as template time series characteristic variables; searching for a third number of time series characteristic variables with the highest similarity for each of the template time series characteristic variables in the historical data KMeans clustering model as reference time series characteristic variables; calculating the reference power generation efficiency corresponding to each of the reference time series characteristic variables; determining the target power generation efficiency according to all the reference power generation efficiencies; and evaluating the power generation data of the day to be evaluated according to the target power generation efficiency to obtain a power generation amount evaluation result. The present invention uses historical power generation records to form a KMeans clustering model. When it is necessary to evaluate the power generation data of a day, the historical power generation data with the highest similarity can be searched from the historical power generation records, and the power generation efficiency of the power generation data of the day to be evaluated can be deduced according to the power generation efficiency of these historical power generation data to complete the evaluation. Compared with other related technologies, the present invention can achieve high-accuracy power generation amount evaluation only by using a small amount of historical power generation records, helping power station managers make more reasonable operation decisions; at the same time, the present invention also greatly reduces the requirement for the amount of historical power generation data, enabling the method of the present invention to be applicable to old power stations with missing historical data and newly built power stations with insufficient historical data, greatly broadening the applicable scope of the power generation amount evaluation method.
[0165] The present invention also provides a training device for a historical data KMeans clustering model, including:
[0166] A memory for storing a computer program;
[0167] A processor, which is used to implement the steps of the training method of the historical data KMeans clustering model described above when executing the computer program. The training method of the historical data KMeans clustering model provided by the present invention includes receiving historical power generation records; the historical power generation records include multiple pieces of historical daily power generation data; converting all the historical daily power generation data into time series feature variables corresponding to dates; selecting the first number of time series feature variables from all the time series feature variables as clustering centers, and performing clustering processing on the remaining time series feature variables, so that each remaining time series feature variable has a corresponding clustering center, and the Euclidean distance from the remaining time series feature variable to the corresponding clustering center is not greater than the Euclidean distance from the remaining time series feature variable to the other clustering centers; classifying the clustering centers and the remaining time series feature variables corresponding to the clustering centers into the same cluster, and performing KMeans clustering iteration on the clustering centers within the same cluster until the clustering centers of all clusters remain unchanged between two consecutive iterations, obtaining a historical data KMeans clustering model. The present invention uses historical power generation records to form a KMeans clustering model. When it is necessary to evaluate daily power generation data, the historical daily power generation data with the highest similarity can be searched from the historical power generation records, and the power generation efficiency of the to-be-evaluated daily power generation data can be deduced based on the power generation efficiency of these historical daily power generation data to complete the evaluation. Compared with other related technologies, the present invention can achieve high-accuracy power generation evaluation only by using a small amount of historical power generation records, helping power station managers make more reasonable operation decisions; at the same time, the present invention also greatly reduces the requirement for the amount of historical power generation data, enabling the method of the present invention to be applicable to old power stations with missing historical data and newly built power stations with insufficient historical data, greatly broadening the applicable range of the power generation evaluation method.
[0168] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned power generation amount evaluation methods based on KMeans clustering and / or the steps of the training method of the historical data KMeans clustering model are implemented. The power generation amount evaluation method based on KMeans clustering provided by the present invention includes receiving power generation data of a day to be evaluated; converting the power generation data of the day to be evaluated into feature variables of the day to be evaluated; in a pre-trained historical data KMeans clustering model, selecting a first number of time series feature variables; the clusters to which the selected time series feature variables belong are different, and the historical data KMeans clustering model has the first number of clusters; calculating the similarity between each of the time series feature variables and the feature variables of the day to be evaluated, and selecting a second number of time series feature variables with the highest similarity as template time series feature variables; searching for a third number of time series feature variables with the highest similarity for each of the template time series feature variables in the historical data KMeans clustering model as reference time series feature variables; calculating the reference power generation efficiency corresponding to each of the reference time series feature variables; determining the target power generation efficiency according to all the reference power generation efficiencies; and evaluating the power generation data of the day to be evaluated according to the target power generation efficiency to obtain a power generation amount evaluation result. The present invention uses historical power generation records to form a KMeans clustering model. When it is necessary to evaluate the daily power generation data, the historical daily power generation data with the highest similarity can be searched from the historical power generation records, and the power generation efficiency of the power generation data of the day to be evaluated can be deduced according to the power generation efficiency of these historical daily power generation data to complete the evaluation. Compared with other related technologies, the present invention can achieve high-accuracy power generation amount evaluation only by using a small amount of historical power generation records, helping power station managers make more reasonable operation decisions; at the same time, the present invention also greatly reduces the requirement for the amount of historical power generation data, enabling the method of the present invention to be applicable to old power stations with missing historical data and newly built power stations with insufficient historical data, greatly broadening the applicable scope of the power generation amount evaluation method.
[0169] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0170] It should be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0171] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0172] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0173] The above has introduced in detail the power generation evaluation method, device, and the training method and device of the historical data KMeans clustering model based on KMeans clustering provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A power generation evaluation method based on KMeans clustering, characterized in that: include: Receive daily power generation data to be evaluated; Converting the power generation data of the day to be evaluated into characteristic variables of the day to be evaluated; In a pre-trained historical data KMeans clustering model, a first number of time series feature variables are selected; the selected time series feature variables belong to different clusters, and the historical data KMeans clustering model has the first number of clusters; Calculating the similarity between each of the time series characteristic variables and the characteristic variable of the day to be evaluated, and selecting a second number of time series characteristic variables with the highest similarity as the model time series characteristic variables; Searching for the third number of time series feature variables with the highest similarity for each of the template time series feature variables in the historical data KMeans clustering model as reference time series feature variables; Calculating the reference power generation efficiency corresponding to each of the reference time series characteristic variables; Determine the target power generation efficiency according to all reference power generation efficiencies; The daily power generation data to be evaluated is evaluated according to the target power generation efficiency to obtain a power generation evaluation result.
2. The power generation evaluation method based on KMeans clustering according to claim 1, characterized in that: Based on all reference power generation efficiencies, the target power generation efficiency is determined including: The average value of all reference power generation efficiencies is calculated as the target power generation efficiency.
3. The power generation evaluation method based on KMeans clustering according to claim 1, characterized in that: The second number and / or the third number ranges from 4 to 8, both inclusive.
4. The power generation evaluation method based on KMeans clustering according to claim 1, characterized in that: After receiving the daily power generation data to be evaluated, it also includes: Normalizing the daily power generation data to be evaluated; Correspondingly, the data corresponding to the time series feature variables in the historical data KMeans clustering model are normalized data.
5. A power generation evaluation device based on KMeans clustering, characterized in that: include: A first receiving module, used for receiving daily power generation data to be evaluated; A transformation module, used for converting the power generation data of the day to be evaluated into characteristic variables of the day to be evaluated; A variable selection module, used for selecting a first number of time series feature variables in a pre-trained historical data KMeans clustering model; the selected time series feature variables belong to different clusters, and the historical data KMeans clustering model has the first number of clusters; A similarity module, used for calculating the similarity between each of the time series characteristic variables and the characteristic variable of the day to be evaluated, and selecting a second number of time series characteristic variables with the highest similarity as the model time series characteristic variables; A reference variable module, used for searching the third number of time series feature variables with the highest similarity for each of the template time series feature variables in the historical data KMeans clustering model as reference time series feature variables; A reference power generation module, used to calculate the reference power generation efficiency corresponding to each of the reference time series characteristic variables; A target power generation module, used to determine a target power generation efficiency according to all reference power generation efficiencies; The evaluation module is used to evaluate the daily power generation data to be evaluated according to the target power generation efficiency to obtain a power generation evaluation result.
6. A training method for a historical data KMeans clustering model, characterized in that: The training method of the historical data KMeans clustering model is used in the power generation evaluation method based on KMeans clustering as claimed in any one of claims 1 to 4, comprising: Receiving historical power generation records; the historical power generation records include a plurality of historical daily power generation data; Convert all historical daily power generation data into time series characteristic variables corresponding to the date; Selecting the first number of time series feature variables from all time series feature variables as cluster centers, and performing clustering processing on the remaining time series feature variables, so that each remaining time series feature variable has a corresponding cluster center, and the Euclidean distance from the remaining time series feature variable to the corresponding cluster center is not greater than the Euclidean distance from the remaining time series feature variable to the remaining cluster centers; The cluster centers and the remaining time series feature variables corresponding to the cluster centers are classified into the same cluster, and KMeans clustering iterations are performed on the cluster centers in the same cluster until the cluster centers of all clusters remain unchanged between two consecutive iterations, thereby obtaining a historical data KMeans clustering model.
7. The method for training a historical data KMeans clustering model according to claim 6, characterized in that: Before converting all historical daily power generation data into time series characteristic variables corresponding to the date, it also includes: Determine that an entry in the historical power generation record lacks power generation data; the entry lacking power generation data is historical daily power generation data with blank entries; The entry with missing power generation data is removed from the historical power generation record.
8. The method for training a historical data KMeans clustering model according to claim 6, characterized in that: The entries of the historical daily power generation record include daily equivalent hours; Before converting all historical daily power generation data into time series characteristic variables corresponding to the date, it also includes: Determine in turn whether the daily equivalent hours of each piece of historical daily power generation data exceeds the preset theoretical upper limit of sunshine; When there is historical daily power generation data in which the daily equivalent hours exceed the theoretical upper limit of sunshine, the historical daily power generation data in which the daily equivalent hours exceed the theoretical upper limit of sunshine is regarded as data error power generation data; The power generation data with data errors is removed from the historical power generation records.
9. The method for training a historical data KMeans clustering model according to claim 6, characterized in that: The method for determining the first quantity includes: The first quantity is obtained by the following formula: m=n / k; Wherein, m is the first number, n is the total number of historical daily power generation data included in the historical power generation record; k is a preset grouping coefficient; The grouping coefficient ranges from 5 to 20, inclusive.
10. A training device for a historical data KMeans clustering model, characterized in that: The training device of the historical data KMeans clustering model is used in the power generation evaluation method based on KMeans clustering as claimed in any one of claims 1 to 4, comprising: A second receiving module is used to receive historical power generation records; the historical power generation records include a plurality of historical daily power generation data; The date variable conversion module is used to convert all historical daily power generation data into time series characteristic variables corresponding to the date; A preliminary clustering module, used for selecting the first number of time series feature variables from all time series feature variables as cluster centers, and performing clustering processing on the remaining time series feature variables, so that each remaining time series feature variable has a corresponding cluster center, and the Euclidean distance from the remaining time series feature variable to the corresponding cluster center is not greater than the Euclidean distance from the remaining time series feature variable to the remaining cluster centers; The clustering iteration module is used to classify the cluster centers and the remaining time series feature variables corresponding to the cluster centers into the same cluster, and perform KMeans clustering iteration on the cluster centers in the same cluster until the cluster centers of all clusters remain unchanged between two consecutive iterations, thereby obtaining a KMeans clustering model of historical data.