Photovoltaic output data similar daily clustering method and system based on integrated clustering algorithm
Through the combination of integrated clustering algorithm and Hungarian algorithm, the instability and information loss of photovoltaic output data are solved, and the efficient similar daily clustering of photovoltaic output data is achieved, which improves the prediction accuracy and management efficiency of data.
Patent Information
- Application Number
- CN202510433173.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the random volatility and meteorological characteristics of photovoltaic output data lead to unstable clustering results. Traditional single clustering algorithms are susceptible to initial parameters and noise easily cause cluster center drift, and cluster label arrangement and combination explosions lead to information loss.
The integrated clustering algorithm is adopted, combined with the base clusterer and the Hungarian algorithm, and the photovoltaic output data is diversified and initial clustering is carried out, and the co-coordination matrix construction and label alignment mechanism is constructed to achieve similar daily clustering of photovoltaic output data.
It improves the stability of photovoltaic output data clustering results and the comparability of cross-cluster results, provides standardized co-coordinated matrix input, and improves the prediction accuracy and management efficiency of photovoltaic output data.
Smart Images

Figure CN120408229A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of photovoltaic power generation data processing, and in particular to a method and system for clustering photovoltaic output data on similar days based on an integrated clustering algorithm, as well as a corresponding computer terminal and computer-readable storage medium. Background Art
[0002] As one of the world's most abundant and promising renewable energy sources, solar energy plays a vital role in the clean energy sector, and photovoltaic power generation has rapidly become a focal point for solar energy applications. However, the intermittent and volatile nature of solar power leads to poor stability in photovoltaic power generation, complicating grid management and making accurate photovoltaic power forecasting a bottleneck for its large-scale application. Accurate photovoltaic power forecasting is crucial for enhancing system reliability and flexibility in power system planning and decision-making. Clustering of photovoltaic output data plays a crucial role in photovoltaic power forecasting, improving prediction accuracy, simplifying prediction models, revealing underlying patterns, assisting in outlier detection, and optimizing energy management.
[0003] Photovoltaic output data is not only affected by various meteorological factors but also exhibits significant random fluctuations. Effective preprocessing of raw photovoltaic data is a technical bottleneck that urgently needs to be addressed in this field. Currently, no descriptions or reports of technologies similar to the present invention have been found, and no similar data has been collected domestically or internationally. Summary of the Invention
[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method and system for clustering photovoltaic output data on similar days based on an integrated clustering algorithm, and also provides a corresponding computer terminal and computer-readable storage medium.
[0005] According to one aspect of the present invention, a method for clustering photovoltaic output data on similar days based on an integrated clustering algorithm is provided, comprising:
[0006] selecting a base clusterer, and performing clustering processing on the photovoltaic output data set based on the base clusterer to obtain a single clustering result;
[0007] The Hungarian algorithm is used to unify the classification labels of different single clustering results to obtain the aligned single clustering results;
[0008] Based on the aligned single clustering result, a co-cooperation matrix is constructed, which is used to represent the frequency of samples being assigned to the same cluster in all clustering results;
[0009] The co-correlation matrix is used as a similarity matrix between samples, and the co-correlation matrix is clustered to achieve the final clustering division of the photovoltaic output data set, thereby obtaining a data set after similar day clustering.
[0010] According to another aspect of the present invention, there is provided a similar-day clustering system for photovoltaic output data based on an integrated clustering algorithm, including:
[0011] A single clustering module, which is used to select a base clusterer and perform clustering processing on a photovoltaic output data set based on the base clusterer to obtain a single clustering result;
[0012] A single result alignment module, which uses the Hungarian algorithm to perform uniform processing on the classification labels of different single clustering results to obtain the aligned single clustering result;
[0013] A co-covariance matrix module, which constructs a co-covariance matrix based on the aligned single clustering result, and the co-covariance matrix is used to represent the frequency of samples being assigned to the same cluster in all clustering results;
[0014] A clustering division module, which is used to use the co-covariance matrix as a similarity matrix between samples, perform clustering on the co-covariance matrix, realize the final clustering division of the photovoltaic output data set, and obtain a data set after similar-day clustering.
[0015] According to a third aspect of the present invention, there is provided a computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method described above in the present invention, or run the system described above in the present invention.
[0016] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute the method described above in the present invention, or run the system described above in the present invention.
[0017] Due to the adoption of the above technical solutions, compared with the prior art, the present invention has at least one of the following beneficial effects:
[0018] The present invention provides a method and system for similar-day clustering of photovoltaic output data based on an integrated clustering algorithm. In view of the characteristics such as the random volatility of photovoltaic output data, by using heterogeneous base clusterers to perform diversified initial clustering on photovoltaic output curves, the problems that the traditional single clustering algorithm is easily affected by initial parameters resulting in unstable results and data noise is likely to cause the drift of clustering centers are solved, and the effect of improving the stability of the clustering results of photovoltaic output data is achieved.
[0019] The present invention provides a method and system for clustering similar days of photovoltaic output data based on an integrated clustering algorithm. Through the label alignment mechanism of the Hungarian algorithm, the problem of loss of consensus information of photovoltaic output data caused by the explosion of permutations and combinations of traditional clustering labels is solved, thereby achieving the effect of improving the comparability of results across clusterers and providing standardized input for the subsequent construction of a consensus matrix for photovoltaic output data. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0021] Figure 1 The figure is a workflow diagram of a photovoltaic output data similar day clustering method based on an integrated clustering algorithm in a preferred embodiment of the present invention.
[0022] Figure 2 The figure is a schematic diagram of the component modules of a photovoltaic output data similar day clustering system based on an integrated clustering algorithm in a preferred embodiment of the present invention.
[0023] Figure 3 This is a power generation curve diagram under three types of weather modes in a verification example of the present invention. DETAILED DESCRIPTION
[0024] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process. It should be noted that those skilled in the art may make various modifications and improvements without departing from the scope of the present invention, and these modifications and improvements fall within the scope of protection of the present invention.
[0025] Photovoltaic output data is not only affected by various meteorological characteristics but also exhibits significant random fluctuations. To better understand and predict the changing patterns of photovoltaic output, an embodiment of the present invention provides a method for clustering photovoltaic output data on similar days based on an integrated clustering algorithm. This method clusters photovoltaic output data using a clustering algorithm based on meteorological characteristics to identify output patterns under different meteorological characteristics.
[0026] Specifically, if Figure 1 As shown, the photovoltaic output data similar day clustering method based on the integrated clustering algorithm provided in this embodiment may include:
[0027] S1, select a base clusterer, perform clustering on the photovoltaic output dataset based on the base clusterer, and obtain a single clustering result;
[0028] S2, using the Hungarian algorithm to unify the classification labels of different single clustering results to obtain the aligned single clustering results;
[0029] S3. Based on the aligned single clustering result, construct a co-association matrix, which is used to represent the frequency that samples are assigned to the same cluster in all clustering results.
[0030] S4. Use the co-association matrix as the similarity matrix between samples, perform clustering on the co-association matrix, achieve the final clustering division of the photovoltaic output dataset, and obtain the dataset after clustering of similar days.
[0031] In some preferred embodiments, before selecting the base clusterer, the above method may further include:
[0032] S0. Select clustering features, including:
[0033] S01. Analyze the core features related to photovoltaic power based on Pearson, Spearman, and Kendall correlation coefficients, including global radiation and diffuse radiation.
[0034] S02. Based on the obtained core features, construct time series features based on daily statistical summaries; divide the time series features by day, and calculate the mean, peak, and standard deviation statistics for each day.
[0035] S03. Use gradient boosting decision tree and random forest algorithms to select the time series features, measure the contributions of different features, and finally normalize the scores respectively to screen the time series features and obtain the required clustering features.
[0036] In some preferred embodiments, for the above S1, when selecting the base clusterer and performing clustering processing on the photovoltaic output dataset based on the base clusterer to obtain a single clustering result, it may further include:
[0037] The base clusterers include: K-Medoids algorithm, spectral clustering algorithm, AGNES (agglomerative nesting) algorithm, and BIRCH (balanced iterative reducing and clustering using hierarchies) algorithm; where:
[0038] S11. The K-Medoids algorithm minimizes the sum of the distances from points within the cluster to the cluster center by selecting the cluster center point, as shown in Equation (1):
[0039]
[0040] In the formula, m i is the updated cluster center of the i-th cluster, C i is the i-th cluster, x is the candidate cluster center, and x j is the sample in cluster C i and d(xj , x) is a distance function that calculates the distance between x j and x, k is the number of clusters, d(x j , m i ) is a distance function that calculates the distance between x j and m i ;
[0041] S12. In the spectral clustering algorithm, by constructing a similarity matrix and a Laplacian matrix, and using eigenvalue decomposition to extract the main eigenvectors, the data point clustering information is mapped in a low-dimensional space, as shown in Equation (2):
[0042] L = D - W (7)
[0043] In the formula, W is an n×n symmetric matrix representing the similarity between nodes, D is the degree matrix representing the connection strength between nodes, and L is the Laplacian matrix;
[0044] S13. The AGNES algorithm adopts a bottom-up hierarchical clustering method and forms a hierarchical structure by gradually merging the most similar clusters; among them:
[0045] The core clustering merge formula of the AGNES algorithm is divided into the shortest distance H min , the longest distance H max and the average distance H avg , as shown in Equations (3) to (5):
[0046]
[0047]
[0048]
[0049] In the formula, C i , C j are clusters, e and f are any two points in the cluster, and dist(e, f) is a distance calculation function used to measure the distance between e and f;
[0050] S14. The BIRCH algorithm realizes data clustering by constructing and refining a clustering feature tree; among them:
[0051] S141. Construct a clustering feature tree;
[0052] S142. Add all samples in the photovoltaic output dataset to the clustering feature tree respectively, and calculate the radius of the leaf node after adding the new sample;
[0053] S143. According to the radius of the obtained leaf node, merge the sample into the corresponding cluster and update all clustering feature triples on the clustering feature tree path;
[0054] S144. Treat the clusters represented by each leaf node in the clustering feature tree as data points, perform secondary clustering, and optimize the clustering feature tree.
[0055] S145. Define the photovoltaic output data set as D = (x1, x2, …, x n ). Through feature engineering, select the feature subset related to the photovoltaic output, and use the selected base clusterer to cluster the data. Each base clusterer is set to generate m cluster classes, and finally obtain different single clustering results.
[0056] In some preferred embodiments, the above S141, constructing the clustering feature tree, may further include:
[0057] Determine three parameters B, L, and T of the clustering feature tree; where B is the maximum number of CFs of the internal node; L is the maximum number of CFs of the leaf node; T is the maximum sample radius threshold of each CF of the leaf node; CF = (N, L S , S S ) represents the clustering feature triple; N represents the number of samples in each CF; L S represents the sum vector of the feature values of each sample point in this CF; S s represents the sum of the squares of the feature values of the samples in this CF; L S and S s are calculated according to formulas (6) and (7) as follows:
[0058]
[0059]
[0060] In the formula: x n represents the nth sample vector; x nm represents the mth feature value of the nth sample point; M represents the number of features of the data in the sample point.
[0061] In some preferred embodiments, the above S142, adding all samples in the photovoltaic output data set into the clustering feature tree and calculating the radius of the leaf node after adding the new sample, may further include:
[0062] Read all samples into the clustering feature tree in sequence, find the leaf node A closest to the new sample, and calculate the radius R of the leaf node A after adding the new sample. The calculation method is as shown in formula (8):
[0063]
[0064] In the formula, x n is the nth sample vector, and x0 is the initial sample vector.
[0065] In some preferred embodiments, for the above S143, according to the radius of the obtained leaf node, merging the sample into the corresponding cluster and updating all the clustering feature triples on the clustering feature tree path may further include:
[0066] S1431, if the radius R of the obtained leaf node ≤ T, then merge the sample into the cluster and update all the clustering feature triples CF on the CF-tree path of the clustering feature tree, and perform the following steps:
[0067] Check whether the parent node needs to be split from bottom to top. If it needs to be split, then use the leaf node splitting method until the root node;
[0068] S1432, if the radius R of the obtained leaf node > T, then perform the following steps:
[0069] Judge whether the number of CFs of leaf node A reaches the maximum L; if it is less than L, then create a new CF node to store the new sample and update all the CF tuples on the CF-tree path; otherwise, select the two CFs with the farthest hypersphere radius distance among all the CF tuples in leaf node A as the new leaf nodes A' and A'' to replace the original leaf node A, and put the CF tuples and new sample tuples in the original leaf node into the new leaf nodes A' and A'' in order of distance.
[0070] In some preferred embodiments, for the above S2, using the Hungarian algorithm to perform consistency processing on the classification labels of different single clustering results may further include:
[0071] Taking the classification label of a certain single clustering result as a benchmark, using the Hungarian algorithm to align the classification labels of other single clustering results.
[0072] In some preferred embodiments, for the above S3, based on the aligned single clustering results, constructing a co-association matrix may further include:
[0073] For the given k clustering results C1, C2, …, C k , each element CA(i, j) of the co-association matrix CA is as shown in the formula:
[0074]
[0075] In the formula, φ k represents the label mapping function of the k-th clustering result; Ι(·) represents the indicator function, which takes the value of 1 when the condition is satisfied and 0 otherwise; CA(i, j) represents the frequency that sample x i and sample x j are assigned to the same cluster in all clustering results, and the value range is [0, 1].
[0076] In some preferred embodiments, step S4 may further include:
[0077] The BIRCH algorithm is used to cluster the co-association matrix to obtain the final clustering result. For the specific algorithm, reference can be made to the relevant steps in S14, which will not be elaborated here.
[0078] Based on the same inventive concept, an embodiment of the present invention further provides a photovoltaic output data similar-day clustering system based on an integrated clustering algorithm.
[0079] Specifically, as Figure 2 shown, the photovoltaic output data similar-day clustering system based on the integrated clustering algorithm provided by this embodiment may include:
[0080] A single clustering module, which is used to select a base clusterer and perform clustering processing on the photovoltaic output data set based on the base clusterer to obtain a single clustering result;
[0081] A single result alignment module, which uses the Hungarian algorithm to perform consistency processing on the classification labels of different single clustering results to obtain the aligned single clustering result;
[0082] A co-association matrix module, which constructs a co-association matrix based on the aligned single clustering results. The co-association matrix is used to represent the frequency of samples being assigned to the same cluster in all clustering results;
[0083] A clustering division module, which is used to use the co-association matrix as the similarity matrix between samples, cluster the co-association matrix, and realize the final clustering division of the photovoltaic output data set to obtain the data set after similar-day clustering.
[0084] The working content of each functional module of the photovoltaic output data similar-day clustering system based on the integrated clustering algorithm provided by this embodiment to implement the corresponding functions will be further described in detail below.
[0085] The single clustering module selects four clustering algorithms, namely K-Medoids, spectral clustering, AGNES (agglomerative nesting), and BIRCH (balanced iterative reducing and clustering using hierarchies), as the base clusterer for modeling. Its working principle is as follows:
[0086] The K-Medoids method minimizes the sum of the distances from the points within the cluster to the center by selecting the cluster center points, effectively reducing the influence of outliers and enhancing the robustness of clustering, as shown in Equation (1).
[0087]
[0088] where: m i is the optimal median of the i-th cluster, C i is the i-th cluster, x is the candidate median, x j is, d(x j , x) is the distance function, calculating the distance between x j and x, k is, d(x j , m i ) is the distance function, calculating the distance between x j and m i .
[0089] Spectral clustering maps the cluster information of data points in a low-dimensional space by constructing a similarity matrix and a Laplacian matrix and using eigenvalue decomposition, and is suitable for processing complex or non-spherical data sets, as shown in Equation (2).
[0090] L = D - W (12)
[0091] where: W is an n×n symmetric matrix representing the similarity between nodes, D is the degree matrix representing the connection strength between nodes, and L is the Laplacian matrix.
[0092] AGNES adopts a bottom-up hierarchical clustering method and forms a hierarchical structure by gradually merging the most similar clusters. BIRCH realizes the efficient processing of data by constructing and refining a clustering feature tree. The core clustering merging formulas of AGNES are divided into the shortest distance, the longest distance, and the average distance, as shown in Equations (3) to (5).
[0093]
[0094]
[0095]
[0096] where: In the formula, C i , C j are clusters, e and f are any two points in the cluster, and dist(e, f) is the distance calculation function used to measure the distance between e and f.
[0097] The single result alignment module further uses the Hungarian algorithm to solve the label inconsistency problem in the integrated clustering modeling stage according to the selected base classifier.
[0098] The co-covariance matrix module constructs a co-covariance matrix co-occurrence matrix based on the aligned single clustering result.
[0099] The clustering division module uses the BIRCH algorithm to cluster the co-occurrence matrix to obtain the final clustering result of the required photovoltaic output data set.
[0100] It should be noted that the steps in the method provided by the present invention can be implemented by using the corresponding components in the system, etc. Those skilled in the art can refer to the technical solution of the system to implement the step flow of the method, or refer to the technical solution of the method to implement the composition of the system. That is, the embodiments in the system and the embodiments in the method can be understood as preferred examples of each other and will not be elaborated here.
[0101] Next, a verification example is combined to illustrate the data clustering effect of the technical solution provided by the above embodiments of the present invention.
[0102] The simulation data used in this verification example is sourced from the operation data of a certain solar power station in Australia. The data acquisition period is 5 minutes, covering multi-dimensional features such as power output and meteorological parameters. Since the influence degrees of different features on the photovoltaic power generation are different, in this verification example, the correlation coefficients between each feature and the power generation are calculated, and the features with stronger correlations are selected as the inputs of the model. Specifically, the calculation results of the correlation coefficients are shown in Table 1. Based on this, the key features are selected, and the original data is reconstructed into samples on a daily basis. Considering that there is no power output during the period without sunlight in photovoltaic power generation, the night data is excluded in this verification example. Finally, each constructed sample contains 157 groups of data collected every 5 minutes between 6:00 and 19:00 every day. The experimental hardware environment is based on an Intel(R) Xeon(R) Gold 6234 CPU @ 3.30 GHz processor, 256 GB of memory, and two GeForce RTX 3090 graphics processing units (GPUs). The operating system is windows, supporting CUDA 11.1 and the corresponding CUDNN acceleration library. The method used in the experiment is implemented based on the open-source Pytorch framework.
[0103] Table 1 Correlation coefficients between power generation and meteorological features
[0104]
[0105] To verify the superiority of the clustering effect of the technical solution provided by the above embodiments of the present invention, this verification example compares the performances of five clustering algorithms, including AGNES, BIRCH, spectral clustering, K-Medoids, and the clustering method proposed in the above embodiments of the present invention. The clustering results are quantitatively analyzed through three evaluation indexes: silhouette coefficient, CH index, and DB index. The specific results are shown in Table 2.
[0106] Table 2
[0107]
[0108] As can be seen from Table 2, the clustering method proposed in the above embodiments of the present invention performs excellently in the clustering performance evaluation, specifically reflected in obtaining good results in two key indicators: the silhouette coefficient and the CH index. Among them, the silhouette coefficient is 0.4240, which is improved by 0.0045, 0.0323, 0.0073, and 0.0087 compared with AGNES, BIRCH, spectral clustering, and K-Medoids respectively. In terms of the CH index, the clustering ensemble model is 572.9346, which is improved by 41.4596, 53.1139, 79.6755, and 35.7804 compared with AGNES, BIRCH, spectral clustering, and K-Medoids respectively. In terms of the DB index, the result of the model in this paper is 0.8886. Although this value is slightly higher than 0.8843 of BIRCH, it is decreased by 0.036, 0.1092, and 0.0243 compared with AGNES, spectral clustering, and K-Medoids respectively.
[0109] Based on the clustering method provided in the above embodiments of the present invention, this verification instance further performs a clustering analysis on the photovoltaic power generation data to explore the influence of different meteorological characteristics on the photovoltaic output. According to meteorological characteristics such as temperature, relative humidity, and daily precipitation, combined with the local climate characteristics of Australia, the data is divided into three types of typical weather patterns: the first type is sunny days, characterized by low relative humidity, zero daily precipitation, and stable total inclined-plane irradiance; the second type is cloudy days, characterized by slightly higher relative humidity, less daily precipitation, and large fluctuations in the total inclined-plane irradiance; the third type is rainy days, characterized by higher daily precipitation and relative humidity, and at the same time, a significant decrease in the total inclined-plane irradiance. From the power generation power curve, all three types of weather patterns show a typical peak-shaped distribution, as Figure 3 shown. Since the power station is located in the desert area of Australia and the solar radiation is strong, a relatively high power generation can still be maintained even on cloudy days. The peak values of photovoltaic output on sunny days and cloudy days are similar, but the output power fluctuates more significantly on cloudy days. The rainfall in this area is less, and the proportion of rainy-day samples is only 5%, while the proportions of cloudy-day and sunny-day samples are 45% and 50% respectively.
[0110] An embodiment of the present invention further provides a computer terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method of any one of the above embodiments of the present invention, or, run the system of any one of the above embodiments of the present invention.
[0111] Optionally, a memory for storing programs; the memory may include volatile memory (e.g., random-access memory, such as static random-access memory (SRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), etc.); the memory may also include non-volatile memory, such as flash memory. The memory is used to store computer programs (such as application programs and functional modules for implementing the above methods), computer instructions, etc. The above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.
[0112] A processor for executing the computer programs stored in the memory to implement each step in the method or each module in the system described in the above embodiments. For specific details, please refer to the relevant descriptions in the foregoing method and system embodiments.
[0113] The processor and the memory can be of an independent structure or an integrated structure integrated together. When the processor and the memory are of an independent structure, the memory and the processor can be coupled and connected through a bus.
[0114] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute the method of any one of the above embodiments of the present invention, or to run the system of any one of the above embodiments of the present invention.
[0115] Among them, the computer-readable medium includes computer storage media and communication media, where the communication media includes any medium facilitating the transfer of computer programs from one place to another. The storage medium can be any available medium accessible by a general or special-purpose computer. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. Additionally, the ASIC can be located in a user device. Of course, the processor and the storage medium can also exist as discrete components in a communication device.
[0116] The photovoltaic output data similar-day clustering method and system based on the integrated clustering algorithm provided in the above embodiments of the present invention cluster the photovoltaic output data through a clustering algorithm based on meteorological characteristics to identify the output patterns under different meteorological characteristics, and thus better understand and predict the change law of photovoltaic output. The multi-clusterer integration strategy effectively solves the limitations of a single algorithm in dealing with the non-stationarity of photovoltaic output data and improves the quality and robustness of the clustering results. The multi-cluster integration strategy can effectively divide the photovoltaic output data with multi-attribute characteristics and achieve an accurate description of the association between complex meteorology and photovoltaic output.
[0117] Matters not described in detail in the above embodiments of the present invention are all well-known technologies in the art.
[0118] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A method for clustering similar days of photovoltaic output data based on an integrated clustering algorithm, characterized in that, Including: Selecting a base clusterer, clustering a photovoltaic output dataset based on the base clusterer to obtain a single clustering result; Using the Hungarian algorithm to unify the classification labels of different single clustering results to obtain the aligned single clustering results; Based on the aligned single clustering results, constructing a co-covariance matrix, which is used to represent the frequency of samples being assigned to the same cluster in all clustering results; Taking the co-covariance matrix as the similarity matrix between samples, clustering the co-covariance matrix to achieve the final clustering division of the photovoltaic output dataset, and obtaining the dataset after clustering of similar days.
2. The method for clustering similar days of photovoltaic output data based on an integrated clustering algorithm according to claim 1, wherein The selecting of the base clusterer and clustering the photovoltaic output dataset based on the base clusterer to obtain a single clustering result includes: The base clusterer includes: K-Medoids algorithm, spectral clustering algorithm, AGNES algorithm and BIRCH algorithm; where: The K-Medoids algorithm minimizes the sum of the distances from the points in the cluster to the cluster center point by selecting the cluster center point, as shown in Equation (1): Where m i is the updated cluster center of the i-th cluster, C i is the i-th clustering cluster, x is the candidate cluster center, x j is the sample in the cluster C i , d(x j , x) is the distance function, calculating the distance between x j and x, k is the number of clustering clusters, d(x j , m i ) is the distance function, calculating the distance between x j and m i ; The spectral clustering algorithm constructs a similarity matrix and a Laplacian matrix, and uses eigenvalue decomposition to extract the main eigenvectors to map the data point cluster information in a low-dimensional space, as shown in Equation (2): L = D - W (2) In the formula, W is an n×n symmetric matrix representing the similarity between each node, D is the degree matrix representing the connection strength between nodes, and L is the Laplacian matrix; The AGNES algorithm uses a bottom-up hierarchical clustering method to gradually merge the most similar clustering clusters to form a hierarchical structure; where: The core clustering merging formula of the AGNES algorithm is divided into the shortest distance H min , the longest distance H max and the average distance H avg , as shown in Equations (3) to (5): where C i , C j are clustering clusters, e and f are any two points in the clustering cluster, and dist(e, f) is a distance calculation function used to measure the distance between e and f; The BIRCH algorithm realizes the clustering process of data by constructing and refining a clustering feature tree; where: Constructing a clustering feature tree; Adding all samples in the photovoltaic output dataset into the clustering feature tree respectively, and calculating the radius of the leaf node after adding the new sample; According to the obtained radius of the leaf node, merging the sample into the corresponding cluster and updating all clustering feature triples on the path of the clustering feature tree; Regarding the cluster represented by each leaf node in the clustering feature tree as a data point, and performing secondary clustering to optimize the clustering feature tree; Define the photovoltaic output dataset as D = (x1, x2, …, x n ), filter out the feature subset related to the photovoltaic output through feature engineering, and use the selected base clusterer to cluster the data, and each base clusterer is set to generate m cluster classes, and finally obtain different single clustering results.
3. The method for clustering similar days of photovoltaic output data based on an integrated clustering algorithm according to claim 2, characterized in that, The constructing of the clustering feature tree includes: Determine the three parameters B, L, and T of the clustering feature tree; where B is the maximum number of CFs of internal nodes; L is the maximum number of CFs of leaf nodes; T is the maximum sample radius threshold of each CF of leaf nodes; CF = (N, L S , S S ) represents the clustering feature triple; N represents the number of samples in each CF; L S represents the sum vector of the feature values of the sample points in this CF; S s represents the sum of the squares of the feature values of the samples in this CF; L S and S s are calculated as shown in equations (6) and (7): where: x n represents the nth sample vector; x nm represents the mth eigenvalue of the nth sample point; M represents the number of features of the data in the sample points.
4. The method for clustering similar days of photovoltaic output data based on an integrated clustering algorithm according to claim 2, wherein The adding of all samples in the photovoltaic output dataset into the clustering feature tree respectively and calculating the radius of the leaf node after adding the new sample includes: Reading all samples into the clustering feature tree in sequence, finding the leaf node A closest to the new sample, and calculating the radius R of the leaf node A after adding the new sample. The calculation method is as shown in Equation (8): where x n is the nth sample vector and x0 is the initial sample vector.
5. The method for clustering similar days of photovoltaic output data based on the integrated clustering algorithm according to claim 2, wherein The merging of the sample into the corresponding cluster according to the obtained radius of the leaf node and updating all clustering feature triples on the path of the clustering feature tree includes: If the obtained radius R of the leaf node ≤ T, then merging the sample into the cluster and updating all clustering feature triples CF on the path of the clustering feature tree CF, and performing the following steps: Checking whether the parent node needs to be split from bottom to top. If it needs to be split, then using the leaf node splitting method until the root node; If the obtained radius R of the leaf node > T, then performing the following steps: Determine whether the CF number of leaf node A reaches the maximum L; if it is less than L, create a new CF node to store the new sample and update all CF tuples on the CF-tree path; otherwise, select the two CFs with the farthest hypersphere radius distance among all CF tuples in leaf node A as the new leaf nodes A' and A'' to replace the original leaf node A, and place the CF tuples and new sample tuples in the original leaf node into the new leaf nodes A' and A'' in order of distance.
6. The method for clustering similar days of photovoltaic output data based on an integrated clustering algorithm according to any one of claims 2-5, characterized in that The method for selecting the clustering features includes: Analyze the core features related to photovoltaic power, including global radiation and diffuse radiation, based on Pearson, Spearman, and Kendall correlation coefficients. Based on the obtained core features, construct time series features based on daily statistical summaries; divide the time series features by day, and calculate the mean, peak, and standard deviation statistics for each day. Use the gradient boosting decision tree and random forest algorithms to select the time series features, and finally normalize the scores respectively by measuring the contributions of different features to screen the time series features to obtain the required clustering features.
7. The method for clustering similar days of photovoltaic output data based on the integrated clustering algorithm according to claim 1, wherein The method for unifying the classification labels of different single clustering results using the Hungarian algorithm includes: Taking the classification label of a certain single clustering result as a benchmark, use the Hungarian algorithm to align the classification labels of other single clustering results.
8. The method for clustering similar days of photovoltaic output data based on an integrated clustering algorithm according to claim 1, characterized in that, Based on the aligned single clustering results, construct a co-association matrix, including: For the given k clustering results C1, C2, …, C k , each element CA(i, j) of the co-association matrix CA is shown as in the formula: where φ k represents the label mapping function of the k-th clustering result; Ι(·) represents the indicator function, which takes the value of 1 when the condition holds and 0 otherwise; CA(i, j) represents the sample x i and the sample x j are assigned to the same cluster in all clustering results, and the value range is [0, 1].
9. The method for clustering similar days of photovoltaic output data based on an integrated clustering algorithm according to claim 1, wherein Use the BIRCH algorithm to cluster the co-association matrix to obtain the final clustering result.
10. A photovoltaic output data similar-day clustering system based on an integrated clustering algorithm, characterized in that, Include: A single clustering module, which is used to select a base clusterer and perform clustering processing on the photovoltaic output dataset based on the base clusterer to obtain a single clustering result. A single result alignment module, which uses the Hungarian algorithm to unify the classification labels of different single clustering results to obtain the aligned single clustering results. A co-association matrix module, which constructs a co-association matrix based on the aligned single clustering results. The co-association matrix is used to represent the frequency of samples being assigned to the same cluster in all clustering results. A clustering division module, which is used to use the co-association matrix as a similarity matrix between samples, cluster the co-association matrix, and implement the final clustering division of the photovoltaic output dataset to obtain the dataset after clustering similar days.
11. A computer terminal, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it can be used to execute the method described in any one of claims 1-9, or run the system described in claim 10.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it can be used to execute the method described in any one of claims 1-9, or run the system described in claim 10.