Weather type classification method based on photovoltaic power generation symbol sequence histogram clustering
By processing and clustering photovoltaic power output data and reclassifying weather types, the error problem caused by relying on weather forecast data in photovoltaic power generation forecasting is solved, thereby improving the accuracy and applicability of photovoltaic power generation forecasting.
Patent Information
- Application Number
- CN202310068883.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-06
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-02-06
AI Technical Summary
Current photovoltaic power generation forecasts rely on weather forecast data from meteorological service providers, resulting in large errors and insufficient spatial resolution, which affects the accuracy and precision of photovoltaic output forecasts, especially in distributed photovoltaic power plants where there is a lack of accurate descriptions of local weather types.
By performing detrending reconstruction, differential calculation, symbolization, and character encoding on photovoltaic power output data, a symbol sequence histogram is generated. Unsupervised clustering algorithms are then used to mine meteorological information from historical photovoltaic power output data, and weather types are reclassified.
It improves the accuracy and robustness of photovoltaic power generation forecasts, and provides more accurate information on similar days, especially suitable for describing local weather types for distributed photovoltaic power plants.
Smart Images

Figure CN115982601B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of photovoltaic new energy and relates to a weather type classification method based on histogram clustering of photovoltaic power generation symbol sequences. Background Technology
[0002] Photovoltaic power generation, as one of the main forms of solar energy utilization, has the advantages of zero pollution and renewability, and has become an important way to solve the problem of carbon emissions. However, solar energy is a naturally fluctuating energy source, and the output of photovoltaic power generation is greatly affected by weather conditions, especially cloud cover changes, resulting in intermittent and uncertain power supply patterns. These characteristics not only bring new challenges to the safe and stable operation of the power system, but also place higher demands on power flexibility and regulation capabilities. With the increasing penetration rate of photovoltaic power generation in power grids and microgrids, the prediction of photovoltaic power generation output has become an important research topic in the field of photovoltaic applications.
[0003] Since photovoltaic power generation is closely related to cloud cover, i.e., weather conditions such as sunny, rainy, cloudy, and overcast, effective classification of historical day weather types is crucial for obtaining accurate similar days and is also a key factor in improving the accuracy of photovoltaic power generation forecasts. However, in current photovoltaic power output forecasting research, studies on weather conditions that directly affect the accuracy of similar days are often neglected, and most of them directly use weather forecast data provided by meteorological service providers. This presents two problems: (1) errors inherent in public weather forecasts themselves. (2) Public weather forecasts classify weather types based on the proportion of cloud cover across the entire sky, so they can generally only provide rough forecasts for a large geographical area, and their spatial resolution is insufficient for photovoltaic power generation forecasting. This means that the training sample labels for similar days in photovoltaic power output forecasting are unreliable, which will seriously affect the accuracy of subsequent photovoltaic power output forecasts. Summary of the Invention
[0004] Current photovoltaic (PV) power generation forecasting largely relies on meteorological data to determine similar days, which to some extent limits further improvements in forecast accuracy. Experimental studies show that the output power fluctuation trend of PV modules is basically consistent with the fluctuation trend of solar radiation intensity. It is generally accepted that solar radiation intensity is closely related to weather types defined by cloud cover. Based on this background, this invention proposes a novel strategy for PV power generation forecasting: inversely calculating the corresponding weather type from historical PV output data. This aims to change the current situation where the selection of similar days in PV output forecasting generally relies on weather forecast data, thereby improving the accuracy and robustness of PV output forecasting.
[0005] This invention proposes a weather type classification method based on histogram clustering of photovoltaic power generation symbol sequences, comprising the following steps:
[0006] Step (1) Calculate the fluctuation score for the sign histogram features of the original photovoltaic power output sequence and the sequence after removing the fixed trend component, and obtain the fluctuation score sequence.
[0007] Step (2) Based on the fluctuation score sequence, cluster the symbol sequence histogram to obtain the optimal boundary and classification rules.
[0008] Step (3) Reclassify the weather type of the sample day according to the optimal boundary and classification rules.
[0009] The beneficial effects of this invention are as follows:
[0010] 1. This invention provides a weather type mapping method based on historical photovoltaic power output fluctuation analysis. Based on the meteorological sensitivity of photovoltaic power generation, an unsupervised clustering algorithm is used to mine the meteorological information implicit in historical photovoltaic power output. The corresponding meteorological conditions are then deduced from the historical photovoltaic power output data, enabling a reclassification of historical weather types and providing more accurate similarity day information for photovoltaic power generation forecasting.
[0011] 2. The historical weather types obtained by clustering in this invention are closely related to the solar radiation conditions in the sky, and are obtained based on the characteristics of historical photovoltaic data. They are for the prediction of photovoltaic power generation days, and are particularly targeted at the description of local weather types for distributed photovoltaic power stations. Attached Figure Description
[0012] Figure 1 A flowchart for histogram clustering of symbol sequences and classification of weather types. Detailed Implementation
[0013] By performing detrending reconstruction, difference calculation, symbolization, word encoding, and symbol histogram feature extraction on the photovoltaic power output sequence, the original photovoltaic power output sequence X can be obtained. (m,d) and sequence R after removing fixed trend components (m,d) symbol sequence SX (m,d) and SR (m,d) And finally obtain the symbol sequence histogram HX. (m,d) and HR (m,d) The detrended reconstruction process specifically involves: for a specific photovoltaic power generation system, recording the output power during the effective photovoltaic power generation period on a daily basis, and denoted by X. (m,d) It is represented as shown in equation (1);
[0014]
[0015] In the formula, i, d, and m represent the sampling period number and the date and month in which it falls, respectively; N represents the total number of sampling points per day; This represents the average photovoltaic power generation on day d of month m and within the i-th sampling period;
[0016] An approximate clear sky sequence is obtained by combining points one by one to replace the ideal clear sky sequence; let... The maximum value of the average photovoltaic power generation within the i-th sampling period of all days in the m-th month is obtained by combining these maximum values into an approximate clear sky sequence for the m-th month. As shown in equation (2):
[0017]
[0018] from Extracting the detrended term r from the decomposition that reflects instantaneous changes in meteorological conditions i (m,d) As shown in equation (3):
[0019]
[0020] Then the sequence R after removing the fixed trend component (m,d) Represented as:
[0021]
[0022] The difference calculation is as follows:
[0023] For X (m,d) and R (m,d) First-order difference calculations are performed separately to obtain the difference time series ΔX of the feature sequence. (m,d) and △R (m,d) Hereinafter referred to as "difference sequence"; with ΔX (m,d) For example, the calculations are shown in equations (5) and (6);
[0024]
[0025]
[0026] Where i, d, and m represent the sampling period number and its corresponding date and month, respectively; △R (m,d) The calculation process is the same as above.
[0027] Symbolization specifically refers to: using △X (m,d) Taking the symbolization of as an example, let the difference sequence ΔX (m,d) Based on the degree of fluctuation, it is divided into M levels, that is, the amplitude space Ω of the difference sequence is divided into M subspaces Ω on the Lebesgue measure. n ,in Then, M characters are used to assign values to the elements of the difference sequence, thereby discretizing the difference sequence into a string;
[0028] Using the principle of equal probability, the set of natural numbers is used as the assignment symbol to assign values to the elements of the difference sequence, as shown in equation (7).
[0029] sx i (m,d) ={n|△x i (m,d) ∈Ω n} (7)
[0030] Among them, Ω n The range of values is represented by n, which is a natural number and takes the value range [0, M-1]. i, d, and m represent the sampling period number and the date and month in which it falls, respectively.
[0031] Through the above time-discrete and space-continuous mapping, △X (m,d) Discretization generates a sequence of symbols, using SX. (m,d) As shown in equation (8);
[0032] SX (m,d) =(sx1) (m,d) ,sx2 (m,d) ,…,sx N-1 (m,d) (8)
[0033] Similarly, the difference sequence ΔR can be obtained. (m,d) symbol sequence SR (m,d) .
[0034] The specific process of character encoding is as follows: encoding the symbol sequence SR (m,d) The resulting characters are encoded, and the character encoding value is recorded as follows: As shown in equation (9);
[0035]
[0036] L is the optimal word length in the sense of minimizing Shannon entropy, and l, d and m represent the sampling period number and the date and month in which it is located, respectively.
[0037] Thus, SX is obtained. (m,d) The character encoding sequence VX (m,d) As shown in equation (10);
[0038] VX (m,d) ={vx1 (m,d) vx2 (m,d) ,…,vx N-L+1 (m,d)} (10)
[0039] Similarly, constituting SR (m,d) VR character encoding sequence (m,d) .
[0040] Specifically, the histogram of the corresponding symbol sequence is obtained by: statistical word encoding sequence VX (m,d) The frequency of occurrence of different encoding values in the character sequence is determined, and the character encoding sequence is sorted from smallest to largest according to the character encoding value, resulting in the symbol sequence histogram HX as shown in equation (11). (m,d) Discrete distributed;
[0041]
[0042] Where k is the word code value corresponding to the histogram component, k∈[0,M] L -1], PX k This indicates the frequency of occurrence of the corresponding character encoding value k in the samples of the current sample day; {} (m,d) This represents the histogram of the symbol sequence for the current sample day being the dth day of month m; similarly, VR is obtained. (m,d) The histogram distribution of the symbol sequence.
[0043] The specific implementation steps of this invention are as follows:
[0044] Step (1) Perform a histogram HX of the symbol sequence. (m,d) and HR (m,d) Calculate the volatility scores separately to obtain the volatility score sequences QX and QR.
[0045] This invention introduces the concept of a fluctuation score, characterizing the degree of fluctuation in photovoltaic power output time series, using a daily unit. Since the magnitude of the symbol sequence's character encoding value reflects the intensity of fluctuation within a character-length time period, the fluctuation score proposed in this invention is calculated based on the character encoding. The fluctuation score is defined as the sum of the products of the encoding value and the frequency in the symbol histogram.
[0046] Using the symbol sequence histogram HX (m,d) Fluctuation score qx (m,d) For example, the calculation of is shown in equation (1):
[0047]
[0048] Among them, PX k This represents the frequency of occurrence of the corresponding character encoding value k for all samples on the current sample day (day d of month m). For example, if HX (m,d) In the text, the character encoding value 12 appears 3 times, then PX 12 =3; L is the optimal word length in the sense of minimizing Shannon entropy, and M is the number of fluctuation levels in the symbolization process.
[0049] Calculate the volatility score for all days in the sample set. Let T be the number of all different volatility scores that have occurred. x All qx (m,d)Rearrange them in ascending order to form X (m,d) The fluctuation score sequence QX is shown in equation (2).
[0050] QX={qx1,qx2,…qx t ,…,qx T} (2)
[0051] In the formula, qx1 is the minimum fluctuation score; qx T This represents the maximum fluctuation score; thus, the fluctuation of the characteristic sequence for a sample day can be quantified by the fluctuation score for that day. The larger the fluctuation score, the more drastic the fluctuation in photovoltaic output on that day. Similarly, R can be obtained. (m,d) The fluctuation score sequence QR.
[0052] Step (2) Based on the fluctuation score sequences QX and QR, respectively analyze the histogram HX of the symbol sequence. (m,d) and HR (m,d) Clustering is performed to obtain the optimal boundary and classification rules.
[0053] The core idea of symbol sequence histogram clustering based on fluctuation scores is to find the optimal cluster boundary of fluctuation scores, divide QX and QR into several labeled subclasses, and realize unsupervised clustering of the symbol sequence histograms corresponding to these subclasses.
[0054] This invention, referencing the principles of clustering, iterates through a moving boundary approach to find the optimal cluster boundary for the histogram sequence. This process involves two parameters: a clustering effectiveness evaluation index and the number of clusters.
[0055] This invention uses the silhouette coefficient (SC) as an evaluation index for the effectiveness of symbolic histogram clustering. SC∈[-1,1], where -1 represents a very poor clustering and 1 represents a perfect clustering; this invention uses dynamic time curvature distance to measure the distance between histograms.
[0056] Considering that the weather types affecting photovoltaic output mainly include sunny, cloudy / rainy, and overcast, the number of weather type labels is 3. This invention classifies the fluctuation pattern of each feature sequence into two types. That is, it is necessary to find the cluster boundaries and cluster each feature sequence into two categories.
[0057] Let qx2, ..., qx be the values respectively. T-1 Using T-2 score points as boundaries, we obtain all segmentation schemes and their corresponding SCs. The boundary where SC reaches its maximum value is taken as the optimal boundary. Let the optimal boundaries for QX and QR be qx and qR, respectively. * and qr *Then, as shown in equations (3) and (4), the subclasses to which the fluctuation scores belong can be labeled respectively.
[0058]
[0059]
[0060] Step (3) Reclassify the weather type of the sample day according to the optimal boundary and classification rules.
[0061] Based on the optimal boundary and clustering labels, the weather type classification rules set in this invention are shown in Table 1.
[0062] Table 1. Clustering labels for weather types based on fluctuation scores.
[0063]
[0064] As mentioned above, A represents drastic fluctuations, F represents moderate fluctuations, and X... (m,d) R reflects the overall intraday fluctuation characteristics. (m,d) This reflects the instantaneous fluctuation characteristics. According to this rule, Table 1 shows that on sunny days (A,F), intraday fluctuations are large, while instantaneous fluctuations are small; on cloudy days (A,A), both intraday and instantaneous fluctuations are large; and on rainy days (F,A), intraday fluctuations are gradual, while instantaneous fluctuations are large. Since (F,F) represents a relatively gradual intraday and instantaneous fluctuation, this invention is based on R... (m,d) The fluctuation characteristics classify it as a sunny day.
[0065] Based on the above methods for extracting fluctuation features from time series data, the complete process for weather type classification based on symbolic sequence histogram clustering proposed in this invention is attached. Figure 1 As shown.
Claims
1. A weather type classification method based on photovoltaic power generation symbol sequence histogram clustering, characterized in that, The method specifically includes the following steps: Step (1) Calculate the fluctuation score for the sign histogram features of the original photovoltaic power output sequence and the sequence after removing the fixed trend component, respectively, to obtain the fluctuation score sequences QX and QR; Using days as the unit, the concept of fluctuation score is introduced to characterize the degree of fluctuation in the photovoltaic power output time series. Since the size of the symbol sequence word code value reflects the intensity of fluctuation within a word-length time period, the proposed fluctuation score is calculated based on the word code. The fluctuation score is defined as the sum of the products of the code value and the frequency in the symbol histogram. Using the symbol sequence histogram HX (m,d) Fluctuation score qx (m,d) For example, the calculation of is shown in equation (1): Among them, PX k L represents the frequency of occurrence of the corresponding character encoding value k for all samples on the current day; L is the symbol sequence SX. (m,d) The optimal word length, M is the number of fluctuation levels during the symbolization process, and X... (m,d) This represents the output power during the effective period of photovoltaic power generation, where d and m represent the date and month of the sampling, respectively; Calculate the volatility score for all days in the sample set; let T be the number of all different volatility scores that have occurred. x ; Put all qx (m,d) Rearrange them in ascending order to form X (m,d) The fluctuation score sequence QX is shown in equation (2); QX={qx1,qx2,…qx t ,…,qx T } (2) In the formula, qx1 is the minimum fluctuation score; qx T This represents the maximum fluctuation score; thus, the fluctuation of the characteristic sequence of a sample day can be quantified by the fluctuation score of that day. The larger the fluctuation score, the more drastic the fluctuation in photovoltaic output on that day. Similarly, R can be obtained. (m,d) The fluctuation score sequence QR; Step (2) Based on the fluctuation score sequences QX and QR, cluster the symbol sequence histogram to obtain the optimal boundary and classification rules; By finding the optimal cluster boundary of the fluctuation score, QX and QR are divided into several labeled subclasses, and unsupervised clustering of the symbol sequence histograms corresponding to these subclasses is achieved. Step (3) Reclassify the weather type of the sample day according to the optimal boundary and classification rules.
2. The weather type classification method based on photovoltaic power generation symbol sequence histogram clustering according to claim 1, characterized in that: Based on the fluctuation score sequences QX and QR, the histogram HX of the sign sequence is analyzed respectively. (m,d) and HR (m,d) Perform clustering; specifically: Following the principle of clustering, the process iterates by moving the boundary to find the optimal cluster boundary of the histogram sequence; this process involves two parameters: the clustering effectiveness evaluation index and the number of clusters. The silhouette coefficient (SC) is used as an evaluation index for the effectiveness of symbolic histogram clustering. SC∈[-1,1], where -1 represents a very poor clustering and 1 represents a perfect clustering; the distance between histograms is measured using dynamic time-bending distance; Considering that the weather types affecting photovoltaic output mainly include sunny, cloudy and rainy, and partly cloudy, the number of weather type labels is 3; the fluctuation pattern of each feature sequence is divided into 2 types; that is, it is necessary to find the cluster boundaries and cluster each feature sequence into two categories respectively; Let qx2, ..., qx be the values respectively. T-1 Using the T-2 score points as boundaries, we obtain all segmentation schemes and their corresponding SCs; the boundary where SC reaches its maximum value is the optimal boundary; let the optimal boundaries for QX and QR be qx and qR respectively. * and qr * Then, as shown in equations (3) and (4), the subclasses to which the fluctuation scores belong are labeled respectively; A represents drastic fluctuations, while F represents moderate fluctuations.
3. The weather type classification method based on photovoltaic power generation symbol sequence histogram clustering according to claim 1, characterized in that: The weather types of the sample days were reclassified based on the optimal boundary and classification rules; specifically: Based on the optimal boundary and clustering labels, the weather type classification rules are set as shown in Table 1. Table 1. Clustering labels for weather types based on fluctuation scores. As mentioned above, A represents drastic fluctuations, F represents moderate fluctuations, and X... (m,d) R reflects the overall intraday fluctuation characteristics. (m,d) Reflecting instantaneous fluctuation characteristics; according to this rule, Table 1 shows that: on sunny days corresponding to (A,F), intraday fluctuations are large, while instantaneous fluctuations are small; on cloudy days corresponding to (A,A), both intraday and instantaneous fluctuations are large; on rainy days corresponding to (F,A), intraday fluctuations are gradual, but instantaneous fluctuations are large; since (F,F) represents a situation where both intraday and instantaneous fluctuations are relatively gradual, according to R... (m,d) The fluctuation characteristics classify it as a sunny day.
Citation Information
Patent Citations
Photovoltaic power interval prediction method combining neural network and parameter estimation
CN108985965A
Photovoltaic power probability estimation method and system based on copula optimization
CN115099511A