Photovoltaic power prediction method based on improved signal decomposition and deep learning fusion
By improving the method of integrating signal decomposition and deep learning, combining light intensity and photovoltaic power historical data, and using the GRU network to predict photovoltaic power, the problems of inaccurate prediction results and equipment dependence in existing technologies are solved, and efficient and accurate medium- and long-term photovoltaic power prediction is achieved.
Patent Information
- Application Number
- CN202510595779.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-12
AI Technical Summary
Existing photovoltaic power prediction methods are easily affected by noise and data loss in medium- and long-term forecasts, resulting in inaccurate prediction results. They also rely on meteorological observation equipment, which limits their scope of application.
A method based on the fusion of improved signal decomposition and deep learning is adopted. Light intensity data and photovoltaic power historical data are collected through feature engineering. The Gated Recurrent Unit (GRU) network is used for photovoltaic power prediction. Combined with spatial image analysis and intrinsic mode function (IMF) decomposition, the dependence on meteorological equipment is reduced.
The accuracy and reliability of photovoltaic power prediction are improved, the amount of calculation is reduced, and stable predictions can be achieved in the medium and long term without the need for additional observation equipment.
Smart Images

Figure CN120633898A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of photovoltaic power prediction, and specifically relates to a photovoltaic power prediction method based on improved signal decomposition and deep learning fusion. Background Art
[0002] Photovoltaic power generation is a method of generating electricity by directly converting solar radiation into electricity through the photovoltaic effect. It is an environmentally friendly energy source, and the photovoltaic power generation industry is an emerging industry that meets the current low-carbon needs of the times. However, due to the limitations of photovoltaic power generation itself, the output power of photovoltaic power generation systems is unstable. This is because the system output power varies synchronously with changes in light intensity, and the larger the photovoltaic power generation system, the greater the fluctuation in the system output power. If photovoltaic power generation systems are directly integrated into the power grid, it will inevitably cause excessive fluctuations in the power transmission capacity of the grid, placing enormous pressure on the load regulation of the power system and even causing serious damage to the power system. This inherent limitation of photovoltaic power generation has greatly hindered the further development of the photovoltaic power generation industry.
[0003] One solution to this problem is to identify patterns in PV power variation over time from historical PV power data, allowing for the appropriate planning of power system load adjustments. This has led to the emergence of PV power forecasting technology. Current PV power forecasting methods can be broadly categorized into two types based on the models used: one uses meteorological data models to predict future solar intensity, while the other uses historical PV power data to analyze PV power trends and predict future PV power changes. The former, while essentially a weather forecast, suffers from high computational complexity and relatively low reliability. The latter provides better short-term predictions, but medium- and long-term predictions are susceptible to significant errors due to noise or missing data in the data model. Some forecasting methods have also emerged that combine these two models. However, most of these methods require supporting meteorological observation and data processing equipment, making their practical application difficult due to equipment limitations. Summary of the Invention
[0004] The purpose of the present invention is to provide a photovoltaic power prediction method based on the fusion of improved signal decomposition and deep learning. This method uses two models, light intensity data collected through feature engineering and intrinsic mode function (IMF) obtained by decomposing historical photovoltaic power data, to analyze the patterns. The light intensity data and IMF are then input into a gated recurrent unit (GRU) network in succession, and the medium- and long-term photovoltaic power prediction results are obtained through the GRU network.
[0005] The present invention adopts the following technical solutions:
[0006] A photovoltaic power prediction method based on improved signal decomposition and deep learning fusion includes the following steps:
[0007] Step 1. Take multiple airspace images of the area to be predicted over a period of time. The time interval between each shot is equal, and the optical parameters for all airspace images are set to the same. Record the time of each shot and the photovoltaic power generation data of the area to be predicted during the shooting period.
[0008] Step 2. Divide each spatial image into different regions based on image brightness, record the brightness levels of each region, and then calculate the area ratio of each region in each image;
[0009] Step 3. Decompose the time-varying curve of photovoltaic power generation using a signal decomposition algorithm, and decompose the power curve into multiple intrinsic mode functions (IMFs) and a residual;
[0010] Step 4. Based on the proportion of the areas of different brightness in all spatial images, summarize the characteristics of the overall spatial illumination intensity over time, and reconstruct the data of the illumination intensity over time based on the obtained characteristics;
[0011] Step 5. Analyze the reconstructed light intensity data, multiple IMFs, and residuals through a neural network to provide a photovoltaic power prediction result.
[0012] Prior art photovoltaic power prediction methods often use a single model based on meteorological data or historical photovoltaic power data. The output is easily affected by noise in the model data, making it difficult to guarantee accurate results in medium- and long-term forecasts. The present method utilizes a dual-model analysis strategy based on both light intensity data and historical photovoltaic power data. These two models can mutually verify their timing characteristics within a neural network, thus preventing the influence of random factors such as noise on the results. Furthermore, the present method only requires the collection of light intensity data, making it simpler than other methods that utilize meteorological data and significantly reducing the computational effort required during data processing.
[0013] Further optimization, the spatial image area division and area ratio calculation in Step 2 specifically include the following steps:
[0014] Step 2.1. Convert the color space used to measure spatial image color to LAB. Each pixel in the image is assigned a unique serial number. Determine the color of each pixel, that is, the corresponding L-axis, A-axis, and B-axis coordinates of the pixel in LAB space, and record them in the form of (l, a, b).
[0015] Step 2.2. Use K-means clustering analysis method to classify all pixels in the image. The specific steps of classification are:
[0016] Step 2.2.1. Determine the number of cluster centers k, the center threshold K, and the iteration threshold E, and then randomly select k pixels in the image as the initial cluster centers; the number of iterations is recorded as e, the initial value of e is 0, and the cluster center is represented by C i , i is the cluster center number, 1≤i≤k;
[0017] Step 2.2.2. The total number of pixels in the image other than the cluster center is n, and the other pixels are represented by P j , j is the pixel number, 1≤j≤n, take a cluster center C i , calculate C i With each P j The distance D Ci-Pj And record, the calculation formula is:
[0018]
[0019] where l Ci 、a Ci 、b Ci C i The corresponding L-axis, A-axis, and B-axis coordinates in the LAB color space, l Pj 、a Pj 、b Pj P j The corresponding L-axis, A-axis, and B-axis coordinates in the LAB color space; when all P j Completed with the current C i After calculating the distance, select another C i Continue calculating until all C i All calculations are completed;
[0020] Step 2.2.3. Compare P j With each cluster center C i The distance between them, find the shortest distance and the C corresponding to the shortest distance i , and P j The closest C i Associated, all P j You need to find the C associated with it i , and then each C i and all P associated with it j into the same category;
[0021] Step 2.2.4. Calculate the average of the L-axis, A-axis, and B-axis coordinates of all pixels in the same class, take the point corresponding to the average of the three axes as the centroid of the class, and calculate the centroid and the cluster center C of the class. i The distance d between the center of mass and C i If the distance is not greater than the center threshold, that is, if the condition d≤K is met, the classification result is output; if not, the number of iterations is compared to see if it reaches the iteration threshold, that is, e=E. If it is met, the classification result is output; if it is still not met, the e value is increased by 1, and the centroid is used as the new cluster center C i , and then repeat the above process from Step 2.2.2;
[0022] Step 2.3. a in all cluster center coordinates Ci Replace with a constant, b Ci It is also replaced by a constant, and the cluster center after coordinate replacement is recorded as C` i ; Set all C` i Sort by the L-axis coordinate value, each C` i The corresponding serial number is recorded as m, and then each C` i The corresponding coordinates are regarded as a characteristic color, and the characteristic color is recorded as SIF m ;
[0023] Step 2.4. Calculate other pixels P in the image j The distance from each characteristic color in the L-axis direction, find the minimum distance and the SIF corresponding to the minimum distance m , and then use the SIF corresponding to the minimum distance to determine the color of the pixel m Replace; each pixel P in the image j All colors require a SIF m replace;
[0024] Step 2.5. Each SIF m The range covered in the image is regarded as a region, the number of pixels in each region is counted, and then the proportion of the region in the image is calculated by dividing the number of pixels in the region by the total number of pixels in the image.
[0025] The purpose of regionalizing a spatial image is to determine the area occupied by each portion of the image with different brightness levels, thereby deriving the overall light intensity in the spatial domain at the time of capture. Using the K-means cluster analysis method to classify all pixels in the image, pixels with similar colors can be grouped into the same category, and the most representative pixels in the same category are output as the final cluster centers. Using the brightness of these cluster centers as the characteristic color brightness can better reflect the overall brightness characteristics of the image. The image's color space is converted to LAB space because LAB space is a color model that represents a color through brightness and chromaticity. The L axis of a point in the space represents the brightness of the color, the A axis represents the chromaticity from green to red, and the B axis represents the chromaticity from blue to yellow. For the method of the present invention, only the brightness dimension of the image is required to reflect the light intensity, and chromaticity can be ignored. Therefore, the A-axis and B-axis coordinates of all characteristic colors are constants. The colors of all pixels are then replaced with the most similar characteristic colors based on their brightness similarity. This allows the image to be divided into several regions of different brightness, thereby reflecting the overall light intensity. Another benefit of using the K-means clustering method is that the number of cluster centers, k, and therefore the number of characteristic colors, can be determined by the operator. A greater number of characteristic colors in an image provides a more accurate representation of the overall spatial brightness, but this also increases the computational complexity of subsequent data analysis. Allowing the operator to determine the number of characteristic colors helps achieve a balance between accuracy and computational complexity.
[0026] For further optimization, the decomposition of the power curve in Step 3 specifically includes the following steps:
[0027] Step 3.1. Generate r Gaussian white noises, which are curves connected by multiple noise distribution points. The curve takes time as the independent variable and power as the dependent variable. The distribution shapes of all Gaussian white noise curves are different, and the duration of all curves is not shorter than the duration of the power curve. The multiple noise distribution points are evenly distributed on the time scale and Gaussian distributed with the mean at position 0 on the power scale. The r white noises are recorded as {noise1,…,noise r};
[0028] Step 3.2. Compare the r white noises with the power curve S to be decomposed q Superposition, q is the number of iterations, q = 0 in the initial power curve; the superposition of the power curve and noise is represented by S q-s =S q +noise s , s is the white noise number, 1≤s≤r, S q-s is the curve formed by superposition, with a total of r curves;
[0029] Step 3.3. For all S q-sPerform Empirical Mode Decomposition (EMD), each S q-s Correspondingly, an intrinsic mode function is generated, expressed as IMF q-s ; for all IMF q-s The IMF is obtained by averaging the q-ave For any time t within the time range corresponding to the power curve, all IMFs exist q-s The arithmetic mean of the corresponding values at time t and the IMF q-ave The relationship in which the corresponding values are equal at time t;
[0030] Step 3.4. Determine the IMF q-ave Are the following conditions met?
[0031] a. In the entire curve, the number of extreme points and zero-crossing points is the same or differs by only one;
[0032] b. The upper envelope formed by the maximum point of the curve and the lower envelope formed by the minimum point are symmetrical about the time axis;
[0033] c. The number of extreme points in the curve is not less than twice the number of days included in the time range corresponding to the initial power curve;
[0034] If all three conditions are met at the same time, then S q+1 =S q -IMF q-ave The new power curve to be decomposed is obtained by processing, and then the above process is repeated from Step 3.2; if the three conditions are not met at the same time, the power curve to be decomposed S in this iterative calculation is q as residuals;
[0035] Step 3.5. Smooth the residuals using the Exponential Moving Average (EMA) method. The specific steps are:
[0036] Step 3.5.1. Take multiple sampling points on the residual, T is the sampling point number, all sampling points are arranged in the order of sampling time and the sampling time interval between two adjacent points is equal, and record the corresponding time value and power value F of each sampling point T ;
[0037] Step 3.5.2. Set all F T Substitute into the EMA method calculation formula in order, the calculation formula is as follows
[0038] G T =βG T-1 +)1-β)F T(2)
[0039] Among them G T is the power value corresponding to the Tth sampling point after processing by the EMA method, G0 = 0, β is the smoothing coefficient;
[0040] The time values of all sampling points are retained, and the power values are replaced by the corresponding G T The data points are formed, and the curve formed by these data points is the residual after smoothing.
[0041] The signal decomposition method described in Step 3 is an improvement on the Ensemble Empirical Mode Decomposition (EEMD) method. The difference between the EEMD method and the EMD method is that the EMD method decomposes the initial signal directly, while the EEMD method first adds Gaussian white noise to the initial signal before decomposing it. The EMD decomposition steps are as follows: First, the upper envelope of the curve is formed by interpolating the maximum points of the signal curve, and the lower envelope is formed by interpolating the minimum points. The upper and lower envelopes are then integrated and averaged to obtain the average line. The definition of the integrated average is the same as described in Step 3.3. The average line is then subtracted from the initial signal to generate another curve. This curve is then judged to determine whether it meets the IMF definition. If so, it is considered an IMF. The decomposition process is then repeated, using the IMF as the signal curve, until the generated curve does not meet the IMF definition. The process stops when a curve that does not meet the IMF definition is found, and this curve is used as the residual. In this way, multiple IMFs and a residual are obtained through the EMD decomposition method. The basic idea of the EMD method is that all complex waveforms are composed of the superposition of simple waveforms, and therefore all complex waveforms can be decomposed into multiple simple waveform components. Traditional transformation methods have difficulty capturing the frequency variation characteristics of complex curves. However, by decomposing complex curves into multiple IMFs, each containing an oscillation mode, and then performing modal analysis on each IMF, the frequency characteristics of the simple waveforms are derived, and then superimposed to obtain the frequency characteristics of the complex curve.
[0042] However, this is only an ideal state. In actual operation, the IMF decomposed by the EMD method may have problems such as modal aliasing. In this case, it is difficult for modal analysis to obtain accurate results, so the EEMD method has emerged. The EEMD method will add different Gaussian white noises to the signal curve to be decomposed each time. The superposition of Gaussian white noises can increase the distance between the extreme points on the same side of the time axis in the signal curve, avoiding the modal aliasing problem caused by the extreme points with similar distances being ignored in the construction process of the envelope. Because the specific positions of the noise distribution points in Gaussian white noise are random, and all noise distribution points are Gaussian distributed with a mean of 0, the waveform after the superposition of enough Gaussian white noises can be considered to be eventually canceled. The EEMD method also brings a problem, that is, the number of IMFs that need to be decomposed for a signal curve is huge, which is bound to cause lengthy calculation time, and the waveforms of the first few IMFs will be greatly affected by the noise. In the method of the present invention, multiple Gaussian white noises are generated at one time, multiple curves are superimposed, and then the multiple curves are decomposed by the EMD method to obtain IMFs. q-s , and then to all IMF q-s The integrated average is the IMF q-ave This approach can also reduce the number of generated IMFs while offsetting the influence of Gaussian white noise on the waveform. This not only saves calculation time but also ensures that all IMF waveforms are not affected by noise.
[0043] Among the three IMF judgment conditions in Step 3.4, two conditions a and b are conventional conditions, and condition c is unique to the method of the present invention. Taking into account that the power curve processed by the method of the present invention is synchronously related to the light intensity, the natural phenomenon of sunrise and sunset within a day determines that the light intensity first rises and then falls, and stabilizes at 0 at night. In this way, the light intensity change curve over time will have at least one maximum point and one minimum point in a day. If the number of extreme points in the power curve is less than twice the number of days included in the time range corresponding to the power curve, it means that the waveform change of the power curve has been distorted. At this time, the IMF obtained by continuing to decompose can no longer reflect the true change trend of the initial power curve, and it is meaningless. The setting of condition c is also to save time and avoid wasting time due to meaningless calculations.
[0044] The purpose of smoothing the residual in Step 3.5 is to filter out the influence of Gaussian white noise on the residual waveform. q-ave The influence of noise is so small that it can be ignored, but the power curve S q After several IMF q-aveThe subtraction may be affected by noise, resulting in a larger waveform amplitude. Since the residual will later be input into the GRU network for analysis, it needs to be filtered to prevent noise from affecting the final result. The EMA method processes data by giving higher weight to current input data and lower weight to older input data. This allows the output value at each moment in the output curve to be correlated with all previous input values, resulting in a smoother output curve relative to the input curve. The specific degree of smoothness of the output curve is determined by the smoothing coefficient β, which is an empirical constant.
[0045] Further optimization, the reconstruction of the light intensity variation data in Step 4 specifically includes the following steps:
[0046] Step 4.1. According to the various SIFs in the spatial image obtained in Step 2 m The proportion of each SIF in the entire image m How the proportion of images changes over time;
[0047] Step 4.2. Determine a characteristic color as the benchmark characteristic color, and then calculate each SIF by Grey Relational Analysis (GRA) method. m The correlation between the trend of the image proportion of the reference characteristic color image proportion over time and the trend of the image proportion of the reference characteristic color image over time; then a correlation threshold is taken and the SIF with a correlation higher than the threshold is m Retain SIFs whose correlation is below the threshold m Eliminate;
[0048] Step 4.3. Then save the various SIFs m Each SIF m The corresponding image proportion change data over time is regarded as a group of data. The multiple groups of data are reduced in dimension by the principal component analysis (PCA) method, and the multiple groups of data are converted into a set of comprehensive indicators. The comprehensive indicators are the reconstructed light intensity change data over time.
[0049] After processing in Step 2, all spatial images are composed of only a few characteristic colors. The baseline characteristic color can be selected as a characteristic color with a significantly larger area in images taken during the strongest lighting conditions over several consecutive days. The image percentage of each characteristic color is then divided by the baseline characteristic color's area percentage. The result is the correlation ratio for each characteristic color. If a characteristic color has a corresponding area in the previous image but not in the next, the correlation ratio for that characteristic color in the next image is recorded as 0. After all images have been processed, the average correlation ratio for each characteristic color is calculated. The result is the correlation between that characteristic color and the baseline characteristic color. By applying a threshold to the correlation data, characteristic colors with a low correlation between the image percentage change trend and the light intensity change trend can be eliminated, thereby improving the authenticity of the obtained light intensity data.
[0050] The dimensionality reduction of the PCA method refers to converting multiple sets of data into one or a few sets of comprehensive indicators while retaining the features. In the method of the present invention, the data on the change of the image proportion of each characteristic color over time is a set of data. The ultimate goal is to convert multiple sets of data into a set of comprehensive indicators. Assuming that a total of u spatial domain images are taken, and there are v characteristic colors screened and retained by the GRA method, then the image proportion data of all characteristic colors can form a sample matrix with v rows and u columns. Let the sample matrix be denoted as X, then X can also be expressed as a v-dimensional sample set, where v is the number of dimensions that need to be reduced by the PCA method. The sample set can be expressed as X = {x`1,…,x` u First, each sample is centered so that the mean of the data in the sample is 0; then the covariance matrix XX of the sample set is calculated. T , the covariance matrix can show the correlation between different dimensions; then XX T Perform eigenvalue decomposition to obtain the eigenvalue and the corresponding eigenvector; then find the largest eigenvalue, standardize the eigenvector corresponding to the eigenvalue to obtain the characteristic matrix vector W; finally, multiply each sample in X by W, and the result is the required set of comprehensive indicators.
[0051] Data screening and reconstruction are essential for neural network predictions. Large amounts of data can make it difficult for the neural network to discern data features, or overfit the data, affecting the accuracy of the analysis results. Furthermore, large amounts of data input inevitably increase the amount of computation required for analysis. Data screening and reconstruction save time by reducing the amount of data required for analysis.
[0052] Further optimization, the photovoltaic power prediction in Step 5 specifically includes the following steps:
[0053] Step 5.1. Input the light intensity data over time into the GRU network and train the GRU network to recognize the temporal features contained in the light intensity data.
[0054] Step 5.2. Input each IMF and residual obtained by decomposition into the GRU network for iterative calculation. After obtaining the predicted value of each IMF and residual output, all predicted values are summed to obtain the final predicted value of photovoltaic power.
[0055] The GRU network uses the same gate control principle as the LSTM network, but the GRU network has a relatively simpler structure. The GRU network only contains an update gate and a reset gate. Its specific operation process is shown in the following formula:
[0056] z t =σ(W z ·[h t-1 ,x t ]+b z )
[0057] r t =σ(W r ·[h t-1 ,x t ]+b r )
[0058]
[0059] where z t 、r t 、 h t They are update gate, reset gate, candidate hidden state, and update hidden state respectively. σ is the sigmoid activation function, and b z with b r are update gate bias and reset gate bias respectively, W z 、W r , W are update gate weight matrix, reset gate weight matrix and candidate hidden weight matrix respectively, x t is the input value.
[0060] After the light intensity changes over time data is input into the GRU network, the time-related features in the data will be updated at the gate z t is retained, reset gate r t The weight of noise or other irrelevant features in the data will be reduced, so that the noise and irrelevant features are gradually ignored in the cycle calculation. and update the hidden state h tUnder the premise of retaining the characteristics of the data at the previous moment, new data is introduced into the loop calculation, so that the GRU network can continuously learn the time series characteristics of the data. The method of the present invention selects light intensity data as training data because the time series characteristics of light intensity data are consistent with the time series characteristics of photovoltaic power data, which can calibrate the analysis of IMF or residual time series characteristics. The IMF and residual are analyzed one by one because the IMF only contains a single oscillation mode, its characteristic analysis is simpler, and the analysis results can be more accurate. The residual also contains time series characteristics, so it also needs to be analyzed. Because all IMFs and residuals are gradually decomposed from the initial power curve, the complete power curve prediction result can be obtained by summing the prediction results of each IMF and residual. This approach is more reliable and faster than directly analyzing the initial power curve.
[0061] The beneficial effects of the method of the present invention are:
[0062] 1. The method of the present invention adopts a dual-model analysis strategy based on light intensity data and photovoltaic power history data, which can effectively avoid the influence of accidental factors such as noise on the prediction results and increase the reliability of the prediction results;
[0063] 2. The raw materials required for the analysis of the method of the present invention only include airspace images and photovoltaic power data. No additional professional observation or data processing equipment is required, making the method highly feasible in practical operation.
[0064] 3. The method of the present invention improves upon the existing EEMD signal decomposition method, making the decomposition and calculation of the power curve faster. The screening and reconstruction of the light intensity data and the use of the GRU network are also designed to speed up data analysis, making the method of the present invention highly efficient.
[0065] 4. The method of the present invention can make medium- and long-term predictions of photovoltaic power change trends, and its prediction results are more valuable for reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 Schematic diagram of the overall process of the method of the present invention;
[0067] Figure 2 Schematic diagram of the spatial image area division and ratio calculation process;
[0068] Figure 3 Flowchart of the improved signal decomposition method based on the EEMD method. DETAILED DESCRIPTION
[0069] To make the purpose, technical solution, and advantages of the method of the present invention more clear, the technical solution of the present invention will be clearly and completely described below through specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0070] Example 1:
[0071] A photovoltaic power prediction method based on improved signal decomposition and deep learning fusion includes the following steps:
[0072] Step 1. Take 60 airspace images of the area to be predicted every two hours for five days. All images are taken with the same optical parameters. The photovoltaic power generation data for the area to be predicted is recorded during the five-day period.
[0073] Step 2. Divide each airspace image into different regions based on image brightness, record the brightness levels of each region, and then calculate the area ratio of each region in each image. This includes the following steps:
[0074] Step 2.1. Convert the color space used to measure spatial image color to LAB. Each pixel in the image is assigned a unique serial number. Determine the color of each pixel, that is, the corresponding L-axis, A-axis, and B-axis coordinates of the pixel in LAB space, and record them in the form of (l, a, b).
[0075] Step 2.2. Use the K-means clustering analysis method to classify all pixels in the image. The specific steps of classification are as follows:
[0076] Step 2.2.1. Determine the number of cluster centers to be 10, the center threshold to be 40, and the iteration threshold to be 20. Then randomly select 10 pixels in the image as the initial cluster centers; the number of iterations is recorded as e, the initial value of e is 0, and the cluster center is represented by C i , i is the cluster center number, 1≤i≤10;
[0077] Step 2.2.2. The total number of pixels in the image other than the cluster center is n, and the other pixels are represented by P j , j is the pixel number, 1≤j≤n, take a cluster center C i , calculate C i With each P j The distance D Ci-Pj And record, the calculation formula is:
[0078]
[0079] where l Ci 、a Ci 、b Ci C i The corresponding L-axis, A-axis, and B-axis coordinates in the LAB color space, l Pj 、a Pj 、b Pj P j The corresponding L-axis, A-axis, and B-axis coordinates in the LAB color space; when all P j Completed with the current C i After calculating the distance, select another C i Continue calculating until all C i All calculations are completed;
[0080] Step 2.2.3. Compare P j With each cluster center C i The distance between them, find the shortest distance and the C corresponding to the shortest distance i , and P j The closest C i Associated, all P j You need to find the C associated with it i , and then each C i and all P associated with it j into the same category;
[0081] Step 2.2.4. Calculate the average of the L-axis, A-axis, and B-axis coordinates of all pixels in the same class, take the point corresponding to the average of the three axes as the centroid of the class, and calculate the centroid and the cluster center C of the class. i The distance d between the center of mass and C i If the distance is not greater than the center threshold, that is, if the condition d≤40 is met, the classification result is output; if not, the number of iterations is compared to see if it reaches the iteration threshold, that is, e=20. If it is met, the classification result is output; if it is still not met, the e value is increased by 1, and the centroid is used as the new cluster center C i , and then repeat the above process from Step 2.2.2;
[0082] Step 2.3. a in all cluster center coordinates Ci Uniformly replaced with -2.42, b Ci It is also uniformly replaced with -12.13, and the cluster center after coordinate replacement is recorded as C` i ; Set all C` i Sort by the L-axis coordinate value, each C` i The corresponding serial number is recorded as m, and then each C` iThe corresponding coordinates are regarded as a characteristic color, and the characteristic color is recorded as SIF m ;
[0083] Step 2.4. Calculate other pixels P in the image j The distance from each characteristic color in the L-axis direction, find the minimum distance and the SIF corresponding to the minimum distance m , and then use the SIF corresponding to the minimum distance to determine the color of the pixel m Replace; each pixel P in the image j All colors require a SIF m replace;
[0084] Step 2.5. Each SIF m The range covered in the image is considered as an area, and the number of pixels in each area is counted. Then the proportion of the area in the image is calculated by dividing the number of pixels in the area by the total number of pixels in the image. The overall process of spatial image area division and area proportion calculation is as follows: Figure 2 As shown, the specific data of the LAB space coordinates corresponding to various characteristic colors and the area ratios of different characteristic colors are listed in Figure 2 middle.
[0085] Step 3. Decompose the time-varying curve of photovoltaic power generation using a signal decomposition algorithm, and decompose the power curve into multiple intrinsic mode functions (IMFs) and a residual. The specific steps include:
[0086] Step 3.1. Generate 30 Gaussian white noises. The Gaussian white noise is a curve connected by multiple noise distribution points. The curve takes time as the independent variable and power as the dependent variable. The distribution shapes of all Gaussian white noise curves are different, and the duration of all curves is not shorter than the duration of the power curve. The multiple noise distribution points are evenly distributed on the time scale and Gaussian distributed with the mean at position 0 on the power scale. The 30 white noises are recorded as {noise1,…,noise 30};
[0087] Step 3.2. Compare 30 white noises with the power curve S to be decomposed q Superposition, q is the number of iterations, q = 0 in the initial power curve; the superposition of the power curve and noise is represented by S q-s =S q +noise s , s is the white noise number, 1≤s≤30, S q-s There are 30 curves formed by superposition;
[0088] Step 3.3. For all S q-sPerform Empirical Mode Decomposition (EMD). The decomposition steps of the EMD method are: first, q-s The upper envelope of the curve is formed by interpolating the maximum points, and the lower envelope of the curve is formed by interpolating the minimum points. Then the upper and lower envelopes are integrated and averaged to obtain the average line. The definition of integrated average is the same as that described later. Then, S q-s Subtracting the average line generates another curve, which is the generated intrinsic mode function, denoted as IMF q-s ; for all IMF q-s The IMF is obtained by averaging the q-ave For any time t within the time range corresponding to the power curve, all IMFs exist q-s The arithmetic mean of the corresponding values at time t and the IMF q-ave The relationship in which the corresponding values are equal at time t;
[0089] Step 3.4. Determine the IMF q-ave Are the following conditions met?
[0090] d. In the entire curve, the number of extreme points and zero crossing points is the same or differs by only one;
[0091] e. The upper envelope formed by the maximum point of the curve and the lower envelope formed by the minimum point are symmetrical about the time axis;
[0092] f. The number of extreme points in the curve is not less than twice the number of days included in the time range corresponding to the initial power curve;
[0093] If all three conditions are met at the same time, then S q+1 =S q -IMF q-ave The new power curve to be decomposed is obtained by processing, and then the above process is repeated from Step 3.2; if the three conditions are not met at the same time, the power curve to be decomposed S in this iterative calculation is q as residuals;
[0094] Step 3.5. Smooth the residuals using the Exponential Moving Average (EMA) method. The specific steps are:
[0095] Step 3.5.1. Take multiple sampling points on the residual, T is the sampling point number, all sampling points are arranged in the order of sampling time and the sampling time interval between two adjacent points is equal, and record the corresponding time value and power value F of each sampling point T ;
[0096] Step 3.5.2. Set all F T Substitute into the EMA method calculation formula in order, the calculation formula is as follows
[0097] G T =βG T-1 +(1-β)F T (2)
[0098] Among them G T is the power value corresponding to the Tth sampling point after processing by the EMA method, G0=0, β is the smoothing coefficient, which is 0.8 in this embodiment;
[0099] The time values of all sampling points are retained, and the power values are replaced by the corresponding G T The data points are formed, and the curve formed by these data points is the residual after smoothing.
[0100] Step 4. Based on the proportion of the areas of different brightness in all spatial images, summarize the characteristics of the overall spatial illumination intensity over time, and reconstruct the data of the illumination intensity over time based on the obtained characteristics. The specific steps include:
[0101] Step 4.1. According to the various SIFs in the spatial image obtained in Step 2 m The proportion of each SIF in the entire image m How the proportion of images changes over time;
[0102] Step 4.2. Select the characteristic color with the largest area in the image taken at the time of the strongest illumination during the five days as the benchmark characteristic color, and then calculate all other SIFs using the Grey Relational Analysis (GRA) method. m The correlation degree with the benchmark characteristic color in the image proportion over time is first calculated by dividing the proportion of the corresponding area of each characteristic color in the same image by the proportion of the benchmark characteristic color area. The result is the correlation ratio of each characteristic color. If a certain characteristic color has a corresponding area in the previous image but not in the next image, the correlation ratio of the characteristic color in the next image is recorded as 0. After all images are calculated, the average value of the correlation ratio of each characteristic color is calculated. The result is the correlation degree between the characteristic color and the benchmark characteristic color. Then, a correlation threshold is taken, and the SIF with a correlation degree higher than the threshold is added. m Retain SIFs whose correlation is below the threshold m After elimination, there are four characteristic colors that are finally retained;
[0103] Step 4.3. Then save the various SIFs m Each SIFm The corresponding image proportion changes over time are regarded as a set of data. The Principal Component Analysis (PCA) method is used to reduce the dimensionality of multiple data sets and transform them into a comprehensive indicator. A total of 60 spatial images were taken, and there are 4 characteristic colors that are retained after screening by the GRA method. Then the image proportion data of all characteristic colors can form a sample matrix with 4 rows and 60 columns. Let the sample matrix be denoted as X, then X can also be expressed as a v-dimensional sample set, where v is the number of dimensions that need to be reduced by the PCA method. The sample set can be expressed as X = {x`1,…,x` 60 First, each sample is centered so that the mean of the data in the sample is 0; then the covariance matrix XX of the sample set is calculated. T , the covariance matrix can show the correlation between different dimensions; then XX T Perform eigenvalue decomposition to obtain the eigenvalue and the corresponding eigenvector; then find the largest eigenvalue and normalize the eigenvector corresponding to the eigenvalue to obtain the characteristic matrix vector W; finally, multiply each sample in X by W to obtain the desired comprehensive index, which is the reconstructed light intensity change data over time;
[0104] Step 5. Analyze the reconstructed light intensity data, multiple IMFs, and residuals through a neural network to provide a PV power prediction result. This includes the following steps:
[0105] Step 5.1. Input the light intensity data over time into the GRU network and train the GRU network to recognize the temporal features contained in the light intensity data.
[0106] Step 5.2. Input each IMF and residual obtained by decomposition into the GRU network for iterative calculation. After obtaining the predicted value of each IMF and residual output, all predicted values are summed to obtain the final predicted value of photovoltaic power.
Claims
1. A photovoltaic power prediction method based on improved signal decomposition and deep learning fusion, characterized in that: The following steps are involved: Step 1. Take multiple airspace images of the area to be predicted over a period of time. The time interval between each shot is equal, and the optical parameters for all airspace images are set to the same. Record the time of each shot and the photovoltaic power generation data of the area to be predicted during the shooting period. Step 2. Divide each spatial image into different regions based on image brightness, record the brightness levels of each region, and then calculate the area ratio of each region in each image; Step 3. Decompose the time-varying curve of photovoltaic power generation using a signal decomposition algorithm, and decompose the power curve into multiple IMFs and a residual; Step 4. Based on the proportion of the areas of different brightness in all spatial images, summarize the characteristics of the overall spatial illumination intensity over time, and reconstruct the data of the illumination intensity over time based on the obtained characteristics; Step 5. Analyze the reconstructed light intensity data, multiple IMFs, and residuals through a neural network to provide a photovoltaic power prediction result.
2. The photovoltaic power prediction method based on improved signal decomposition and deep learning fusion according to claim 1, characterized in that: The spatial image region division and region ratio calculation in Step 2 specifically include the following steps: Step 2.
1. Convert the color space used to measure spatial image color to LAB. Each pixel in the image is assigned a unique serial number. Determine the color of each pixel, that is, the corresponding L-axis, A-axis, and B-axis coordinates of the pixel in LAB space, and record them in the form of (l, a, b). Step 2.
2. Use the K-means clustering method to classify all pixels in the image. Each class contains a cluster center. Step 2.
3. The A-axis coordinates of all cluster centers are replaced by a constant, and the B-axis coordinates are also replaced by a constant. The cluster center after coordinate replacement is recorded as C` i ; Set all C` i Sort by the L-axis coordinate value, each C` i The corresponding serial number is recorded as m, and then each C` i The corresponding coordinates are regarded as a characteristic color, and the characteristic color is recorded as SIF m ; Step 2.
4. Calculate the distance between each characteristic color and the pixels other than the cluster center in the image along the L axis, and find the minimum distance and the SIF corresponding to the minimum distance. m , and then use the SIF corresponding to the minimum distance to calculate the color of the pixel m Replacement; the color of each pixel in the image needs to be replaced with a SIF m replace; Step 2.
5. Each SIF m The range covered in the image is regarded as a region, the number of pixels in each region is counted, and then the proportion of the region in the image is calculated by dividing the number of pixels in the region by the total number of pixels in the image.
3. The photovoltaic power prediction method based on improved signal decomposition and deep learning fusion according to claim 1, characterized in that: The decomposition of the power curve in Step 3 specifically includes the following steps: Step 3.
1. Generate r Gaussian white noises, which are curves connected by multiple noise distribution points. The curve takes time as the independent variable and power as the dependent variable. The distribution shapes of all Gaussian white noise curves are different, and the duration of all curves is not shorter than the duration of the power curve. The multiple noise distribution points are evenly distributed on the time scale and Gaussian distributed with the mean at position 0 on the power scale. The r white noises are recorded as {noise1,…,noise r }; Step 3.
2. Compare the r white noises with the power curve S to be decomposed q Superposition, q is the number of iterations, q = 0 in the initial power curve; the superposition of the power curve and noise is represented by S q-s =S q +noise s , s is the white noise number, 1≤s≤r, S q-s is the curve formed by superposition, with a total of r curves; Step 3.
3. For all S q-s Perform Empirical Mode Decomposition (EMD), each S q-s Correspondingly, an intrinsic mode function is generated, expressed as IMF q-s ; for all IMF q-s The IMF is obtained by averaging the q-ave For any time t within the time range corresponding to the power curve, all IMFs exist q-s The arithmetic mean of the corresponding values at time t and the IMF q-ave The relationship in which the corresponding values are equal at time t; Step 3.
4. Determine the IMF q-ave Are the following conditions met? a. In the entire curve, the number of extreme points and zero-crossing points is the same or differs by only one; b. The upper envelope formed by the maximum point of the curve and the lower envelope formed by the minimum point are symmetrical about the time axis; c. The number of extreme points in the curve is not less than twice the number of days included in the time range corresponding to the initial power curve; If all three conditions are met at the same time, then S q+1 =S q -IMF q-ave The new power curve to be decomposed is obtained by processing, and then the above process is repeated from Step 3.2; if the three conditions are not met at the same time, the power curve to be decomposed S in this iterative calculation is q as residuals; Step 3.
5. Smooth the residuals using the Exponential Moving Average (EMA) method. The specific steps are as follows: Step 3.5.
1. Take multiple sampling points on the residuals, where T is the sampling point number. All sampling points are arranged in the order of sampling time, and the sampling time intervals between adjacent points are equal. Record the time value and power value F corresponding to each sampling point. T ; Step 3.5.
2. Set all F T Substitute into the EMA method calculation formula in order, the calculation formula is as follows G T =βG T-1 +(1-β)F T (2) Among them G T is the power value corresponding to the Tth sampling point after processing by the EMA method, G0 = 0, β is the smoothing coefficient; The time values of all sampling points are retained, and the power values are replaced by the corresponding G T The data points are formed, and the curve formed by these data points is the residual after smoothing.
4. The photovoltaic power prediction method based on improved signal decomposition and deep learning fusion according to claim 1, characterized in that: The reconstruction of the light intensity variation data in Step 4 specifically includes the following steps: Step 4.
1. According to the various SIFs in the spatial image obtained in Step 2 m The proportion of each SIF in the entire image m How the proportion of images changes over time; Step 4.
2. Determine a characteristic color as the benchmark characteristic color, and then calculate each SIF by Grey Relational Analysis (GRA) method. m The correlation between the trend of the image proportion of the reference characteristic color image proportion over time and the trend of the image proportion of the reference characteristic color image over time; then a correlation threshold is taken and the SIF with a correlation higher than the threshold is m Retain SIFs whose correlation is below the threshold m Eliminate; Step 4.
3. Then save the various SIFs m Each SIF m The corresponding image proportion change data over time is regarded as a group of data. The multiple groups of data are reduced in dimension by the principal component analysis (PCA) method, and the multiple groups of data are converted into a set of comprehensive indicators. The comprehensive indicators are the reconstructed light intensity change data over time.
5. The photovoltaic power prediction method based on improved signal decomposition and deep learning fusion according to claim 1, characterized in that: The photovoltaic power prediction in Step 5 specifically includes the following steps: Step 5.
1. Input the light intensity data over time into the GRU network and train the GRU network to recognize the temporal features contained in the light intensity data. Step 5.
2. Input each IMF and residual obtained by decomposition into the GRU network for iterative calculation. After obtaining the predicted value of each IMF and residual output, all predicted values are summed to obtain the final predicted value of photovoltaic power.