Data fusion photovoltaic generation power prediction method and system
By collecting and processing feature vectors from satellite cloud images and radar echo images, a fused feature matrix is constructed and input into the prediction model, which solves the problem of incomplete meteorological data in photovoltaic power generation prediction and improves prediction accuracy.
Patent Information
- Application Number
- CN202512042200.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-17
AI Technical Summary
Existing photovoltaic power generation prediction methods rely on a single data source, making it difficult to comprehensively reflect the overall meteorological conditions of the photovoltaic power station area. In particular, when clouds move quickly, local monitoring data cannot reflect changes in cloud distribution in a timely manner, resulting in a large deviation between the prediction results and the actual power generation.
Satellite cloud images and radar echo images of the photovoltaic area are collected. The first feature vector and the second feature vector are obtained through pixel segmentation and feature extraction. A fused feature matrix is constructed and input into a preset power prediction model to calculate the power generation.
It effectively reduces prediction errors in complex scenarios such as extreme weather and rapid cloud movement, providing more reliable technical support for stable grid dispatch and renewable energy consumption.
Smart Images

Figure CN121884276A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of photovoltaic technology, and in particular to a data fusion method and system for predicting photovoltaic power generation. Background Technology
[0002] Currently, most existing photovoltaic (PV) power generation forecasting methods rely on a single type of data source. Some methods only build statistical models based on historical power generation data. These models cannot effectively capture the impact of sudden changes in meteorological conditions on power generation, and the prediction error often increases significantly during extreme weather or seasonal transitions. Other methods attempt to combine temperature and humidity data collected by ground-based meteorological stations for forecasting, but the monitoring range of these stations is limited and cannot fully reflect the overall meteorological conditions of the PV power plant area. Especially when cloud cover moves rapidly, local monitoring data cannot reflect changes in cloud distribution across the entire PV area in a timely manner, leading to significant deviations between the predicted results and the actual power generation. Summary of the Invention
[0003] The purpose of this invention is to at least partially solve one of the technical problems existing in the prior art.
[0004] To achieve the above objectives, the present invention provides a data fusion method for predicting photovoltaic power generation, comprising the following steps:
[0005] Satellite cloud images and radar echo images of the photovoltaic area are collected, and pixel segmentation and feature extraction are performed on the satellite cloud images and radar echo images to obtain the first feature vector and the second feature vector.
[0006] Construct a fused feature matrix based on the first and second eigenvectors;
[0007] The fused feature matrix is input into a preset power prediction model to calculate the power generation output value of the photovoltaic power generation.
[0008] Furthermore, satellite cloud images and radar echo images of the photovoltaic area were collected, including:
[0009] The geographical coordinate range of the photovoltaic area is determined, and the data collection boundary range is set based on the geographical coordinate range;
[0010] Meteorological data is collected from the photovoltaic area based on the aforementioned collection boundary range to obtain satellite cloud images, and radar scans are performed on the photovoltaic area based on the aforementioned collection boundary range to obtain radar echo images.
[0011] Furthermore, meteorological data is collected from the photovoltaic area based on the aforementioned collection boundary range to obtain satellite cloud images, including:
[0012] The geographic coordinate range is converted into satellite image acquisition coordinates to obtain the converted coordinate range. Based on the converted coordinate range, a data acquisition command is sent to the meteorological satellite to collect meteorological data of the photovoltaic area, and an initial satellite cloud image is obtained.
[0013] The initial satellite cloud image is subjected to image correction processing to eliminate image distortion, resulting in a corrected cloud image. The corrected cloud image is then filtered at the pixel level to extract satellite cloud images corresponding to the acquisition boundary range.
[0014] Furthermore, the photovoltaic area is scanned by radar based on the acquisition boundary range to obtain a radar echo image, including:
[0015] The radar scanning parameters are set for the acquisition boundary range to determine the frequency and angle range of the radar scan, and the set scanning parameters are obtained. Based on the set scanning parameters, the radar equipment is controlled to scan the photovoltaic area to obtain initial radar echo data containing precipitation particle reflection intensity and distance information.
[0016] The initial radar echo data is parsed and processed to convert it into a visual image format to obtain a preliminary echo image. The preliminary echo image is then subjected to noise filtering to eliminate interference signals, thus obtaining the final radar echo image.
[0017] Furthermore, pixel segmentation and feature extraction are performed on the satellite cloud image and radar echo image to obtain a first feature vector and a second feature vector, including:
[0018] The satellite cloud image is converted to grayscale to obtain a grayscale cloud image, and the grayscale cloud image is then segmented using a threshold to obtain cloud pixel regions.
[0019] Texture features are extracted from the cloud pixel region to obtain cloud texture features. Based on the cloud texture features, cloud types are classified to obtain a first classification result including cumulus, stratus and cirrus. The first classification result is then quantized and encoded to obtain a first feature vector.
[0020] The radar echo image is subjected to intensity grading processing to obtain a graded echo map, and edge detection is performed on the graded echo map to identify echo edge regions. Shape features are extracted from the echo edge regions to obtain echo shape features.
[0021] Precipitation type is determined based on the echo shape characteristics to obtain a second classification result. The second classification result is then quantized and encoded to obtain a second feature vector.
[0022] Furthermore, threshold segmentation is performed on the grayscale cloud image to obtain cloud pixel regions, including:
[0023] The grayscale cloud map is statistically analyzed to obtain a grayscale histogram, and the peak and valley distribution of the grayscale histogram is analyzed. The initial segmentation threshold is determined based on the peak and valley distribution.
[0024] The grayscale cloud image is initially segmented based on the initial segmentation threshold to obtain a preliminary segmentation region. Then, connectivity analysis is performed on the preliminary segmentation region to merge adjacent pixel regions with similar grayscale values to obtain a cloud pixel region.
[0025] Furthermore, determining the initial segmentation threshold based on the peak and trough distribution includes:
[0026] Peak values are extracted from the peak and trough distribution to obtain the main peak and trough positions. Based on the main peak and trough positions, the gray-level interval sequence of adjacent peaks and troughs is calculated, wherein the gray-level interval sequence includes the gray-level difference between the cloud layer and the background.
[0027] The maximum interval is detected in the gray-level interval sequence to obtain the position of the maximum gray-level span, and the initial segmentation threshold is determined based on the position of the maximum gray-level span.
[0028] Furthermore, a fused feature matrix is constructed based on the first and second feature vectors, including:
[0029] The cloud impact quantification is performed on the first feature vector to obtain the first weight, and the precipitation impact quantification is performed on the second feature vector to obtain the second weight;
[0030] The first feature vector is weighted based on the first weight to obtain a weighted first vector; the second feature vector is weighted based on the second weight to obtain a weighted second vector; the weighted first vector and the weighted second vector are then concatenated column by column to obtain a fused feature matrix.
[0031] Furthermore, the fused feature matrix is input into a preset power prediction model to calculate the power generation output value of the photovoltaic power generation, including:
[0032] The fused feature matrix is input into a preset power prediction model; wherein the power prediction model includes a feature encoding layer, a temporal correlation layer, and a power mapping layer;
[0033] The fused feature matrix is input into the feature coding layer for layer-by-layer nonlinear transformation to obtain the coding feature vector. Based on the coding feature vector, the attenuation of the theoretical irradiance of the photovoltaic module is calculated to obtain the irradiance attenuation coefficient.
[0034] The encoded feature vector and the irradiation attenuation coefficient are input into the temporal association layer to perform state transfer calculations between time steps to obtain the temporal association vector.
[0035] The time-series correlation vector is input into the power mapping layer for weighted summation to obtain the initial power value; and the initial power value is subjected to upper and lower limit constraints based on the installed capacity of the photovoltaic power station to obtain the photovoltaic power output value.
[0036] The present invention also provides a data fusion photovoltaic power generation prediction system, comprising:
[0037] The acquisition module is used to acquire satellite cloud images and radar echo images of the photovoltaic area, and to perform pixel segmentation and feature extraction on the satellite cloud images and radar echo images to obtain the first feature vector and the second feature vector.
[0038] The fusion module is used to construct a fused feature matrix based on the first feature vector and the second feature vector;
[0039] The calculation module is used to input the fused feature matrix into a preset power prediction model to calculate the power generation and obtain the photovoltaic power output value.
[0040] This invention provides a data fusion-based photovoltaic power generation prediction method, comprising the following steps: acquiring satellite cloud images and radar echo images of the photovoltaic area, and performing pixel segmentation and feature extraction on the satellite cloud images and radar echo images to obtain a first feature vector and a second feature vector; constructing a fusion feature matrix based on the first feature vector and the second feature vector; inputting the fusion feature matrix into a preset power prediction model to calculate the power generation output value of the photovoltaic power generation. This method solves the technical problem that traditional technologies cannot fully reflect the overall meteorological conditions of the photovoltaic power station area, especially when the cloud layer moves quickly, and local monitoring data cannot reflect the changes in cloud distribution in the entire photovoltaic area in a timely manner, resulting in a large deviation between the prediction results and the actual power generation. It effectively reduces prediction errors in complex scenarios such as extreme weather and rapid cloud movement, and provides a more reliable technical guarantee for stable grid dispatch and efficient consumption of new energy. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1This is a schematic diagram of the steps of a photovoltaic power generation prediction method based on data fusion in one embodiment of the present invention;
[0043] Figure 2 This is a schematic diagram of a photovoltaic power generation prediction system based on data fusion in one embodiment of the present invention;
[0044] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0045] The embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0046] The following describes in detail, with reference to the accompanying drawings, a data fusion method for predicting photovoltaic power generation according to an embodiment of the present invention.
[0047] Figure 1 This invention provides a data fusion-based photovoltaic power generation prediction method, comprising the following steps:
[0048] Step S1: Collect satellite cloud images and radar echo images of the photovoltaic area, and perform pixel segmentation and feature extraction on the satellite cloud images and radar echo images to obtain the first feature vector and the second feature vector.
[0049] Specifically, we first focus on the photovoltaic area, collecting satellite cloud images and radar echo images. After collection, we don't use them directly; we first process these two types of images—specifically, pixel segmentation. For example, when dealing with satellite cloud images, we separate cloud areas and clear-sky areas corresponding to different grayscale values, clearly distinguishing between thick and thin cloud pixel blocks; the same applies to radar echo images, where we break down strong and weak echo areas based on echo intensity pixel values. After pixel segmentation, we extract features from these two types of processed images. For example, from the segmented areas of satellite cloud images, we extract features such as cloud coverage area and average grayscale value, organizing them into a first feature vector; from the segmented areas of radar echo images, we extract features such as average echo intensity and echo movement speed, organizing them into a second feature vector.
[0050] Step S2: Construct a fused feature matrix based on the first feature vector and the second feature vector.
[0051] Specifically, after obtaining the first and second eigenvectors, they cannot be directly combined; some adaptation processing is required first. For example, the first eigenvector contains parameters such as cloud coverage area and average gray value, while the second eigenvector contains the average echo intensity and echo velocity. The dimensions of the two types of vectors must first be adjusted to be consistent—for instance, if the first eigenvector is 1×8, the second eigenvector is also processed to 1×8, ensuring that the parameters at each corresponding position reflect the meteorological conditions at the same time. After adjustment, the two vectors are stacked column by column, for example, the first eigenvector as the first column of the matrix and the second eigenvector as the second column, thus forming a fused feature matrix. If the feature weights are different, a weighting coefficient is added to different features before stacking; for example, cloud coverage area, which has a greater impact, is assigned more weight to ensure that the matrix can integrate information from both types of features.
[0052] Step S3: Input the fused feature matrix into the preset power prediction model to calculate the power generation and obtain the photovoltaic power output value.
[0053] Specifically, the next step after constructing the fusion feature matrix is to input it into the preset power prediction model. However, before inputting, it's necessary to confirm that the matrix format matches the model requirements. For example, if the model requires row vector input, the matrix must be converted to the corresponding format to avoid calculation errors due to format incompatibility. After input, the model will calculate based on the correlation between meteorological features and power generation power trained within it—for example, when encountering situations with large cloud coverage and high echo intensity in the matrix, the model will calculate the corresponding lower power generation power based on past data. Once the model completes its calculations, the final result is the photovoltaic power generation output value. This value will also be compared with the model's preset accuracy threshold to ensure the reliability of the result.
[0054] In a specific embodiment, satellite cloud images and radar echo images of the photovoltaic area are collected, including:
[0055] The geographical coordinate range of the photovoltaic area is determined, and the data collection boundary range is set based on the geographical coordinate range;
[0056] Meteorological data is collected from the photovoltaic area based on the aforementioned collection boundary range to obtain satellite cloud images, and radar scans are performed on the photovoltaic area based on the aforementioned collection boundary range to obtain radar echo images.
[0057] Specifically, to collect satellite cloud images and radar echo images of a photovoltaic (PV) area, the first step is to determine the geographical coordinates of that area. For example, if a PV power station is located between 30°15′ and 30°20′ north latitude and 120°05′ and 120°10′ east longitude, this latitude and longitude range is clearly defined; this is the determined geographical coordinate range. After determining this range, the data collection doesn't begin directly. Instead, the collection boundary is set based on it—generally, it's expanded slightly outward by 1-2 kilometers from the determined geographical coordinate range. For example, the area mentioned earlier would be set to 30°14′ and 30°21′ north latitude and 120°04′ and 120°11′ east longitude. This helps avoid missing meteorological data at the boundaries.
[0058] Once the data collection boundary is set, two things can be done based on this range: one is to collect meteorological data on the photovoltaic area to obtain satellite cloud images. This usually involves calling the real-time data interface of a meteorological satellite, specifying the data collection boundary range just set, and the satellite will transmit back the cloud distribution image within this range, which is the required satellite cloud image; the other is to perform radar scanning on the photovoltaic area based on the same data collection boundary range. For example, by using a nearby meteorological radar station, the radar scanning range is adjusted to match the data collection boundary range. The image obtained after scanning, reflecting the distribution of precipitation particles, is the desired radar echo image.
[0059] In a specific embodiment, meteorological data is collected from the photovoltaic area based on the collection boundary range to obtain a satellite cloud image, including:
[0060] The geographic coordinate range is converted into satellite image acquisition coordinates to obtain the converted coordinate range. Based on the converted coordinate range, a data acquisition command is sent to the meteorological satellite to collect meteorological data of the photovoltaic area, and an initial satellite cloud image is obtained.
[0061] The initial satellite cloud image is subjected to image correction processing to eliminate image distortion, resulting in a corrected cloud image. The corrected cloud image is then filtered at the pixel level to extract satellite cloud images corresponding to the acquisition boundary range.
[0062] Specifically, to collect meteorological data for the photovoltaic area based on the previously set collection boundary and obtain satellite cloud images, the first step is to process the coordinates. This involves converting the geographic coordinates of the collection boundary into satellite imagery coordinates that the meteorological satellite can recognize. This conversion process requires specialized coordinate transformation tools. For example, geographic coordinates such as 30°14′-30°21′ N and 120°04′-120°11′ E are converted into row and column coordinates for satellite images, such as the X-axis range of 1200-1500 and the Y-axis range of 800-1100 for a specific meteorological satellite. This is the converted coordinate range. With this range, a data collection command is sent to the meteorological satellite, explicitly specifying the converted coordinate range. After receiving the command, the satellite will photograph the corresponding area and transmit the initial satellite cloud image.
[0063] After obtaining the initial satellite cloud image, it cannot be used directly; image correction processing must be performed first. For example, due to the angle during satellite imaging, the image edges may be stretched and distorted. In this case, radiometric correction and geometric correction methods are used to eliminate these distortions. The processed image is the corrected cloud image. Next, pixel-level filtering is performed on the corrected cloud image. For example, it is checked whether the coordinates of each pixel are within the pixel range corresponding to the previously acquired boundary. Pixels outside the range are removed, leaving only the qualified parts. This extracts the satellite cloud image corresponding to the acquired boundary range.
[0064] In a specific embodiment, a radar scan is performed on the photovoltaic area based on the acquisition boundary range to obtain a radar echo image, including:
[0065] The radar scanning parameters are set for the acquisition boundary range to determine the frequency and angle range of the radar scan, and the set scanning parameters are obtained. Based on the set scanning parameters, the radar equipment is controlled to scan the photovoltaic area to obtain initial radar echo data containing precipitation particle reflection intensity and distance information.
[0066] The initial radar echo data is parsed and processed to convert it into a visual image format to obtain a preliminary echo image. The preliminary echo image is then subjected to noise filtering to eliminate interference signals, thus obtaining the final radar echo image.
[0067] Specifically, to perform radar scanning on the photovoltaic area based on the previously determined acquisition boundary range and ultimately obtain radar echo images, the first step is to set the radar scanning parameters. These parameters must be determined according to the acquisition boundary range. For example, if the previously set boundary is 30°14′-30°21′N and 120°04′-120°11′E, then the radar scanning angle range must cover this area. This could be set with an azimuth angle from 80° to 110° and an elevation angle from 0.5° to 3°. For the frequency, commonly used for S-band radar, it would be set to around 2.8GHz. These determined frequency and angle ranges constitute the post-scanning parameters. With these parameters, the radar equipment is used to control the radar, scanning the photovoltaic area. The scanned data, which reflects the intensity of precipitation particle reflection and the distance of particles to the radar, constitutes the initial radar echo data.
[0068] After obtaining the initial radar echo data, data analysis processing is required. Since the data at this stage is still raw electrical signal data, specialized analysis tools are needed to convert it into visual image formats such as PNG and TIFF. After conversion, the preliminary echo image is obtained. However, the preliminary echo image often contains clutter interference, such as signals reflected from ground objects. In this case, Kalman filtering or mean filtering methods are used to filter out noise and remove these interfering signals. The resulting clear image is the desired radar echo image.
[0069] In a specific embodiment, pixel segmentation and feature extraction are performed on satellite cloud images and radar echo images to obtain a first feature vector and a second feature vector, including:
[0070] The satellite cloud image is converted to grayscale to obtain a grayscale cloud image, and the grayscale cloud image is then segmented using a threshold to obtain cloud pixel regions.
[0071] Texture features are extracted from the cloud pixel region to obtain cloud texture features. Based on the cloud texture features, cloud types are classified to obtain a first classification result including cumulus, stratus and cirrus. The first classification result is then quantized and encoded to obtain a first feature vector.
[0072] The radar echo image is subjected to intensity grading processing to obtain a graded echo map, and edge detection is performed on the graded echo map to identify echo edge regions. Shape features are extracted from the echo edge regions to obtain echo shape features.
[0073] Precipitation type is determined based on the echo shape characteristics to obtain a second classification result. The second classification result is then quantized and encoded to obtain a second feature vector.
[0074] Specifically, to perform pixel segmentation and feature extraction on satellite cloud images and radar echo images, and ultimately obtain the first and second feature vectors, we must begin with the processing of the satellite cloud images. The first step is to perform grayscale processing on the satellite cloud images, converting the original color cloud images into images with only black, white, and gray gradients. For example, the blue clear sky areas and white cloud areas in the cloud image are mapped to grayscale values from 0 to 255 according to their brightness. This results in a grayscale cloud image. Next, thresholding is performed on the grayscale cloud image. Generally, a suitable grayscale threshold is chosen, such as grouping pixels with grayscale values greater than 150 into one category and those less than 150 into another. This separates the pixels corresponding to the clouds from other background pixels, and the segmented portion becomes the cloud pixel region.
[0075] Once the cloud pixel regions are identified, the next step is to extract their texture features. A common method is to calculate the gray-level co-occurrence matrix (GLCM) between pixels and extract parameters such as contrast and correlation; these parameters constitute the cloud texture features. Then, the clouds are categorized based on these texture features. For example, high contrast and coarse texture indicate cumulus clouds, low contrast and uniform texture indicate stratus clouds, and fine texture and high gray-level values indicate cirrus clouds. The resulting classification, encompassing cumulus, stratus, and cirrus clouds, is the first classification result. This first classification result is then quantized and encoded—for example, assigning a 1 to cumulus, a 2 to stratus, and a 3 to cirrus—and combined with the previously extracted texture feature parameters to form an ordered set of values; this is the first feature vector.
[0076] After processing the satellite cloud image, let's look at the radar echo image. First, we need to perform intensity grading, which involves classifying the echoes according to their intensity values. For example, values between 0-10 dBz are classified as weak echoes, 10-30 dBz as medium echoes, and values above 30 dBz as strong echoes. Each grade is labeled with a different color, resulting in a graded echo map. Next, we perform edge detection on the graded echo map, typically using the Canny or Sobel operator, to find the boundaries between the echo region and the background region. These boundary regions are the echo edge regions. Then, we extract shape features from the echo edge regions, such as calculating the perimeter and area of the edges, or determining whether the edges are circular or irregular. These are the echo shape features.
[0077] Finally, the precipitation type is determined based on the echo shape characteristics. For example, thunderstorms are characterized by irregular edges and large areas of strong echoes, while light rain is characterized by regular edges and uniform areas of medium echoes. This determination is the second classification result. The second classification result is then quantized and encoded, for example, thunderstorms are encoded as 1 and light rain as 2. Combined with the previously obtained shape feature parameters, this is organized into another set of ordered values, which is the second feature vector.
[0078] In a specific embodiment, threshold segmentation is performed on the grayscale cloud image to obtain cloud pixel regions, including:
[0079] The grayscale cloud map is statistically analyzed to obtain a grayscale histogram, and the peak and valley distribution of the grayscale histogram is analyzed. The initial segmentation threshold is determined based on the peak and valley distribution.
[0080] The grayscale cloud image is initially segmented based on the initial segmentation threshold to obtain a preliminary segmentation region. Then, connectivity analysis is performed on the preliminary segmentation region to merge adjacent pixel regions with similar grayscale values to obtain a cloud pixel region.
[0081] Specifically, as mentioned earlier, thresholding the grayscale cloud image to obtain cloud pixel regions requires starting with grayscale value statistics. First, count the grayscale value of each pixel in the grayscale cloud image, for example, count the number of pixels corresponding to each grayscale value from 0 to 255. Then, create a graph with the grayscale value on the horizontal axis and the number of pixels on the vertical axis; this is a grayscale histogram. After obtaining the grayscale histogram, focus on its peak and trough distribution—generally, a grayscale cloud histogram will have two main peaks: one corresponding to the background (e.g., a clear sky area with lower grayscale values), and the other corresponding to the clouds (with higher grayscale values). The lowest point between the two peaks is the trough. For example, if the histogram has a peak at grayscale value 50 (background) and another at 180 (clouds), with a trough at grayscale value 120, then the grayscale value 120 corresponding to this trough can be used as the initial segmentation threshold. This method of selecting a threshold is much more accurate than arbitrarily setting a threshold.
[0082] With an initial segmentation threshold, we use it to perform preliminary segmentation of the grayscale cloud image. For example, if the initial threshold is set to 120, then all pixels with grayscale values greater than 120 in the grayscale cloud image are grouped into one category, and those less than or equal to 120 are grouped into another. These two parts are the preliminary segmented regions—the parts with grayscale values greater than 120 are most likely cloud layers, but there may be scattered small pixel blocks (such as noise points in the background) mixed in. At this point, we need to perform connectivity analysis on the preliminary segmented regions, using a connected component labeling algorithm, such as the 8-neighborhood connectivity method, to determine the relationship between those small pixel blocks and their surrounding pixels: if a small pixel block is adjacent to a large region (such as a large area of high grayscale values), and their grayscale values are not much different (for example, both are between 120 and 255), then this small pixel block is merged with the large region; if a small pixel block is surrounded by low grayscale background pixels, and its grayscale values are closer to the background, then it is classified as background. After merging in this way, the remaining complete and continuous high grayscale value area is the cloud pixel area we need. This process can avoid mistaking noise for clouds and make the segmentation results more accurate.
[0083] In a specific embodiment, determining the initial segmentation threshold based on the peak and trough distribution includes:
[0084] Peak values are extracted from the peak and trough distribution to obtain the main peak and trough positions. Based on the main peak and trough positions, the gray-level interval sequence of adjacent peaks and troughs is calculated, wherein the gray-level interval sequence includes the gray-level difference between the cloud layer and the background.
[0085] The maximum interval is detected in the gray-level interval sequence to obtain the position of the maximum gray-level span, and the initial segmentation threshold is determined based on the position of the maximum gray-level span.
[0086] Specifically, as mentioned earlier, the initial segmentation threshold needs to be determined based on the peak and trough distribution of the grayscale histogram. This requires peak extraction first. This involves identifying the main peak and trough positions from the peak and trough distribution. For example, in the histogram example, the peak position corresponding to the background is at grayscale value 50, the peak position corresponding to the cloud is at grayscale value 180, and the trough position in the middle is at grayscale value 120. These are the extracted main peak and trough positions. Next, the grayscale interval sequence between adjacent peaks and troughs is calculated based on these positions. This means calculating the grayscale value difference between each adjacent peak and trough—for example, the difference between the background peak (50) and the middle trough (120) is 70, and the difference between the middle trough (120) and the cloud peak (180) is 60. The [70, 60] range formed by these two differences is the grayscale interval sequence. Here, 70 and 60 actually contain the grayscale difference information between the cloud and the background.
[0087] With the grayscale interval sequence in hand, the next step is to perform maximum interval detection. This involves finding the interval with the largest value in the sequence. In the previous sequence, 70 is larger than 60, so the interval corresponding to 70 is the maximum grayscale span, located in the range from the background peak (50) to the middle trough (120). Then, the initial segmentation threshold is determined based on the location of this maximum grayscale span. Usually, the point in the middle of this span that best distinguishes the two types of regions is selected. For example, the trough position corresponding to the maximum span, 120, can be used directly, or an intermediate value can be selected within this span. However, in practice, selecting the trough position is more common, so 120 is used as the initial segmentation threshold here. This threshold can separate the grayscale regions of the clouds and the background to the greatest extent.
[0088] In a specific embodiment, a fused feature matrix is constructed based on the first feature vector and the second feature vector, including:
[0089] The cloud impact quantification is performed on the first feature vector to obtain the first weight, and the precipitation impact quantification is performed on the second feature vector to obtain the second weight;
[0090] The first feature vector is weighted based on the first weight to obtain a weighted first vector; the second feature vector is weighted based on the second weight to obtain a weighted second vector; the weighted first vector and the weighted second vector are then concatenated column by column to obtain a fused feature matrix.
[0091] Specifically, to construct a fused feature matrix based on the previously obtained first and second eigenvectors, weights must first be calculated for each vector. First, the first eigenvector, derived from satellite cloud images, needs to be quantified for cloud impact. For example, considering the cloud type in the first eigenvector, if cumulus is predominant (previously encoded as 1), its impact on photovoltaic power generation is moderate, so a first weight of 0.4 is given; if stratus is predominant (encoded as 2), with strong shading and a significant impact, a first weight of 0.6 is given. For example, if a set of first eigenvectors corresponds mainly to stratus, the calculated first weight would be 0.6. Next, the second eigenvector, derived from radar echo images, needs to be quantified for precipitation impact. For example, if the second eigenvector shows thunderstorms (encoded as 1), its impact on power generation is significant, so a second weight of 0.7 is given; if it shows light rain (encoded as 2), its impact is minor, so a second weight of 0.3 is given. Assuming this set of second eigenvectors corresponds to light rain, the second weight would be 0.3. After calculating the weights, a weighted calculation is performed. Taking the first feature vector as an example, suppose it is [2, 5, 8] (2 represents stratus, 5 is texture contrast, and 8 is the coverage area parameter). Multiplying it by the first weight of 0.6, each value is calculated as follows: 2 × 0.6 = 1.2, 5 × 0.6 = 3, 8 × 0.6 = 4.8, resulting in [1.2, 3, 4.8], which is the weighted first vector. Now consider the second feature vector, for example, [2, 4, 6] (2 represents light rain, 4 is edge perimeter, and 6 is shape complexity). Multiplying it by the second weight of 0.3, we get 2 × 0.3 = 0.6, 4 × 0.3 = 1.2, 6 × 0.3 = 1.8, resulting in [0.6, 1.2, 1.8], which is the weighted second vector. Finally, these two weighted vectors are concatenated column-wise, with the weighted first vector as the first column and the weighted second vector as the second column, forming... This refers to the desired fusion feature matrix.
[0092] In a specific embodiment, the fused feature matrix is input into a preset power prediction model to calculate the power generation output value of the photovoltaic power generation, including:
[0093] The fused feature matrix is input into a preset power prediction model; wherein the power prediction model includes a feature encoding layer, a temporal correlation layer, and a power mapping layer;
[0094] The fused feature matrix is input into the feature coding layer for layer-by-layer nonlinear transformation to obtain the coding feature vector. Based on the coding feature vector, the attenuation of the theoretical irradiance of the photovoltaic module is calculated to obtain the irradiance attenuation coefficient.
[0095] The encoded feature vector and the irradiation attenuation coefficient are input into the temporal association layer to perform state transfer calculations between time steps to obtain the temporal association vector.
[0096] The time-series correlation vector is input into the power mapping layer for weighted summation to obtain the initial power value; and the initial power value is subjected to upper and lower limit constraints based on the installed capacity of the photovoltaic power station to obtain the photovoltaic power output value.
[0097] Specifically, the next step after constructing the fused feature matrix is to input it into the pre-defined power prediction model. This model is not a single structure; it contains a feature encoding layer, a temporal correlation layer, and a power mapping layer, each processing its own data. First, the fused feature matrix is fed into the feature encoding layer, where a layer-by-layer nonlinear transformation is performed. For example, using the previously obtained fused feature matrix... In the first layer, the ReLU activation function is used, and each value in the matrix is substituted into... The calculation involves substituting 1.2 into the matrix to get 1.2, 0.6 to get 0.6, 3 to get 3, 1.2 to get 1.2, 4.8 to get 4.8, and 1.8 to get 1.8. After two such transformations, the resulting [1.5, 2.1, 3.3] is the encoded feature vector. This encoded feature vector is then used to calculate the theoretical irradiance attenuation coefficient of the photovoltaic module. Assuming the theoretical irradiance is 1000W / ㎡, a larger value in the encoded feature vector indicates a greater impact from cloud cover or precipitation. For example, using the formula "attenuation coefficient = 1 - (mean of encoded feature vector / 10)", the mean is (1.5 + 2.1 + 3.3) / 3 = 2.3, so the attenuation coefficient is 1 - 2.3 / 10 = 0.77. After calculating the attenuation coefficient, the encoded feature vector and the irradiance attenuation coefficient are input into the time-series correlation layer, where state transfer calculations between time steps are performed. For example, the current encoded feature vector is [1.5, 2.1, 3.3] with an attenuation coefficient of 0.77, and the previous state vector is [1.3, 1.9, 3.0]. Using the gating mechanism in the LSTM network, the current data and the previous state are combined, and the calculated [1.45, 2.02, 3.18] is the temporal correlation vector. Finally, the temporal correlation vector is put into the power mapping layer, where a weighted summation is performed. Assuming that the weights of each vector element in the power mapping layer are 0.2, 0.3, and 0.5, we use 1.45×0.2 + 2.02×0.3 + 3.18×0.5 to calculate: 1.45×0.2 = 0.29, 2.02×0.3 = 0.606, and 3.18×0.5 = 1.59. Adding these together, 0.29 + 0.606 + 1.59 = 2.486 kW, which is the initial power value. However, this value cannot be used directly. It must be constrained by upper and lower limits based on the installed capacity of the photovoltaic power station. For example, if the installed capacity of the power station is 5kW, the initial power value cannot exceed 5kW or be lower than 0. 2.486kW is within this range, so the final value of 2.486kW is the photovoltaic power output value.
[0098] Additionally, it should be noted that during the training data preparation phase, historical operating data from the target photovoltaic power station for more than 12 consecutive months was collected as training samples. Each sample set includes two parts: input data and label data. The input data includes the fused feature matrix of satellite cloud image and radar echo image corresponding to that moment, and the label data is the actual power generation record value of the photovoltaic power station at that moment. The collected raw data underwent preprocessing, including removing outliers caused by sensor malfunctions, filling in missing data using linear interpolation, and normalizing various feature data to unify their numerical range to between 0 and 1. The preprocessed data was then randomly divided into training, validation, and test sets in an 8:1:1 ratio.
[0099] In the model structure setting phase, the feature encoding layer adopts a three-layer fully connected neural network structure, with 128 neurons in the first layer, 64 in the second layer, and 32 in the third layer. ReLU is used as the activation function between each layer to achieve nonlinear transformation. The temporal correlation layer uses a single-layer LSTM network structure with a hidden layer dimension of 64 to capture the time-series dependence of photovoltaic power generation. The power mapping layer uses a single-layer fully connected network with the same input dimension as the output dimension of the temporal correlation layer, and an output dimension of 1, corresponding to a single predicted power value.
[0100] During model training, mean squared error (MSE) was used as the loss function to calculate the deviation between the predicted and actual power values. The Adam optimizer was used for parameter updates, with an initial learning rate of 0.001, which decreased to 0.5 times its original value after every 50 training epochs. The batch size was set to 32, and the maximum number of training epochs was set to 500. An early stopping strategy was employed during training: training was stopped when the validation set loss no longer decreased for 20 consecutive epochs, and the model parameters corresponding to the minimum validation set loss were saved as the final model.
[0101] During the model validation phase, the trained model is tested on a test set to verify its prediction accuracy. The root mean square error (RMSE) and mean absolute percentage error (MASE) between the predicted and actual values are calculated. When the RMSE is less than 5% of the installed capacity and the MASE is less than 10%, the model is considered to have completed training and is ready for use.
[0102] Through the above construction and training process, those skilled in the art can reproduce the complete training process of the power prediction model according to the data preparation method, model structure parameters, training configuration and verification standards described in the specification.
[0103] Furthermore, the feature encoding layer processes the fused feature matrix through two steps: nonlinear transformation and dimensionality compression. Taking the aforementioned fused feature matrix as an example, this matrix is in 3x2 matrix form:
[0104] The feature encoding layer processes the fused feature matrix through two steps: nonlinear transformation and dimensionality compression. Taking the fused feature matrix in the specification as an example, this matrix... It is in the form of a 3x2 matrix:
[0105] The first column [1.2, 3, 4.8] corresponds to satellite cloud image features, and the second column [0.6, 1.2, 1.8] corresponds to radar echo image features. Before inputting the feature encoding layer, this 3x2 matrix is first expanded row-wise into a one-dimensional vector [1.2, 0.6, 3, 1.2, 4.8, 1.8] with a dimension of 6.
[0106] The first fully connected layer performs a nonlinear transformation on the unfolded one-dimensional vector. This layer has 6 neurons, the weight matrix is initialized to the identity matrix, the bias term is initialized to 0, and the activation function is ReLU. Since all elements in the input vector are positive, the values remain unchanged after ReLU activation, and the output of the first layer is still [1.2, 0.6, 3, 1.2, 4.8, 1.8].
[0107] The second fully connected layer performs dimensionality compression on the output of the first layer. This layer has 3 neurons, compressing the 6-dimensional vector into a 3-dimensional vector. This layer uses a sparsely connected weight matrix structure, where each output neuron is connected to only two specific input elements. The specific connection method and weight parameters are as follows:
[0108] The first output neuron connects the first input element (1, 2) and the fourth input element (1, 2), with a weight of 0.5 for each and a bias term of 0.3.
[0109] Calculation process: 1.2 × 0.5 + 1.2 × 0.5 + 0.3 = 0.6 + 0.6 + 0.3 = 1.5
[0110] The second output neuron connects the second input element (0.6) and the fifth input element (4.8), with a weight of 0.5 for each and a bias term of -0.6.
[0111] Calculation process: 0.6 × 0.5 + 4.8 × 0.5 + (-0.6) = 0.3 + 2.4 - 0.6 = 2.1
[0112] The third output neuron connects the third input element (3) and the sixth input element (1.8), with a weight of 0.5 for both and a bias term of 0.9.
[0113] Calculation process: 3 × 0.5 + 1.8 × 0.5 + 0.9 = 1.5 + 0.9 + 0.9 = 3.3
[0114] After the above two transformations, the encoded feature vector [1.5, 2.1, 3.3] is obtained.
[0115] The specific values of the aforementioned weight parameters and bias terms are the optimal parameters obtained after model training, and are not pre-set manually. During model training, historical photovoltaic power generation data is used as training samples, and the weights and biases of each layer are continuously adjusted through the backpropagation algorithm to minimize the mean square error between the predicted power value and the actual power value, ultimately converging to obtain the aforementioned parameter values.
[0116] The above describes a data fusion-based photovoltaic power generation prediction method according to an embodiment of the present invention. The following describes a data fusion-based photovoltaic power generation prediction system according to an embodiment of the present invention. Please refer to [link / reference]. Figure 2 One embodiment of the photovoltaic power generation prediction system based on data fusion according to the present invention includes:
[0117] The acquisition module 21 is used to acquire satellite cloud images and radar echo images of the photovoltaic area, and to perform pixel segmentation and feature extraction on the satellite cloud images and radar echo images to obtain the first feature vector and the second feature vector.
[0118] Fusion module 22 is used to construct a fused feature matrix based on the first feature vector and the second feature vector;
[0119] The calculation module 23 is used to input the fused feature matrix into a preset power prediction model to calculate the power generation and obtain the photovoltaic power output value.
[0120] In this embodiment, the specific implementation of each module in the above system embodiment is described in the above method embodiment, and will not be repeated here.
Claims
1. A method for predicting photovoltaic power generation through data fusion, characterized in that, Includes the following steps: Satellite cloud images and radar echo images of the photovoltaic area are collected, and pixel segmentation and feature extraction are performed on the satellite cloud images and radar echo images to obtain the first feature vector and the second feature vector. Construct a fused feature matrix based on the first and second eigenvectors; The fused feature matrix is input into a preset power prediction model to calculate the power generation output value of the photovoltaic power generation.
2. The photovoltaic power generation prediction method based on data fusion according to claim 1, characterized in that, Collect satellite cloud images and radar echo images of the photovoltaic area, including: The geographical coordinate range of the photovoltaic area is determined, and the data collection boundary range is set based on the geographical coordinate range; Meteorological data is collected from the photovoltaic area based on the aforementioned collection boundary range to obtain satellite cloud images, and radar scans are performed on the photovoltaic area based on the aforementioned collection boundary range to obtain radar echo images.
3. The photovoltaic power generation prediction method based on data fusion according to claim 2, characterized in that, Meteorological data is collected from the photovoltaic area based on the aforementioned collection boundary range to obtain satellite cloud images, including: The geographic coordinate range is converted into satellite image acquisition coordinates to obtain the converted coordinate range. Based on the converted coordinate range, a data acquisition command is sent to the meteorological satellite to collect meteorological data of the photovoltaic area and obtain an initial satellite cloud image. The initial satellite cloud image is subjected to image correction processing to eliminate image distortion, resulting in a corrected cloud image. The corrected cloud image is then filtered at the pixel level to extract satellite cloud images corresponding to the acquisition boundary range.
4. The photovoltaic power generation prediction method based on data fusion according to claim 2, characterized in that, Based on the aforementioned acquisition boundary range, a radar scan of the photovoltaic area is performed to obtain a radar echo image, including: The radar scanning parameters are set for the acquisition boundary range to determine the frequency and angle range of the radar scan, and the set scanning parameters are obtained. Based on the set scanning parameters, the radar equipment is controlled to scan the photovoltaic area to obtain initial radar echo data containing information on precipitation particle reflection intensity and distance. The initial radar echo data is parsed and converted into a visual image format to obtain a preliminary echo image. The preliminary echo image is then subjected to noise filtering to eliminate interference signals, thus obtaining the final radar echo image.
5. The photovoltaic power generation prediction method based on data fusion according to claim 1, characterized in that, Pixel segmentation and feature extraction are performed on satellite cloud images and radar echo images to obtain a first feature vector and a second feature vector, including: The satellite cloud image is converted to grayscale to obtain a grayscale cloud image, and the grayscale cloud image is then segmented using a threshold to obtain cloud pixel regions. Texture features are extracted from the cloud pixel region to obtain cloud texture features. Based on the cloud texture features, cloud types are classified to obtain a first classification result including cumulus, stratus and cirrus. The first classification result is then quantized and encoded to obtain a first feature vector. The radar echo image is subjected to intensity grading processing to obtain a graded echo map, and edge detection is performed on the graded echo map to identify echo edge regions. Shape features are extracted from the echo edge regions to obtain echo shape features. Precipitation type is determined based on the echo shape characteristics to obtain a second classification result. The second classification result is then quantized and encoded to obtain a second feature vector.
6. The photovoltaic power generation prediction method based on data fusion according to claim 5, characterized in that, Threshold segmentation is performed on the grayscale cloud image to obtain cloud pixel regions, including: The grayscale cloud map is statistically analyzed to obtain a grayscale histogram, and the peak and valley distribution of the grayscale histogram is analyzed. The initial segmentation threshold is determined based on the peak and valley distribution. The grayscale cloud image is initially segmented based on the initial segmentation threshold to obtain a preliminary segmentation region. Then, connectivity analysis is performed on the preliminary segmentation region to merge adjacent pixel regions with similar grayscale values to obtain a cloud pixel region.
7. The photovoltaic power generation prediction method based on data fusion according to claim 6, characterized in that, Determining the initial segmentation threshold based on the peak and trough distribution includes: Peak values are extracted from the peak and trough distribution to obtain the main peak and trough positions. Based on the main peak and trough positions, the gray-level interval sequence of adjacent peaks and troughs is calculated, wherein the gray-level interval sequence includes the gray-level difference between the cloud layer and the background. The maximum interval is detected in the gray-level interval sequence to obtain the position of the maximum gray-level span, and the initial segmentation threshold is determined based on the position of the maximum gray-level span.
8. The photovoltaic power generation prediction method based on data fusion according to claim 1, characterized in that, A fused feature matrix is constructed based on the first and second eigenvectors, including: The cloud impact quantification is performed on the first feature vector to obtain the first weight, and the precipitation impact quantification is performed on the second feature vector to obtain the second weight; The first feature vector is weighted based on the first weight to obtain a weighted first vector; the second feature vector is weighted based on the second weight to obtain a weighted second vector; the weighted first vector and the weighted second vector are then concatenated column by column to obtain a fused feature matrix.
9. The photovoltaic power generation prediction method based on data fusion according to claim 1, characterized in that, The fused feature matrix is input into a preset power prediction model to calculate the power generation output value of the photovoltaic power generation, including: The fused feature matrix is input into a preset power prediction model; wherein the power prediction model includes a feature encoding layer, a temporal correlation layer, and a power mapping layer; The fused feature matrix is input into the feature coding layer for layer-by-layer nonlinear transformation to obtain the coding feature vector. Based on the coding feature vector, the attenuation of the theoretical irradiance of the photovoltaic module is calculated to obtain the irradiance attenuation coefficient. The encoded feature vector and the irradiation attenuation coefficient are input into the temporal association layer to perform state transfer calculations between time steps to obtain the temporal association vector. The time-series correlation vector is input into the power mapping layer for weighted summation to obtain the initial power value; and the initial power value is subjected to upper and lower limit constraints based on the installed capacity of the photovoltaic power station to obtain the photovoltaic power output value.
10. A data fusion photovoltaic power generation prediction system, characterized in that, A photovoltaic power generation prediction method for performing data fusion as described in any one of claims 1 to 9, comprising: The acquisition module is used to acquire satellite cloud images and radar echo images of the photovoltaic area, and to perform pixel segmentation and feature extraction on the satellite cloud images and radar echo images to obtain the first feature vector and the second feature vector. The fusion module is used to construct a fused feature matrix based on the first feature vector and the second feature vector; The calculation module is used to input the fused feature matrix into a preset power prediction model to calculate the power generation and obtain the photovoltaic power output value.