Joint prediction method for distributed new energy power generation device

By clustering analysis of the spatial coordinates and timing data of new energy power generation devices, key meteorological factors are screened, and time-series data sequences are generated, the problem of low prediction accuracy caused by the sensitivity of new energy power generation devices to natural resources is solved, and more efficient joint power prediction is achieved.

CN120237648AActive Publication Date: 2025-07-01STATE GRID ZHEJIANG ELECTRIC POWER CO LTD NINGBO POWER SUPPLY CO

Patent Information

Application Number
CN202510726967.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-01
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The prior art fails to effectively consider the sensitivity of new energy power generation devices to different natural resources, resulting in low joint power prediction accuracy.

Method used

Cluster analysis is carried out based on the spatial coordinates and time-sequential power generation data of new energy power generation devices, regional data clusters are determined, correlation coefficients and mutual information are calculated to screen key meteorological factors, and time-sequential data sequences are generated to predict future power.

Benefits of technology

The combined power prediction accuracy of distributed new energy power generation devices has been significantly improved, and weather factors that have a weak or no impact on new energy power generation devices have been eliminated, which has improved prediction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120237648A_ABST
    Figure CN120237648A_ABST
Patent Text Reader

Abstract

The invention discloses a joint prediction method for a distributed new energy power generation device, and belongs to the technical field of new energy power generation power prediction, and the method comprises the steps: collecting the space coordinates, time sequence power generation data, output data, meteorological characteristic data, output power, and meteorological data of the new energy power generation device; performing clustering analysis based on the space coordinates and the time sequence power generation data to determine a regional data cluster; calculating a correlation coefficient and mutual information according to the regional data cluster, the output data and the meteorological characteristic data; screening meteorological factors based on correlation coefficients and mutual information to obtain key meteorological factors; generating a time sequence data sequence corresponding to the new energy power generation device based on the key meteorological factors, the output power and the meteorological data; predicting the future power of the corresponding new energy power generation device based on the time sequence data sequence; the method overcomes the problem of low joint power prediction accuracy caused by the fact that the sensitivity of a new energy power generation device to natural resources is not considered in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of new energy power generation prediction, and specifically to a combined prediction method for distributed new energy power generation devices. Background Art

[0002] With the development of new energy power generation technology, the proportion of green power sources represented by new energy in the power grid is increasing. Although new energy power generation devices are widely distributed in different regions, the installation locations of new energy power generation devices have similar or even the same natural resources, such as wind resources, solar energy resources, etc. These natural resources have certain periodic changes over time. Considering the characteristic that new energy power generation devices are greatly affected by natural resources, when natural resources change, the power of the corresponding new energy power generation devices also changes accordingly, making it difficult to predict the power generation of new energy power generation devices. Despite the specific spatio-temporal complementary characteristics of the power of different types and regions of new energy power generation devices, due to the instability of the power generation, it is difficult to truly complement each other.

[0003] Chinese Patent, Publication No.: CN119226702A, Publication Date: December 31, 2024, discloses a high-precision combined wind-solar power prediction method, including: A. First, collect wind power generation data and photovoltaic power generation data respectively; B. Preprocess the collected wind power generation data and photovoltaic power generation data; C. Extract features from the preprocessed power generation data; D. Then, according to the spatio-temporal correlation of wind-solar power generation and the load curve, introduce a temporal attention mechanism, and construct a combined wind-solar power generation prediction model based on the improved GRU-CNN algorithm by leveraging the temporal feature mining ability of the GRU network and the spatial feature mining ability of the CNN network; E. Conduct model training, verification, and testing to achieve combined prediction of wind-solar power generation; F. Finally, based on high-precision meteorological data, evaluate the prediction accuracy performance of different methods through testing and form an optimization strategy to obtain the best prediction method; however, this invention does not consider the different sensitivities of different types of new energy power generation devices to different natural resources, resulting in low accuracy in the prediction of combined power. Summary of the Invention

[0004] The object of the present invention is to address the problem that the prior art does not consider the sensitivity of new energy power generation devices to natural resources, resulting in low accuracy in the prediction of combined power. A combined prediction method for distributed new energy power generation devices is proposed. Cluster analysis is performed based on the spatial coordinates and time-series power generation data of new energy power generation devices to divide regions and determine regional data clusters. Key meteorological factors are obtained by screening meteorological factors that have a great impact on new energy power generation devices according to the regional data clusters, output data, and meteorological characteristic data. A time-series data sequence is generated based on the key meteorological factors, output power, and meteorological data to predict the future power of new energy power generation devices, significantly improving the accuracy of combined power prediction for distributed new energy power generation devices.

[0005] In a first aspect, a technical solution provided in an embodiment of the present invention is a combined prediction method for a distributed new energy power generation device, including the following steps: Collect the spatial coordinates, time-series power generation data, output data, meteorological characteristic data, output power, and meteorological data of the new energy power generation device; Perform clustering analysis based on the spatial coordinates and time-series power generation data to determine regional data clusters; Calculate the correlation coefficient and mutual information according to the regional data clusters, output data, and meteorological characteristic data; Screen meteorological factors based on the correlation coefficient and mutual information to obtain key meteorological factors; Generate a time-series data sequence corresponding to the new energy power generation device based on the key meteorological factors, output power, and meteorological data; Predict the future power of the corresponding new energy power generation device based on the time-series data sequence.

[0006] In this solution, clustering analysis is performed based on the spatial coordinates and time-series power generation data of the new energy power generation device to determine regional data clusters, and new energy power generation devices with similar geographical distributions and similar or even the same power generation modes are classified into the same category, thus providing strong support for subsequent unified prediction; secondly, weather is a factor affecting the power generation power of new energy power generation devices, and the change of the weather itself leads to complex changes in the power generation power, making it difficult to predict. Calculate the correlation coefficient and mutual information according to the regional data clusters, output data, and meteorological characteristic data, and depict the highly non-linear relationship and complex mapping relationship between weather factors and the power of new energy power generation devices through the correlation coefficient and mutual information; in addition, different new energy power generation devices have different sensitivities to different weather factors, that is, when different weather factors act on the same new energy power generation device, the impact on the power generation power of the new energy power generation device is different. Screen meteorological factors based on the correlation coefficient and mutual information to obtain key meteorological factors, eliminate weather factors that have little or no impact on the corresponding new energy power generation device, and thus propose a large amount of useless data, effectively improving the efficiency of power prediction; however, at this time, only the category of the new energy power generation device and the key weather factors affecting the new energy power generation device are obtained. Although the highly non-linear relationship and complex mapping relationship between weather factors and the power of new energy power generation devices are depicted, the change of the power generation power of the new energy power generation device over time series with weather factors is not clear. Then, generate a time-series data sequence corresponding to the new energy power generation device based on the key meteorological factors, output power, and meteorological data, obtain the change of the power generation power of the new energy power generation device over time series with weather factors, and predict the future power of the corresponding new energy power generation device through the change, significantly improving the combined power prediction accuracy of distributed new energy power generation devices.

[0007] Preferably, the specific process of clustering analysis based on spatial coordinates and time-series power generation data to determine regional data clusters is as follows: A1. Combine the spatial coordinates and time-series power generation data to obtain a cluster data set, and determine the distance threshold, minimum number of points, and time-series data weight based on the data characteristics of the cluster data set; A2. Perform weighted calculation based on the time-series data weight and the cluster data set to obtain the cluster Euler distance, and construct a neighborhood based on the cluster Euler distance and the distance threshold; A3. Determine the core points based on the neighborhood and the minimum number of points. If the number of data points in the neighborhood is less than the minimum number of points, mark the cluster data corresponding to the neighborhood as ordinary points. If the number of data points in the neighborhood is greater than or equal to the minimum number of points, mark the cluster data corresponding to the neighborhood as core points; A4. Perform cluster division based on the core points and the neighborhoods corresponding to the core points to obtain regional data clusters.

[0008] In this solution, due to the different spatial positions and possible different power generation modes of new energy power generation devices, the spatial coordinates and time-series power generation data are combined to obtain a cluster data set, combining two different types of data into a new type of data, which is convenient for overall clustering analysis. Secondly, in order to optimize the clustering analysis effect, the distance threshold, minimum number of points, and time-series data weight are determined based on the data characteristics of the cluster data set; the Euler formula is improved to a weighted Euler formula using the time-series data weight, and the cluster data set is input into the weighted Euler formula for weighted calculation to obtain the Euler distance between data, that is, the cluster Euler distance. Then, based on a certain data, all data with an Euler distance less than or equal to the distance threshold from the data is marked as points within the neighborhood of the certain data, thus establishing a neighborhood; based on the neighborhood and the minimum number of points, core points and ordinary points are judged, and then cluster division can be performed based on the core points and the neighborhoods corresponding to the core points to obtain regional data clusters. If data is a core point, and data is within the neighborhood of data , then it is said that data is density-reachable from data . For the core point data , all data that is density-reachable from data is grouped into the same cluster to obtain regional data clusters. At the same time, data that does not belong to any cluster will also be obtained, and these data can be marked as noise points.

[0009] Preferably, in A1, the specific process of combining the spatial coordinates and time-series power generation data to obtain a cluster data set is as follows: Statistically eliminate outliers in the time-series power generation data based on the standard score method, and fill in the missing values in the time-series power generation data based on the linear interpolation method to obtain complete time-series power generation data; Normalize the spatial coordinates to obtain normalized coordinates, and normalize the complete time-series power generation data to obtain normalized power generation data; Sort out the normalized coordinates and normalized power generation data and merge them to obtain a cluster dataset.

[0010] In this solution, the spatial coordinates are specifically a set of coordinate points, and the time-series power generation data is specifically a sequence with a length of time ; Adopt the Z-score method, that is, the standard score method, to filter out the outliers in the time-series power generation data, so that only missing values exist in the time-series power generation data, and adopt the linear interpolation method to fill in the missing values according to the linear relationship of the time-series power generation data to obtain the complete time-series power generation data; Secondly, in order to eliminate the geographical differences of new energy power generation devices in different regions, Use the normalization formula to normalize the spatial coordinates to obtain normalized coordinates. In order to eliminate the power magnitude differences of new energy power generation devices in different regions, Use the normalization formula to normalize the complete time-series power generation data to obtain normalized power generation data. It should be noted that when normalizing the spatial coordinates, the abscissa and ordinate of the spatial coordinates need to be normalized separately. Finally, based on the new energy power generation device, the corresponding normalized coordinates and the corresponding normalized power generation data are merged into the same sequence, that is, the abscissa of the normalized coordinates is used as the first item of the sequence, the ordinate of the normalized coordinates is used as the second item of the sequence, and the abscissa of the normalized coordinates and the ordinate of the normalized coordinates are inserted into the sequence of the normalized power generation data.

[0011] Preferably, after A4 is completed, it is also necessary to calculate the silhouette coefficient to verify the regional data clusters. The corresponding silhouette coefficient formula is specifically: ; In the formula, is the silhouette coefficient of the th data point, is the average distance from the th data point to other data points in the same regional data cluster, is the average distance from the th data point to other regional data clusters.

[0012] In this solution, the closer the silhouette coefficient is to 1, the better the clustering effect of the corresponding th data point is proved. While ensuring the clustering effect through the silhouette coefficient , further eliminate the data with poor clustering effect, that is, effectively filter out the noise data.

[0013] Preferably, the specific process of calculating the correlation coefficient and mutual information based on the regional data cluster, output data, and meteorological characteristic data is as follows: B1. Screen the output data based on the regional data cluster to obtain regional output data, and screen the meteorological characteristic data based on the regional data cluster to obtain regional meteorological characteristic data; B2. Standardize the regional meteorological characteristic data to obtain standardized meteorological characteristic data, and organize the standardized meteorological characteristic data and regional output data to obtain regional samples; B3. Calculate the correlation coefficient and mutual information according to the regional samples and the cross-correlation formula.

[0014] In this solution, the output data and meteorological characteristic data of the new energy power generation devices in the corresponding region are obtained based on the regional data cluster. It should be noted that the purpose of the regional data cluster is to divide the new energy power generation devices into regions, and then obtain the corresponding output data and meteorological characteristic data of the new energy power generation devices. Therefore, it can be achieved by screening the already collected corresponding data or collecting the corresponding data of specific new energy power generation devices, etc.; in order to eliminate the scale difference of the meteorological characteristic data, the standardized formula is used to standardize the regional meteorological characteristic data to obtain standardized meteorological characteristic data, and the standardized meteorological characteristic data and regional output data are organized to obtain regional samples. Specifically, both the standardized meteorological characteristic data and the regional output data are in the form of sequences. The rank of each data in the sequence can be counted, as well as the joint probability distribution and marginal probability distribution of the standardized meteorological characteristic data and the regional output data. The rank, joint probability distribution, and marginal probability distribution are organized to obtain regional samples, and the regional samples are input into the cross-correlation formula to calculate the correlation coefficient and mutual information.

[0015] Preferably, in B3, the cross-correlation formula is specifically: ; ; In the formula, is the correlation coefficient between the th meteorological factor corresponding to the regional sample and the new energy output , is the number of regional samples, is the rank of the th regional sample corresponding to the th meteorological factor , is the rank of the th regional sample corresponding to the new energy output , is the mutual information between the th meteorological factor corresponding to the regional sample and the new energy output . is the th meteorological factor and the joint probability distribution of new energy output . is the th meteorological factor 's marginal probability distribution, is the marginal probability distribution of new energy output .

[0016] Preferably, the screening model for obtaining key meteorological factors by screening meteorological factors based on correlation coefficients and mutual information is specifically: ; In the formula, is the key index corresponding to the th meteorological factor , is the weight coefficient, is the th meteorological factor and the correlation coefficient between new energy output , is the th meteorological factor and the mutual information between new energy output .

[0017] Preferably, the specific process of generating a time series data sequence of the corresponding new energy power generation device based on key meteorological factors, output power, and meteorological data is as follows: C1. Screen meteorological data based on key meteorological factors to obtain key meteorological data, and screen output power based on the regional data cluster corresponding to the key meteorological factors to obtain key output power; C2. Integrate the time series of key meteorological data and the time series of key output power to obtain input data; Synchronously, perform similarity analysis on key meteorological data to obtain meteorological similarity features, and perform similarity analysis on the geographical location of the corresponding new energy power generation device to obtain geographical similarity features; C3. Generate a heterogeneous graph based on input data, meteorological similarity features, and geographical similarity features; C4. Extract spatial embedding features in the heterogeneous graph based on the time series of input data to obtain feature vectors, and organize and splice all the feature vectors to obtain a time series data sequence.

[0018] In this solution, meteorological data is screened based on key meteorological factors, and the output power is screened based on the regional data clusters corresponding to the key meteorological factors. The screening process is to delete unnecessary data and optimize the volume and accuracy of the data to be processed. It should be noted that the purpose is to delete unnecessary data, which can be achieved by deleting the corresponding data that has been collected, or by deleting the corresponding data collection devices. To solve the problems of dimensional differences and missing values in the key meteorological data and key output power, methods such as normalization formulas and missing value filling are used for integration to obtain input data. At the same time, to capture the key influencing factors and regional sharing factors, similarity analysis is performed on the key meteorological data and geographical locations respectively, and the in-depth relationships between meteorology, climate, geographical location, etc. and the new energy power generation device are analyzed. Then, based on the input data, meteorological similarity features, and geographical similarity features, a heterogeneous graph that meets the requirements of the heterogeneous graph neural network is generated. The heterogeneous graph is input into the heterogeneous graph neural network. Specifically, based on the time series of the input data, the corresponding node feature matrix in the heterogeneous graph is input into the heterogeneous graph neural network, and the spatial embedding features are extracted through multiple graph convolutions of the heterogeneous graph neural network to obtain feature vectors. The feature vectors are output in the form of a matrix, and the feature vectors at each time step are concatenated to reconstruct a time series data sequence.

[0019] Preferably, in C2, the specific process of integrating the input data based on the time series of the key meteorological data and the time series of the key output power is as follows: Count the missing values of the key meteorological data and fill them to obtain complete key meteorological data, and normalize the complete key meteorological data to obtain normalized meteorological data; Count the missing values of the key output power and fill them to obtain complete key output power, and normalize the complete key output power to obtain normalized output power; Align the time series of the normalized meteorological data with the time series of the normalized output power to obtain a unified time series; Generate input data based on the unified time series, normalized meteorological data, and normalized output power.

[0020] In this solution, historical data of the same period is selected based on the periodicity of the meteorological data to fill the missing values of the meteorological data, and it is normalized using the normalization formula. The linear interpolation method is used to fill the missing values of the output power according to the linear relationship of the output power, and it is normalized using the normalization formula. Align the time series of the normalized meteorological data with the time series of the normalized output power to obtain a unified time series. Segment the same time series based on the sampling time interval, and organize the normalized meteorological data and normalized output power based on the segmented same time series to obtain input data.

[0021] Preferably, the specific process of predicting the future power of the corresponding new energy power generation device based on the time series data sequence is as follows: Gradually update the hidden state based on the time series data sequence to obtain a hidden state sequence, and use the hidden state sequence as the initial state to generate a prediction sequence; Generate a power output matrix based on the prediction sequence to obtain the future power of the corresponding new energy power generation device.

[0022] In this solution, the time series data sequence is input into the time series prediction model. The encoder of the time series prediction model encodes the time series data sequence to gradually update the hidden state, obtaining a hidden state sequence, and uses the hidden state sequence as the initial state of the decoder of the time series prediction model to gradually generate a prediction sequence within the future time period. The power output matrix formed by the prediction sequence is essentially the future power of the corresponding new energy power generation device.

[0023] Advantages of the present invention: (1) This application performs clustering analysis based on the spatial coordinates and time series power generation data of new energy power generation devices to determine regional data clusters, divides new energy power generation devices with similar geographical distributions and similar or even the same power generation modes into a unified category, realizes the precise division of new energy power generation devices under complex non-linear spatial relationships, and uses the silhouette coefficient to optimize the clustering quality, effectively filtering out noise data; (2) This application calculates the correlation coefficient and mutual information based on the regional data cluster, output data, and meteorological characteristic data, depicts the highly non-linear relationship and complex mapping relationship between weather factors and the power of new energy power generation devices through the correlation coefficient and mutual information, and screens meteorological factors based on the correlation coefficient and mutual information to obtain key meteorological factors, eliminating weather factors that have little or no impact on the corresponding new energy power generation devices, effectively improving the efficiency of power prediction; (3) This application generates a time series data sequence of the corresponding new energy power generation device based on the key meteorological factors, output power, and meteorological data, obtains the change in the power generation power of the new energy power generation device over time series with weather factors, clarifies the change in the power generation power of the new energy power generation device over time series with weather factors, and predicts the future power of the corresponding new energy power generation device through the change, significantly improving the joint power prediction accuracy of distributed new energy power generation devices. Description of the Drawings

[0024] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious. The drawings are only for the purpose of showing the preferred embodiments and are not considered as limiting the present invention. Moreover, throughout the drawings, the same reference symbols are used to represent the same components.

[0025] Figure 1Schematic flowchart of a combined prediction method for a distributed new energy power generation device of the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only the best embodiments of the present invention, which are only used to explain the present invention and do not limit the protection scope of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0027] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the operations (or steps) as sequential processes, many of the operations (or steps) can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but there can also be additional steps not included in the drawings; the process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0028] Embodiment 1: As Figure 1 shown, this embodiment provides a combined prediction method for a distributed new energy power generation device, including the following steps: Collect the spatial coordinates, time-series power generation data, output data, meteorological characteristic data, output power, and meteorological data of the new energy power generation device; Perform clustering analysis based on the spatial coordinates and time-series power generation data to determine regional data clusters; A1. Merge the spatial coordinates and time-series power generation data to obtain a cluster data set, and determine the distance threshold, minimum number of points, and time-series data weight based on the data characteristics of the cluster data set; Specifically, statistically eliminate the outliers in the time-series power generation data based on the standard score method, and fill the missing values in the time-series power generation data based on the linear interpolation method to obtain complete time-series power generation data; Normalize the spatial coordinates to obtain normalized coordinates, and normalize the complete time-series power generation data to obtain normalized power generation data; Sort out the normalized coordinates and normalized power generation data and merge them to obtain a cluster data set; A2. Perform weighted calculation based on the time-series data weight and the cluster data set to obtain the cluster Euler distance, and construct a neighborhood based on the cluster Euler distance and the distance threshold; A3. Determine the core points based on the domain and the minimum number of points. If the number of data points in the domain is less than the minimum number of points, mark the cluster data corresponding to the domain as ordinary points. If the number of data points in the domain is greater than or equal to the minimum number of points, mark the cluster data corresponding to the domain as core points; A4. Perform cluster division based on the core points and the domains corresponding to the core points to obtain regional data clusters.

[0029] In this embodiment, due to the different spatial positions of new energy power generation devices and the possible different power generation modes, the spatial coordinates and time-series power generation data are merged and processed to obtain a cluster data set, which combines two different types of data into a new type of data, facilitating overall clustering analysis. Specifically, the spatial coordinates are a set of coordinate points, and the th spatial coordinate can be expressed as , and the time-series power generation data is a sequence with a length of time , which can be expressed as . represents the th new energy power generation device corresponding to the th spatial coordinate. The Z-score method, that is, the standard score method, is used to filter out the outliers in the time-series power generation data, so that only missing values exist in the time-series power generation data, and the linear interpolation method is used to fill the missing values according to the linear relationship of the time-series power generation data to obtain complete time-series power generation data. Secondly, in order to eliminate the geographical differences of new energy power generation devices in different regions, the normalization formula is used to normalize the spatial coordinates to obtain normalized coordinates. In order to eliminate the power magnitude differences of new energy power generation devices in different regions, the normalization formula is used to normalize the complete time-series power generation data to obtain normalized power generation data. The normalization formula is specifically: ; In the formula, is the th data after normalization, including the abscissa of the normalized coordinates, the ordinate of the normalized coordinates, or the time-series power data, is the th data participating in the normalization, is the minimum value of the data participating in the normalization, is the maximum value of the data participating in normalization; it should be noted that when normalizing the spatial coordinates, the abscissa and ordinate of the spatial coordinates need to be normalized separately, and finally, based on the new energy power generation device, the corresponding normalized coordinates and the corresponding normalized power generation data are combined into the same sequence, that is, the abscissa of the normalized coordinates is used as the first item of the sequence, the ordinate of the normalized coordinates is used as the second item of the sequence, and the abscissa of the normalized coordinates and the ordinate of the normalized coordinates are inserted into the sequence of the normalized power generation data to obtain the cluster dataset, and the cluster dataset can be expressed as , where is the abscissa of the normalized coordinates of the th new energy power generation device, is the ordinate of the normalized coordinates of the th new energy power generation device, is the time-series power generation data of the th new energy power generation device.

[0030] Secondly, in order to optimize the effect of clustering analysis, a distance threshold , minimum number of points and time-series data weight are determined based on the data characteristics of the cluster dataset; using the time-series data weight , the Euler formula is improved to a weighted Euler formula, and the specific weighted Euler formula is:[[]] ; where is the Euler distance between the th cluster data and the th cluster data, is the abscissa of the normalized coordinates corresponding to the th cluster data, is the abscissa of the normalized coordinates corresponding to the th cluster data, is the ordinate of the normalized coordinates corresponding to the th cluster data, is the ordinate of the normalized coordinates corresponding to the th cluster data, is the length of the time-series power generation data, is the time-series power generation data corresponding to the th cluster data, is the time-series power generation data corresponding to the th cluster data.

[0031] Input the cluster dataset into the weighted Euler formula for weighted calculation to obtain the Euler distance between data, that is, the cluster Euler distance. Then, based on a certain piece of data, all data with an Euler distance less than or equal to the distance threshold from the data is marked as points within the domain of the certain piece of data, thereby establishing a domain, which can be expressed as: ; In the formula, represents the domain of the th cluster data, represents the cluster data in the domain , is the cluster dataset, is the cluster data and the cluster data is the Euler distance between them, is the distance threshold.

[0032] Based on the domain and the minimum number of points, determine the core points and ordinary points. The judgment condition for the core points can be expressed as: ; In the formula, represents the domain of the th cluster data, is the minimum number of points; Furthermore, based on the core points and the domains corresponding to the core points, cluster division can be performed to obtain regional data clusters. If the data is a core point and the data is within the domain of the data , then it is said that the density of the data is reachable from the data . For the core point data , all data with a density reachable from the data is grouped into the same cluster to obtain regional data clusters. At the same time, data that does not belong to any cluster will also be obtained, and these data can be marked as noise points.

[0033] In one embodiment, after the above A4 is completed, it is also necessary to calculate the silhouette coefficient to verify the regional data clusters. The specific silhouette coefficient formula is: ; In the formula, is the silhouette coefficient of the th data point, is the average distance from the th data point to other data points in the same regional data cluster, is the average distance from the th data point to other regional data clusters.

[0034] In this embodiment, a corresponding silhouette coefficient threshold is preset. When the silhouette coefficient is less than or equal to the silhouette coefficient threshold, it indicates that the corresponding data is noise data. When the silhouette coefficient is greater than the silhouette coefficient threshold, the clustering effect is judged. At this time, the closer the silhouette coefficient is to 1, the better the clustering effect is proved.

[0035] Calculate the correlation coefficient and mutual information according to the regional data cluster, output data, and meteorological characteristic data; B1. Screen the output data based on the regional data cluster to obtain regional output data, and screen the meteorological characteristic data based on the regional data cluster to obtain regional meteorological characteristic data; B2. Perform standard processing on the regional meteorological characteristic data to obtain standard meteorological characteristic data, and organize the standard meteorological characteristic data and regional output data to obtain regional samples; B3. Calculate the correlation coefficient and mutual information according to the regional samples and the cross-correlation formula; The specific cross-correlation formula is: ; ; In the formula, is the correlation coefficient between the th meteorological factor corresponding to the regional sample and the new energy output , is the number of regional samples, is the rank of the th regional sample corresponding to the th meteorological factor , is the rank of the th regional sample corresponding to the new energy output , is the mutual information between the th meteorological factor corresponding to the regional sample and the new energy output , is the th meteorological factor and the joint probability distribution of the new energy output , is the th meteorological factor 's marginal probability distribution, is the marginal probability distribution of the new energy output .

[0036] In this embodiment, the th meteorological factor and the th meteorological factor are essentially the same. The difference is that the meteorological factor The quantity is not unique, and the meteorological factors are the meteorological factors The sum of, that is, the meteorological factors are all meteorological factors This set of meteorological factors, the new energy output and the new energy output are essentially the same. The difference is that the new energy output The quantity is not unique, and the new energy output is the new energy output The sum of, that is, the new energy output is the set of all new energy outputs, including different types of new energy outputs; based on the regional data cluster, the output data and meteorological characteristic data of the new energy power generation devices in the corresponding region are obtained. It should be noted that the purpose of the regional data cluster is to divide the new energy power generation devices into regions, and then obtain the output data and meteorological characteristic data corresponding to the new energy power generation devices. Therefore, it can be achieved by means of screening the already collected corresponding data or collecting the corresponding data of specific new energy power generation devices, etc.; in order to eliminate the scale difference of the meteorological characteristic data, a standardization formula is used to standardize the regional meteorological characteristic data to obtain the standard meteorological characteristic data. The specific standardization formula is: ; In the formula, is the th regional meteorological characteristic data, is the th standard meteorological characteristic data, is the mean value of the th meteorological factor, is the th standard deviation of the meteorological factor; and the standardized standard meteorological characteristic data has a mean value of 0 and a standard deviation of 1, which is applicable to Gaussian distribution data.

[0037] The standard meteorological characteristic data and the regional output data are sorted to obtain a regional sample. Specifically, the standard meteorological characteristic data and the regional output data are in the form of a sequence. The rank of each data in the sequence can be counted, as well as the joint probability distribution and marginal probability distribution of the standard meteorological characteristic data and the regional output data. The rank, joint probability distribution and marginal probability distribution are sorted to obtain a regional sample, and the regional sample is input into the cross-correlation formula to calculate the correlation coefficient and mutual information.

[0038] Based on the correlation coefficient and mutual information, the meteorological factors are screened to obtain the key meteorological factors; Specifically, the corresponding screening model is specifically: ; In the formula, is the key index corresponding to the nth meteorological factor, is the weight coefficient, is the correlation coefficient between the nth meteorological factor and the new energy output , is the mutual information between the nth meteorological factor and the new energy output.

[0039] Generate a time series data sequence of the corresponding new energy power generation device based on the key meteorological factors, output power and meteorological data; C1. Screen the meteorological data based on the key meteorological factors to obtain the key meteorological data, and screen the output power based on the regional data cluster corresponding to the key meteorological factors to obtain the key output power; C2. Integrate the time series of the key meteorological data and the time series of the key output power to obtain the input data; Specifically, count the missing values of the key meteorological data and fill them to obtain the complete key meteorological data, and normalize the complete key meteorological data to obtain the normalized meteorological data; Count the missing values of the key output power and fill them to obtain the complete key output power, and normalize the complete key output power to obtain the normalized output power; Align the time series of the normalized meteorological data and the time series of the normalized output power to obtain a unified time series; Generate the input data based on the unified time series, normalized meteorological data and normalized output power; Synchronously, perform similarity analysis on the key meteorological data to obtain meteorological similarity features, and perform similarity analysis on the geographical locations of the corresponding new energy power generation devices to obtain geographical similarity features; C3. Generate a heterogeneous graph based on the input data, meteorological similarity features and geographical similarity features; C4. Extract the spatial embedding features in the heterogeneous graph based on the time series of the input data to obtain feature vectors, and organize and splice all the feature vectors to obtain the time series data sequence.

[0040] In this embodiment, meteorological data is screened based on key meteorological factors, and the output power is screened based on the regional data clusters corresponding to the key meteorological factors. The screening process is to delete unnecessary data and optimize the volume and accuracy of the data to be processed. It should be noted that the purpose is to delete unnecessary data, which can be achieved by deleting the corresponding data that has been collected, or by deleting the corresponding data collection devices. To solve the problems of dimensional differences and missing values in the key meteorological data and key output power, methods such as normalization formulas and missing value filling are used for integration to obtain input data. Specifically, based on the periodicity of the meteorological data, historical data of the same period is selected to fill the missing values of the meteorological data, and the normalization formula is used to perform normalization processing on it. The linear interpolation method is used to fill the missing values of the output power according to the linear relationship of the output power, and the normalization formula is used to perform normalization processing on it. Align the time series of the normalized meteorological data with the time series of the normalized output power to obtain a unified time series, segment the same time series based on the sampling time interval, and sort the normalized meteorological data and normalized output power based on the segmented same time series to obtain input data. At the same time, in order to capture key influencing factors and regional sharing factors, similarity analysis is performed on the key meteorological data and geographical locations respectively, and the in-depth relationships between meteorology, climate, geographical location, etc. and new energy power generation devices are analyzed. Then, based on the input data, meteorological similarity features, and geographical similarity features, a heterogeneous graph that meets the requirements of the heterogeneous graph neural network is generated. The heterogeneous graph can be expressed as: ; In the formula, is the heterogeneous graph, is the cluster site of the cluster dataset corresponding to the new energy power generation device, is the edge set composed of multiple relationships ; is the set of relationships ; is the normalized adjacency matrix of the edge; The essence of the similarity analysis is to calculate the weights of each edge. The weights can be calculated according to the similarity function. The specific similarity function is: ; In the formula, is the feature of the node corresponding to the cluster data , and , where is the th screened meteorological factor, is the feature of the node corresponding to the cluster data , is the feature of the node corresponding to the cluster data , is the cluster data Corresponding nodes and cluster data The similarity of the corresponding nodes, is the cluster data The set of adjacent nodes to which the corresponding nodes are connected by a relationship connected.

[0041] Input the heterogeneous graph into a heterogeneous graph neural network. Specifically, based on the time series of the input data, input the corresponding node feature matrix in the heterogeneous graph into the heterogeneous graph neural network, and extract the spatial embedding features through multiple layers of graph convolution of the heterogeneous graph neural network to obtain feature vectors. The feature vectors are output in the form of a matrix. The update formula corresponding to the multiple layers of graph convolution is: ; In the formula, is the node feature matrix of the th layer, is the weight matrix corresponding to the relationship ; is the activation function; Subsequently, splice the feature vectors of each time step and reconstruct them into a time series data sequence.

[0042] Predict the future power of the corresponding new energy power generation device based on the time series data sequence; Specifically, gradually update the hidden state based on the time series data sequence to obtain a hidden state sequence, and use the hidden state sequence as the initial state to generate a prediction sequence; Generate a power output matrix based on the prediction sequence to obtain the future power of the corresponding new energy power generation device.

[0043] In this embodiment, input the time series data sequence into a time series prediction model. The encoder of the time series prediction model encodes the time series data sequence to gradually update the hidden state and obtain a hidden state sequence. The specific step-by-step update formula corresponding to the encoder is: ; ; ; ; ; ; In the formula, represents the forget gate corresponding to the encoder time step ; is the activation function, is the forget gate weight, is the feature dimension corresponding to the time step ; is the time step input of the corresponding time series data is the bias term of the forget gate is the encoder time step corresponding input gate is the input gate weight is the input gate bias term is the encoder time step corresponding candidate gate is the activation function of the candidate gate is the candidate gate weight is the candidate gate bias term is the encoder time step corresponding state update is the encoder time step corresponding state update is the encoder time step corresponding output gate is the output gate weight is the output gate bias term is the encoder time step corresponding hidden state

[0044] Using the hidden state sequence as the initial state of the decoder of the time series prediction model to gradually generate the prediction sequence within the future time period, the power output matrix formed by the prediction sequence is essentially the future power of the corresponding new energy power generation device, and the prediction model corresponding to the decoder is specifically: ; ; In the formula, LSTM represents the decoder is the time step corresponding hidden sequence is the time step corresponding updated state sequence is the time step corresponding decoder output is the time step corresponding output is the fully connected mapping layer of the decoder, used to map the decoder hidden state to the predicted power value

[0045] In addition, before formally predicting the combined power, it is also necessary to use the mean square error function to compare the entire prediction sequence with the true value, and update the parameters of the heterogeneous graph neural network and the time series prediction model according to the comparison result using the chain rule of backpropagation. Specifically, the gradient descent algorithm is used for the update, and the gradient descent algorithm is specifically: ; ; In the formula, is the parameter vector of the heterogeneous graph neural network or the time series prediction model, is the time step corresponding actual output power matrix, is the time step corresponding predicted output power matrix, is the total number of time steps, is the learning rate, is the mean square error function gradient with respect to the parameter.

[0046] This embodiment at least has the following substantial effects: (1) This embodiment conducts clustering analysis based on the spatial coordinates and time-series power generation data of new energy power generation devices to determine regional data clusters, divides new energy power generation devices with similar geographical distributions and similar or even the same power generation patterns into the same category, realizes the precise division of new energy power generation devices under complex non-linear spatial relationships, and optimizes the clustering quality using the silhouette coefficient to effectively filter out noise data; (2) This embodiment calculates the correlation coefficient and mutual information based on the regional data clusters, output data, and meteorological characteristic data, depicts the highly non-linear relationship and complex mapping relationship between weather factors and the power of new energy power generation devices through the correlation coefficient and mutual information, and screens out key meteorological factors based on the correlation coefficient and mutual information, eliminating weather factors that have weak or even no impact on the corresponding new energy power generation devices, effectively improving the efficiency of power prediction; (3) This embodiment generates a time-series data sequence of the corresponding new energy power generation device based on the key meteorological factors, output power, and meteorological data, obtains the change in the power generation power of the new energy power generation device over time with weather factors, clarifies the change in the power generation power of the new energy power generation device over time with weather factors, and predicts the future power of the corresponding new energy power generation device through the change, significantly improving the joint power prediction accuracy of distributed new energy power generation devices.

[0047] The above specific implementation manners are the preferred implementation manners of the present invention, which do not limit the specific implementation scope of the present invention. The scope of the present invention includes but is not limited to this specific implementation manner. Any equivalent changes made according to the shape, structure, and method of the present invention are within the protection scope of the present invention.

Claims

1. A combined prediction method for a distributed new energy power generation device, characterized in that, It includes the following steps: Collect the spatial coordinates, time-series power generation data, output data, meteorological characteristic data, output power, and meteorological data of the new energy power generation device; Perform clustering analysis based on the spatial coordinates and time-series power generation data to determine regional data clusters; Calculate the correlation coefficient and mutual information according to the regional data clusters, output data, and meteorological characteristic data; Screen the meteorological factors based on the correlation coefficient and mutual information to obtain the key meteorological factors; Generate a time-series data sequence corresponding to the new energy power generation device based on the key meteorological factors, output power, and meteorological data; Predict the future power of the corresponding new energy power generation device based on the time-series data sequence.

2. The combined prediction method of a distributed new energy power generation device according to claim 1, characterized in that The specific process of performing clustering analysis based on the spatial coordinates and time-series power generation data to determine regional data clusters is as follows: A1. Combine and process the spatial coordinates and time-series power generation data to obtain a cluster data set, and determine the distance threshold, minimum number of points, and time-series data weight based on the data characteristics of the cluster data set; A2. Perform weighted calculation based on the time-series data weight and the cluster data set to obtain the cluster Euler distance, and construct a neighborhood based on the cluster Euler distance and the distance threshold; A3. Judge and determine the core points based on the neighborhood and the minimum number of points. If the number of data points in the neighborhood is less than the minimum number of points, mark the cluster data corresponding to the neighborhood as ordinary points. If the number of data points in the neighborhood is greater than or equal to the minimum number of points, mark the cluster data corresponding to the neighborhood as core points; A4. Perform cluster division based on the core points and the neighborhoods corresponding to the core points to obtain regional data clusters.

3. The combined prediction method of a distributed new energy power generation device according to claim 2, wherein In A1, the specific process of combining and processing the spatial coordinates and time-series power generation data to obtain a cluster data set is as follows: Statistically eliminate the outliers in the time-series power generation data based on the standard score method, and fill the missing values in the time-series power generation data based on the linear interpolation method to obtain complete time-series power generation data; Perform normalization processing on the spatial coordinates to obtain normalized coordinates, and perform normalization processing on the complete time-series power generation data to obtain normalized power generation data; Sort out the normalized coordinates and normalized power generation data and combine them to obtain a cluster data set.

4. The combined prediction method of a distributed new energy power generation device according to claim 2, characterized in that, After A4 is completed, it is also necessary to calculate the silhouette coefficient to verify the regional data clusters. The specific silhouette coefficient formula is as follows: ; In the formula, is the silhouette coefficient of the -th data point, is the average distance from the -th data point to other data points in the data cluster of the same region, is the average distance from the -th data point to the data clusters in other regions.

5. The combined prediction method of a distributed new energy power generation device according to claim 1, characterized in that The specific process of calculating the correlation coefficient and mutual information according to the regional data clusters, output data, and meteorological characteristic data is as follows: B1. Screen the output data based on the regional data clusters to obtain regional output data, and screen the meteorological characteristic data based on the regional data clusters to obtain regional meteorological characteristic data; B2. Perform standard processing on the regional meteorological characteristic data to obtain standard meteorological characteristic data, and sort out the standard meteorological characteristic data and the regional output data to obtain regional samples; B3. Calculate the correlation coefficient and mutual information according to the regional samples and the cross-correlation formula.

6. The combined prediction method of a distributed new energy power generation device according to claim 5, characterized in that In B3, the cross-correlation formula is specifically: ; ; In the formula, is the correlation coefficient between the th meteorological factor corresponding to the regional sample and the new energy output , is the number of regional samples, is the rank of the th meteorological factor corresponding to the th regional sample, is the rank of the new energy output corresponding to the th regional sample, is the mutual information between the th meteorological factor corresponding to the regional sample and the new energy output , is the joint probability distribution of the th meteorological factor and the new energy output , is the marginal probability distribution of the th meteorological factor , is the marginal probability distribution of the new energy output . is the th meteorological factor , is the marginal probability distribution of the new energy output.

7. A combined prediction method for a distributed new energy power generation device according to claim 1, characterized in that The specific screening model for screening the meteorological factors based on the correlation coefficient and mutual information to obtain the key meteorological factors is as follows: ; In the formula, is the key index corresponding to the th meteorological factor, is the weight coefficient, is the correlation coefficient between the th meteorological factor and the new energy output , and is the mutual information between the th meteorological factor and the new energy output . th meteorological factor and the new energy output .

8. The combined prediction method of a distributed new energy power generation device according to claim 1, characterized in that The specific process of generating a time-series data sequence corresponding to the new energy power generation device based on the key meteorological factors, output power, and meteorological data is as follows: C1. Screen the meteorological data based on the key meteorological factors to obtain key meteorological data, and screen the output power based on the regional data clusters corresponding to the key meteorological factors to obtain key output power; C2. Integrate the time series of key meteorological data and the time series of key output power to obtain input data; Synchronously, perform similarity analysis on the key meteorological data to obtain meteorological similarity features, and perform similarity analysis on the geographical locations of the corresponding new energy power generation devices to obtain geographical similarity features; C3. Generate a heterogeneous graph based on the input data, meteorological similarity features, and geographical similarity features; C4. Extract spatial embedding features in the heterogeneous graph based on the time series of the input data to obtain feature vectors, and organize and splice all the feature vectors to obtain a time series data sequence.

9. A combined prediction method for a distributed new energy power generation device according to claim 8, characterized in that In C2, the specific process of integrating the time series of key meteorological data and the time series of key output power to obtain input data is as follows: Count the missing values of the key meteorological data and fill them to obtain complete key meteorological data, and normalize the complete key meteorological data to obtain normalized meteorological data; Count the missing values of the key output power and fill them to obtain complete key output power, and normalize the complete key output power to obtain normalized output power; Align the time series of the normalized meteorological data with the time series of the normalized output power to obtain a unified time series; Generate input data based on the unified time series, normalized meteorological data, and normalized output power.

10. The combined prediction method of a distributed new energy power generation device according to claim 1, characterized in that, The specific process of predicting the future power of the corresponding new energy power generation device based on the time series data sequence is as follows: Gradually update the hidden state based on the time series data sequence to obtain a hidden state sequence, and use the hidden state sequence as the initial state to generate a prediction sequence; Generate a power output matrix based on the prediction sequence to obtain the future power of the corresponding new energy power generation device.

Citation Information

Patent Citations

  • High-precision wind-solar combined power prediction method

    CN119226702A

  • Meteorological-driven new energy generation power prediction method

    CN116090635A

  • Distributed photovoltaic short-term power prediction method based on fusion clustering and VQC-LSTM

    CN117498319A

  • Photovoltaic short-term generation power prediction method based on weather clustering and LSTM combined model

    CN117791595A

  • Photovoltaic output prediction method and device based on multi-source data, and storage medium

    CN118399406A

Cited By

  • Electric power time sequence data fusion method for power grid planning optimization

    CN120705820A