A joint prediction method for distributed new energy power generation devices

By clustering analysis of the spatial coordinates and timing data of new energy power generation devices, key meteorological factors are screened, and time-series data sequences are generated, the problem of low prediction accuracy caused by the sensitivity of new energy power generation devices to natural resources is solved, and efficient joint power prediction is achieved.

CN120237648BActive Publication Date: 2025-08-12STATE GRID ZHEJIANG ELECTRIC POWER CO LTD NINGBO POWER SUPPLY CO
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510726967.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-08-12
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The prior art fails to effectively consider the sensitivity of new energy power generation devices to different natural resources, resulting in low joint power prediction accuracy.

Method used

Cluster analysis is carried out based on the spatial coordinates and time-sequential power generation data of new energy power generation devices, regional data clusters are determined, correlation coefficients and mutual information are calculated to screen key meteorological factors, and time-sequential data sequences are generated to predict future power.

Benefits of technology

The combined power prediction accuracy of distributed new energy power generation devices has been significantly improved, the precise division of new energy power generation devices and the filtering of noise data has been realized, and the prediction efficiency has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120237648B_ABST
    Figure CN120237648B_ABST
Patent Text Reader

Abstract

The present invention discloses a joint prediction method for distributed new energy power generation devices, which belongs to the technical field of new energy power generation prediction, comprising: collecting spatial coordinates, time-series power generation data, output data, meteorological characteristic data, output power, and meteorological data of the new energy power generation devices; performing cluster analysis based on the spatial coordinates and time-series power generation data to determine regional data clusters; calculating correlation coefficients and mutual information based on the regional data clusters, output data, and meteorological characteristic data; screening meteorological factors based on the correlation coefficients and mutual information to obtain key meteorological factors; generating a time-series data sequence corresponding to the new energy power generation device based on the key meteorological factors, output power, and meteorological data; and predicting the future power of the corresponding new energy power generation device based on the time-series data sequence. The present invention overcomes the problem that the prior art does not consider the sensitivity of the new energy power generation device to natural resources, resulting in low accuracy in joint power prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of new energy power generation prediction technology, and in particular to a joint prediction method for distributed new energy power generation devices. Background Art

[0002] With the development of new energy power generation technology, the proportion of green power sources represented by new energy in the power grid is getting higher and higher. Although new energy power generation devices are widely distributed and regionally different, the installation locations of new energy power generation devices have similar or even the same natural resources, such as wind resources, solar energy resources, etc. These natural resources have certain periodic changes in time. Combined with the characteristics of new energy power generation devices themselves being greatly affected by natural resources, when natural resources change, the power of the corresponding new energy power generation devices also changes accordingly, making it difficult to predict the power generation power of new energy power generation devices. Although the power of new energy power generation devices of different types and regions has the characteristics of temporal and spatial complementarity, it is difficult to truly complement each other due to the instability of power generation.

[0003] Chinese patent, publication number: CN119226702A, publication date: December 31, 2024, discloses a high-precision wind-solar combined power prediction method, including: A. first, separately collect wind power generation data and photovoltaic power generation data; B. pre-process the collected wind power generation data and photovoltaic power generation data; C. extract features from the pre-processed power generation data; D. then, based on the spatiotemporal correlation and load curve of wind and solar power generation, introduce a temporal attention mechanism, and use the temporal feature mining capability of the GRU network and the spatial feature mining capability of the CNN network to construct a wind-solar combined power prediction model based on the improved GRU-CNN algorithm; E. carry out model training, verification and testing to realize the combined prediction of wind and solar power generation; F. finally, based on high-precision meteorological data, evaluate the prediction accuracy performance of different methods through testing, and form an optimization strategy to obtain the best prediction method; however, this invention does not take into account the different sensitivities of different types of new energy power generation devices to different natural resources, resulting in low prediction accuracy of combined power. Summary of the Invention

[0004] The purpose of the present invention is to address the problem that the existing technology does not take into account the sensitivity of new energy power generation devices to natural resources, resulting in low accuracy of joint power prediction; a joint prediction method for distributed new energy power generation devices is proposed, which divides the region into regions based on cluster analysis of the spatial coordinates and time-series power generation data of the new energy power generation devices to determine regional data clusters, and screens meteorological factors that have a great impact on the new energy power generation devices according to the regional data clusters, output data and meteorological characteristic data to obtain key meteorological factors, and generates a time-series data sequence based on the key meteorological factors, output power and meteorological data to predict the future power of the new energy power generation devices, thereby significantly improving the joint power prediction accuracy of distributed new energy power generation devices.

[0005] In a first aspect, a technical solution provided in an embodiment of the present invention is a joint prediction method for distributed new energy power generation devices, comprising the following steps:

[0006] Collect spatial coordinates, time-series power generation data, output data, meteorological characteristic data, output power, and meteorological data of new energy power generation devices;

[0007] Cluster analysis is performed based on spatial coordinates and time series power generation data to determine regional data clusters;

[0008] Calculate correlation coefficients and mutual information based on regional data clusters, output data, and meteorological characteristic data;

[0009] Meteorological factors are screened based on correlation coefficient and mutual information to obtain key meteorological factors;

[0010] Generate a time series data sequence corresponding to the new energy power generation device based on key meteorological factors, output power and meteorological data;

[0011] Predict the future power of corresponding new energy power generation devices based on time series data.

[0012] In this solution, cluster analysis is performed based on the spatial coordinates and time-series power generation data of the new energy power generation device to determine the regional data cluster, and the new energy power generation devices with similar geographical distribution and similar or even identical power generation modes are divided into a unified category, thereby providing strong support for subsequent unified predictions; secondly, weather is a factor that affects the power generation of new energy power generation devices. Changes in the weather itself lead to complex changes in power generation, which is difficult to predict. The correlation coefficient and mutual information are calculated based on the regional data cluster, output data and meteorological characteristic data. The highly nonlinear relationship and complex mapping relationship between weather factors and the power of new energy power generation devices are depicted through the correlation coefficient and mutual information; in addition, different new energy power generation devices have different sensitivities to different weather factors, that is, when different weather factors act on the same new energy power generation device, the impact on the power generation of the new energy power generation device is different. Based on the correlation coefficient And mutual information is used to screen meteorological factors to obtain key meteorological factors, eliminate weather factors that have little or no impact on the corresponding new energy power generation devices, and then propose a large amount of useless data, which effectively improves the efficiency of power forecasting; however, at this time, only the category of new energy power generation devices and key weather factors that affect new energy power generation devices are obtained. Although the highly nonlinear relationship and complex mapping relationship between weather factors and the power of new energy power generation devices are described, the changes in the power generation of new energy power generation devices in the time series with weather factors are not clear. Based on the key meteorological factors, output power and meteorological data, a time series data sequence of the corresponding new energy power generation device is generated, and the changes in the power generation of new energy power generation devices in the time series with weather factors are obtained. The future power of the corresponding new energy power generation device is predicted through the changes, which significantly improves the joint power prediction accuracy of distributed new energy power generation devices.

[0013] Preferably, the specific process of determining regional data clusters by cluster analysis based on spatial coordinates and time-series power generation data is as follows:

[0014] A1. Merge the spatial coordinates and time-series power generation data to obtain a cluster data set, and determine the distance threshold, minimum number of points, and time-series data weight based on the data characteristics of the cluster data set;

[0015] A2. Perform weighted calculation based on the time series data weight and the cluster data set to obtain the cluster Euler distance, and construct a domain based on the cluster Euler distance and distance threshold;

[0016] A3. Determine the core point based on the domain and the minimum number of points. If the number of data points in the domain is less than the minimum number of points, mark the cluster data corresponding to the domain as a normal point. If the number of data points in the domain is greater than or equal to the minimum number of points, mark the cluster data corresponding to the domain as a core point.

[0017] A4. Clustering is performed based on core points and their corresponding areas to obtain regional data clusters.

[0018] In this solution, since the spatial positions of new energy power generation devices are different and the power generation modes may also be different, the spatial coordinates and time-series power generation data are merged to obtain a cluster data set, and the two different data are merged into a new data to facilitate overall cluster analysis. Secondly, in order to optimize the effect of cluster analysis, the distance threshold, the minimum number of points and the time-series data weight are determined based on the data characteristics of the cluster data set; the Euler formula is improved into a weighted Euler formula using the time-series data weight, and the cluster data set is input into the weighted Euler formula for weighted calculation to obtain the Euler distance between data, namely the cluster Euler distance. Then, based on a certain data statistic, all data whose Euler distance to the data is less than or equal to the distance threshold are marked as points within the certain data field, thereby establishing a field; based on the field and the minimum number of points, the core points and ordinary points are judged, and then clustering can be performed based on the core points and the fields corresponding to the core points to obtain regional data clusters. If the data As the core point, data In data In the field of The density of the data can be , for the core point data , all densities available to the data The data are classified into the same cluster to obtain the regional data cluster. At the same time, data that does not belong to any cluster will also be obtained, which can be marked as noise points.

[0019] Preferably, in A1, the specific process of merging the spatial coordinate and time-series power generation data to obtain the cluster data set is:

[0020] The outliers in the time series power generation data are statistically removed based on the standard score method, and the missing values in the time series power generation data are filled in based on the linear interpolation method to obtain the complete time series power generation data;

[0021] Normalizing the spatial coordinates to obtain normalized coordinates, and normalizing the complete time series power generation data to obtain normalized power generation data;

[0022] The normalized coordinates and normalized power generation data are sorted and merged to obtain the cluster dataset.

[0023] In this solution, the spatial coordinates are specifically a set of coordinate points, and the time series power generation data are specifically a set of lengths of time. The Z-score method, that is, the standard score method, is used to filter out the abnormal values in the time series power generation data, so that only missing values exist in the time series power generation data, and the linear interpolation method is used to fill the missing values according to the linear relationship of the time series power generation data to obtain the complete time series power generation data; secondly, in order to eliminate the geographical differences of new energy power generation devices in different regions, the The normalization formula is used to normalize the spatial coordinates to obtain normalized coordinates. In order to eliminate the differences in power levels of new energy power generation devices in different regions, the normalization formula is used to normalize the spatial coordinates to obtain normalized coordinates. The normalization formula normalizes the complete time series power generation data to obtain normalized power generation data. It should be noted that when normalizing the spatial coordinates, the horizontal coordinate and the vertical coordinate of the spatial coordinate need to be separated to perform normalization calculations. Finally, based on the new energy power generation device, the corresponding normalized coordinates and the corresponding normalized power generation data are merged into the same series, that is, the horizontal coordinate of the normalized coordinate is used as the first item of the series, and the vertical coordinate of the normalized coordinate is used as the second item of the series, and the horizontal coordinate and the vertical coordinate of the normalized coordinate are inserted into the series of normalized power generation data.

[0024] Preferably, after A4 is completed, the silhouette coefficient needs to be calculated to verify the regional data cluster. The corresponding silhouette coefficient formula is specifically as follows:

[0025] ;

[0026] Where, For the The silhouette coefficient of the data points, For the The average distance from a data point to other data points in the same area data cluster, For the The average distance from a data point to other regional data clusters.

[0027] In this scheme, the silhouette coefficient The closer it is to 1, the more it corresponds to The better the clustering effect of the data points, the better the While ensuring the clustering effect, data with poor clustering effect is further eliminated, that is, noise data is effectively filtered out.

[0028] Preferably, the specific process of calculating the correlation coefficient and mutual information based on the regional data cluster, output data and meteorological characteristic data is as follows:

[0029] B1. Filtering the output data based on the regional data cluster to obtain regional output data, and filtering the meteorological characteristic data based on the regional data cluster to obtain regional meteorological characteristic data;

[0030] B2. Perform standard processing on the regional meteorological characteristic data to obtain standard meteorological characteristic data, and organize the standard meteorological characteristic data and regional output data to obtain regional samples;

[0031] B3. Calculate the correlation coefficient and mutual information based on the regional samples and the cross-correlation formula.

[0032] In this solution, the output data and meteorological characteristic data of the new energy power generation device in the corresponding area are obtained based on the regional data cluster. It should be noted that the purpose of the regional data cluster is to divide the new energy power generation device into regions, and then obtain the output data and meteorological characteristic data corresponding to the new energy power generation device. Therefore, it can be achieved by screening the corresponding data that have been collected or collecting the corresponding data of specific new energy power generation devices. In order to eliminate the scale difference of the meteorological characteristic data, the regional meteorological characteristic data is processed by a standardized formula to obtain standard meteorological characteristic data, and the standard meteorological characteristic data and regional output data are sorted to obtain regional samples. Specifically, the standard meteorological characteristic data and regional output data are in the form of series. The rank of each data in the series, as well as the joint probability distribution and marginal probability distribution of the standard meteorological characteristic data and the regional output data can be counted, and the rank, joint probability distribution and marginal probability distribution are sorted to obtain regional samples. The regional samples are input into the cross-correlation formula to calculate the correlation coefficient and mutual information.

[0033] Preferably, in B3, the cross-correlation formula is specifically:

[0034] ;

[0035] ;

[0036] Where, The regional sample corresponds to Meteorological factors With new energy output The correlation coefficient of is the number of regional samples, For the The regional samples correspond to Meteorological factors rank, For the Regional samples corresponding to new energy output rank, The regional sample corresponds to Meteorological factors With new energy output The mutual information of For the Meteorological factors With new energy output The joint probability distribution of For the Meteorological factors The marginal probability distribution of Contribute to new energy The marginal probability distribution of .

[0037] Preferably, the screening model corresponding to the key meteorological factors obtained by screening meteorological factors based on correlation coefficient and mutual information is specifically:

[0038] ;

[0039] Where, For the Meteorological factors The corresponding key indicators are is the weight coefficient, For the Meteorological factors With new energy output The correlation coefficient of For the Meteorological factors With new energy output mutual information.

[0040] Preferably, the specific process of generating the time series data sequence corresponding to the new energy power generation device based on key meteorological factors, output power and meteorological data is as follows:

[0041] C1. Filtering meteorological data based on key meteorological factors to obtain key meteorological data, and filtering output power based on regional data clusters corresponding to the key meteorological factors to obtain key output power;

[0042] C2, integrating the time series of key meteorological data and the time series of key output power to obtain input data;

[0043] Simultaneously, similarity analysis is performed on key meteorological data to obtain meteorological similarity features, and similarity analysis is performed on the geographical locations of corresponding new energy power generation devices to obtain geographical similarity features;

[0044] C3, generating heterogeneous graphs based on input data, meteorological similarity features, and geographical similarity features;

[0045] C4. Extract the spatial embedding features in the heterogeneous graph based on the time series of the input data to obtain feature vectors, and organize and concatenate all the feature vectors to obtain the time series data sequence.

[0046] In this solution, meteorological data is screened based on key meteorological factors, and output power is screened based on regional data clusters corresponding to key meteorological factors. The screening process is to delete unnecessary data and optimize the volume and accuracy of the data to be processed. It should be noted that its purpose is to delete unnecessary data, which can be achieved by deleting the corresponding data that has been collected, or by deleting the corresponding data collection equipment. In order to solve the dimensional differences and missing values of the key meteorological data and key output power, normalization formulas, missing value filling and other methods are used to integrate the input data. At the same time, in order to capture key influencing factors and regional shared factors, Similarity analysis is performed on key meteorological data and geographical locations respectively, and the deep relationship between meteorology, climate, geographical location and new energy power generation devices is analyzed. Then, a heterogeneous graph that meets the requirements of the heterogeneous graph neural network is generated based on the input data, meteorological similarity features and geographical similarity features, and the heterogeneous graph is input into the heterogeneous graph neural network. Specifically, based on the time series of the input data, the corresponding node feature matrix in the heterogeneous graph is input into the heterogeneous graph neural network, and the spatial embedding features are extracted through the multi-layer graph convolution of the heterogeneous graph neural network to obtain a feature vector. The feature vector is output in the form of a matrix, and the feature vector of each time step is spliced to reconstruct a time series data sequence.

[0047] Preferably, in C2, the specific process of integrating the time series of key meteorological data and the time series of key output power to obtain input data is:

[0048] Counting missing values of key meteorological data and filling them to obtain complete key meteorological data, and normalizing the complete key meteorological data to obtain normalized meteorological data;

[0049] Counting missing values of key output power and filling them to obtain complete key output power, and normalizing the complete key output power to obtain normalized output power;

[0050] Align the time series of normalized meteorological data with the time series of normalized output power to obtain a unified time series;

[0051] The input data is generated based on unified time series, normalized meteorological data, and normalized output power.

[0052] In this scheme, based on the periodicity of meteorological data, historical data of the same period are selected to fill the missing values of meteorological data, and the normalization formula is used to normalize them. Linear interpolation is used to fill the missing values of output power according to the linear relationship of output power, and the normalization formula is used to normalize them. The time series of normalized meteorological data and the time series of normalized output power are aligned to obtain a unified time series, the same time series is segmented based on the sampling time interval, and the normalized meteorological data and normalized output power are sorted based on the segmented same time series to obtain input data.

[0053] Preferably, the specific process of predicting the future power of the corresponding new energy power generation device based on the time series data sequence is:

[0054] Based on the time series data sequence, the hidden state is gradually updated to obtain the hidden state sequence, and the hidden state sequence is used as the initial state to generate the prediction sequence;

[0055] The power output matrix is generated based on the prediction sequence to obtain the future power of the corresponding new energy power generation device.

[0056] In this solution, the time series data sequence is input into the time series prediction model. The encoder of the time series prediction model encodes the time series data sequence to realize the gradual update of the hidden state, obtains the hidden state sequence, and uses the hidden state sequence as the initial state of the time series prediction model decoder to gradually generate the prediction sequence in the future time period. The power output matrix formed by the prediction sequence is actually the future power of the corresponding new energy power generation device.

[0057] Beneficial effects of the present invention:

[0058] (1) This application determines regional data clusters based on the spatial coordinates and time-series power generation data of new energy power generation devices through cluster analysis, and divides new energy power generation devices with similar geographical distribution and similar or even identical power generation modes into a unified category, thereby achieving accurate division of new energy power generation devices under complex nonlinear spatial relationships, and using the silhouette coefficient to optimize clustering quality and effectively filter out noise data;

[0059] (2) This application calculates the correlation coefficient and mutual information based on regional data clusters, output data and meteorological characteristic data, and uses the correlation coefficient and mutual information to depict the highly nonlinear relationship and complex mapping relationship between weather factors and the power of new energy power generation devices. Based on the correlation coefficient and mutual information, the application selects meteorological factors to obtain key meteorological factors, eliminates weather factors that have little or no effect on the corresponding new energy power generation devices, and effectively improves the efficiency of power forecasting;

[0060] (3) This application generates a time series data sequence corresponding to the new energy power generation device based on key meteorological factors, output power and meteorological data, obtains the power generation change of the new energy power generation device in the time series accompanied by weather factors, clarifies the power generation change of the new energy power generation device in the time series accompanied by weather factors, and predicts the future power of the corresponding new energy power generation device through the said change, thereby significantly improving the joint power prediction accuracy of distributed new energy power generation devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Other features, objects, and advantages of the present invention will become more apparent upon reading the detailed description of the non-limiting embodiments made with reference to the following drawings. The drawings are for the purpose of illustrating preferred embodiments only and are not to be construed as limiting the present invention. Like reference characters are used throughout the drawings to designate like parts.

[0062] Figure 1 The figure is a flow chart of a joint prediction method for distributed renewable energy power generation devices according to the present invention. DETAILED DESCRIPTION

[0063] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation method described herein is only an optimal embodiment of the present invention, which is only used to explain the present invention and does not limit the scope of protection of the present invention. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0064] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the operations (or steps) as sequential processes, many of the operations (or steps) therein can be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but can also have additional steps not included in the figures; the process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0065] Example 1:

[0066] like Figure 1 As shown, this embodiment provides a joint prediction method for distributed new energy power generation devices, including the following steps:

[0067] Collect spatial coordinates, time-series power generation data, output data, meteorological characteristic data, output power, and meteorological data of new energy power generation devices;

[0068] Cluster analysis is performed based on spatial coordinates and time series power generation data to determine regional data clusters;

[0069] A1. Merge the spatial coordinates and time-series power generation data to obtain a cluster data set, and determine the distance threshold, minimum number of points, and time-series data weight based on the data characteristics of the cluster data set;

[0070] Specifically, outliers in the time series power generation data are counted and eliminated based on the standard score method, and missing values in the time series power generation data are filled in based on the linear interpolation method to obtain complete time series power generation data;

[0071] Normalizing the spatial coordinates to obtain normalized coordinates, and normalizing the complete time series power generation data to obtain normalized power generation data;

[0072] Arrange the normalized coordinates and normalized power generation data and merge them to obtain the cluster data set;

[0073] A2. Perform weighted calculation based on the time series data weight and the cluster data set to obtain the cluster Euler distance, and construct a domain based on the cluster Euler distance and distance threshold;

[0074] A3. Determine the core point based on the domain and the minimum number of points. If the number of data points in the domain is less than the minimum number of points, mark the cluster data corresponding to the domain as a normal point. If the number of data points in the domain is greater than or equal to the minimum number of points, mark the cluster data corresponding to the domain as a core point.

[0075] A4. Clustering is performed based on core points and their corresponding areas to obtain regional data clusters.

[0076] In this embodiment, since the spatial positions of the new energy power generation devices are different and the power generation modes may also be different, the spatial coordinates and the time series power generation data are merged to obtain a cluster data set, and the two different data are merged into a new data to facilitate the overall cluster analysis; specifically, the spatial coordinates are a set of coordinate points, The spatial coordinates can be expressed as The time series power generation data is specifically a set of time length The sequence can be expressed as , Indicates the The space coordinates corresponding to the New energy power generation devices; the Z-score method, that is, the standard score method, is used to filter out abnormal values in the time series power generation data, so that only missing values exist in the time series power generation data, and the linear interpolation method is used to fill the missing values according to the linear relationship of the time series power generation data to obtain complete time series power generation data; secondly, in order to eliminate the geographical differences of new energy power generation devices in different regions, the The normalization formula is used to normalize the spatial coordinates to obtain normalized coordinates. In order to eliminate the differences in power levels of new energy power generation devices in different regions, the normalization formula is used to normalize the spatial coordinates to obtain normalized coordinates. The normalization formula normalizes the complete time series power generation data to obtain normalized power generation data. The normalization formula is as follows:

[0077] ;

[0078] Where, For the The normalized data includes the horizontal coordinate of the normalized coordinate, the vertical coordinate of the normalized coordinate or the time series power data. For the The data involved in normalization, is the minimum value of the normalized data, is the maximum value of the normalized data; it should be noted that when the spatial coordinates are normalized, the abscissa and ordinate of the spatial coordinates need to be separated to perform normalization calculations. Finally, based on the new energy power generation device, the corresponding normalized coordinates and the corresponding normalized power generation data are merged into the same series, that is, the abscissa of the normalized coordinates is taken as the first item of the series, and the ordinate of the normalized coordinates is taken as the second item of the series. The abscissa and ordinate of the normalized coordinates are inserted into the series of normalized power generation data to obtain a cluster data set, which can be expressed as , where For the The horizontal coordinate of the normalized coordinate of each new energy power generation device, For the The vertical coordinate of the normalized coordinate of each new energy power generation device, For the Time series power generation data of new energy power generation devices.

[0079] Secondly, in order to optimize the effect of cluster analysis, the distance threshold is determined based on the data characteristics of the cluster data set. , minimum number of points Time series data weights ; Using time series data weight The Euler formula is improved into the weighted Euler formula, which is specifically:

[0080] ;

[0081] Where, For the Cluster data and The Euler distance between cluster data, For the The horizontal coordinate of the normalized coordinate corresponding to the cluster data, For the The horizontal coordinate of the normalized coordinate corresponding to the cluster data, For the The vertical coordinate of the normalized coordinate corresponding to each cluster data, For the The vertical coordinate of the normalized coordinate corresponding to each cluster data, is the length of the time series power generation data, For the The time series power generation data corresponding to the cluster data, For the The time series power generation data corresponding to the cluster data.

[0082] The cluster data set is input into the weighted Euler formula for weighted calculation to obtain the Euler distance between data, i.e., the cluster Euler distance. Then, based on a certain data, all data whose Euler distance to the data is less than or equal to the distance threshold are counted and all the data are marked as points within the certain data domain, thereby establishing a domain. The domain can be expressed as:

[0083] ;

[0084] Where, Indicates the The field of cluster data, Representing clustered data in the domain , For cluster dataset, For cluster data With cluster data The Euler distance of is the distance threshold.

[0085] The core points and common points are determined based on the domain and the minimum number of points. The determination condition of the core points can be expressed as:

[0086] ;

[0087] Where, Indicates the The field of cluster data, is the minimum number of points;

[0088] Then, clustering can be performed based on core points and their corresponding areas to obtain regional data clusters. As the core point, data In data In the field of The density of the data can be , for the core point data , all densities available to the data The data are classified into the same cluster to obtain the regional data cluster. At the same time, data that does not belong to any cluster will also be obtained, which can be marked as noise points.

[0089] In one embodiment, after A4 is completed, it is necessary to calculate the silhouette coefficient to verify the regional data cluster. The corresponding silhouette coefficient formula is specifically as follows:

[0090] ;

[0091] Where, For the The silhouette coefficient of the data points, For the The average distance from a data point to other data points in the same area data cluster, For the The average distance from a data point to other regional data clusters.

[0092] In this embodiment, a corresponding silhouette coefficient threshold is pre-set. When the silhouette coefficient is less than or equal to the silhouette coefficient threshold, it indicates that the corresponding data is noise data. When the silhouette coefficient is greater than the silhouette coefficient threshold, the clustering effect is judged. At this time, the closer the silhouette coefficient is to 1, the better the clustering effect.

[0093] Calculate correlation coefficients and mutual information based on regional data clusters, output data, and meteorological characteristic data;

[0094] B1. Filtering the output data based on the regional data cluster to obtain regional output data, and filtering the meteorological characteristic data based on the regional data cluster to obtain regional meteorological characteristic data;

[0095] B2. Perform standard processing on the regional meteorological characteristic data to obtain standard meteorological characteristic data, and organize the standard meteorological characteristic data and regional output data to obtain regional samples;

[0096] B3. Calculate the correlation coefficient and mutual information based on regional samples and the cross-correlation formula;

[0097] The cross-correlation formula is specifically:

[0098] ;

[0099] ;

[0100] Where, The regional sample corresponds to Meteorological factors With new energy output The correlation coefficient of is the number of regional samples, For the The regional samples correspond to Meteorological factors rank, For the Regional samples corresponding to new energy output rank, The regional sample corresponds to Meteorological factors With new energy output The mutual information of For the Meteorological factors With new energy output The joint probability distribution of For the Meteorological factors The marginal probability distribution of Contribute to new energy The marginal probability distribution of .

[0101] In this embodiment, Meteorological factors With the Meteorological factors The essence is the same, the difference is that the meteorological factors The number of meteorological factors is not unique. It is meteorological factors The sum of meteorological factors All meteorological factors This type of meteorological factors, new energy output With new energy output The essence is the same, the difference is that new energy output The number of new energy output is not unique. It is new energy output The sum of new energy output It is the collection of all new energy outputs, including different types of new energy outputs; based on the regional data cluster, the output data and meteorological characteristic data of the new energy power generation device in the corresponding area are obtained. It should be noted that the purpose of the regional data cluster is to divide the new energy power generation device into regions, and then obtain the output data and meteorological characteristic data corresponding to the new energy power generation device. Therefore, it can be achieved by screening the corresponding data that has been collected or collecting the corresponding data of a specific new energy power generation device. In order to eliminate the scale difference of the meteorological characteristic data, the regional meteorological characteristic data is processed by a standardization formula to obtain the standard meteorological characteristic data. The standardization formula is specifically as follows:

[0102] ;

[0103] Where, For the Regional meteorological characteristic data, For the Standard meteorological characteristic data, For the The mean value of the meteorological factors, For the The standard deviation of each meteorological factor; and the standard meteorological characteristic data after standardization has a mean of 0 and a standard deviation of 1, which is suitable for Gaussian distribution data.

[0104] Standard meteorological characteristic data and regional output data are sorted to obtain regional samples. Specifically, the standard meteorological characteristic data and regional output data are in the form of series. The rank of each data in the series, as well as the joint probability distribution and marginal probability distribution of the standard meteorological characteristic data and the regional output data can be counted. The ranks, joint probability distribution and marginal probability distribution are sorted to obtain regional samples, and the regional samples are input into the cross-correlation formula to calculate the correlation coefficient and mutual information.

[0105] Meteorological factors are screened based on correlation coefficient and mutual information to obtain key meteorological factors;

[0106] Specifically, the corresponding screening model is:

[0107] ;

[0108] Where, For the Meteorological factors The corresponding key indicators are is the weight coefficient, For the Meteorological factors With new energy output The correlation coefficient of For the Meteorological factors With new energy output mutual information.

[0109] Generate a time series data sequence corresponding to the new energy power generation device based on key meteorological factors, output power and meteorological data;

[0110] C1. Filtering meteorological data based on key meteorological factors to obtain key meteorological data, and filtering output power based on regional data clusters corresponding to the key meteorological factors to obtain key output power;

[0111] C2, integrating the time series of key meteorological data and the time series of key output power to obtain input data;

[0112] Specifically, missing values of key meteorological data are counted and filled to obtain complete key meteorological data, and the complete key meteorological data are normalized to obtain normalized meteorological data;

[0113] Counting missing values of key output power and filling them to obtain complete key output power, and normalizing the complete key output power to obtain normalized output power;

[0114] Align the time series of normalized meteorological data with the time series of normalized output power to obtain a unified time series;

[0115] Generating input data based on unified time series, normalized meteorological data, and normalized output power;

[0116] Simultaneously, similarity analysis is performed on key meteorological data to obtain meteorological similarity features, and similarity analysis is performed on the geographical locations of corresponding new energy power generation devices to obtain geographical similarity features;

[0117] C3, generating heterogeneous graphs based on input data, meteorological similarity features, and geographical similarity features;

[0118] C4. Extract the spatial embedding features in the heterogeneous graph based on the time series of the input data to obtain feature vectors, and organize and concatenate all the feature vectors to obtain the time series data sequence.

[0119] In this embodiment, meteorological data is screened based on key meteorological factors, and output power is screened based on regional data clusters corresponding to key meteorological factors. The screening process is to delete unnecessary data and optimize the volume and accuracy of the data to be processed. It should be noted that its purpose is to delete unnecessary data, which can be achieved by deleting the corresponding data that has been collected, and can also be achieved by deleting the corresponding data collection equipment. In order to solve the problems of dimensional differences and missing values in the key meteorological data and key output power, normalization formulas, missing value filling and other methods are used to integrate and obtain input data. Specifically, based on the periodicity of meteorological data, historical data of the same period are selected to fill the missing values of meteorological data, and normalization formulas are used to normalize them. Linear The interpolation method fills the missing values of output power according to the linear relationship of output power and normalizes it using the normalization formula; the time series of normalized meteorological data and the time series of normalized output power are aligned to obtain a unified time series, the same time series is segmented based on the sampling time interval, and the normalized meteorological data and normalized output power are sorted based on the segmented same time series to obtain input data; at the same time, in order to capture key influencing factors and regional shared factors, similarity analysis is performed on key meteorological data and geographical location respectively, and the deep relationship between meteorology, climate, geographical location, etc. and new energy power generation devices is analyzed. Then, based on the input data, meteorological similarity features and geographical similarity features, a heterogeneous graph that meets the requirements of the heterogeneous graph neural network is generated. The heterogeneous graph can be expressed as:

[0120] ;

[0121] Where, is a heterogeneous graph, The cluster site of the cluster dataset corresponding to the new energy power generation device, Due to multiple relationships The set of edges that form For the relationship A collection of is the normalized adjacency matrix of edges;

[0122] The essence of the similarity analysis is to calculate the weight of each edge. The weight can be calculated according to the similarity function. The similarity function is specifically:

[0123] ;

[0124] Where, For cluster data The characteristics of the corresponding node, and ,in For the After screening, the meteorological factors For the corresponding cluster data The characteristics of the node, For the corresponding cluster data The characteristics of the node, For cluster data Corresponding node and cluster data The similarity of the corresponding nodes, For cluster data Corresponding nodes are related A set of connected adjacent nodes.

[0125] The heterogeneous graph is input into the heterogeneous graph neural network. Specifically, based on the time series of the input data, the corresponding node feature matrix in the heterogeneous graph is input into the heterogeneous graph neural network. The spatial embedding features are extracted through the multi-layer graph convolution of the heterogeneous graph neural network to obtain a feature vector. The feature vector is output in the form of a matrix. The update formula corresponding to the multi-layer graph convolution is:

[0126] ;

[0127] Where, For the The node feature matrix of the layer, For the relationship The corresponding weight matrix, is the activation function;

[0128] Subsequently, the feature vectors of each time step are concatenated and reconstructed to form a time series data sequence.

[0129] Predict the future power of the corresponding new energy power generation device based on the time series data sequence;

[0130] Specifically, the hidden state is gradually updated based on the time series data sequence to obtain the hidden state sequence, and the hidden state sequence is used as the initial state to generate the prediction sequence;

[0131] The power output matrix is generated based on the prediction sequence to obtain the future power of the corresponding new energy power generation device.

[0132] In this embodiment, the time series data sequence is input into the time series prediction model. The encoder of the time series prediction model encodes the time series data sequence to implement the gradual update of the hidden state, thereby obtaining a hidden state sequence. The gradual update formula corresponding to the encoder is specifically:

[0133] ;

[0134] ;

[0135] ;

[0136] ;

[0137] ;

[0138] ;

[0139] Where, represents the encoder time step The corresponding forget gate, is the activation function, is the forget gate weight, is the time step The corresponding feature dimension, is the time step The corresponding time series data input, is the bias term of the forget gate, is the encoder time step The corresponding input gate, is the input gate weight, is the input gate bias term, is the encoder time step The corresponding candidate gate, is the activation function of the candidate gate, is the candidate gate weight, is the candidate gate bias term, is the encoder time step The corresponding status update, is the encoder time step The corresponding status update, is the encoder time step The corresponding output gate, is the output gate weight, is the output gate bias term, is the encoder time step The corresponding hidden state.

[0140] The hidden state sequence is used as the initial state of the time series prediction model decoder to gradually generate a prediction sequence in the future time period. The power output matrix formed by the prediction sequence is essentially the future power of the corresponding new energy power generation device. The prediction model corresponding to the decoder is specifically:

[0141] ;

[0142] ;

[0143] Where LSTM represents the decoder, is the time step The corresponding hidden sequence, is the time step The corresponding update state sequence, is the time step The corresponding decoder output is, is the time step The corresponding output is, is the fully connected mapping layer of the decoder, which is used to map the decoder hidden state to the predicted power value.

[0144] In addition, before formally predicting the joint power, it is necessary to use the mean square error function The entire predicted sequence is compared with the true value, and the parameters of the heterogeneous graph neural network and time series prediction model are updated based on the comparison results using the chain rule of back propagation. Specifically, the gradient descent algorithm is used for updating. The gradient descent algorithm is specifically as follows:

[0145] ;

[0146] ;

[0147] Where, is the parameter vector of the heterogeneous graph neural network or time series prediction model, is the time step The corresponding actual output power matrix, is the time step The corresponding predicted output power matrix, is the total number of time steps, is the learning rate, is the mean square error function Gradient with respect to the parameters.

[0148] This embodiment has at least the following substantial effects:

[0149] (1) This embodiment performs cluster analysis based on the spatial coordinates and time-series power generation data of new energy power generation devices to determine regional data clusters, classifying new energy power generation devices with similar geographical distribution and similar or even identical power generation modes into a unified category, achieving accurate classification of new energy power generation devices under complex nonlinear spatial relationships, and using the silhouette coefficient to optimize clustering quality and effectively filter out noise data;

[0150] (2) This embodiment calculates the correlation coefficient and mutual information based on the regional data cluster, output data and meteorological characteristic data. The correlation coefficient and mutual information are used to describe the highly nonlinear relationship and complex mapping relationship between weather factors and the power of new energy power generation devices. Based on the correlation coefficient and mutual information, meteorological factors are screened to obtain key meteorological factors, and weather factors that have little or no effect on the corresponding new energy power generation devices are eliminated, effectively improving the efficiency of power forecasting.

[0151] (3) This embodiment generates a time series data sequence corresponding to the new energy power generation device based on key meteorological factors, output power and meteorological data, obtains the power generation change of the new energy power generation device accompanied by weather factors in the time series, clarifies the power generation change of the new energy power generation device accompanied by weather factors in the time series, and predicts the future power of the corresponding new energy power generation device through the said change, thereby significantly improving the joint power prediction accuracy of distributed new energy power generation devices.

[0152] The above specific embodiments are preferred embodiments of the present invention and are not intended to limit the specific implementation scope of the present invention. The scope of the present invention includes but is not limited to the specific embodiments. All equivalent changes made in accordance with the shape, structure, and method of the present invention are within the scope of protection of the present invention.

Claims

1. A joint prediction method for distributed new energy power generation devices, characterized in that: The following steps are involved: Collect spatial coordinates, time-series power generation data, output data, meteorological characteristic data, output power, and meteorological data of new energy power generation devices; Cluster analysis is performed based on spatial coordinates and time series power generation data to determine regional data clusters; Calculate the correlation coefficient and mutual information based on the regional data cluster, output data, and meteorological characteristic data, which includes the following sub-steps: B1. Filtering the output data based on the regional data cluster to obtain regional output data, and filtering the meteorological characteristic data based on the regional data cluster to obtain regional meteorological characteristic data; B2. Perform standard processing on the regional meteorological characteristic data to obtain standard meteorological characteristic data, and organize the standard meteorological characteristic data and regional output data to obtain regional samples; B3. Calculate the correlation coefficient and mutual information based on regional samples and the cross-correlation formula; Meteorological factors are screened based on correlation coefficient and mutual information to obtain key meteorological factors; Generating a time series data sequence corresponding to a new energy power generation device based on key meteorological factors, output power, and meteorological data specifically includes the following sub-steps: C1. Filtering meteorological data based on key meteorological factors to obtain key meteorological data, and filtering output power based on regional data clusters corresponding to the key meteorological factors to obtain key output power; C2, integrating the time series of key meteorological data and the time series of key output power to obtain input data; Simultaneously, similarity analysis is performed on key meteorological data to obtain meteorological similarity features, and similarity analysis is performed on the geographical locations of corresponding new energy power generation devices to obtain geographical similarity features; C3, generating heterogeneous graphs based on input data, meteorological similarity features, and geographical similarity features; C4. Extract spatial embedding features from the heterogeneous graph based on the time series of the input data to obtain feature vectors, and then organize and concatenate all feature vectors to obtain a time series data sequence. Predict the future power of corresponding new energy power generation devices based on time series data.

2. The method for joint prediction of distributed renewable energy power generation devices according to claim 1, characterized in that: The specific process of performing cluster analysis based on spatial coordinates and time-series power generation data to determine regional data clusters is as follows: A1. Merge the spatial coordinates and time-series power generation data to obtain a cluster data set, and determine the distance threshold, minimum number of points, and time-series data weight based on the data characteristics of the cluster data set; A2. Perform weighted calculation based on the time series data weight and the cluster data set to obtain the cluster Euler distance, and construct a domain based on the cluster Euler distance and distance threshold; A3. Determine the core point based on the domain and the minimum number of points. If the number of data points in the domain is less than the minimum number of points, mark the cluster data corresponding to the domain as a normal point. If the number of data points in the domain is greater than or equal to the minimum number of points, mark the cluster data corresponding to the domain as a core point. A4. Clustering is performed based on core points and their corresponding areas to obtain regional data clusters.

3. The joint prediction method of a distributed new energy power generation device according to claim 2, characterized in that: In A1, the specific process of merging the spatial coordinate and time-series power generation data to obtain the cluster data set is as follows: The outliers in the time series power generation data are statistically removed based on the standard score method, and the missing values in the time series power generation data are filled in based on the linear interpolation method to obtain the complete time series power generation data; Normalizing the spatial coordinates to obtain normalized coordinates, and normalizing the complete time series power generation data to obtain normalized power generation data; The normalized coordinates and normalized power generation data are sorted and merged to obtain the cluster dataset.

4. The joint prediction method of a distributed new energy power generation device according to claim 2, characterized in that: After A4 is completed, the silhouette coefficient needs to be calculated to verify the regional data cluster. The corresponding silhouette coefficient formula is as follows: ; Where, For the The silhouette coefficient of the data points, For the The average distance from a data point to other data points in the same area data cluster, For the The average distance from a data point to other regional data clusters.

5. The joint prediction method of a distributed renewable energy power generation device according to claim 1, characterized in that: In B3, the cross-correlation formula is specifically: ; ; Where, The regional sample corresponds to Meteorological factors With new energy output The correlation coefficient of is the number of regional samples, For the The regional samples correspond to Meteorological factors rank, For the Regional samples corresponding to new energy output rank, The regional sample corresponds to Meteorological factors With new energy output The mutual information of For the Meteorological factors With new energy output The joint probability distribution of For the Meteorological factors The marginal probability distribution of Contribute to new energy The marginal probability distribution of .

6. The method for joint prediction of distributed renewable energy power generation devices according to claim 1, characterized in that: The screening model corresponding to the key meteorological factors obtained by screening meteorological factors based on correlation coefficient and mutual information is specifically: ; Where, For the Meteorological factors The corresponding key indicators are is the weight coefficient, For the Meteorological factors With new energy output The correlation coefficient of For the Meteorological factors With new energy output mutual information.

7. The joint prediction method of distributed renewable energy power generation devices according to claim 1, characterized in that: In C2, the specific process of integrating the time series of key meteorological data and the time series of key output power to obtain input data is as follows: Counting missing values of key meteorological data and filling them to obtain complete key meteorological data, and normalizing the complete key meteorological data to obtain normalized meteorological data; Counting missing values of key output power and filling them to obtain complete key output power, and normalizing the complete key output power to obtain normalized output power; Align the time series of normalized meteorological data with the time series of normalized output power to obtain a unified time series; The input data is generated based on unified time series, normalized meteorological data, and normalized output power.

8. The method for joint prediction of distributed renewable energy power generation devices according to claim 1, characterized in that: The specific process of predicting the future power of the corresponding new energy power generation device based on the time series data sequence is as follows: Based on the time series data sequence, the hidden state is gradually updated to obtain the hidden state sequence, and the hidden state sequence is used as the initial state to generate the prediction sequence; The power output matrix is generated based on the prediction sequence to obtain the future power of the corresponding new energy power generation device.

Citation Information

Patent Citations

  • High-precision wind-solar combined power prediction method

    CN119226702A

  • Photovoltaic short-term generation power prediction method based on weather clustering and LSTM combined model

    CN117791595A

  • Photovoltaic output prediction method and device based on multi-source data, and storage medium

    CN118399406A

  • Distributed photovoltaic power station generation power prediction method and device

    CN119482360A