Regional new energy power prediction method and device

The geographical and climatic sub-regions are divided by HDBSCAN and K-means algorithms, combined with the LSTM-XGBoost hybrid model, the weights are calculated dynamically and the total power prediction results are generated, which solves the problems of high computational complexity and low prediction accuracy in regional new energy power prediction, especially in cloudy weather, which significantly improves the prediction accuracy and efficiency.

CN120373525APending Publication Date: 2025-07-25CHAOYANG POWER SUPPLY COMPANY OF STATE GRID LIAONING ELECTRIC POWER SUPPLY +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510378181.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has high computational complexity in regional new energy power prediction, and it is difficult to take into account both spatial and temporal correlation and multi-type resource collaborative prediction. In addition, traditional clustering algorithms lead to unreasonable sub-region division, affecting prediction accuracy.

Method used

The HDBSCAN algorithm is used to divide the geoclimate sub-regions, and the proxy site group is screened through the K-means algorithm, combined with the LSTM-XGBoost hybrid model for prediction, dynamically calculate the weights and generate the total power prediction results of the region through the STCN network.

Benefits of technology

It reduces the computational complexity, improves prediction accuracy, especially in cloudy weather, reduces the impact of spatial heterogeneity, and improves the robustness and efficiency of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373525A_ABST
    Figure CN120373525A_ABST
Patent Text Reader

Abstract

The invention provides a regional new energy power prediction method and device. The method comprises the following steps: firstly, collecting static and dynamic data of a new energy station in a region; constructing a geographic climate feature vector based on the geographic position and the meteorological data, and dividing geographic climate sub-regions by using an HDBSCAN algorithm; constructing output feature vectors according to the dynamic data of the sub-regions, and screening agent station groups by using a K-means algorithm; calculating a dynamic weight in combination with a historical prediction error, an output characteristic and an installed capacity of the agent station; predicting the power of the proxy station by using the L-X hybrid model in combination with the dynamic weight, the weather and the power data; calculating the predicted power of each sub-region based on the predicted power of the proxy station and the dynamic weight; and finally, integrating the predicted powers and geographic positions of all the sub-regions to obtain region total power prediction. Spatial heterogeneity influence is reduced through hybrid clustering, weather and power association is captured by using an L-X model, and calculation complexity is reduced through a proxy station mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of power calculation, and particularly relates to a method and device for predicting the power of regional new energy sources. Background Art

[0002] Regional renewable energy power stations are widely distributed, and their power output is affected by many factors such as geographical location, meteorological conditions, and equipment types, resulting in high data dimensions and complex features.

[0003] Traditional prediction methods usually build models based on the historical power data of a single power station. However, in the face of regional aggregation scenarios, directly building a full-scale model will cause the computational complexity to increase exponentially. Although some studies have tried to simplify the input through feature dimensionality reduction or data fusion, these methods have not been fully adapted in the energy field, especially lacking the ability to dynamically model the spatio-temporal correlation between power stations.

[0004] In addition, existing regional power prediction models mostly rely on a single machine learning algorithm, making it difficult to balance the long-term dependence of time series and the non-linear relationship of features.

[0005] At the same time, when traditional clustering algorithms divide power station groups, they often lead to unreasonable sub-region division due to fixed preset cluster numbers or ignoring geographical and climatic differences, thereby affecting the prediction accuracy.

[0006] In addition, existing research mostly focuses on a single energy type, lacking support for collaborative prediction of multiple types of resources, and the weight allocation is often based on static indicators, ignoring the dynamic contributions of power stations. Summary of the Invention

[0007] The purpose of this application is to overcome the above-mentioned defects in the prior art and provide a method and device for predicting the power of regional new energy sources.

[0008] This application provides a method for predicting the power of regional new energy sources, including:

[0009] Obtain the static data and dynamic data of each new energy power station in the region, where the static data includes the geographical location coordinates and installed capacity of the power station, and the dynamic data includes historical meteorological data, historical power data, and historical prediction errors;

[0010] Based on the geographical location coordinates and the historical meteorological data, construct a geographical and climatic feature vector;

[0011] Based on the geographical and climatic feature vector, determine geographical and climatic sub-regions through the HDBSCAN algorithm;

[0012] According to the dynamic data in the geographical and climatic sub-regions, construct an output power feature vector;

[0013] Based on the output power feature vector, screen the proxy power station group through the K-means algorithm;

[0014] Calculate the dynamic weight coefficient based on the historical prediction error, output characteristic vector, and installed capacity of the proxy station group;

[0015] Construct an LSTM-XGBoost hybrid prediction model according to the dynamic weight coefficient, the historical meteorological data, and the historical power data, and generate the predicted power of the proxy station group;

[0016] Calculate the predicted power of each geographical climate sub-region based on the predicted power of the proxy station group and the dynamic weight coefficient;

[0017] Generate the regional total power prediction result according to the predicted power of all geographical climate sub-regions and the geographical location coordinates.

[0018] Optionally, the parameter configuration of the HDBSCAN algorithm includes:

[0019] The minimum cluster size of the HDBSCAN algorithm is set to 5;

[0020] The geographical climate characteristic vector includes the longitude and latitude of the station, the historical average irradiance, and the historical average wind speed.

[0021] Optionally, the K-means algorithm includes:

[0022] The number of clusters of the K-means algorithm analyzes the inflection point of the sum of squared errors within the cluster with the change of the number of clusters through the elbow method, and selects the optimal number of clusters.

[0023] Optionally, the hyperparameters of the dynamic weight coefficient are set as the installed capacity weight coefficient α = 0.4, the output correlation coefficient weight coefficient β = 0.3, and the historical prediction error weight coefficient γ = 0.3.

[0024] Optionally, the structure of the LSTM-XGBoost hybrid prediction model includes:

[0025] The LSTM network is configured with 2 hidden layers, 64 neurons in each layer, and the Dropout rate is 0.2;

[0026] The XGBoost model uses the Huber loss function and adds an L2 regularization term to suppress overfitting.

[0027] This application also provides a regional new energy power prediction device, including:

[0028] An acquisition module that acquires the static data and dynamic data of each new energy station in the region, where the static data includes the geographical location coordinates and installed capacity of the station, and the dynamic data includes historical meteorological data, historical power data, and historical prediction errors;

[0029] A climate module that constructs a geographical climate feature vector based on the geographical location coordinates and the historical meteorological data;

[0030] A partitioning module that determines geographical climate sub-regions based on the geographical climate feature vector through the HDBSCAN algorithm;

[0031] An output module that constructs an output feature vector according to the dynamic data within the geographical climate sub-region; a station group module that filters proxy station groups based on the output feature vector through the K-means algorithm;

[0032] A weight module that calculates dynamic weight coefficients based on the historical prediction errors, output feature vectors, and installed capacities of the proxy station groups;

[0033] A prediction module that constructs an LSTM-XGBoost hybrid prediction model according to the dynamic weight coefficients, the historical meteorological data, and the historical power data, and generates the predicted power of the proxy station groups;

[0034] A power module that calculates the predicted power of each geographical climate sub-region based on the predicted power of the proxy station groups and the dynamic weight coefficients;

[0035] A result module that generates a regional total power prediction result according to the predicted powers of all geographical climate sub-regions and the geographical location coordinates.

[0036] Optionally, the parameter configuration of the HDBSCAN algorithm includes:

[0037] The minimum cluster size of the HDBSCAN algorithm is set to 5;

[0038] The geographical climate feature vector includes the longitude and latitude of the station, the historical average irradiance, and the historical average wind speed.

[0039] Optionally, the K-means algorithm includes:

[0040] The number of clusters of the K-means algorithm analyzes the inflection point of the sum of squared errors within the cluster with the change of the number of clusters through the elbow method, and selects the optimal number of clusters.

[0041] Optionally, the hyperparameters of the dynamic weight coefficients are set as the installed capacity weight coefficient α = 0.4, the output correlation coefficient weight coefficient β = 0.3, and the historical prediction error weight coefficient γ = 0.3.

[0042] Optionally, the structure of the LSTM-XGBoost hybrid prediction model includes:

[0043] The LSTM network is configured with 2 hidden layers, each layer having 64 neurons, and the Dropout rate is 0.2;

[0044] The XGBoost model adopts the Huber loss function and adds an L2 regularization term to suppress overfitting.

[0045] The beneficial effects of this application are as follows:

[0046] This application provides a method for predicting the power of regional new energy, including: obtaining the static data and dynamic data of each new energy power station in the region, where the static data includes the geographical location coordinates and installed capacity of the power station, and the dynamic data includes historical meteorological data, historical power data, and historical prediction errors; constructing a geographical climate feature vector based on the geographical location coordinates and the historical meteorological data; determining geographical climate sub-regions through the HDBSCAN algorithm based on the geographical climate feature vector; constructing an output feature vector according to the dynamic data within the geographical climate sub-region; screening a proxy power station group through the K-means algorithm based on the output feature vector; calculating a dynamic weight coefficient based on the historical prediction errors, output feature vector, and installed capacity of the proxy power station group; constructing an LSTM-XGBoost hybrid prediction model according to the dynamic weight coefficient, the historical meteorological data, and the historical power data to generate the predicted power of the proxy power station group; calculating the predicted power of each geographical climate sub-region based on the predicted power of the proxy power station group and the dynamic weight coefficient; generating a regional total power prediction result according to the predicted power of all geographical climate sub-regions and the geographical location coordinates. Through the hybrid clustering mechanism, this application divides the power stations into more reasonable sub-regions, thereby reducing the impact of spatial heterogeneity on prediction. By adopting the LSTM-XGBoost hybrid prediction model, it integrates time series modeling and feature importance analysis, and more accurately captures the meteorological and power correlation in the time series. Through the proxy power station mechanism, this application screens representative power stations as proxy power stations according to the output characteristics within the sub-region, reducing the computational complexity. Brief Description of the Drawings

[0047] Figure 1 is a schematic diagram of the regional new energy power prediction process in this application;

[0048] Figure 2 is a schematic diagram of the K-means elbow rule curve in this application;

[0049] Figure 3 is a schematic diagram of the STCN network structure in this application;

[0050] Figure 4 is a schematic diagram of the comparison of prediction results under cloudy weather in this application. Detailed Embodiments

[0051] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it can be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, the embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0052] Please refer to Figure 1 As shown, the present application provides a method for predicting the power of new energy in a region, including:

[0053] S101. Obtain the static data and dynamic data of each new energy power station in the region. The static data includes the geographical location coordinates and installed capacity of the power station, and the dynamic data includes historical meteorological data, historical power data, and historical prediction errors.

[0054] The input data for regional aggregated power prediction includes: static data and dynamic data.

[0055] The static data includes:

[0056] 1. The geographical location coordinates of the power station (x i , y i ), where i = 1, 2,..., N (N is the total number of power stations).

[0057] 2. The installed capacity C i (unit: MW).

[0058] 3. The encoding of the equipment type (such as photovoltaic, wind power, energy storage) into a one-hot vector T i ∈ {0, 1} K (K is the total number of energy types).

[0059] The dynamic data includes:

[0060] 1. Historical power data (T is the historical time step).

[0061] 2. Historical meteorological data, including: meteorological observation data: irradiance G i,t , wind speed V i,t , temperature T i,t ; numerical weather prediction (NWP) data: meteorological prediction values for the next H hours

[0062] Preprocess the above data.

[0063] Missing value filling: Adopt the spatio-temporal collaborative interpolation method to fill in the missing meteorological or power data based on the spatio-temporal correlation of adjacent power stations. For the missing value X i,t of power station i at time t, its estimated value is:

[0064]

[0065] where is the set of geographically adjacent stations, is the time - proximity window, and the weights and are determined by the Gaussian kernel function.

[0066] Feature normalization: For dynamic data, quantile normalization is adopted to eliminate the dimensional differences of different stations. Let a certain feature sequence of station i be {x i,t}}, and its normalized value is:

[0067]

[0068] where Q p (x i ) represents the p - quantile of the feature data of station i.

[0069] S102. Based on the geographical location coordinates and the historical meteorological data, construct a geo - climate feature vector;

[0070] For each station i, extract the following geo - climate feature vector:

[0071]

[0072] where are the latitude and longitude coordinates, and the last two items are the historical average irradiance and wind speed.

[0073] S103. Based on the geo - climate feature vector, determine the geo - climate sub - regions through the HDBSCAN algorithm;

[0074] Adopt the density - based hierarchical clustering algorithm HDBSCAN, and the expression is as follows:

[0075] d mreach (i,j)=max{core k (i),core k (j),d(i,j)}

[0076] where core i (i) is the distance from point i to its k - th nearest neighbor, and d(i,j) is the Euclidean distance.

[0077] Through the minimum spanning tree (MST) to segment the clustering, finally generate M geo - climate sub - regions {S1,S2,…,S M}.

[0078] S104. Based on the dynamic data within the geo - climate sub - regions, construct a power output feature vector;

[0079] For station i ∈ S within sub - region S m m ​, extract the output feature vector:

[0080]

[0081] where: σ(P i ) is the standard deviation of the power sequence of substation i; ρ(P i , P total ) is the correlation coefficient between the power of substation i and the total regional power; the third term is the installed capacity ratio.

[0082] S105. Screen the proxy substation group through the K-means algorithm based on the output feature vector;

[0083] Perform K-means clustering on F i power , with the goal of minimizing the within-cluster sum of squares error, and the expression is as follows:

[0084]

[0085] The central substation of each cluster is selected as the proxy substation, and finally K m proxy substations are retained in each sub-region (K m << |S m |).

[0086] S106. Calculate the dynamic weight coefficient based on the historical prediction error, output feature vector and installed capacity of the proxy substation group;

[0087] For the proxy substation k in the sub-region S m , its weight w m,k consists of three parts:

[0088]

[0089] where: α, β, γ are adjustable hyperparameters (default α = 0.4, β = 0.3, γ = 0.3); MAE k is the historical prediction mean absolute error of substation k.

[0090] S107. Construct an LSTM-XGBoost hybrid prediction model according to the dynamic weight coefficient, the historical meteorological data and the historical power data, and generate the predicted power of the proxy substation group;

[0091] LSTM time series feature extraction:

[0092] Input: the meteorological sequence of proxy substation k and the historical power sequence

[0093] LSTM unit calculation:

[0094] f t = σ(W f · [h t-1 , x t + b f )

[0095] i t = σ(W i · [h t-1 , x t + b i )

[0096] ot = σ(W o · [h t-1 , x t + b o )

[0097] c t = f t ⊙ c t-1 + i t ⊙ tanh(W c · [h t-1 , x t + b c )

[0098] h. t = o t ⊙ tanh(c t )

[0099] where x t is the input feature at time step t, and h t is the hidden state.

[0100] XGBoost Feature Importance Weighting:

[0101] Input: LSTM hidden state h t and static feature T k ;

[0102] XGBoost Objective Function:

[0103]

[0104] where l is the Huber loss function, is the regularization term.

[0105] Output: Predicted power P of proxy station k k,t .

[0106] S108. Calculate the predicted power of each geographical and climatic sub-region based on the predicted power of the proxy station group and the dynamic weight coefficient;

[0107] The predicted power of the sub-region is the weighted sum of the predicted values of the proxy power stations:

[0108]

[0109] S109. Generate the total regional power prediction result based on the predicted power of all geographical and climatic sub-regions and the geographical location coordinates.

[0110] Input: The prediction sequences of all sub-regions and their geographical location matrices (Construct an adjacency matrix based on the HDBSCAN clustering result).

[0111] Alternately stack graph convolution (GCN) and temporal convolution (TCN):

[0112]

[0113] where is the adjacency matrix with self-connections, is the degree matrix, and W (l) is the trainable parameter.

[0114] Output: The predicted value P of the total regional power total,t .

[0115] In this application, the HDBSCAN algorithm is used to perform preliminary clustering on the power stations, and homogeneous sub-regions are divided based on geographical location and installed capacity to reduce the impact of geographical and climatic differences on prediction. For the power stations within each sub-region, secondary clustering is performed through the k-means algorithm, and a group of proxy power stations is selected according to historical output characteristics (such as output volatility and correlation coefficient with the total regional power) to participate in the prediction instead of all power stations.

[0116] Based on the installed capacity, historical prediction accuracy, and correlation coefficient with the total regional power of the proxy power stations, the weight coefficients are dynamically allocated.

[0117] For each sub-region, train the LSTM-XGBoost hybrid model: LSTM module: Capture the correlation between meteorology and power in the time series. XGBoost module: Optimize the feature importance ranking and improve the prediction robustness.

[0118] Sum the predicted powers of the proxy power stations in each sub-region by weight to obtain the predicted power of the sub-region. The spatio-temporal correlation between sub-regions is fused through the spatio-temporal convolutional network (STCN) to generate the total regional power prediction result.

[0119] Please refer to Figures 2 to 4As shown in the figure, in this embodiment, a certain area in East China is taken as an example. There are 50 distributed photovoltaic power stations in the target area, with a total installed capacity of 320 MW and a time resolution of 15 minutes. The data covers the whole year of 2023 and includes historical power, meteorological observation and numerical weather prediction (NWP) data.

[0120] Data source and parameter configuration:

[0121] Static data: The geographical location (latitude and longitude) of the power station is collected by a GPS device with an accuracy of ±0.001°; installed capacity range: 1 MW to 20 MW, discrete distribution; equipment type: monocrystalline silicon photovoltaic module (encoded as [1,0]), thin-film photovoltaic (encoded as [0,1]).

[0122] Dynamic data: Historical power data: Obtained from the SCADA system and filtered for outliers (removing outliers exceeding 120% of the installed capacity); meteorological data: irradiance (W / m 2 ), temperature (°C), cloud cover (%) observed values provided by the regional meteorological station; NWP data: 72-hour forecast from the European Centre for Medium-Range Weather Forecasts (ECMWF) with a spatial resolution of 0.1°×0.1°.

[0123] For missing irradiance data, weighted interpolation of neighboring power stations is used. Suppose the irradiance of power station i is missing at time t, and its interpolation formula is:

[0124]

[0125] where d ij is the Euclidean distance between power stations i and j, and D = 50 km is the attenuation radius.

[0126] Feature normalization: Power data is normalized by installed capacity:

[0127]

[0128] Meteorological data is normalized to the [0,1] interval using Min-Max normalization.

[0129] Hybrid clustering and proxy power station screening:

[0130] Geographical climate clustering (HDBSCAN):

[0131] Parameter settings: minimum cluster size (min_cluster_size) = 5; feature vector where is the annual average irradiance, is the annual average wind speed.

[0132] Clustering result: 50 power stations are divided into 4 sub-regions (S1 - S4), and the cluster sizes are 12, 15, 10, and 13 respectively.

[0133] Output feature secondary clustering (K-means):

[0134] Feature construction: For sub-region S1 (12 stations), extract feature vectors:

[0135] F i power = [σ(P i ), ρ(P i , P total ), C i

[0136] where σ(P i ) is calculated as the coefficient of variation of the power sequence (standard deviation / mean)

[0137] Determination of the number of clusters:

[0138] As Figure 2 shown, the elbow method is used to select the optimal number of clusters KK. When K = 3K = 3, the rate of decrease in the sum of squared errors within the cluster (SSE) slows down significantly.

[0139] Selection of proxy stations: For each cluster, select the 2 stations closest to the cluster center as proxy stations. The S1 sub-region finally retains 6 proxy stations (a 50% reduction from the original 12).

[0140] Dynamic weighting and model training

[0141] For the proxy stations k1 - k6 in the S1 sub-region, the dynamic weights are calculated as follows:

[0142]

[0143] Ck: The installed capacity of k1 is 18MW (total capacity of the S1 sub-region is 45MW), so the first term is 0.4 x (18 / 45) = 0.16; Pk: The correlation coefficient between k1 and the total regional power is 0.92, so the second term is 0.3 x 0.92 = 0.276; MAEk: The historical prediction MAE of k1 is 1.2MW, so the third term is 0.3 x (1 / 1.2) = 0.25; The total weight ω1,k1 = 0.16 + 0.276 + 0.25 = 0.686 (adjusted to 0.228 after normalization)

[0144] LSTM-XGBoost hybrid model training:

[0145] LSTM network structure, input layer: 3 features (normalized irradiance, temperature, cloud cover); hidden layer: 2 layers of LSTM, 64 neurons in each layer; Dropout rate: 0.2.

[0146] ​Output layer: Power prediction for the next 4 hours (16 time steps).

[0147] Upscaling prediction and aggregation:

[0148] Sub-region power prediction: For sub-region S1, the predicted powers of proxy stations k1 - k6 are weighted and summed:

[0149]

[0150] STCN global aggregation:

[0151] Adjacency matrix construction: Based on the HDBSCAN clustering results, define the adjacency relationship between sub-regions:

[0152]

[0153] where d(S m , S n ) is the distance between the centroids of sub-regions.

[0154] Spatio-temporal convolutional layer design:

[0155] Temporal convolution: 1D convolution kernel size = 3, stride = 1, output channels = 32; Spatial convolution; Number of channels in the graph convolutional layer (GCN) = 32, activation function: ReLU; Output layer: Fully connected layer mapping to the total regional power.

[0156] Verification and result analysis

[0157] Experimental settings:

[0158] Training set: Data from January to October 2023; Test set: Data from November to December 2023; Comparison benchmarks: Traditional full-scale LSTM, fixed clustering + XGBoost.

[0159] Performance metrics:

[0160]

[0161]

[0162] Under cloudy weather (irradiance fluctuation > 200 W / m2 / h), the MAE is reduced by 26.7% compared with the benchmark method; In sunny scenarios, the error is stable at 4.2 - 5.1 MW, and the volatility is reduced by 40%.

[0163] The proxy station mechanism reduces the training data volume by 58%, and the hybrid model parallel training strategy shortens the iteration time by 32%.

[0164] This application also provides a regional new energy power prediction device, including:

[0165] An acquisition module that acquires the static data and dynamic data of each new energy power station in the area. The static data includes the geographical location coordinates and installed capacity of the power station, and the dynamic data includes historical meteorological data, historical power data, and historical prediction errors.

[0166] A climate module that constructs a geographical climate feature vector based on the geographical location coordinates and the historical meteorological data.

[0167] A division module that determines geographical climate sub-regions based on the geographical climate feature vector through the HDBSCAN algorithm.

[0168] An output module that constructs an output feature vector according to the dynamic data within the geographical climate sub-region; a station group module that filters the proxy power station group based on the output feature vector through the K-means algorithm.

[0169] A weight module that calculates the dynamic weight coefficient based on the historical prediction errors, output feature vectors, and installed capacity of the proxy power station group.

[0170] A prediction module that constructs an LSTM-XGBoost hybrid prediction model according to the dynamic weight coefficient, the historical meteorological data, and the historical power data, and generates the predicted power of the proxy power station group.

[0171] A power module that calculates the predicted power of each geographical climate sub-region based on the predicted power of the proxy power station group and the dynamic weight coefficient.

[0172] A result module that generates the regional total power prediction result according to the predicted power of all geographical climate sub-regions and the geographical location coordinates.

[0173] Further, the parameter configuration of the HDBSCAN algorithm includes:

[0174] The minimum cluster size of the HDBSCAN algorithm is set to 5.

[0175] The geographical climate feature vector includes the longitude and latitude of the power station, the historical average irradiance, and the historical average wind speed.

[0176] Further, the K-means algorithm includes:

[0177] The number of clusters of the K-means algorithm analyzes the inflection point of the sum of squared errors within the cluster with the change of the number of clusters through the elbow method, and selects the optimal number of clusters.

[0178] Further, the hyperparameters of the dynamic weight coefficient are set as the installed capacity weight coefficient α = 0.4, the output correlation coefficient weight coefficient β = 0.3, and the historical prediction error weight coefficient γ = 0.3.

[0179] Furthermore, the structure of the LSTM-XGBoost hybrid prediction model includes:

[0180] The LSTM network is configured with 2 hidden layers, each layer having 64 neurons, and a Dropout rate of 0.2;

[0181] The XGBoost model uses the Huber loss function and adds an L2 regularization term to suppress overfitting.

[0182] The above description of the embodiments is to enable those of ordinary skill in the art to understand and apply the present invention. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art based on the disclosure of the present invention should fall within the protection scope of the present invention.

Claims

1. A method for predicting the power of new energy in a region, characterized in that, Including: Obtain the static data and dynamic data of each new energy power station in the area. The static data includes the geographical location coordinates and installed capacity of the power station, and the dynamic data includes historical meteorological data, historical power data, and historical prediction errors; Based on the geographical location coordinates and the historical meteorological data, construct a geographical climate feature vector; Based on the geographical climate feature vector, determine geographical climate sub-regions through the HDBSCAN algorithm; According to the dynamic data within the geographical climate sub-regions, construct an output feature vector; Based on the output feature vector, screen the proxy power station group through the K-means algorithm; Based on the historical prediction errors, output feature vectors, and installed capacity of the proxy power station group, calculate the dynamic weight coefficient; According to the dynamic weight coefficient, the historical meteorological data, and the historical power data, construct an LSTM-XGBoost hybrid prediction model to generate the predicted power of the proxy power station group; Based on the predicted power of the proxy power station group and the dynamic weight coefficient, calculate the predicted power of each geographical climate sub-region; According to the predicted power of all geographical climate sub-regions and the geographical location coordinates, generate the regional total power prediction result.

2. The method according to claim 1, wherein The parameter configuration of the HDBSCAN algorithm includes: The minimum cluster size of the HDBSCAN algorithm is set to 5; The geographical climate feature vector includes the longitude and latitude of the power station, historical average irradiance, and historical average wind speed.

3. The method according to claim 1, characterized in that The K-means algorithm includes: For the K-means algorithm, the number of clusters is analyzed by the elbow method to find the inflection point of the change of the sum of squared errors within the cluster with the number of clusters, and the optimal number of clusters is selected.

4. The method according to claim 1, characterized in that, The hyperparameters of the dynamic weight coefficient are set as the installed capacity weight coefficient α = 0.4, the output-related coefficient weight coefficient β = 0.3, and the historical prediction error weight coefficient γ = 0.

3.

5. The method according to claim 1, characterized in that, The structure of the LSTM-XGBoost hybrid prediction model includes: The LSTM network is configured with 2 hidden layers, 64 neurons in each layer, and a Dropout rate of 0.2; The XGBoost model uses the Huber loss function and adds an L2 regularization term to suppress overfitting.

6. A regional new energy power prediction device, characterized in that Including: An acquisition module that obtains the static data and dynamic data of each new energy power station in the area. The static data includes the geographical location coordinates and installed capacity of the power station, and the dynamic data includes historical meteorological data, historical power data, and historical prediction errors; A climate module that constructs a geographical climate feature vector based on the geographical location coordinates and the historical meteorological data; A division module that determines geographical climate sub-regions through the HDBSCAN algorithm based on the geographical climate feature vector; An output module that constructs an output feature vector according to the dynamic data within the geographical climate sub-regions; A power station group module that screens the proxy power station group through the K-means algorithm based on the output feature vector; A weight module that calculates the dynamic weight coefficient based on the historical prediction errors, output feature vectors, and installed capacity of the proxy power station group; A prediction module, based on the dynamic weight coefficient, the historical meteorological data, and the historical power data, constructs an LSTM-XGBoost hybrid prediction model to generate the predicted power of the proxy station group; A power module, based on the predicted power of the proxy station group and the dynamic weight coefficient, calculates the predicted power of each geographical climate sub-region; A result module, based on the predicted power of all geographical climate sub-regions and the geographical location coordinates, generates a regional total power prediction result.

7. The device according to claim 6, wherein The parameter configuration of the HDBSCAN algorithm includes: The minimum cluster size of the HDBSCAN algorithm is set to 5; The geographical climate feature vector includes the longitude and latitude of the station, the historical average irradiance, and the historical average wind speed.

8. The device according to claim 6, characterized in that The K-means algorithm includes: For the K-means algorithm, the number of clusters is determined by analyzing the inflection point of the sum of squared errors within clusters with respect to the number of clusters using the elbow method to select the optimal number of clusters.

9. The device according to claim 6, characterized in that, The hyperparameters of the dynamic weight coefficient are set as the installed capacity weight coefficient α = 0.4, the output-related coefficient weight coefficient β = 0.3, and the historical prediction error weight coefficient γ = 0.

3.

10. The device according to claim 6, characterized in that, The structure of the LSTM-XGBoost hybrid prediction model includes: The LSTM network is configured with 2 hidden layers, each layer having 64 neurons, and a Dropout rate of 0.2; The XGBoost model uses the Huber loss function and adds an L2 regularization term to suppress overfitting.

Citation Information

Cited By

  • Method for predicting output of new energy station

    CN120744385A

  • A method for predicting the output of new energy power plants

    CN120744385B

  • Wind power plant power prediction method and device based on space-time collaboration and storage medium

    CN121529519A