Photovoltaic power prediction method and device based on dual-mode hybrid analysis, and electronic equipment

By combining adaptive noise-complete ensemble empirical mode decomposition and cluster analysis with the prediction results of BiLSTM and SVR models, the accuracy problem of photovoltaic power prediction under complex meteorological conditions is solved, and more efficient photovoltaic power generation prediction is achieved.

CN120933924APending Publication Date: 2025-11-11SHAOGUAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511075620.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing photovoltaic power prediction technologies are ineffective under complex weather conditions, making it difficult to accurately predict photovoltaic power generation.

Method used

A dual-mode hybrid analysis approach is adopted, which decomposes photovoltaic power data into multiple sets of intrinsic mode functions through adaptive noise complete set empirical mode decomposition and cluster analysis. BiLSTM model and SVR model are used for prediction respectively, and the prediction results are fused to improve accuracy and reliability.

Benefits of technology

It improves the accuracy and reliability of photovoltaic power prediction, can handle complex photovoltaic power variation patterns, reduces the uncertainty of single model prediction, and enhances the stability of prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120933924A_ABST
    Figure CN120933924A_ABST
Patent Text Reader

Abstract

The invention provides a photovoltaic power prediction method and device based on dual-mode hybrid analysis and electronic equipment, and relates to the technical field of power systems and photovoltaic power generation. The method comprises the following steps: acquiring a current photovoltaic power sequence; carrying out adaptive noise complete set empirical mode decomposition on the current photovoltaic power sequence to obtain multiple groups of intrinsic mode functions; based on a preset clustering algorithm and a Bayesian information criterion, performing clustering processing on the plurality of groups of intrinsic mode functions to obtain a plurality of data sub-pools; for each data sub-pool, inputting the plurality of intrinsic mode functions in the data sub-pool into a pre-trained BiLSTM model and a pre-trained SVR model to obtain a first prediction result and a second prediction result; performing fusion processing on the first prediction result and the second prediction result to obtain a target prediction result; and determining a photovoltaic power prediction sequence according to the target prediction results of all the data sub-pools. The method is used for improving the accuracy of photovoltaic power prediction and the reliability of a prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of power systems and photovoltaic power generation technology, and in particular to a photovoltaic power prediction method, device and electronic equipment based on dual-mode hybrid analysis. Background Technology

[0002] With increasing environmental awareness and growing demand for sustainable energy, renewable energy is gradually replacing traditional fossil fuels and becoming the mainstream of energy development. Photovoltaic energy, as a clean, efficient, safe, and large-scale renewable energy source, is seeing its share in the power grid increase year by year, playing a significant role in optimizing the energy structure and ensuring energy security. However, photovoltaic power generation is significantly affected by weather conditions; fluctuations in weather patterns can lead to instability in photovoltaic power generation, posing a significant challenge to the power grid. Therefore, accurately predicting the power generation of photovoltaic power plants has become one of the core issues for the industry's development.

[0003] Currently, photovoltaic power prediction technologies mainly fall into three categories: using regression models or neural networks to predict photovoltaic power, using satellite remote sensing technology to predict photovoltaic power, and using numerical weather prediction and other methods to describe atmospheric conditions to predict photovoltaic power.

[0004] However, these photovoltaic power prediction technologies still suffer from poor prediction results under complex weather conditions. Summary of the Invention

[0005] The photovoltaic power prediction method, apparatus, and electronic equipment based on dual-mode hybrid analysis provided in this application are intended to address the problem that existing photovoltaic power prediction technologies still produce poor prediction results under complex weather conditions.

[0006] In a first aspect, embodiments of this application provide a photovoltaic power prediction method based on dual-mode hybrid analysis, comprising:

[0007] Obtain the current photovoltaic power sequence, which includes multiple daily photovoltaic power data.

[0008] Adaptive noise-complete set empirical mode decomposition is performed on the current photovoltaic power sequence to obtain multiple sets of intrinsic mode functions. Each set of intrinsic mode functions includes multiple intrinsic mode functions decomposed from a single day's photovoltaic power data.

[0009] Based on a pre-defined clustering algorithm and Bayesian information criteria, multiple sets of intrinsic mode functions are clustered to obtain multiple data sub-pools. Each data sub-pool includes: a cluster label for a cluster, all daily photovoltaic power data corresponding to the cluster label, and multiple intrinsic mode functions corresponding to each daily photovoltaic power data.

[0010] For each data sub-pool, multiple intrinsic mode functions from the data sub-pool are input into the pre-trained BiLSTM model and SVR model respectively to obtain the first prediction result and the second prediction result.

[0011] The first and second prediction results are fused to obtain the target prediction result.

[0012] Based on the target prediction results of all data sub-pools, a photovoltaic power prediction sequence is determined.

[0013] In one possible implementation, the preset clustering algorithm is the X-means algorithm; based on the preset clustering algorithm and the Bayesian information criterion, multiple sets of intrinsic mode functions are clustered to obtain multiple data sub-pools, including: extracting features from multiple sets of intrinsic mode functions to obtain a target feature matrix; and clustering the target feature matrix based on the X-means algorithm and the Bayesian information criterion to obtain multiple data sub-pools.

[0014] In one possible implementation, feature extraction is performed on multiple sets of intrinsic mode functions to obtain a target feature matrix, including: for each set of intrinsic mode functions, feature extraction is performed on multiple intrinsic mode functions corresponding to the decomposition of daily photovoltaic power data to obtain energy features; based on the daily average photovoltaic irradiance and energy features, a feature sub-matrix corresponding to the daily photovoltaic power data is obtained; and based on the feature sub-matrix corresponding to all daily photovoltaic power data, a target feature matrix is ​​obtained.

[0015] In one possible implementation, clustering is performed on the target feature matrix based on the X-means algorithm and Bayesian information criterion to obtain multiple data sub-pools. This includes: initializing the number of clusters; dividing each cluster into two sub-clusters to obtain the first sub-cluster center and the second sub-cluster center; determining the first criterion value of the cluster, the second criterion value of the first sub-cluster center, and the third criterion value of the second sub-cluster center based on the Bayesian information criterion; when the second criterion value is less than the first criterion value, the first sub-cluster center is taken as the new cluster center; and when the third criterion value is less than the first criterion value, the second sub-cluster center is taken as the new cluster center; based on the updated cluster centers, the updated number of clusters is obtained, and the steps of dividing each cluster into two sub-clusters to obtain the first sub-cluster center and the second sub-cluster center are repeated until the updated number of clusters meets the preset iteration conditions, thereby obtaining the target number of clusters and multiple data sub-pools.

[0016] In one possible implementation, after clustering multiple sets of intrinsic mode functions based on a preset clustering algorithm and Bayesian information criteria to obtain multiple data sub-pools, the method further includes: for each data sub-pool, performing frequency division on multiple intrinsic mode functions in each daily photovoltaic power data to obtain high-frequency intrinsic mode functions and mid-to-low-frequency intrinsic mode functions for each daily photovoltaic power data; determining the optimization parameters of variational mode decomposition based on the whale swarm algorithm and the high-frequency intrinsic mode functions of each daily photovoltaic power data; and, according to the optimization parameters, processing the high-frequency intrinsic mode functions of each daily photovoltaic power data... Variational mode decomposition is performed to obtain multiple narrowband sub-modes. The low- and mid-frequency intrinsic mode functions and multiple narrowband sub-modes of the daily photovoltaic power data are normalized to obtain the target mode data of the data sub-pool. Accordingly, for each data sub-pool, the multiple intrinsic mode functions in the data sub-pool are input into the pre-trained BiLSTM model and SVR model to obtain the first prediction result and the second prediction result, including: for each data sub-pool, the target mode data in the data sub-pool are input into the pre-trained BiLSTM model and SVR model to obtain the first prediction result and the second prediction result.

[0017] In one possible implementation, the first prediction result and the second prediction result are fused to obtain the target prediction result, including: fusion processing of the first prediction result and the second prediction result based on a fusion coefficient to obtain the target prediction result, wherein the fusion coefficient is obtained by particle swarm optimization and is used to determine the contribution weight of each model in the target prediction result.

[0018] In one possible implementation, before inputting multiple intrinsic mode functions (EMFs) from each data sub-pool into a pre-trained BiLSTM model and an SVR model to obtain a first prediction result and a second prediction result, the method further includes: acquiring historical photovoltaic power sequences, which include multiple daily photovoltaic power data acquired within a preset historical time period; performing adaptive noise complete set empirical mode decomposition (AMD) on the historical photovoltaic power sequences to obtain multiple sets of historical EMFs; clustering the multiple sets of historical EMFs based on a preset clustering algorithm and Bayesian information criteria to obtain multiple historical data sub-pools; constructing a training set and a validation set based on all historical data sub-pools; and initializing a particle swarm based on a particle swarm optimization (PSO) algorithm, wherein each particle includes the first model parameters of the BiLSTM model. The system uses the second model parameters of the SVR model and initial fusion coefficients to fuse the prediction results of the BiLSTM and SVR models. For each particle, the BiLSTM and SVR models are trained separately based on the training set and the model parameters in the particle to obtain the first and second historical prediction results. Based on the initial fusion coefficients in the particle, the first and second historical prediction results are fused to obtain the target historical prediction result. Based on the target historical prediction result and the corresponding validation samples in the validation set, the loss function value is obtained. Based on the loss function values ​​of all particles, the particle swarm is optimized until the preset loss function converges. Based on the optimal particle in the particle swarm, the trained BiLSTM and SVR models and the fusion coefficients are determined.

[0019] Secondly, embodiments of this application provide a photovoltaic power prediction device based on dual-mode hybrid analysis, comprising:

[0020] The acquisition module is used to acquire the current photovoltaic power sequence, which includes multiple daily photovoltaic power data.

[0021] The decomposition module is used to perform adaptive noise complete set empirical mode decomposition on the current photovoltaic power sequence to obtain multiple sets of intrinsic mode functions. Each set of intrinsic mode functions includes multiple intrinsic mode functions decomposed from a single day's photovoltaic power data.

[0022] The processing module is used to perform clustering processing on multiple sets of intrinsic mode functions based on a preset clustering algorithm and Bayesian information criteria to obtain multiple data sub-pools. Each data sub-pool includes: a cluster label of a cluster, all daily photovoltaic power data corresponding to the cluster label, and multiple intrinsic mode functions corresponding to each daily photovoltaic power data.

[0023] The prediction module is used to input multiple intrinsic mode functions from each data sub-pool into the pre-trained BiLSTM model and SVR model to obtain the first prediction result and the second prediction result.

[0024] The fusion module is used to fuse the first prediction result and the second prediction result to obtain the target prediction result;

[0025] The determination module is used to determine the photovoltaic power prediction sequence based on the target prediction results of all data sub-pools.

[0026] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;

[0027] The memory stores instructions that the computer executes;

[0028] The processor executes computer execution instructions stored in memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0029] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed, are used to implement the first aspect and / or various possible implementations of the first aspect.

[0030] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed, implements the first aspect and / or various possible implementations of the first aspect.

[0031] The photovoltaic power prediction method, apparatus, and electronic device based on dual-mode hybrid analysis provided in this application acquire the current photovoltaic power sequence and perform adaptive noise complete ensemble empirical mode decomposition to reveal the intrinsic characteristics of the data. Then, a preset clustering algorithm and Bayesian information criterion are used to cluster the multiple sets of intrinsic mode functions obtained from the decomposition into multiple data sub-pools. For each data sub-pool, a pre-trained BiLSTM model and an SVR model are used for prediction. Finally, the prediction results of the two models are fused to obtain the target prediction result and determine the photovoltaic power prediction sequence accordingly. By utilizing adaptive noise complete ensemble empirical mode decomposition and cluster analysis, the intrinsic characteristics and patterns of photovoltaic power data can be revealed, grouping data with similar characteristics into one category, providing a foundation for subsequent targeted prediction. The dual-model prediction strategy can fully utilize the advantages of the BiLSTM and SVR models to improve prediction accuracy. Simultaneously, fusing the prediction results of the BiLSTM and SVR models comprehensively considers the prediction information of both models, reducing the uncertainty of single-model prediction and improving the reliability of the prediction results. Attached Figure Description

[0032] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0033] Figure 1 A flowchart illustrating the photovoltaic power prediction method based on dual-mode hybrid analysis provided in this application. Figure 1 ;

[0034] Figure 2 A flowchart illustrating the photovoltaic power prediction method based on dual-mode hybrid analysis provided in this application. Figure 2 ;

[0035] Figure 3 A schematic diagram of the photovoltaic power prediction device based on dual-mode hybrid analysis provided in this application;

[0036] Figure 4 A schematic diagram of the structure of the electronic device provided in this application.

[0037] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0038] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0039] First, let me explain the terms used in this application:

[0040] IXWV: refers to a comprehensive data processing framework based on ICEEMDAN-X-means-WOA-VMD. ICEEMDAN is an adaptive noise complete set empirical mode decomposition, an advanced method for signal processing that can adaptively decompose nonlinear and non-stationary signals and effectively extract intrinsic mode functions (IMFs) from signals.

[0041] X-means is an improved version of the K-means clustering algorithm. It can automatically determine the optimal number of clusters and improve the clustering effect by optimizing the clustering process.

[0042] WOA (Whale Optimization Algorithm) is an optimization algorithm inspired by the behavior of whale groups in nature. It is used to solve global optimization problems and is characterized by fast convergence speed and strong optimization ability.

[0043] Variational Mode Decomposition (VMD) is a non-stationary signal processing technique that decomposes complex signals into a series of modal components with specific center frequencies, which helps in further signal analysis and processing.

[0044] BIC (Bayesian Information Criterion): This refers to the Bayesian Information Criterion, a scoring criterion used for model selection. It is used to evaluate the merits of different clustering models, for example, to determine whether a cluster should be split into two sub-clusters.

[0045] PSO (Particle Swarm Optimization) is an optimization algorithm based on swarm intelligence, inspired by the foraging behavior of flocks of birds or schools of fish. In PSO, each particle represents a potential solution in the problem space, and through information sharing and cooperation among particles, the entire particle swarm is guided towards the optimal solution. The PSO algorithm has advantages such as simplicity, few parameters, and fast convergence.

[0046] BiLSTM (Bidirectional Long Short-Term Memory): refers to a bidirectional long short-term memory network, which is a deep learning model that combines long short-term memory networks and bidirectional processing mechanisms. It can simultaneously consider the forward and backward information of sequence data, capture long-term dependencies in the sequence, and improve the accuracy and robustness of sequence modeling.

[0047] SVR (Support Vector Regression) is a regression analysis method based on Support Vector Machines (SVM) used to solve regression problems. It works by finding an optimal hyperplane to fit the data points, minimizing the sum of the distances from all data points to this hyperplane (within a certain error range). Its goal is to find a function on a given dataset that can predict the value of the target variable as accurately as possible.

[0048] In existing technologies, while regression models and neural networks perform well under stable weather conditions, their predictive effectiveness is less than ideal when the weather changes rapidly. When using satellite remote sensing technology for photovoltaic power prediction, there is usually a certain error between satellite data and actual ground-based meteorological data. When using methods such as numerical weather prediction to describe atmospheric conditions for photovoltaic power prediction, the highly chaotic nature of atmospheric conditions makes it difficult to obtain accurate prediction results. Therefore, existing photovoltaic power prediction technologies still suffer from poor prediction results under complex meteorological conditions.

[0049] To address the aforementioned issues, this application provides a photovoltaic power prediction method, apparatus, and electronic device based on dual-mode hybrid analysis. The method acquires the current photovoltaic power sequence and performs adaptive noise complete set empirical mode decomposition (EMD). Then, it uses a pre-defined clustering algorithm and Bayesian information criteria to cluster the multiple sets of intrinsic mode functions obtained from the decomposition into multiple data sub-pools. This reveals the inherent characteristics and patterns of photovoltaic power data, grouping data with similar characteristics into one category, providing a foundation for subsequent targeted predictions. For each data sub-pool, a pre-trained BiLSTM model and an SVR model are used for prediction, fully utilizing the advantages of both models to improve prediction accuracy. The prediction results from the two models are fused to obtain the target prediction result, which is then used to determine the photovoltaic power prediction sequence. This comprehensively considers the prediction information of both models, reducing the uncertainty of single-model predictions and improving the reliability of the prediction results. Therefore, this method improves the accuracy of photovoltaic power prediction and enhances the reliability of the prediction results through refined data processing and a dual-model fusion strategy. It can handle complex photovoltaic power variation patterns and has broad applicability and scalability.

[0050] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0051] The photovoltaic power prediction method based on dual-mode hybrid analysis provided in this application can be executed by a computing device such as a server or server cluster. The server can be a mobile phone, computer, tablet, or other similar device. This application embodiment does not impose any particular restrictions on the implementation method of the execution subject, as long as the execution subject can obtain the current photovoltaic power sequence, which includes multiple daily photovoltaic power data; perform adaptive noise complete set empirical mode decomposition on the current photovoltaic power sequence to obtain multiple sets of intrinsic mode functions, wherein each set of intrinsic mode functions includes multiple intrinsic mode functions corresponding to a single day photovoltaic power data; based on a preset clustering algorithm and Bayesian information criterion, perform clustering processing on the multiple sets of intrinsic mode functions to obtain multiple data sub-pools, wherein each data sub-pool includes: a cluster label of a cluster, and all the single day photovoltaic power data corresponding to the cluster label and multiple intrinsic mode functions corresponding to each single day photovoltaic power data; for each data sub-pool, input the multiple intrinsic mode functions in the data sub-pool into a pre-trained BiLSTM model and an SVR model respectively to obtain a first prediction result and a second prediction result; perform fusion processing on the first prediction result and the second prediction result to obtain a target prediction result; determine the photovoltaic power prediction sequence based on the target prediction results of all data sub-pools. In one example, the method can be applied to the field of new energy power prediction; further, it can be applied to photovoltaic power generation systems, which include computing devices such as servers or server clusters.

[0052] Figure 1 A flowchart illustrating the photovoltaic power prediction method based on dual-mode hybrid analysis provided in this application. Figure 1 The execution entity of this method can be a server storing a photovoltaic power prediction method based on dual-mode hybrid analysis or other servers. This embodiment does not impose any particular limitations. Figure 1 As shown, the method may include:

[0053] S101. Obtain the current photovoltaic power sequence, which includes multiple daily photovoltaic power data.

[0054] The current photovoltaic power sequence can be a sequence of power data output by a photovoltaic power generation system arranged in chronological order within a specific time period, including multiple daily photovoltaic power data points. In some examples, these data points can be collected in real time by power measurement equipment (such as meters, sensors, etc.) installed in the photovoltaic power station, and can include power readings at hourly, 15-minute, or shorter time intervals.

[0055] Optionally, after data collection, the collected data needs to be initially screened and cleaned to remove outliers and noisy data; the cleaned data should be organized into a sequence according to time order to facilitate subsequent processing.

[0056] S102. Perform adaptive noise complete set empirical mode decomposition on the current photovoltaic power sequence to obtain multiple sets of intrinsic mode functions. Each set of intrinsic mode functions includes multiple intrinsic mode functions decomposed from a single day's photovoltaic power data.

[0057] In this step, each intrinsic mode function (IMF) represents a specific frequency component in the signal. Through adaptive noise complete set empirical mode decomposition, the complex variation patterns in the photovoltaic power sequence can be decomposed into multiple simple and easy-to-process components, which helps to reveal the intrinsic characteristics and laws of the data.

[0058] In some examples, adaptive white noise is added during the decomposition process to improve its stability and accuracy. The resulting IMF can reflect different dynamic characteristics in the photovoltaic power sequence, such as periodic variations and trends. Furthermore, the decomposed IMF can be visualized for a more intuitive understanding of the data's features.

[0059] S103. Based on the preset clustering algorithm and Bayesian information criterion, multiple sets of intrinsic mode functions are clustered to obtain multiple data sub-pools. Each data sub-pool includes: a cluster label of a cluster, all daily photovoltaic power data corresponding to the cluster label, and multiple intrinsic mode functions corresponding to each daily photovoltaic power data.

[0060] Furthermore, based on a pre-defined clustering algorithm and the Bayesian information criterion, intrinsic mode functions with similar characteristics are grouped into one category for targeted prediction in the future. Clustering simplifies data analysis and improves prediction efficiency. In practical applications, clustering algorithms such as K-means and X-means can be used. The Bayesian information criterion evaluates the clustering effect by considering the model's degrees of freedom and likelihood function, selecting the optimal number of clusters as the number that minimizes the Bayesian information criterion value.

[0061] S104. For each data sub-pool, input the multiple intrinsic mode functions in the data sub-pool into the pre-trained BiLSTM model and SVR model respectively to obtain the first prediction result and the second prediction result.

[0062] In this step, the IMF from each data sub-pool is fed into two pre-trained prediction models: a BiLSTM model to capture long-term dynamic characteristics and an SVR model to capture local smoothness characteristics.

[0063] Furthermore, the IMFs in each data sub-pool are formatted to suit the input model. For BiLSTM models, the data typically needs to be formatted as time series, while for SVR models, the data needs to be formatted as feature vectors (simultaneously, IMFs can be standardized to improve the model's prediction accuracy). The formatted IMFs are then input into the BiLSTM model. Through its bidirectional structure, BiLSTM can simultaneously consider past and future information of the data, thereby generating the first prediction result for the time series, which is a prediction of the future trend of the input IMFs. At the same time, the IMFs are input as feature vectors into the SVR model. SVR finds the optimal hyperplane, performs regression analysis, and outputs the second prediction result, which is a regression prediction based on the input IMFs.

[0064] S105. The first prediction result and the second prediction result are fused to obtain the target prediction result.

[0065] Fusion processing refers to combining the prediction results of multiple models to obtain more accurate predictions. For example, weighted averaging, voting mechanisms, or other fusion strategies can be used to combine the prediction results of BiLSTM and SVR to obtain more accurate predictions. Furthermore, post-processing, such as smoothing and outlier detection, can be applied to the fused results to improve the reliability of the predictions.

[0066] S106. Based on the target prediction results of all data sub-pools, determine the photovoltaic power prediction sequence.

[0067] Among them, the photovoltaic power prediction sequence refers to a sequence that reflects the trend of photovoltaic power change over a period of time, which is composed of the target prediction results of multiple data sub-pools. This sequence is the final output of the prediction results and provides an important reference for the operation and scheduling of photovoltaic power generation systems.

[0068] In some examples, the prediction results of all data sub-pools are aggregated to form a complete photovoltaic power prediction sequence. Aggregation can be achieved through simple concatenation, weighted averaging, or other more complex strategies. For example, the target prediction results of all data sub-pools can be combined into a sequence in chronological order; if a soft clustering method (such as fuzzy C-means) was used during clustering, the prediction results can be weighted and aggregated based on the membership degrees obtained from the clustering process; in hard clustering scenarios, the arithmetic mean of the prediction results of all samples within a sub-pool can be used to aggregate the prediction results.

[0069] The photovoltaic power prediction method based on dual-mode hybrid analysis provided in this application improves the accuracy and reliability of photovoltaic power prediction through steps such as adaptive noise complete ensemble empirical mode decomposition, cluster analysis, dual-model prediction, and result fusion. This method can handle complex photovoltaic power variation patterns, reveal the inherent characteristics and laws of the data, and provide strong support for the operation and scheduling of photovoltaic power generation systems. Simultaneously, by combining pre-trained BiLSTM and SVR models and employing flexible fusion strategies, the advantages of both models can be fully utilized, reducing the uncertainty of single-model prediction and improving the stability and reliability of the prediction results.

[0070] Based on the above embodiments, the preset clustering algorithm is the X-means algorithm; the method described in S103 for clustering multiple sets of intrinsic mode functions based on the preset clustering algorithm and Bayesian information criterion to obtain multiple data sub-pools may include: extracting features from multiple sets of intrinsic mode functions to obtain a target feature matrix; and clustering the target feature matrix based on the X-means algorithm and Bayesian information criterion to obtain multiple data sub-pools.

[0071] The X-means algorithm can automatically determine the optimal number of clusters without prior specification. It finds the optimal clustering structure by iteratively splitting and evaluating clusters, using the Bayesian Information Criterion (BIC) to guide the selection of the number of clusters.

[0072] Feature extraction is the process of extracting key information that represents the characteristics of raw data, essentially transforming IMFs (Integrated Factors) into feature vectors suitable for clustering. Feature extraction can include various methods, such as calculating the statistical properties (mean, variance, etc.), time-domain properties (duration, rate of change, etc.), and frequency-domain properties (spectral distribution, dominant frequency, etc.) of IMFs, or extracting energy features, entropy features, etc. These features should adequately reflect the differences and similarities between IMFs. For example, the energy percentage of each IMF can be calculated as a feature.

[0073] Furthermore, during clustering, the target feature matrix is ​​first input into the X-means algorithm. The algorithm attempts different numbers of clusters and calculates the BIC value for each number of clusters. Then, the number of clusters that minimizes the BIC value is selected as the optimal number of clusters, and IMFs with similar characteristics are grouped into the same data sub-pool. Each data sub-pool will contain a cluster label for one cluster and all IMFs corresponding to that cluster label.

[0074] For example, feature extraction is performed on each group of IMFs, and their energy percentage is calculated as a feature; the features of all IMFs are combined into a target feature matrix; the target feature matrix is ​​input into the X-means algorithm; the algorithm starts with the initially assumed number of clusters (e.g., 2 clusters), attempts to split and evaluate the clusters; according to the BIC criterion, the algorithm automatically determines the optimal number of clusters (e.g., 5 clusters); IMFs with similar characteristics are grouped into the same data sub-pool, resulting in 5 data sub-pools.

[0075] By using the X-means algorithm and feature extraction, automatic clustering of multiple intrinsic mode functions can be achieved. This method can automatically determine the optimal number of clusters based on the inherent characteristics of the data, improving the accuracy and flexibility of clustering. This allows IMFs with similar characteristics to be grouped into the same data sub-pool, providing a more accurate and targeted data foundation for subsequent prediction modeling. This helps to optimize model performance and improve the accuracy and reliability of predictions.

[0076] Based on the above embodiments, the method for extracting features from multiple sets of intrinsic mode functions to obtain a target feature matrix may include: for each set of intrinsic mode functions, extracting features from multiple intrinsic mode functions corresponding to the decomposition of daily photovoltaic power data to obtain energy features; obtaining feature sub-matrix corresponding to daily photovoltaic power data based on the daily average irradiance and energy features; and obtaining the target feature matrix based on the feature sub-matrix corresponding to all daily photovoltaic power data.

[0077] The energy characteristics are extracted from multiple intrinsic mode functions (IMFs) derived from the daily photovoltaic power data and are used to characterize the energy level of each IMF. Energy characteristics can be obtained by calculating the sum of squares or root mean square values ​​of the IMFs, reflecting the distribution of the signal across different frequency components. Furthermore, the energy proportion of each IMF—the ratio of its energy to the total energy—can be considered to further characterize the importance of the IMFs in the overall signal.

[0078] The average daily irradiance of a photovoltaic system refers to the average energy of solar radiation emanating from a unit area within a day, usually expressed in watts per square meter. Daily average irradiance data can be obtained through weather stations or irradiance sensors integrated into the photovoltaic power generation system. The acquired irradiance data is then averaged to obtain the daily average value.

[0079] In one example, for each IMF group, the root mean square value of each IMF is calculated as an energy feature; the energy share of each IMF is calculated, i.e., the ratio of the energy of each IMF to the total energy; the daily average solar irradiance data is obtained; the daily average solar irradiance, the root mean square value of each IMF, and the energy share are combined to form a feature submatrix. For example, if there are 3 IMFs, the feature submatrix might be a 4-dimensional vector (including irradiance, IMF1 energy feature, IMF2 energy feature, and IMF3 energy feature); the feature submatrixes corresponding to all daily solar power data are combined to form a complete target feature matrix. The number of rows in this matrix equals the number of daily solar power data points, and the number of columns equals the dimension of each feature submatrix.

[0080] Furthermore, if needed, other features related to photovoltaic power, such as daily maximum irradiance, daily minimum irradiance, and irradiance variation rate, can be extracted and added to the feature submatrix. The construction method of the feature submatrix can be adjusted according to specific needs and data availability.

[0081] By incorporating the daily average solar irradiance into the feature submatrix, the intrinsic characteristics and external influencing factors of solar power data can be captured more comprehensively. This provides the model with more comprehensive input information, enabling it to consider not only the intrinsic characteristics of historical power data but also the impact of external environmental factors when making predictions. Furthermore, by combining irradiance data, the model can better understand the relationship between solar power output and environmental conditions. This additional information helps the model make more accurate predictions under different weather conditions, especially in situations of drastic weather changes.

[0082] Based on the above embodiments, a method for clustering the target feature matrix to obtain multiple data sub-pools based on the X-means algorithm and Bayesian information criterion may include: initializing the number of clusters; for each cluster, dividing the cluster into two sub-clusters to obtain the first sub-cluster center and the second sub-cluster center; based on the Bayesian information criterion, determining the first criterion value of the cluster, the second criterion value of the first sub-cluster center, and the third criterion value of the second sub-cluster center; when the second criterion value is less than the first criterion value, taking the first sub-cluster center as the new cluster center; and when the third criterion value is less than the first criterion value, taking the second sub-cluster center as the new cluster center; based on the updated cluster centers, obtaining the updated number of clusters, and re-executing the steps of dividing the cluster into two sub-clusters for each cluster to obtain the first sub-cluster center and the second sub-cluster center, until the updated number of clusters meets the preset iteration conditions, thereby obtaining the target number of clusters and multiple data sub-pools.

[0083] In this embodiment, an initial assumption about the number of clusters needs to be given before clustering begins. This number can be determined based on experience, data characteristics, or simple trial and error. For example, one can try starting with a smaller number of clusters (such as 2 or 3), or use some heuristics (such as the elbow rule) to estimate a reasonable initial number of clusters.

[0084] Furthermore, each current cluster is further divided into two subclusters, and the centroids (means) of these two subclusters are calculated. The split points can be selected randomly or based on a distance metric (such as Euclidean distance). The subcluster centers can be obtained by calculating the mean of all points within each subcluster.

[0085] The BIC value (i.e., the criterion value) can be obtained by calculating the likelihood function and the number of parameters of the model. For each cluster and its subclusters, the corresponding BIC value is calculated. The smaller the BIC value, the better the clustering structure achieves between fitting the data and complexity. Therefore, if the BIC value of a subcluster is less than the BIC value of the original cluster, it means that the split is beneficial, and the subcluster center is retained as the new cluster center.

[0086] Iteration conditions can be set according to specific needs and data characteristics. For example, a maximum number of iterations can be set, or iteration can stop when the BIC value changes less than a certain threshold after several consecutive iterations. The final number of clusters is the target number of clusters, and each cluster corresponds to a data sub-pool.

[0087] For example, assuming the initial number of clusters is 2; the two current clusters are split into four sub-clusters (each original cluster is split into two sub-clusters); the center point of the four sub-clusters is calculated; the BIC value (first criterion value) of the original two clusters is calculated; the BIC value of the four sub-clusters (second criterion value and third criterion value, BIC value of each original cluster corresponding to the two sub-clusters) is calculated; for each original cluster, the BIC value of its sub-clusters is compared with the BIC value of the original cluster; if the BIC value of the sub-cluster is smaller, the center of the sub-cluster is selected as the new cluster center; based on the updated cluster center, the splitting and BIC value calculation steps are re-executed; this process is repeated until the preset iteration conditions are met (such as reaching the maximum number of iterations or the change in BIC value is less than a threshold); the final number of clusters is the target number of clusters, and each cluster corresponds to a data sub-pool.

[0088] Clustering using the X-means algorithm automatically adjusts the number of clusters based on the inherent characteristics of the data, finding the optimal clustering scheme through iterative segmentation and evaluation of the cluster structure. This improves the accuracy and flexibility of clustering, providing a more accurate and targeted data foundation for subsequent photovoltaic power prediction modeling. This clustering process allows for a better understanding and analysis of the inherent characteristics of photovoltaic power data, optimizing the performance of prediction models, and ultimately improving the operational efficiency and energy management level of photovoltaic power plants.

[0089] Based on the above embodiments, after clustering multiple sets of intrinsic mode functions (IMFs) using a preset clustering algorithm and Bayesian information criterion as described in S103 to obtain multiple data sub-pools, the method may further include: for each data sub-pool, performing frequency division on multiple IMFs in each daily photovoltaic power data to obtain high-frequency IMFs and mid-to-low-frequency IMFs for each daily photovoltaic power data; determining optimization parameters for variational mode decomposition based on the whale swarm algorithm and the high-frequency IMFs of each daily photovoltaic power data; and, according to the optimization parameters, performing frequency division on the high-frequency IMFs of each daily photovoltaic power data. The intrinsic mode functions are subjected to variational mode decomposition to obtain multiple narrowband sub-modes. The low- and mid-frequency intrinsic mode functions and multiple narrowband sub-modes of the daily photovoltaic power data are normalized to obtain the target mode data of the data sub-pool. Accordingly, for each data sub-pool, the multiple intrinsic mode functions in the data sub-pool are input into the pre-trained BiLSTM model and SVR model to obtain the first prediction result and the second prediction result, including: for each data sub-pool, the target mode data in the data sub-pool are input into the pre-trained BiLSTM model and SVR model to obtain the first prediction result and the second prediction result.

[0090] In this embodiment, frequency partitioning involves dividing the IMF into high-frequency and mid-to-low-frequency components based on its frequency characteristics. High-frequency IMFs typically contain rapidly changing signal components, while mid-to-low-frequency IMFs contain slower-changing signal components. In some examples, the high-frequency and mid-to-low-frequency components are determined by the energy percentage of the IMF, where IMFs with an energy percentage exceeding 2% are considered high-frequency intrinsic mode functions.

[0091] The whale pod algorithm updates the positions of candidate solutions by simulating whale behaviors such as predation, encirclement, and bubble-web formation. When searching for VMD optimization parameters (such as the number of modes K and the penalty factor α), the parameter space can be regarded as the search range of the whales, and the optimal combination of parameters is found through algorithm iteration.

[0092] Furthermore, a pod of whales is initialized, with each whale representing a possible parameter combination (K, α). An optimization objective function is defined, for example, a bi-objective function of minimizing bandwidth and sample entropy. For each whale (i.e., parameter combination), VMD is used to decompose the high-frequency intrinsic mode function, resulting in multiple narrowband sub-modes. The sum of the bandwidths and the sum of the sample entropies of the decomposed sub-modes are calculated as the value of the bi-objective function. A smaller sum of bandwidth indicates a more concentrated sub-mode in the frequency domain, while a smaller sum of sample entropies indicates a smoother sub-mode in the time domain. Based on the value of the bi-objective function, the fitness of each whale is evaluated. Fitness can be calculated using fitness functions in multi-objective optimization algorithms, such as weighted sums or Pareto dominance. The whales' positions and velocities are updated to gradually converge the pod towards the optimal solution. During the iteration process, the whales simulate behaviors such as predation, encirclement, and bubble netting to explore the parameter space and find the optimal solution. When the maximum number of iterations is reached or other stopping conditions are met, the optimal parameter combination (K, α) is obtained. This parameter combination will enable the sub-modes after VMD decomposition to achieve optimal bandwidth and sample entropy at the same time.

[0093] By employing the whale swarm algorithm to search for optimal VMD parameters with the dual objectives of minimizing bandwidth and sample entropy, it is possible to ensure that the sub-modes after VMD decomposition are more concentrated in the frequency domain and smoother in the time domain. This helps improve the accuracy and reliability of subsequent prediction models, because purer and more discernible sub-modes can better reflect the intrinsic characteristics of photovoltaic power data.

[0094] Normalization can be performed by scaling the data to a specific range (e.g., [0,1] or [-1,1]) to eliminate the influence of different dimensions and orders of magnitude on the prediction model. This can be achieved through linear or nonlinear transformations. For example, min-max normalization can be used to scale the data to the [0,1] range. For low-to-mid-frequency IMFs and narrowband submodes, normalization can be performed separately to obtain the target mode data. Finally, the normalized target mode data is input into pre-trained BiLSTM and SVR models to obtain the prediction results. BiLSTM is used to capture long-term dynamic characteristics, and SVR is used to capture local smoothing characteristics.

[0095] By employing frequency partitioning, whale swarm algorithm optimization of VMD parameters, VMD decomposition, and normalization, more accurate decomposition and processing of IMFs within each data sub-pool can be achieved, resulting in cleaner and more discernible input data. By inputting this target modal data into pre-trained BiLSTM and SVR models, more accurate long-term dynamics and local smoothing characteristics can be captured, thereby improving the accuracy and reliability of photovoltaic power prediction.

[0096] Based on the above embodiments, the method described in S105 for fusing the first prediction result and the second prediction result to obtain the target prediction result may include: fusing the first prediction result and the second prediction result based on the fusion coefficient to obtain the target prediction result, wherein the fusion coefficient is obtained by particle swarm optimization and is used to determine the contribution weight of each model in the target prediction result.

[0097] The fusion coefficient is specifically used to determine the weights of the BiLSTM and SVR models in the target prediction result. It can be a number between 0 and 1, representing the proportion of each model in the final prediction result. For example, if the fusion coefficient of the BiLSTM model is 0.7 and the fusion coefficient of the SVR model is 0.3, then the final prediction result will be 70% of the prediction result of the BiLSTM model plus 30% of the prediction result of the SVR model.

[0098] The fusion coefficients obtained through particle swarm optimization enable a more accurate fusion of the prediction results from the BiLSTM and SVR models. This allows for better utilization of the advantages of different models, improving overall prediction performance and thus enhancing the accuracy of photovoltaic power prediction.

[0099] Based on the above embodiments, before inputting multiple intrinsic mode functions from each data sub-pool into the pre-trained BiLSTM model and SVR model to obtain the first prediction result and the second prediction result, as described in S104, the method may further include: acquiring historical photovoltaic power sequences, which include multiple daily photovoltaic power data acquired within a preset historical time period; performing adaptive noise complete set empirical mode decomposition on the historical photovoltaic power sequences to obtain multiple sets of historical intrinsic mode functions; performing clustering processing on the multiple sets of historical intrinsic mode functions based on a preset clustering algorithm and Bayesian information criteria to obtain multiple historical data sub-pools; constructing a training set and a validation set based on all historical data sub-pools; and initializing a particle swarm based on a particle swarm optimization algorithm, wherein each particle includes the first [presumably a component] of the BiLSTM model. The model parameters, the second model parameters of the SVR model, and the initial fusion coefficients are used to fuse the prediction results of the BiLSTM model and the SVR model. For each particle, the BiLSTM model and the SVR model are trained separately according to the training set and the model parameters in the particle to obtain the first historical prediction result and the second historical prediction result. Based on the initial fusion coefficients in the particle, the first historical prediction result and the second historical prediction result are fused to obtain the target historical prediction result. Based on the target historical prediction result and the corresponding validation samples in the validation set, the loss function value is obtained. Based on the loss function values ​​of all particles, the particle swarm is optimized until the preset loss function converges. Based on the optimal particle in the particle swarm, the trained BiLSTM model and SVR model, as well as the fusion coefficients, are determined.

[0100] In this embodiment, the historical photovoltaic power sequence refers to photovoltaic power data collected within a certain time period in the past. This data is used to train and validate the prediction model. The data processing methods for the historical photovoltaic power sequence, such as adaptive noisy complete ensemble empirical mode decomposition and clustering, are similar to those for the current photovoltaic power sequence in the above embodiments, and will not be repeated here.

[0101] The training set is used to train the model, and the validation set is used to evaluate the model's performance. Historical data sub-pools can be randomly allocated to the training and validation sets, or they can be divided chronologically to ensure the representativeness of the training and validation sets.

[0102] Initializing the particle swarm can refer to creating a set of particles, each representing a set of possible BiLSTM model parameters (such as the number of hidden units, number of layers, learning rate, etc.), SVR model parameters (such as kernel function type, penalty factor, etc.), and initial fusion coefficients.

[0103] The fusion process can be a simple weighted average, and the loss function can be mean squared error, mean absolute error, etc. In the particle swarm optimization process, particles gradually converge towards the optimal solution. The update rule of the PSO algorithm can be used to adjust the position and velocity of the particles to find the combination of model parameters and fusion coefficients that minimizes the loss function.

[0104] By combining the PSO algorithm, this method enables joint training and optimization of BiLSTM and SVR models. It automatically adjusts model parameters and fusion coefficients based on historical data to find the model combination that minimizes prediction error. Furthermore, joint training and optimization better leverage the advantages of different models, improving overall prediction performance and adapting to the needs of different photovoltaic power plants and data characteristics.

[0105] Figure 2 A flowchart illustrating the photovoltaic power prediction method based on dual-mode hybrid analysis provided in this application. Figure 2 ,like Figure 2 As shown, in this embodiment... Figure 1 Based on the examples, taking a photovoltaic power generation system, such as a photovoltaic electric field or photovoltaic power station, as an example, the photovoltaic power prediction method based on dual-mode hybrid analysis is described in detail. This method includes:

[0106] Step 1: Use ICEEMDAN for noise reduction, and then use X-means to adaptively segment homogeneous day patterns.

[0107] First, ICEEMDAN is used to perform one-time noise reduction and empirical mode decomposition on the original photovoltaic power sequence to obtain multiple IMFs (such as IMF1, IMF2, ..., IMFn). For each daily sample, IMF energy features are extracted and the average irradiance of the day is added to form a feature matrix. Then, X-means is used to automatically determine the number of clusters according to the BIC criterion, and these features are clustered to form several homogeneous data sub-pools (such as SE1, SE2, ..., SEn) for subsequent modeling.

[0108] Further, step one includes: (1) ICEEMDAN injects adaptive white noise into the photovoltaic historical power data (or current photovoltaic power data, or photovoltaic / load sequence, etc.) and iterates EMD to output several IMFs; (2) calculates the IMF energy vector and photovoltaic daily average amount of each sample to obtain the feature matrix; (3) uses X-means (BIC automatically selects the number of clusters of k-clustering) to divide the feature space into daily clusters, obtains the membership degree μ(c), and outputs k homogeneous data sub-pools.

[0109] Step 2: Perform VMD with WOA dual-objective optimization on the high-frequency IMF of each cluster.

[0110] Within each sub-pool, only high-frequency IMFs with an energy percentage exceeding 2% are selected for secondary processing; the remaining IMFs are directly retained. For the high-frequency portion, a whale swarm algorithm is used to optimize within a given parameter space, ensuring that the decomposed modes simultaneously possess the narrowest bandwidth and lowest entropy. Based on the search results, VMD is performed to decompose the high-frequency signal into non-overlapping narrowband sub-modes (such as imf1, imf2, ..., imf9), and normalization is performed on each sub-mode to provide a cleaner and more discernible input for the prediction model.

[0111] Further, step two includes: (1) Based on the homogeneous data sub-pool obtained in step one, only high-frequency IMFs with energy greater than 2% are selected for secondary processing, and the rest are directly retained; (2) The Whale Algorithm (WOA) is used to search for (K,α) with the dual objectives of "minimum bandwidth and minimum sample entropy"; (3) Finally, VMD is performed to obtain non-overlapping narrowband sub-modes, and each sub-mode is normalized separately.

[0112] Step 3: Use a BiLSTM and SVR dual model based on PSO optimization for prediction.

[0113] For each cluster subset, BiLSTM is used to capture long-term dynamics, while SVR captures local smoothness. Particle Swarm Optimization (PSO) searches for hyperparameters of the BiLSTM, such as hidden units, number of layers, and learning rate, as well as the fusion coefficients of the two models, and iteratively updates the data using validation set error as the evaluation criterion. Finally, the two outputs of each sub-mode are superimposed with optimal weights (i.e., the fusion coefficients determined by PSO weighted ensemble) to obtain (predicted value 1, predicted value 2, ..., predicted value n). These predicted values ​​are then aggregated at the cluster level (or weighted by membership degree for soft clustering) to obtain the entire power prediction curve.

[0114] Further, step three includes: (1) For each cluster subset, BiLSTM is used to capture long-term dynamics and SVR is used to capture local smoothness; (2) Particle swarm optimization (PSO) searches for the hidden units, number of layers, learning rate and other hyperparameters of BiLSTM, as well as the fusion coefficients of the two models, and iteratively updates them with the validation set error as the evaluation criterion; (3) The two outputs of each sub-mode are superimposed with the optimal weights and then summarized at the cluster level (for soft clustering, the weights are based on membership degree) to obtain the entire power prediction curve.

[0115] The photovoltaic power prediction method based on dual-mode hybrid analysis provided in this application integrates three originally separate signal processing tasks—noise reduction, homogenization, and narrowband decomposition—into a continuous process using IXWV: First, ICEEMDAN is used to suppress spikes and endpoint pseudomodes; then, X-means is used to automatically classify homogeneous day patterns such as weekdays, weekends, or sunny / cloudy days based on BIC; finally, VMD fine segmentation with WOA optimization is performed only on the high-frequency components that still retain peak energy. The entire framework has adaptive capabilities, enabling it to find suitable cluster numbers and decomposition parameters without repeated manual parameter trials; in terms of computational efficiency, by focusing on high-frequency components rather than the entire signal to complete the secondary decomposition, resource consumption is significantly reduced; structurally, the narrowband modes are orthogonal to each other and have ordered energy, allowing direct input to any backend model such as BiLSTM or Transformer, achieving algorithm decoupling; at the generalization level, the framework exhibits robust error suppression and training acceleration effects in distributed photovoltaic scenarios.

[0116] By training and optimizing the BiLSTM-SVR dual model using PSO, the complementary characteristics of temporal deep learning and kernel regression are incorporated into a unified prediction module. First, the BiLSTM is trained in parallel on the same subset of data to capture long-range dependencies and nonlinear dynamics; simultaneously, the SVR is trained to characterize local smoothness and minor fluctuations. The Particle Swarm Optimization (PSO) algorithm performs a dual function: searching for key structures and hyperparameters such as the learning rate of the BiLSTM, and simultaneously adjusting the fusion weights of the BiLSTM and SVR to achieve optimal alignment of the two outputs on the validation set. This global optimization mechanism allows the model to adaptively migrate to different data segments without repeated manual parameter tuning; the complementarity of the two networks effectively avoids performance imbalances in peak or plateau segments inherent in single-model approaches.

[0117] Therefore, the method provided in this application, while ensuring controllable inference delay, also possesses both long-period sensitivity and short-time noise suppression capabilities, thus maintaining stable prediction accuracy in distributed photovoltaic scenarios. This method can effectively improve the accuracy of power prediction for distributed photovoltaic power plants, providing a safer and more reliable guarantee for photovoltaic grid connection; at the same time, by providing prediction accuracy and stability, it enables new energy systems to better cope with the volatility and uncertainty of new energy sources, thereby ensuring a reliable energy supply.

[0118] Figure 3 The schematic diagram of the photovoltaic power prediction device based on dual-mode hybrid analysis provided in this application is as follows: Figure 3 As shown, the photovoltaic power prediction device 30 based on dual-mode hybrid analysis provided in this embodiment includes:

[0119] The acquisition module 301 is used to acquire the current photovoltaic power sequence, which includes multiple daily photovoltaic power data.

[0120] The decomposition module 302 is used to perform adaptive noise complete set empirical mode decomposition on the current photovoltaic power sequence to obtain multiple sets of intrinsic mode functions, wherein each set of intrinsic mode functions includes multiple intrinsic mode functions decomposed corresponding to a single day's photovoltaic power data;

[0121] The processing module 303 is used to perform clustering processing on multiple sets of intrinsic mode functions based on a preset clustering algorithm and Bayesian information criteria to obtain multiple data sub-pools. Each data sub-pool includes: a cluster label of a cluster, all daily photovoltaic power data corresponding to the cluster label, and multiple intrinsic mode functions corresponding to each daily photovoltaic power data.

[0122] The prediction module 304 is used to input multiple intrinsic mode functions in each data sub-pool into the pre-trained BiLSTM model and SVR model respectively to obtain the first prediction result and the second prediction result.

[0123] The fusion module 305 is used to fuse the first prediction result and the second prediction result to obtain the target prediction result;

[0124] The determination module 306 is used to determine the photovoltaic power prediction sequence based on the target prediction results of all data sub-pools.

[0125] In one possible implementation, the processing module 303 can also be used to: extract features from multiple sets of intrinsic mode functions to obtain a target feature matrix; and perform clustering processing on the target feature matrix based on the X-means algorithm and Bayesian information criterion to obtain multiple data sub-pools.

[0126] In one possible implementation, the processing module 303 can also be used to: extract features from multiple intrinsic mode functions corresponding to the decomposition of daily photovoltaic power data for each set of intrinsic mode functions to obtain energy features; obtain the feature sub-matrix corresponding to the daily photovoltaic power data based on the daily average irradiance and energy features; and obtain the target feature matrix based on the feature sub-matrix corresponding to all daily photovoltaic power data.

[0127] In one possible implementation, the processing module 303 can also be used to: initialize the number of clusters for clustering; for each cluster, divide the cluster into two sub-clusters to obtain the first sub-cluster center and the second sub-cluster center; based on the Bayesian information criterion, determine the first criterion value of the cluster, and the second criterion value and the third criterion value of the second sub-cluster center; when the second criterion value is less than the first criterion value, take the first sub-cluster center as the new cluster center; and when the third criterion value is less than the first criterion value, take the second sub-cluster center as the new cluster center; based on the updated cluster centers, obtain the updated number of clusters for clustering, and re-execute the steps of dividing the cluster into two sub-clusters for each cluster to obtain the first sub-cluster center and the second sub-cluster center, until the updated number of clusters meets the preset iteration conditions, thereby obtaining the target number of clusters and multiple data sub-pools.

[0128] In one possible implementation, the processing module 303 can also be used to: for each data sub-pool, perform frequency division on multiple intrinsic mode functions in each daily photovoltaic power data to obtain high-frequency intrinsic mode functions and mid-to-low-frequency intrinsic mode functions of each daily photovoltaic power data; determine the optimization parameters of variational mode decomposition based on the whale swarm algorithm and the high-frequency intrinsic mode functions of each daily photovoltaic power data; perform variational mode decomposition on the high-frequency intrinsic mode functions of each daily photovoltaic power data according to the optimization parameters to obtain multiple narrowband sub-modes; normalize the mid-to-low-frequency intrinsic mode functions and multiple narrowband sub-modes of the daily photovoltaic power data to obtain the target mode data of the data sub-pool; correspondingly, for each data sub-pool, input multiple intrinsic mode functions in the data sub-pool to a pre-trained BiLSTM model and an SVR model respectively to obtain a first prediction result and a second prediction result, including: for each data sub-pool, input the target mode data in the data sub-pool to a pre-trained BiLSTM model and an SVR model respectively to obtain a first prediction result and a second prediction result.

[0129] In one possible implementation, the fusion module 305 can also be used to: fuse the first prediction result and the second prediction result based on the fusion coefficient to obtain the target prediction result, wherein the fusion coefficient is obtained by particle swarm optimization and is used to determine the contribution weight of each model in the target prediction result.

[0130] In one possible implementation, the processing module 303 can also be used to: acquire historical photovoltaic power sequences, which include multiple daily photovoltaic power data acquired within a preset historical time period; perform adaptive noise complete set empirical mode decomposition on the historical photovoltaic power sequences to obtain multiple sets of historical intrinsic mode functions; perform clustering processing on the multiple sets of historical intrinsic mode functions based on a preset clustering algorithm and Bayesian information criteria to obtain multiple historical data sub-pools; construct training sets and validation sets based on all historical data sub-pools; initialize a particle swarm based on a particle swarm optimization algorithm, wherein each particle includes the first model parameters of the BiLSTM model, the second model parameters of the SVR model, and initial fusion coefficients, the initial fusion coefficients being used for... The prediction results of the BiLSTM model and the SVR model are fused. For each particle, the BiLSTM model and the SVR model are trained separately according to the training set and the model parameters in the particle to obtain the first historical prediction result and the second historical prediction result. Based on the initial fusion coefficients in the particle, the first historical prediction result and the second historical prediction result are fused to obtain the target historical prediction result. The loss function value is obtained according to the target historical prediction result and the corresponding validation samples in the validation set. Based on the loss function values ​​of all particles, the particle swarm is optimized until the preset loss function converges. Based on the optimal particle in the particle swarm, the trained BiLSTM model and SVR model, as well as the fusion coefficients, are determined.

[0131] The photovoltaic power prediction device based on dual-mode hybrid analysis provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0132] Figure 4 A schematic diagram of the structure of the electronic device provided in this application. Figure 4 As shown, the electronic device 40 provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the device 40 further includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.

[0133] In a specific implementation, at least one processor 401 executes computer execution instructions stored in memory 402, causing at least one processor 401 to perform the above-described method.

[0134] The specific implementation process of processor 401 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0135] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0136] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0137] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0138] This application also provides a photovoltaic electric field, including a photovoltaic array and a power prediction device, wherein the photovoltaic array is used to convert solar energy into electrical energy, and the power prediction device includes electronic equipment 40.

[0139] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0140] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0141] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0142] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0143] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0144] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0145] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0146] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0147] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0148] It should be understood that the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover but not exclude inclusion. For example, a product or device that includes a series of components is not necessarily limited to those components that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.

[0149] As used in this application, the term "module" means any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.

[0150] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A photovoltaic power prediction method based on dual-mode hybrid analysis, characterized in that, include: Obtain the current photovoltaic power sequence, which includes multiple daily photovoltaic power data. Adaptive noise complete set empirical mode decomposition is performed on the current photovoltaic power sequence to obtain multiple sets of intrinsic mode functions, wherein each set of intrinsic mode functions includes multiple intrinsic mode functions decomposed corresponding to a single day's photovoltaic power data; Based on a preset clustering algorithm and Bayesian information criteria, the multiple sets of intrinsic mode functions are clustered to obtain multiple data sub-pools. Each data sub-pool includes: a cluster label for a cluster, all daily photovoltaic power data corresponding to the cluster label, and multiple intrinsic mode functions corresponding to each daily photovoltaic power data. For each data sub-pool, multiple intrinsic mode functions in the data sub-pool are input into the pre-trained BiLSTM model and SVR model respectively to obtain the first prediction result and the second prediction result; The first prediction result and the second prediction result are fused together to obtain the target prediction result; Based on the target prediction results of all data sub-pools, a photovoltaic power prediction sequence is determined.

2. The method according to claim 1, characterized in that, The preset clustering algorithm is the X-means algorithm; based on the preset clustering algorithm and Bayesian information criteria, the multiple sets of intrinsic mode functions are clustered to obtain multiple data sub-pools, including: Feature extraction is performed on the multiple sets of intrinsic mode functions to obtain the target feature matrix; Based on the X-means algorithm and Bayesian information criterion, the target feature matrix is ​​clustered to obtain multiple data sub-pools.

3. The method according to claim 2, characterized in that, The step of extracting features from the multiple sets of intrinsic mode functions to obtain the target feature matrix includes: For each set of intrinsic mode functions, feature extraction is performed on multiple intrinsic mode functions corresponding to the decomposition of daily photovoltaic power data to obtain energy features; Based on the daily average photovoltaic irradiance and the energy characteristics, the feature sub-matrix corresponding to the daily photovoltaic power data is obtained; The target feature matrix is ​​obtained based on the feature sub-matrices corresponding to all daily photovoltaic power data.

4. The method according to claim 2, characterized in that, The target feature matrix is ​​clustered based on the X-means algorithm and Bayesian information criterion to obtain multiple data sub-pools, including: Initialize the number of clusters for clustering; For each cluster, the cluster is divided into two sub-clusters to obtain the center of the first sub-cluster and the center of the second sub-cluster; Based on Bayesian information criteria, the first criterion value of the cluster, the second criterion value of the center of the first sub-cluster, and the third criterion value of the center of the second sub-cluster are determined. When the second criterion value is less than the first criterion value, the center of the first sub-cluster is taken as the new cluster center; and when the third criterion value is less than the first criterion value, the center of the second sub-cluster is taken as the new cluster center. Based on the updated cluster centers, the updated cluster number is obtained, and the steps of dividing each cluster into two sub-clusters to obtain the first sub-cluster center and the second sub-cluster center are re-executed until the updated cluster number meets the preset iteration conditions, thereby obtaining the target cluster number and the multiple data sub-pools.

5. The method according to any one of claims 1-4, characterized in that, After clustering the multiple sets of intrinsic mode functions based on a preset clustering algorithm and Bayesian information criteria to obtain multiple data sub-pools, the method further includes: For each data sub-pool, the frequency of multiple intrinsic mode functions in each daily photovoltaic power data is divided to obtain the high-frequency intrinsic mode function and the mid-to-low frequency intrinsic mode function of each daily photovoltaic power data. Based on the whale swarm algorithm and the high-frequency intrinsic mode function of each day's photovoltaic power data, the optimization parameters of variational mode decomposition are determined; Based on the optimized parameters, variational mode decomposition is performed on the high-frequency intrinsic mode function of each daily photovoltaic power data to obtain multiple narrowband sub-modes; The low- and mid-frequency intrinsic mode functions of the daily photovoltaic power data and the multiple narrowband sub-modes are normalized to obtain the target mode data of the data sub-pool. Correspondingly, For each data sub-pool, multiple intrinsic mode functions from that sub-pool are input into a pre-trained BiLSTM model and an SVR model to obtain a first prediction result and a second prediction result, including: For each data sub-pool, the target modal data in the data sub-pool are input into the pre-trained BiLSTM model and SVR model respectively to obtain the first prediction result and the second prediction result.

6. The method according to any one of claims 1-4, characterized in that, The process of fusing the first prediction result and the second prediction result to obtain the target prediction result includes: Based on the fusion coefficient, the first prediction result and the second prediction result are fused to obtain the target prediction result. The fusion coefficient is obtained by particle swarm optimization and is used to determine the contribution weight of each model in the target prediction result.

7. The method according to any one of claims 1-4, characterized in that, Before inputting multiple intrinsic mode functions from each data sub-pool into a pre-trained BiLSTM model and an SVR model to obtain a first prediction result and a second prediction result, the method further includes: Acquire historical photovoltaic power sequences, which include multiple daily photovoltaic power data acquired within a preset historical time period; Adaptive noise-complete ensemble empirical mode decomposition is performed on the historical photovoltaic power sequence to obtain multiple sets of historical eigenmode functions; Based on a preset clustering algorithm and Bayesian information criteria, the multiple sets of historical intrinsic mode functions are clustered to obtain multiple historical data sub-pools. Based on all historical data sub-pools, construct training and validation sets; Based on the particle swarm optimization algorithm, the particle swarm is initialized, wherein each particle includes the first model parameters of the BiLSTM model, the second model parameters of the SVR model, and the initial fusion coefficients, which are used to fuse the prediction results of the BiLSTM model and the SVR model. For each particle, the BiLSTM model and the SVR model are trained respectively based on the training set and the model parameters in the particle to obtain the first historical prediction result and the second historical prediction result. Based on the initial fusion coefficient in the particles, the first historical prediction result and the second historical prediction result are fused to obtain the target historical prediction result. The loss function value is obtained based on the historical prediction results of the target and the corresponding validation samples in the validation set; Based on the loss function values ​​of all particles, the particle swarm is optimized until the preset loss function converges. Based on the optimal particle in the particle swarm, the trained BiLSTM model and SVR model, as well as the fusion coefficients, are determined.

8. A photovoltaic power prediction device based on dual-mode hybrid analysis, characterized in that, include: The acquisition module is used to acquire the current photovoltaic power sequence, which includes multiple daily photovoltaic power data. The decomposition module is used to perform adaptive noise complete set empirical mode decomposition on the current photovoltaic power sequence to obtain multiple sets of intrinsic mode functions, wherein each set of intrinsic mode functions includes multiple intrinsic mode functions decomposed corresponding to a single day's photovoltaic power data; The processing module is used to perform clustering processing on the multiple sets of intrinsic mode functions based on a preset clustering algorithm and Bayesian information criteria to obtain multiple data sub-pools. Each data sub-pool includes: a cluster label of a cluster, all daily photovoltaic power data corresponding to the cluster label, and multiple intrinsic mode functions corresponding to each daily photovoltaic power data. The prediction module is used to input multiple intrinsic mode functions from each data sub-pool into a pre-trained BiLSTM model and an SVR model to obtain a first prediction result and a second prediction result. The fusion module is used to fuse the first prediction result and the second prediction result to obtain the target prediction result; The determination module is used to determine the photovoltaic power prediction sequence based on the target prediction results of all data sub-pools.

9. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed, are used to implement the method as described in any one of claims 1-7.

Citation Information

Cited By

  • Photovoltaic power generation power prediction method based on mode perception adaptive deep learning

    CN121216441A