A photovoltaic power medium and long term prediction method and system based on frequency domain driving

CN122418640BActive Publication Date: 2026-09-11XINJIANG PETROLEUM ADMINISTRATION BUREAU +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610886894.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-11
Estimated Expiration
2046-06-18

AI Technical Summary

Technical Problem

一种基于频域驱动的光伏功率中长期预测方法和系统,及其相关技术,以解决现有方法对光伏功率序列多尺度季节性建模能力不足、复杂气候适应性差及高精度模型计算开销大难以实际部署等技术问题或其组合

Benefits of technology

1、本发明通过快速傅里叶变换构造K个半周期核函数,将光伏功率时序数据分解为多组不同尺度的季节子序列,并结合季节因子加权机制对各子序列差异化处理,实现了对日内短周期、昼夜中周期与跨季节长周期特征的同步精细捕捉。相比现有方法采用固定模式统计分解或仅提取单一低频分量的方式,本发明对光伏功率序列中多尺度季节性成分的拟合误差降低25%~35%,中长期(7~30天)预测偏差可控制在15%以内,较传统单一模型预测精度提升40%以上。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122418640B_ABST
    Figure CN122418640B_ABST
Patent Text Reader

Abstract

The application discloses a photovoltaic power medium and long-term prediction method and system based on frequency domain driving, and belongs to the technical field of photovoltaic power prediction. The technical problem to be solved is that the existing method has poor modeling ability for multi-scale seasonality of photovoltaic power sequence, poor adaptability to complex climate, and large calculation overhead of high-precision model, which is difficult to be actually deployed. The technical solution points are as follows: obtaining historical power time series data of a photovoltaic power station; performing fast Fourier transform on the historical power time series data to obtain a seasonal sequence and a trend sequence; sequentially performing embedding processing, downsampling, fast Fourier transform, learnable frequency weight adjustment, inverse Fourier transform and linear mapping on the seasonal sequence to obtain a seasonal prediction result; inputting the trend sequence into a linear prediction model to obtain a trend prediction result; and summing the two to obtain a photovoltaic power medium and long-term prediction value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of photovoltaic power prediction technology, and specifically to a method and system for medium- and long-term photovoltaic power prediction based on frequency domain driving. Background Technology

[0002] As an important component of clean energy, the accuracy of photovoltaic power generation forecasting directly affects the stability of power grid dispatching and the absorption rate of new energy sources. Especially in oil and gas field scenarios, core production equipment such as pumping units and gathering and transportation equipment have high requirements for power supply continuity. The medium- and long-term (7-30 days) forecast results of photovoltaic power are the key basis for formulating power dispatching plans and reducing production energy consumption.

[0003] However, the climate conditions in the oil and gas fields are complex, with frequent alternations of strong sandstorms, short-term strong gusts, cold waves, and extreme diurnal temperature differences, resulting in multi-scale seasonal and periodic fluctuations in photovoltaic power output. Traditional forecasting methods mainly face the following three shortcomings.

[0004] First, the multi-scale seasonal modeling capability is insufficient. Existing time series analysis-based methods (such as ARIMA) mainly model within the time domain, lacking the ability to effectively distinguish multi-scale features such as short intraday cycles, medium diurnal cycles, and long cross-seasonal cycles that coexist in the series. The error rate increases significantly when the prediction period exceeds 7 days. Although statistical decomposition methods such as X-11 can split the series into trend and seasonal components, the decomposition mode is fixed and cannot be dynamically adjusted according to the frequency characteristics of the input data. The multi-scale seasonal fitting error of power series under oil and gas field climatic conditions is generally high.

[0005] Secondly, they exhibit poor adaptability to complex climate conditions. Oil and gas field scenarios are characterized by frequent abrupt climate events such as sandstorms blocking sunlight and cold waves causing sharp drops in component efficiency. Existing models often rely on single data sources or employ static feature selection methods to incorporate meteorological factors, resulting in a dynamic identification accuracy of less than 60% for different climate types. This makes it difficult to internalize the impact of various extreme climates into model parameters during training. Medium- and long-term prediction biases generally exceed 25%, easily leading to delayed power grid dispatch instructions or excessive reserve capacity, increasing energy consumption costs in oilfield production.

[0006] Third, existing high-precision models suffer from high computational overhead and deployment difficulties. Deep learning models such as LSTM have a large number of parameters, and their computation time is long under the limited computing power conditions in oil fields, making it difficult to meet the requirements of minute-level prediction response. Although cascaded models such as ARIMA-LSTM have improved accuracy, their overall architecture is complex, maintenance costs are high, and engineering implementation is difficult.

[0007] Chinese patent application CN113988477A discloses a method, device, and storage medium for short-term photovoltaic power prediction based on machine learning. It discloses a technical solution based on X-11 time series trend decomposition and machine learning component modeling, which decomposes historical power sequences into long-term trend components, seasonal variation components, cyclic variation components, and random fluctuation components and models them separately. This has the technical effect of improving the accuracy of short-term prediction and improving the robustness of traditional single models. However, the following problems still exist: the solution is only applicable to short-term prediction (6 hours to 1 day), the decomposition method is fixed and does not have data adaptation capabilities, the granularity of distinguishing multi-scale seasonal features is limited, no frequency domain analysis mechanism is introduced, and no adaptive design is made for extreme climate scenarios in oil and gas fields, which cannot meet the accuracy requirements of medium and long-term prediction.

[0008] Chinese patent application CN118232324A discloses a method and system for medium- and long-term photovoltaic power prediction that considers periodic characteristics. It presents a technical solution based on DFT extraction of low-frequency periodic components, combined with ARIMA and LSTM modeling, which effectively utilizes meteorological factors and power time-series characteristics to improve the accuracy of medium- and long-term predictions. However, the following problems remain: the frequency domain analysis of this solution is limited to feature extraction, extracting only a single low-frequency periodic component for input into ARIMA, resulting in insufficient multi-resolution discrimination of seasonal features at different scales; meteorological features are directly input into the model after static screening by MIC, failing to dynamically internalize the impact of extreme climates into model weight parameters; the LSTM component has a large number of parameters and long inference time, making real-time deployment difficult in oilfields with limited computing power; furthermore, the solution lacks key mechanisms such as kernel function construction and seasonal factor weighting, resulting in significant deficiencies in its ability to model multi-scale features of power sequences under uncertain climatic conditions in oil and gas fields. Summary of the Invention

[0009] The purpose of this invention is to provide: A frequency-domain driven method and system for medium- and long-term photovoltaic power prediction, and related technologies, are proposed to address the technical problems of existing methods, such as insufficient ability to model multi-scale seasonality of photovoltaic power sequences, poor adaptability to complex climates, and high computational cost of high-precision models that make them difficult to deploy in practice, or a combination thereof.

[0010] Terminology Explanation: It should be understood that the following detailed description is exemplary and for illustrative purposes only, and does not limit the subject matter of the invention in any way. In this invention, the singular is used in conjunction with the plural unless otherwise specified. It should also be noted that, unless otherwise specified, the use of “or” or “or” means “and / or”. Furthermore, the use of the term “comprising” and other forms such as “including,” “containing,” and “contains” are not limiting.

[0011] Unless specifically defined herein, the use of various commercially available products herein employs standard techniques, or is carried out in accordance with methods known in the art or the description of this invention. The techniques and methods described herein can generally be implemented according to conventional methods well known in the art, based on the descriptions in the various summary and more specific documents cited and discussed in this specification.

[0012] The term "medium- to long-term forecast" used in this article refers to photovoltaic power forecasting tasks with a forecasting period of 7 to 30 days, which is different from ultra-short-term (within a few hours) and short-term (within a few days) forecasts. It is mainly used to support the formulation and optimization of oilfield power dispatching plans.

[0013] The term "Fast Fourier Transform" used in this article refers to an algorithm that efficiently converts time-domain signals into frequency-domain representations. It can decompose photovoltaic power time-series data into sinusoidal components of different frequencies and calculate the amplitude value of each component to identify periodic components in the sequence.

[0014] The term "frequency component" used in this paper refers to the signal components corresponding to each discrete frequency obtained after applying a fast Fourier transform to a time-domain sequence. Each frequency component has a corresponding frequency value and amplitude value. The larger the amplitude value, the more significant the contribution of the frequency component to the original sequence.

[0015] The term "kernel function" used in this paper refers to a periodic function constructed with half the period corresponding to a dominant frequency component as the period parameter. It is used to decompose historical power time series data and extract seasonal and trend subsequences at the corresponding scale.

[0016] The term "half-cycle kernel function" used in this paper refers to a kernel function constructed with half of the period pi (i.e., pi / 2) corresponding to the dominant frequency component as the period parameter. Compared with the full-cycle kernel function, the half-cycle kernel function can effectively capture the seasonal characteristics of the sequence while reducing the loss of seasonal detail information.

[0017] The term "seasonal subsequence" used in this paper refers to a subsequence that reflects the seasonal fluctuation pattern at a specific time scale, obtained by decomposing historical power time series data using a kernel function. Seasonal subsequences generated by different kernel functions correspond to periodic characteristics at different scales.

[0018] The term "trend subsequence" used in this paper refers to the subsequence that reflects the long-term trend of the series after decomposing historical power time series data using a kernel function. The trend subsequence does not contain significant nonlinear periodic characteristics and its changes are relatively stable.

[0019] The term "seasonal factor" used in this paper refers to the differentiated weighting coefficient assigned to each seasonal subsequence, which is used to adjust the contribution ratio of seasonal subsequences of different scales when merging. The smaller the period of the kernel function, the larger the corresponding seasonal factor, so as to enhance the model's adaptability to short-period high-volatility components.

[0020] The term "seasonal sequence" used in this paper refers to the merged sequence obtained by multiplying K groups of seasonal subsequences by their corresponding seasonal factors and then summing them. This merged sequence integrates multi-scale seasonal information and serves as the input to the frequency domain weighting module.

[0021] The term "trend sequence" used in this paper refers to the combined sequence obtained by directly adding K trend subsequences, which reflects the overall long-term trend of photovoltaic power and serves as the input to the linear prediction module.

[0022] The terms “value embedding, position embedding, and temporal embedding” used in this paper refer to three types of feature encoding operations performed on the input sequence. Value embedding maps the sequence values ​​to a high-dimensional feature space, position embedding introduces the positional information of each time step in the sequence, and temporal embedding introduces the temporal attribute information corresponding to the timestamp. Together, they enhance the model’s ability to perceive the sequence structure.

[0023] The term "downsampling" used in this paper refers to: sampling the sequence at intervals according to the hyperparameter c, compressing the sequence length to the original length divided by c and rounded down, thereby reducing the computational load while preserving the main frequency information of the sequence.

[0024] The term "learnable weight parameters" used in this paper refers to the frequency component weight coefficients that are automatically and iteratively updated through the gradient descent algorithm during model training. These parameters are used to dynamically adjust the model's attention to different frequency components, enabling the model to adaptively focus on low-frequency components corresponding to the seasonal cycle and suppress high-frequency components that reflect noise.

[0025] The term "inverse Fourier transform" used in this paper refers to the transformation operation that converts the frequency domain representation after frequency weighting back into a time domain sequence. It is the inverse of the fast Fourier transform and is used to restore the frequency domain weighted result to a time domain feature that can be used for subsequent linear mapping.

[0026] The term "linear prediction model" used in this paper refers to a fully connected linear network consisting of an input layer, an intermediate layer, and an output layer connected sequentially. It directly maps trend sequences to the output prediction space through linear transformation. It has few parameters, high computational efficiency, and is suitable for predicting trend components with stable changes.

[0027] The term "MAPE" used in this paper refers to Mean Absolute Percentage Error, which measures the relative error between the predicted value and the true value. This paper uses whether the fluctuation range of MAPE within a preset number of consecutive rounds is lower than a preset threshold as the convergence criterion for model training.

[0028] In a first aspect, the present invention provides a method for medium- and long-term photovoltaic power prediction based on frequency domain driving, comprising: S1: Obtain historical power time-series data of photovoltaic power plants; S2: Perform a fast Fourier transform on the historical power time series data, calculate the amplitude value corresponding to each frequency component, select the K frequency components with the largest amplitude, and construct K kernel functions based on the period of the K frequency components with half period as the parameter. S3: Decompose the historical power time series data using the K kernel functions to obtain K groups of seasonal subsequences and K groups of trend subsequences; assign seasonal factor weights to each group of seasonal subsequences, with the smaller the kernel function period corresponding to the larger the seasonal factor weight; merge the weighted seasonal subsequences to obtain a seasonal sequence, and merge the trend subsequences to obtain a trend sequence. S4: Perform frequency domain weighting processing on the seasonal sequence, specifically: perform value embedding, position embedding and time embedding on the seasonal sequence and then perform downsampling, apply fast Fourier transform to the downsampled sequence to obtain frequency domain representation, set learnable weight parameters for each frequency component and dynamically adjust them, and obtain the seasonal prediction result after inverse Fourier transform and linear mapping. S5: Input the trend sequence into the linear prediction model to obtain the trend prediction result; S6: Sum the seasonal forecast results with the trend forecast results to obtain the medium- and long-term forecast values ​​of photovoltaic power.

[0029] Furthermore: The K kernel functions described in S2 are constructed as follows: the K frequency components with the largest amplitudes are sorted in descending order of amplitude, and for the i-th frequency component... Take its corresponding period Half of it is used as the periodic parameter of the kernel function to construct a system with... Periodic kernel function with periodicity ,in .

[0030] Furthermore: In S3, the specific method for assigning seasonal factor weights to each seasonal subsequence is as follows: the smaller the period of the kernel function, the larger the seasonal factor assigned to the corresponding seasonal subsequence; the larger the period of the kernel function, the smaller the seasonal factor assigned to the corresponding seasonal subsequence; each seasonal subsequence is multiplied by its corresponding seasonal factor and then added together to obtain the seasonal sequence; each trend subsequence is directly added together to obtain the trend sequence.

[0031] Furthermore: the number of samples for downsampling in S4 is controlled by the hyperparameter c, and the length of the sequence after downsampling is the original sequence length divided by c and then rounded down; the learnable weight parameters are automatically updated in each round of training through gradient descent, so that the model assigns higher weights to low-frequency components corresponding to the seasonal cycle and lower weights to high-frequency components reflecting noise during the training process.

[0032] Furthermore, the linear prediction model described in S5 is a fully connected linear network, comprising an input layer, an intermediate layer, and an output layer connected in sequence, which directly maps the trend sequence to the output prediction space through a linear transformation.

[0033] Furthermore, after S1 and before S2, a data preprocessing step is also included: the historical power time series data is checked for missing values ​​and outliers are removed, and the values ​​of each feature parameter are normalized and mapped to the interval [0,1].

[0034] Furthermore: the historical power time-series data includes actual photovoltaic output data, as well as at least one of solar irradiance, ambient temperature, wind speed, dust concentration, and relative humidity; the normalization process employs the minimum-maximum normalization method, and the normalization formula is: ; in, The original data, For the normalized data, and These are the minimum and maximum values ​​of the parameter in the dataset, respectively.

[0035] Furthermore, before S2, a feature screening step is included: based on the Pearson correlation coefficient, the correlation between each meteorological feature parameter and photovoltaic output is analyzed, feature parameters with a correlation coefficient higher than a preset threshold are retained, and feature parameters with a correlation coefficient lower than the preset threshold or with a correlation coefficient higher than a redundancy threshold with the retained feature parameters are removed.

[0036] Furthermore: the prediction period for the medium- and long-term prediction is 7 to 30 days; when training the learnable weight parameters, a sliding window method is used to enhance the training data, and cross-validation is used to optimize the model parameters; the training termination condition is that the fluctuation range of the prediction error MAPE within a consecutive preset number of rounds is lower than a preset threshold.

[0037] Secondly, the present invention provides a frequency-domain driven photovoltaic power medium- and long-term prediction system, applying the prediction method described in the first aspect, including: The data acquisition module is used to acquire historical power time-series data of photovoltaic power plants; The kernel function construction module is used to perform a fast Fourier transform on the historical power time series data, calculate the amplitude value of each frequency component, select the K frequency components with the largest amplitude, and construct K kernel functions based on half of the period corresponding to each frequency component. The sequence decomposition module is used to decompose the historical power time series data using the K kernel functions to obtain K sets of seasonal subsequences and K sets of trend subsequences. Each seasonal subsequence is assigned a seasonal factor weight that is negatively correlated with the period of the corresponding kernel function, and the two subsequences are merged to obtain the seasonal sequence and the trend sequence. The frequency domain weighting module is used to sequentially perform embedding processing, downsampling, fast Fourier transform, frequency component weighting based on learnable weight parameters, inverse Fourier transform, and linear mapping on the seasonal sequence, and output the seasonal prediction result. The linear prediction module is used to input the trend sequence into the linear prediction model and output the trend prediction result; The result fusion module is used to sum the seasonal forecast results and the trend forecast results to output the medium- and long-term forecast values ​​of photovoltaic power.

[0038] The present invention has at least the following beneficial effects: 1. This invention constructs K half-cycle kernel functions using Fast Fourier Transform to decompose photovoltaic power time-series data into multiple sets of seasonal subsequences at different scales. It then employs a seasonal factor weighting mechanism to differentiate each subsequence, achieving simultaneous and precise capture of intraday short-cycle, diurnal medium-cycle, and cross-seasonal long-cycle characteristics. Compared to existing methods that use fixed-pattern statistical decomposition or extract only a single low-frequency component, this invention reduces the fitting error of multi-scale seasonal components in photovoltaic power sequences by 25%–35%, and controls the medium- to long-term (7–30 days) prediction deviation to within 15%, improving prediction accuracy by more than 40% compared to traditional single-model predictions.

[0039] 2. Enhanced adaptability to uncertain climatic conditions. This invention introduces learnable frequency weight parameters into the frequency domain weighting module. During training, the model can automatically adjust its focus on different frequency components based on data distribution, internalizing the influence of different climate types into model weights without requiring manual specification of fixed weight rules. Compared to existing solutions that rely on static feature selection or fixed weight mechanisms, this invention has a stronger dynamic response capability to power fluctuations caused by extreme weather events such as strong dust storms and cold waves, effectively reducing the interference of complex climatic conditions on the accuracy of medium- and long-term predictions.

[0040] 3. This invention uses a fully connected linear network to predict trend sequences, replacing complex nonlinear models such as LSTM with linear mapping to process trend components. The number of model parameters is reduced by more than 50% compared to LSTM, and the response time for a single prediction round can reach the minute level, meeting the real-time deployment needs under the limited computing power conditions of oilfields. It solves the problems of high computational overhead and difficulty in engineering implementation of existing high-precision prediction models. Attached Figure Description

[0041] Figure 1 The flowchart illustrates a frequency-domain driven method for medium- and long-term photovoltaic power prediction.

[0042] Figure 2 This is a schematic diagram of the structure of a frequency-domain driven photovoltaic power medium- and long-term prediction system provided by the present invention.

[0043] Figure 3 This is a schematic diagram of the overall framework of a method in one embodiment of the present invention.

[0044] Figure 4 This is a schematic diagram illustrating the process of time-domain to frequency-domain conversion and kernel function construction in one embodiment of the present invention.

[0045] Figure 5 This is a schematic diagram of a decomposition method incorporating seasonal factors in one embodiment of the present invention.

[0046] Figure 6 This is a schematic diagram of a linear model structure of a method in one embodiment of the present invention. Detailed Implementation

[0047] The following non-limiting embodiments are intended to enable those skilled in the art to gain a more comprehensive understanding of the present invention, but do not limit the invention in any way. The following content is merely an exemplary description of the scope of protection claimed by the present invention, and those skilled in the art can make various changes and modifications to the present invention based on the disclosed content, and such changes should also fall within the scope of protection claimed by the present invention.

[0048] The present invention will be further described below by way of specific embodiments. Unless otherwise specified, all instruments, devices, equipment, reagents, products, etc., used in the embodiments of the present invention are obtained through conventional commercial means.

[0049] Unless otherwise stated, conventional methods within the scope of the art shall be used.

[0050] Example 1 like Figure 1 As shown, this embodiment provides a medium- to long-term photovoltaic power prediction method based on frequency domain driving, including steps S1 to S6.

[0051] Step S1: Obtain historical power time-series data of the photovoltaic power plant. This historical power time-series data forms the data basis for subsequent frequency domain decomposition and model training.

[0052] Step S2 involves performing a Fast Fourier Transform (FFT) on the historical power time-series data to calculate the amplitude value corresponding to each frequency component. The K frequency components with the largest amplitudes are selected, and K kernel functions are constructed based on the periods of these K frequency components, using half-periods as parameters. Specifically, the time-domain sequence contains various seasonal patterns. The FFT transforms the time-domain sequence to the frequency domain, and the K frequency components with the largest amplitudes correspond to the main seasonal information in the sequence. Using half-periods instead of full-periods to construct kernel functions can preserve seasonal details while reducing information loss.

[0053] Step S3 involves decomposing the historical power time-series data using the K kernel functions to obtain K groups of seasonal subsequences and K groups of trend subsequences. Each seasonal subsequence is assigned a seasonal factor weight; the smaller the kernel function period, the larger the corresponding seasonal factor weight. The weighted seasonal subsequences are then merged to obtain a seasonal sequence, and the trend subsequences are merged to obtain a trend sequence. Specifically, subsequences generated with smaller kernel function periods exhibit more volatile fluctuations and are more susceptible to noise interference; assigning a larger seasonal factor enhances the model's adaptability to short-period rapid fluctuations. Conversely, subsequences with larger kernel function periods show more stable fluctuations, and assigning a smaller seasonal factor is sufficient to meet prediction requirements. This weighted merging process fully integrates multi-scale seasonal information.

[0054] Step S4 involves performing frequency domain weighting processing on the seasonal sequence. This includes value embedding, position embedding, and time embedding, followed by downsampling. A Fast Fourier Transform is then applied to the downsampled sequence to obtain its frequency domain representation. Learnable weight parameters are set for each frequency component and dynamically adjusted. After inverse Fourier transform and linear mapping, the seasonal prediction result is obtained. Specifically, seasonal components in the frequency domain are typically concentrated in low-frequency components. By dynamically adjusting the attention given to each frequency component through learnable weight parameters, the model adaptively focuses on the dominant seasonal frequencies during training, improving the prediction accuracy of seasonal components.

[0055] Step S5: Input the trend sequence into the linear prediction model to obtain the trend prediction result. The trend component does not contain complex nonlinear features, and the linear model can efficiently complete the direct mapping of the trend with low computational cost and strong interpretability.

[0056] Step S6: Summate the seasonal forecast results with the trend forecast results to obtain the medium- and long-term forecast values ​​for photovoltaic power. The seasonal and trend components are predicted separately and then summed to restore the original values, forming a closed loop with the sequence decomposition step to ensure the completeness of the forecast results.

[0057] In one specific implementation of this embodiment, the K kernel functions in S2 are constructed as follows: the K frequency components with the largest amplitudes are sorted in descending order of amplitude, and the i-th frequency component... Take its corresponding period Half of it is used as the periodic parameter of the kernel function to construct a system with... Periodic kernel function with periodicity ,in Kernel functions are constructed sequentially after being sorted by amplitude to ensure that the seasonal components that contribute the most to the sequence are captured first; the half-cycle design avoids the blurring of seasonal details by the full-cycle kernel function and preserves the resolution of seasonal features at each scale.

[0058] In one specific implementation of this embodiment, the K kernel functions in S2 are constructed as follows: Let the historical power time series data be... The sampling time interval is ,right After performing a fast Fourier transform, the discrete frequency components are obtained. and its corresponding amplitude First, eliminate the DC component, and then proceed according to amplitude. Select the top K dominant frequency components from largest to smallest, denoted as For the i-th dominant frequency component, its corresponding period is calculated as follows: ; If frequency is indexed by discrete frequency If we express it as such, then its period can also be expressed as: ; In the above formula, N is the length of the input sequence. The sampling time interval, Let be the frequency index corresponding to the i-th dominant frequency component. Then, half of this period is taken as the period parameter of the kernel function, i.e.: ; and with Construct the i-th periodic kernel function for the periodic scale; specifically, the following periodic Gaussian kernel function can be constructed: ; In the above formula, t and s represent different time points in the sequence. This represents the time interval between two points in time. Let be the half-cycle parameter determined by the i-th dominant frequency component. This is the kernel function bandwidth parameter, used to control the smoothness of the kernel function for adjacent period positions. Through the above construction, when two time points are within the period... When the two phases are close, the kernel function value is larger; when the phase difference between the two is large, the kernel function value is smaller, so that the kernel function can highlight the periodic variation pattern corresponding to the i-th dominant frequency.

[0059] In one specific implementation of this embodiment, the method for assigning seasonal factor weights to each seasonal subsequence in S3 is as follows: the smaller the kernel function period, the larger the seasonal factor assigned to the corresponding seasonal subsequence; the larger the kernel function period, the smaller the seasonal factor assigned to the corresponding seasonal subsequence; each seasonal subsequence is multiplied by its corresponding seasonal factor and then summed to obtain the seasonal sequence; each trend subsequence is directly summed to obtain the trend sequence. This weighting strategy enables the model to maintain a strong fitting ability for short-period, high-volatility subsequences, while not weakening the expression of long-period, stationary subsequences, thus achieving a balanced utilization of multi-scale seasonal information.

[0060] In one specific implementation of this embodiment, the embedding process in S4 includes value embedding, positional embedding, and temporal embedding; the number of samples in the downsampling is controlled by the hyperparameter c, and the length of the sequence after downsampling is the original sequence length divided by c and then rounded down; the learnable weight parameters are automatically updated through gradient descent in each training round, so that the model assigns higher weights to low-frequency components corresponding to the seasonal cycle and lower weights to high-frequency components reflecting noise during training. The three types of embedding introduce numerical information, positional information, and temporal information, respectively, enhancing the model's understanding of the sequence structure; the hyperparameter c controls the downsampling rate, retaining the main frequency information while reducing the computational load; the learnable weights automatically focus on the effective frequency components through training, eliminating the need for manually specifying fixed weights and improving the model's adaptability.

[0061] In one specific implementation of this embodiment, the linear prediction model in S5 is a fully connected linear network, comprising an input layer, an intermediate layer, and an output layer connected sequentially. It directly maps the trend sequence to the output prediction space through a linear transformation. The trend components change smoothly and exhibit clear linearity. The fully connected linear network has fewer parameters and higher computational efficiency. Compared to complex models such as LSTM, it can significantly reduce computational overhead while maintaining accuracy, thus meeting the response requirements of real-time prediction.

[0062] In one specific embodiment of this example, a data preprocessing step is included after S1 and before S2: missing value screening and outlier removal are performed on the historical power time series data, and the values ​​of each feature parameter are normalized to map the values ​​of each parameter to the interval [0,1]. Data preprocessing can remove noise data introduced by sensor failure or abnormal acquisition, and normalization eliminates the numerical differences between features of different dimensions, ensuring the stability of subsequent frequency domain analysis and model training.

[0063] In one specific embodiment of this example, the historical power time-series data includes actual photovoltaic output data, and at least one of solar irradiance, ambient temperature, wind speed, dust concentration, and relative humidity; the normalization process employs the minimum-maximum normalization method, and the normalization formula is:

[0064] in, The original data, For the normalized data, and These are the minimum and maximum values ​​of the parameter in the dataset, respectively. The introduction of multidimensional meteorological parameters enables the model to perceive the impact of external climate factors on photovoltaic output; the maximum-minimum normalization formula is concise and computationally efficient, and can unify parameters with significantly different dimensions to the same numerical range, avoiding the dominance of large-scale features in model training.

[0065] In one specific embodiment of this example, a feature selection step is included before S2: Based on Pearson correlation coefficient analysis, the correlation between various meteorological characteristic parameters and photovoltaic output is analyzed; feature parameters with correlation coefficients higher than a preset threshold are retained; and feature parameters with correlation coefficients lower than the preset threshold or with correlation coefficients higher than a redundancy threshold with the retained feature parameters are removed. Feature selection can remove redundant features with weak impact on photovoltaic output, reduce input dimensionality, reduce the computational load of model training, and simultaneously avoid interference from highly collinear features on frequency domain analysis results, thus improving prediction stability.

[0066] In one specific implementation of this embodiment, the prediction period for the medium- to long-term forecast is 7 to 30 days. When training the learnable weight parameters, a sliding window approach is used to augment the training data, and cross-validation is used to optimize the model parameters. The training termination condition is that the fluctuation range of the prediction error (MAPE) within a consecutive preset number of rounds is lower than a preset threshold. The prediction period covers weekly to monthly scheduling requirements. Sliding window augmentation can expand the effective training samples with limited data volume, and cross-validation prevents overfitting. The MAPE fluctuation range is used as the convergence criterion to ensure that the model stops training after achieving stable prediction accuracy on the validation set.

[0067] Example 2 like Figure 2 As shown, this embodiment provides a medium- to long-term photovoltaic power prediction system based on frequency domain driving, applying the prediction method described in Embodiment 1, including a data acquisition module, a kernel function construction module, a sequence decomposition module, a frequency domain weighting module, a linear prediction module, and a result fusion module.

[0068] The data acquisition module obtains historical power time-series data of the photovoltaic power plant, providing a unified data input for subsequent modules.

[0069] The kernel function construction module performs a Fast Fourier Transform on the historical power time series data, calculates the amplitude of each frequency component, selects the K frequency components with the largest amplitudes, and constructs K kernel functions based on half the period corresponding to each frequency component. By constructing kernel functions with the half-period as a parameter, the module achieves precise capture of multi-scale seasonal components in the sequence.

[0070] The sequence decomposition module uses the K kernel functions to decompose the historical power time series data, obtaining K sets of seasonal subsequences and K sets of trend subsequences. Each seasonal subsequence is assigned a seasonal factor weight that is negatively correlated with the period of the corresponding kernel function, and the subsequences are merged to obtain the seasonal sequence and the trend sequence. The seasonal factor weighting integrates the seasonal components at each scale according to their contribution differences, and the trend sequence is output independently for subsequent linear prediction.

[0071] The frequency domain weighting module sequentially performs embedding processing, downsampling, fast Fourier transform, frequency component weighting based on learnable weight parameters, inverse Fourier transform, and linear mapping on the seasonal sequence, outputting the seasonal prediction result. The learnable frequency weights are adaptively updated during training, enabling the module to focus on the seasonal low-frequency components and suppress high-frequency noise interference.

[0072] The linear prediction module inputs the trend sequence into the linear prediction model and outputs the trend prediction result. Linear models have a small number of parameters and high computational efficiency, making them suitable for rapid prediction of trend components with stable changes.

[0073] The results fusion module sums the seasonal forecast results with the trend forecast results and outputs the medium- and long-term forecast values ​​of photovoltaic power, thus completing the complete medium- and long-term forecast process.

[0074] Example 3 The technical solution of this invention will be described in detail below with reference to the actual application scenario of the Tarim Oil and Gas Field. It should be noted that the described embodiments are only some embodiments of this invention, and not all of them. Any other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the protection scope of this invention.

[0075] like Figures 3-6 As shown in the attached diagram, the following is a brief description. Figure 3 The overall framework of the method of this invention mainly includes four parts: frequency domain-driven data decomposition, extraction of multi-scale seasonal features, frequency-weighted features, and processing of trend sequence data.

[0076] Figure 4 The process of converting the time domain to the frequency domain and constructing a kernel function in the method of this invention decomposes a time domain sequence containing multiple seasonal patterns into different frequency domain components through a fast Fourier transform.

[0077] Figure 5 The method of this invention combines the decomposition of seasonal factors, uses kernel function decomposition to generate seasonal subsequences, highlights the seasonality of the corresponding cycle, and finally integrates multi-dimensional climate characteristics into the sequence.

[0078] Figure 6 The linear model structure of the method of this invention, in contrast to complex multi-scale climate characteristics, uses a simple linear model for trend component prediction, which can quickly and effectively process trend components.

[0079] Using this invention for medium- and long-term photovoltaic power prediction involves two steps: the first step is to train the frequency-domain constrained multi-scale seasonal photovoltaic power deep prediction model (MSFC) described in this invention using historical operating data of photovoltaic power plants in oil and gas fields; the second step is to use the trained model, combined with the actual meteorological and photovoltaic operating parameters of the target oil and gas field, to achieve accurate medium- and long-term photovoltaic power prediction.

[0080] I. Description of Implementation Methods in the Model Training Phase First, this embodiment surveyed the operation of photovoltaic power stations in typical oil and gas fields in China, collecting 150 sets of operational data across different seasons, weather conditions (sunny days, dust storms, cold waves, and short-term strong gusts), and photovoltaic array sizes (30MW-100MW). The data covers actual photovoltaic output, multi-source meteorological parameters (solar irradiance, ambient temperature, wind speed, dust concentration, and relative humidity), photovoltaic module status (temperature and conversion efficiency), and terrain-aided data (digital elevation model, DEM), meeting the model training requirements for data diversity.

[0081] For each data point, missing and outlier values ​​were investigated for the input parameters of the frequency domain decomposition module, the adjustment factor of the frequency weighting module, and the relevant indicators of the output results of the linear prediction module. The investigation revealed that the dust concentration monitoring data had the highest proportion of outliers (approximately 10%), mainly due to momentary sensor failure caused by obstruction during severe dust storms. 8% of the photovoltaic module conversion efficiency data was missing, stemming from the lack of real-time monitoring modules on some older modules. Basic meteorological data such as solar irradiance and ambient temperature showed no significant anomalies.

[0082] Data sets with missing or outlier conditions were discarded; 30 sets were discarded, leaving 120 sets for model training. After removing outlier and missing data, the data was normalized, mapping parameters such as photovoltaic output (0-5000kW), solar irradiance (0-1200W / m²), dust concentration (0-1000μg / m³), and ambient temperature (-35℃~+45℃) to the [0,1] interval. The normalization formula is as follows: Where x is the original data, These are the actual minimum and maximum values ​​of the parameter, respectively.

[0083] Pearson correlation coefficient analysis was applied to analyze the correlation between features. The correlation coefficient between solar irradiance and photovoltaic (PV) output was found to be 0.94. Therefore, solar irradiance was retained as the core driving feature in feature selection, and historical PV output values ​​were not introduced as independent features. The correlation coefficient between dust concentration and atmospheric visibility was 0.87, so atmospheric visibility was removed. The correlations between other parameters (wind speed, module temperature, DEM data, etc.) were all below 0.75, and all were retained. The final training dataset includes: an input feature matrix (120 rows, 32 columns, covering meteorological, environmental, module status, and terrain parameters), a frequency feature matrix (120 rows, 15 columns, corresponding to seasonal frequency components at different scales), and a prediction result vector (120 rows, 1 column, representing the predicted daily average peak PV power for the next month).

[0084] After data preprocessing, the MSFC model parameters were initialized: the Empirical Mode Decomposition (EMD) was set to 5 layers, and the Gaussian kernel bandwidth parameter σ was determined to be 0.5 through grid search (range 0.1-1.0); the initial values ​​of the frequency weighting coefficients were allocated according to the contribution of each frequency component (seasonal low-frequency components weighted 0.6, diurnal mid-frequency components weighted 0.3, and random fluctuation high-frequency components weighted 0.1); the learning rate of the linear prediction module was set to 0.001, and the regularization coefficient was 0.0001. During training, sliding window data augmentation (window size 30 days) was used, with 64 data points sampled per batch, 30% of which came from extreme climate samples (dust storm days, cold wave days). The model parameters were optimized through 5-fold cross-validation. When the training iteration reached 1200 rounds, the prediction error (MAPE) converged (fluctuation range <3%), and the model prediction results achieved a 93% fit with the actual power curve, meeting the training termination condition and completing the model training.

[0085] II. Explanation of the results of medium- and long-term prediction of photovoltaic power in the Tarim Oil and Gas Field using the trained model. In the actual application project of this embodiment, a 50MW photovoltaic power station is selected. The power station is equipped with a multi-source meteorological monitoring system (including laser dust sensor, irradiance meter, anemometer, etc.) and data acquisition terminal. It needs to provide medium and long-term photovoltaic power forecast data for the next month (30 days) for three surrounding oil and gas extraction sites (including 20 oil pumping units, gathering and transportation pumps, etc.) to support the formulation of oilfield power dispatch plan. The real-time and historical data of the power station (photovoltaic power data for the same period in the past 3 years, meteorological monitoring data for the past 15 days, current wind speed of 2.5 m / s, irradiance of 650 W / m², dust concentration of 180 μg / m³, and module temperature of 42℃) were input into the trained MSFC model, and the prediction process is as follows: Multi-source data fusion module: integrates meteorological monitoring data (irradiance, temperature, wind speed, dust concentration), historical operation data of photovoltaic power plants (power curves of the same period in the past 3 years) and topographic data (DEM digital elevation model) to construct a standardized input matrix (30 rows and 32 columns, corresponding to the daily feature parameters of the next 30 days). Frequency-domain driven seasonal decomposition: The historical photovoltaic power sequence is decomposed into 5 IMF components and 1 residual term by EMD, where IMF1 (high frequency, 0.1-0.5Hz) corresponds to intraday short-term fluctuations, IMF3 (mid frequency, 0.01-0.1Hz) corresponds to diurnal periodic changes, and IMF5 (low frequency, <0.01Hz) corresponds to seasonal fluctuations. Combining the Gaussian kernel function and seasonal factors (1.2 for high temperature and strong sunshine in summer, and 0.8 for low temperature and weak radiation in winter), the seasonal frequency features in each IMF component are extracted to generate a frequency feature matrix.

[0086] Frequency weighting adjustment: Based on the current dust concentration (180 μg / m³, which is at a moderate level), the frequency weights are dynamically adjusted: the weight of the mid-frequency range (0.02-0.08 Hz), which is sensitive to the impact of dust, is increased from 0.3 to 0.4 to enhance the model's ability to capture diurnal power fluctuations caused by dust; at the same time, the weight of low-frequency seasonal components is maintained at 0.5 and the weight of high-frequency random fluctuations is maintained at 0.1 to ensure stable prediction of long-term trends.

[0087] Linear prediction output: The trained linear model predicts the input data of the fusion frequency characteristics and outputs the daily peak, valley and average daily power generation prediction results for the next 30 days, forming a complete medium and long-term prediction curve. After running this model for one month, the prediction results of the traditional LSTM model, ARIMA model, and VMD-SSA-LSTM hybrid model are compared, and the performance indicators are shown in the table below:

[0088] The results show that the MSFC model of this invention has significant advantages: Better prediction accuracy: MAPE is only 4.2%, which is 2.6 percentage points lower than the LSTM model and 5.3 percentage points lower than the ARIMA model. It accurately captures the seasonality (such as the power drop caused by sandstorm weather on the 15th-20th day) and diurnal periodicity of photovoltaic power in Tarim Oil and Gas Field. Its lightweight characteristics are outstanding: the number of model parameters is only 12,000, far lower than LSTM (85,000) and VMD-SSA-LSTM (103,000), and the single-round prediction time is 1.8s, which is suitable for the limited computing equipment resources in the front line of oil fields; High practical application value: Based on the prediction results of the MSFC model, the oilfield dispatching department optimized the operation plan of the pumping unit and arranged high-load operations (such as wellbore dewaxing) during the peak photovoltaic power period (11:00-15:00), which increased the photovoltaic absorption rate from the original 72% to 89% and reduced the daily grid power purchase cost by 28%, verifying the effectiveness and practicality of the invention under uncertain weather conditions.

[0089] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention do not depart from the essence and scope of the technical solution of the present invention.

Claims

1. A medium- to long-term photovoltaic power prediction method based on frequency domain driving, characterized in that, include: S1: Obtain historical power time-series data of photovoltaic power plants; S2: Perform a fast Fourier transform on the historical power time series data, calculate the amplitude value corresponding to each frequency component, select the K frequency components with the largest amplitude, and construct K kernel functions based on the period of the K frequency components with half period as the parameter. S3: Decompose the historical power time series data using the K kernel functions to obtain K groups of seasonal subsequences and K groups of trend subsequences; assign seasonal factor weights to each group of seasonal subsequences, with the smaller the kernel function period corresponding to the larger the seasonal factor weight; merge the weighted seasonal subsequences to obtain a seasonal sequence, and merge the trend subsequences to obtain a trend sequence. S4: Perform frequency domain weighting processing on the seasonal sequence, specifically: perform value embedding, position embedding and time embedding on the seasonal sequence and then perform downsampling, apply fast Fourier transform to the downsampled sequence to obtain frequency domain representation, set learnable weight parameters for each frequency component and dynamically adjust them, and obtain the seasonal prediction result after inverse Fourier transform and linear mapping. S5: Input the trend sequence into the linear prediction model to obtain the trend prediction result; S6: Sum the seasonal forecast results with the trend forecast results to obtain the medium- and long-term forecast values ​​of photovoltaic power.

2. The method for medium- and long-term photovoltaic power prediction based on frequency domain driving according to claim 1, characterized in that: The K kernel functions described in S2 are constructed as follows: The K frequency components with the largest amplitudes are sorted in descending order of amplitude; for the i-th frequency component… Take its corresponding period Half of it is used as the periodic parameter of the kernel function to construct a system with... Periodic kernel function with periodicity ,in .

3. The method for medium- and long-term photovoltaic power prediction based on frequency domain driving according to claim 1, characterized in that: The specific method for assigning seasonal factor weights to each seasonal subsequence in S3 is as follows: The smaller the period of the kernel function, the larger the seasonal factor assigned to the corresponding seasonal subsequence; the larger the period of the kernel function, the smaller the seasonal factor assigned to the corresponding seasonal subsequence. The seasonal sequence is obtained by multiplying each seasonal subsequence by its corresponding seasonal factor and then summing them. The trend subsequences are directly added together to obtain the trend sequence.

4. The method for medium- and long-term photovoltaic power prediction based on frequency domain driving according to claim 1, characterized in that: The number of samples for downsampling described in S4 is controlled by the hyperparameter c. The length of the sequence after downsampling is the original sequence length divided by c and then rounded down. The learnable weight parameters are automatically updated in each round of training through gradient descent, so that the model assigns higher weights to low-frequency components corresponding to the seasonal cycle and lower weights to high-frequency components reflecting noise during the training process.

5. The method for medium- and long-term photovoltaic power prediction based on frequency domain driving according to claim 1, characterized in that: The linear prediction model described in S5 is a fully connected linear network, which includes an input layer, an intermediate layer and an output layer connected in sequence, and directly maps the trend sequence to the output prediction space through a linear transformation.

6. The method for medium- and long-term photovoltaic power prediction based on frequency domain driving according to claim 1, characterized in that: The process after S1 and before S2 includes a data preprocessing step: checking for missing values ​​and removing outliers in the historical power time series data, and normalizing each feature parameter to map the value of each parameter to the interval [0,1].

7. The method for medium- and long-term photovoltaic power prediction based on frequency domain driving according to claim 6, characterized in that: The historical power time-series data includes actual photovoltaic output data, as well as at least one of solar irradiance, ambient temperature, wind speed, dust concentration, and relative humidity; the normalization process uses the maximum-minimum normalization method, and the normalization formula is: ; in, The original data, For the normalized data, and These are the minimum and maximum values ​​of the parameter in the dataset, respectively.

8. The method for medium- and long-term photovoltaic power prediction based on frequency domain driving according to claim 1, characterized in that: Before S2, a feature selection step is also included: based on the Pearson correlation coefficient, the correlation between each meteorological feature parameter and photovoltaic output is analyzed, feature parameters with a correlation coefficient higher than a preset threshold are retained, and feature parameters with a correlation coefficient lower than the preset threshold or with a correlation coefficient higher than a redundancy threshold with the retained feature parameters are removed.

9. The method for medium- and long-term photovoltaic power prediction based on frequency domain driving according to claim 1, characterized in that: The prediction period for the medium- and long-term forecasts is 7 to 30 days; when training the learnable weight parameters, a sliding window method is used to enhance the training data, and cross-validation is used to optimize the model parameters; The training termination condition is that the fluctuation range of the prediction error MAPE within a preset number of consecutive rounds is lower than a preset threshold.

10. A medium- to long-term photovoltaic power prediction system based on frequency domain driving, employing the prediction method described in any one of claims 1-9, characterized in that, include: The data acquisition module is used to acquire historical power time-series data of photovoltaic power plants; The kernel function construction module is used to perform a fast Fourier transform on the historical power time series data, calculate the amplitude value of each frequency component, select the K frequency components with the largest amplitude, and construct K kernel functions based on half of the period corresponding to each frequency component. The sequence decomposition module is used to decompose the historical power time series data using the K kernel functions to obtain K sets of seasonal subsequences and K sets of trend subsequences. Each seasonal subsequence is assigned a seasonal factor weight that is negatively correlated with the period of the corresponding kernel function, and the two subsequences are merged to obtain the seasonal sequence and the trend sequence. The frequency domain weighting module is used to sequentially perform embedding processing, downsampling, fast Fourier transform, frequency component weighting based on learnable weight parameters, inverse Fourier transform, and linear mapping on the seasonal sequence, and output the seasonal prediction result. The linear prediction module is used to input the trend sequence into the linear prediction model and output the trend prediction result; The result fusion module is used to sum the seasonal forecast results and the trend forecast results to output the medium- and long-term forecast values ​​of photovoltaic power.

Citation Information

Patent Citations

  • Photovoltaic power short-term prediction method and device based on machine learning, and storage medium

    CN113988477A

  • Medium and long term photovoltaic power prediction method and system considering periodic characteristic characterization

    CN118232324A

  • Sub-daylight photovoltaic power generation prediction method and system based on seasonal decomposition and convolutional network

    CN116014722A

  • Electromagnetic spectrum prediction method and system based on fractional Fourier transform

    CN120596903A