Photovoltaic prediction method
By integrating spatiotemporal and time-sequential neural networks, the data heterogeneity and computing efficiency problems in photovoltaic power generation prediction are solved, and high-precision and interpretable photovoltaic power generation prediction is achieved, which is suitable for power system scheduling and edge computing equipment.
Patent Information
- Application Number
- CN202510974776.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-15
AI Technical Summary
The existing photovoltaic power generation prediction methods have shortcomings in data heterogeneity, computing efficiency, physical interpretability and dynamic environmental adaptability, and it is difficult to achieve high-precision and high-reliability photovoltaic power prediction, especially on edge computing devices, which are difficult to meet the needs of real-time and accuracy.
The method of fusing spatiotemporal and time-series neural network is adopted to design the local-global attention mechanism and combine physical regularization terms to achieve efficient feature extraction and prediction through time series regularization, space-time embedding, local-global attention layer, joint attention fusion layer and edge optimization technology.
While ensuring accuracy, the calculation complexity is reduced, and the model runs in real time on embedded GPUs, improving the robustness and interpretability of predictions, meeting the needs of power system scheduling and grid stability, and supporting offline prediction of edge devices.
Smart Images

Figure CN120470547A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of smart grid and energy management, and particularly relates to a photovoltaic prediction method. Background Art
[0002] As the global energy mix gradually shifts toward cleaner and more sustainable energy, photovoltaic power generation (PV) is becoming increasingly prominent as a key renewable energy solution. With PV installed capacity growing at an astonishing rate across China, accurately forecasting PV power generation is crucial for ensuring power system stability, optimizing grid dispatch, and improving system economics. However, PV power generation is inherently intermittent and uncertain, primarily influenced by meteorological factors such as sunlight intensity, ambient temperature, and cloud cover, making precise prediction of its output power difficult. These uncertainties pose significant challenges to power system planning, real-time dispatch, and stable operation, particularly when considering large-scale integration of PV power into existing power grids. This requires not only accurate assessment of PV power generation capacity but also in-depth research on PV generation behavior under varying meteorological conditions to overcome inherent forecasting challenges. Therefore, developing highly accurate and reliable PV power generation forecasting technologies is crucial for advancing PV technology, improving PV power utilization, and ensuring the safe and stable operation of power systems. Furthermore, combining big data analysis, deep learning algorithms, and refined meteorological models can significantly improve the accuracy and reliability of PV power generation forecasts. This not only helps alleviate grid fluctuations caused by the intermittent nature of photovoltaic power generation, but also promotes the efficient use of renewable energy, providing strong technical support for achieving a green transformation of the global energy structure. Developing highly adaptable and accurate photovoltaic power generation prediction methods has far-reaching strategic significance for promoting the development and application of clean energy.
[0003] Existing photovoltaic power forecasting methods can be divided into four categories based on their technological development: prediction methods based on physical models, prediction techniques based on statistical models, prediction methods based on traditional machine learning, deep learning technologies, and comprehensive prediction methods. Among them, prediction methods based on physical models primarily simulate the physical processes of photovoltaic cells to establish an electrical model of the cells, and then combine it with meteorological data to predict power. While this method has clear physical meaning, it has a large number of model parameters, requires high accuracy of meteorological data, and struggles to accurately describe photovoltaic output characteristics under complex meteorological conditions. Prediction methods based on statistical models primarily analyze the statistical and logical relationships between historical photovoltaic power generation data and meteorological data to establish a prediction model. Commonly used statistical models include autoregressive models (AR), moving average models (MA), and autoregressive moving average models (ARMA). These methods are simple and easy to use, but have difficulty handling nonlinear relationships and are highly dependent on historical data. Machine learning-based prediction methods, on the other hand, primarily train neural network-based learning models to learn the complex nonlinear relationships between photovoltaic power generation data and meteorological data and perform power forecasting. Commonly used machine learning models typically include support vector machines (SVMs), random forests (RFs), and neural networks (NNs). These methods have powerful nonlinear mapping capabilities, but model training requires extensive historical data and suffers from poor interpretability. Deep learning technology represents a cutting-edge advancement in photovoltaic forecasting methods, particularly adept at handling complex, nonlinear relationships. These techniques primarily rely on multi-layer neural network architectures, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and their improved versions, long short-term memory (LSTM) and gated recurrent units (GRUs). These models can automatically learn and extract features from large amounts of historical data, eliminating the need for manual feature engineering. In photovoltaic forecasting, deep learning models can effectively capture the dynamic patterns in time series data by learning power generation patterns under varying meteorological conditions, thereby achieving more accurate power forecasts. Furthermore, deep learning models can adapt to varying weather conditions and geographic locations, improving forecast robustness. However, deep learning models typically require extensive training data and are computationally expensive to train. Furthermore, their complex structure leads to poor interpretability. Comprehensive forecasting methods aim to combine the strengths of these different approaches to overcome the limitations of individual approaches. These methods can improve forecasting performance by integrating the advantages of physical models, statistical models, and machine learning or deep learning models. For example, a common strategy is to use physics-based models for preliminary estimates and then use machine learning or deep learning models to correct for errors. This approach not only preserves the clear physical meaning of the physical model but also leverages the powerful nonlinear mapping capabilities of data-driven models to compensate for the poor performance of physical models under complex conditions.Ensemble learning is also an important comprehensive approach. It further improves forecast accuracy and stability by integrating the prediction results of multiple basic models, such as through averaging, weighted averaging, or more complex meta-learning algorithms. These ensemble approaches emphasize the flexible selection and combination of different forecasting techniques and strategies based on the needs of specific application scenarios to achieve optimal forecasting results. This makes them particularly well-suited to addressing the challenges posed by the high variability and uncertainty of PV system output.
[0004] Photovoltaic power generation forecasting requires not only predicting power generation at a single moment in time but also predicting power generation trends over time, maintaining consistency and continuity across time. Currently, photovoltaic power generation forecasting technology is increasingly advantageous in power system scheduling, grid stability control, and power market trading. High-precision forecasting can reduce the risk of grid frequency fluctuations. Day-ahead market trading requires a low root mean square error (RMSE) for 24-hour forecasts (typically <3%). The industry also has certain requirements for edge computing deployment, some of which require millisecond-level real-time inference on low-computing devices, such as the chips built into photovoltaic inverters. Photovoltaic output is affected by heterogeneous data from multiple sources, including meteorological conditions, device status, and spatiotemporal correlations. Traditional models such as LSTM and CNN struggle to effectively integrate cross-modal features. Traditional methods rely on fixed time window modeling, making them difficult to handle sudden output fluctuations caused by meteorological events such as rapid cumulus cloud obstruction. While models like the Transformer can capture long-term dependencies, their global self-attention mechanism results in a computational complexity of O(T) / (T). 2 ), it is difficult to meet the real-time requirements of photovoltaic power generation forecasting, especially in minute-level forecasting scenarios, which are even more challenging. In addition, the existing data-driven deep neural network model architecture is mostly black-box operation, which cannot effectively embed the physical laws of photovoltaic power generation, such as temperature attenuation effect, which easily leads to excessive prediction deviations in extreme weather or equipment aging. In response to the above problems, the existing mainstream methods have obvious shortcomings. For example, pure physical models are usually highly dependent on accurate equipment parameters and environmental measurements, but cannot calibrate dynamic factors such as dust obstruction and component aging in real time; for example, the hybrid model of statistical-machine learning is prone to insufficient spatial resolution of meteorological numerical forecasts, and it is difficult to characterize the impact of local cloud movement on distributed photovoltaics; for example, the encoding and decoding framework contained in the Transformer model consumes a lot of computing resources. For example, predicting a 24-hour minute-level series requires 16GB of video memory, which limits its deployment on the edge side, and its position encoding does not integrate meteorological time series features, resulting in deviations in spatiotemporal correlation modeling. Therefore, exploring and researching efficient photovoltaic power generation prediction methods, automatically realizing accurate prediction of photovoltaic power generation, and ensuring that the prediction results maintain high precision and high reliability in the time dimension have become major issues that need to be urgently addressed in the current photovoltaic power generation application field. Summary of the Invention
[0005] In order to solve at least one technical problem in the above-mentioned background technology, the purpose of the present invention is to provide a new photovoltaic intelligent prediction method that integrates spatiotemporal sparse attention and time series neural network, aiming to solve the problems of data heterogeneity, computational efficiency, physical interpretability and insufficient adaptability to dynamic environments in the existing technology.
[0006] The present invention solves the technical problem by adopting the following technical solutions:
[0007] A photovoltaic prediction method comprises the following steps:
[0008] Step S1, data preprocessing: mainly align the time series of historical power generation data with the meteorological data through the time series regularization strategy, design feature engineering, extract the cumulative change rate of irradiance, and construct the equipment health index, which is input into the input layer;
[0009] Step S2: Based on the preprocessing results of step S1, a spatiotemporal embedding layer is designed to perform spatiotemporal encoding on the data transmitted by the input layer and all meteorological data branches according to the time series characteristics and spatial series characteristics;
[0010] Step S3: Define a neural network module. Design a local-global attention layer in this module and add attention window parameters to capture key information about the spatiotemporal features of the data in the spatiotemporal embedding layer. Then design a network layer to achieve efficient feature transformation and enhancement. In addition, inject meteorological constraints through physical regularization terms to ensure that the model output is more consistent with physical laws.
[0011] Step S4: Design a joint attention fusion layer, in which meteorological cross-attention calculation, gating weight generation, and attention result fusion are completed to effectively integrate key features from different sources and processed by the network layer;
[0012] Step S5, after step S4, designing the output layer, completing feature integration, linear transformation, setting activation function, physical regularization, and finally generating output of photovoltaic power prediction;
[0013] Step S6, edge optimization, compresses this improved neural network model to the memory device through INT8 quantization and model pruning technology, enabling offline prediction.
[0014] Furthermore, in step S1, firstly, any input data I is given in ={P1,P2,P3,……,P m ,R n}, where P i ={d i1 ,d i2 ,…,d im|1≤i≤m,m∈N + ,d im ∈ }, P i is the input photovoltaic data set of column i, m is the number of input photovoltaic feature columns, d im is the mth photovoltaic data value in the i-th column, R n It is a set of all influencing factors n, which supports dynamic addition and mainly includes meteorological data and user behavior patterns.
[0015] Furthermore, in step S1, it is necessary to accurately find the optimal alignment path between the two sequences through the time series regularization strategy and complete the processing of missing values and outliers. The specific method is as follows:
[0016] First, the time series of historical power generation data and meteorological data are constructed into a two-dimensional cost matrix. The elements in the matrix represent the differences between corresponding points in the two series. Secondly, starting from the starting point of the matrix, the cumulative cost matrix is traversed step by step to calculate the cumulative cost matrix. The elements of the cumulative cost matrix represent the optimal path cost from the starting point to the current point. Starting from the end point of the cumulative cost matrix, backtracking along the optimal path is used to find the optimal alignment path between the two time series data. Then, missing values in the aligned data are checked and processed using interpolation filling. Outliers are identified and processed by removing the mean. The cumulative rate of change of irradiance is then calculated using the following formula: , (1)
[0017] Where s is the size of the time window, V is the rate of change of irradiance, The cumulative rate of change of irradiance, which reflects the trend of irradiance change over time, is used to assist in analyzing its impact on power generation. When constructing the equipment health index, irradiance, power generation, temperature, wind speed, wind direction, humidity, equipment operating time, maintenance records, and fault history records are selected as key features. These features are normalized to their maximum and minimum values so that they are on the same scale. The normalized feature set obtained after processing is set to Nor, and the health index is calculated using the following formula function:
[0018] (2)
[0019] where w i is the weight of feature i.
[0020] Furthermore, in step S2, time series feature encoding involves setting multi-column one-dimensional convolution operations with different kernel numbers according to the number of input features to capture local features, and dynamically setting the weights assigned to different time steps; spatial sequence feature encoding mainly extracts the input features of the spatial relationship as spatial sequence feature encoding within the regular range of linear fitting based on the time resolution of the input data and the fluctuation law of sample sampling; then the outputs of the time coding layer and the space coding layer are structurally fused to form a unified spatiotemporal coding structure, and the information feature representation after spatiotemporal coding is output.
[0021] Furthermore, in step S3, the custom neural network module is first initialized and set to accept two parameters, namely the model dimension d_model and the number of attention heads n_heads; a local-global attention layer is created in the module to implement the local-global attention mechanism, and this layer adds an attention window parameter w_size=24; multiple neural network sublayers are created in the network layer; in each neural network sublayer, the input data first undergoes a linear transformation to expand the dimension from d_model to d_model×4, which uses a dynamically learnable radial basis function as the activation function to introduce nonlinear characteristics and enhance physical interpretability; after the activation function, the data undergoes another linear transformation to restore the dimension from d_model×4 back to d_model.
[0022] Furthermore, in step S3, a physical constraint module is created after the model output to implement the physical regularization term. The output of the physical constraint module conforms to the physical laws. The parameter coef=0.1 is set to control the influence of the regularization term on the model output. Then, the output of the network layer is added to the physical regularization term calculated by the physical constraint module, and the meteorological condition constraint is injected to obtain the final output of the physical constraint module.
[0023] Furthermore, in step S4, a joint attention fusion layer must first be designed. In this fusion layer, meteorological cross-attention calculation is first performed. Meteorological data is required as the key and historical power generation data as the query. By comparing the similarity between each power generation data point and all meteorological data points, the degree of attention to different meteorological conditions is determined; then an attention weight matrix is generated to represent the degree of attention of each power generation data point to each meteorological data point; secondly, gated weight generation is performed, and the meteorological cross-attention result obtained in the first step is input into the adaptive kernel network structure together with the local-global attention result of the previous step; by learning the characteristics of the input data, the weight distribution of the three types of attention is generated; then a weight vector is output, which contains the weight values of the three types of attention; followed by attention result fusion, and the weight vector generated in the previous step is used to perform weighted summation of the three types of attention results.
[0024] Furthermore, the key features in step S4 include spatiotemporal data, meteorological data and equipment status data.
[0025] Furthermore, in step S5, the features output by the joint attention fusion layer are further integrated through the output layer, and the input features are dimensionally adjusted to ensure that they match the expected input of the output layer; at the same time, multiple fully connected layers are set, and their weights and bias parameters are learned during the training process to achieve nonlinear transformation of features and map the fused features to the target output space; a linear activation function is used in the fully connected layer, and the above-mentioned physical regularization term is called before the final prediction result is generated, and the output is additionally adjusted to ensure that the prediction result conforms to the physical laws.
[0026] Furthermore, in step S6, the edge optimization method includes:
[0027] Step S61: First, evaluate the size of the current neural network model, including weights and biases, and perform a benchmark test on the uncompressed model, recording performance indicators as a reference; calculate the absolute value of each weight and identify small, unimportant weights; then set a pruning threshold based on the model size and performance requirements, set weights less than the threshold to 0, remove corresponding connections, remove redundant weights and connections, and reconstruct the model structure; evaluate the importance of each layer by removing and testing the performance layer by layer, then remove layers with less impact on performance based on the evaluation results, and gradually iterate the pruning to fine-tune the model;
[0028] Step S62: prepare a small calibration dataset for parameter adjustment during the quantization process; convert the model's weights from floating point numbers to INT8 representation; calculate the scaling factor and zero point for each weight layer to maintain numerical stability; convert the model's activation values from floating point numbers to INT8 representation; reconstruct the model based on the results of pruning and quantization to ensure the correct model structure; then use TensorRT, an inference engine optimized for INT8, and convert the quantized model to a format supported by the edge device, deploy the model to an edge device with 256KB of memory, and perform offline prediction tests on the edge device to verify the performance and accuracy of the model; finally, perform performance analysis and iterative optimization.
[0029] Beneficial effects of the present invention:
[0030] 1. Propose a local-global sparse attention mechanism to reduce the computational complexity from O(T 2 ) is reduced to O(TlogT), while ensuring accuracy, the model can run in real time on embedded GPU; the algorithm structure is designed with joint dynamic fusion based on multi-attention, and the gated KAN network adaptively weights the spatiotemporal, meteorological, and device status attention to improve the efficiency of cross-modal feature fusion.
[0031] 2. An interpretable dual KAN activation function is proposed, and a dynamically learnable radial basis function is introduced to replace ReLU. This introduces nonlinear characteristics to improve physical interpretability and combines it with photovoltaic equation constraints, so that the model satisfies both data-driven and physical consistency, improving the interpretability of the model and making it easier for users to understand and trust the model prediction results. Meteorological perception position coding is integrated, and real-time meteorological parameters such as irradiance and cloud cover are cross-swapped in multiple steps, embedded in position coding, to enhance the model's sensitivity to microclimate changes and improve the robustness of the prediction.
[0032] 3. Through the improved neural network structure and algorithm, the joint attention fusion layer can effectively fuse multi-source heterogeneous data, improve prediction accuracy, and meet the needs of grid stability, power trading economy, etc.; the method of the present invention can support edge optimization version, that is, it supports INT8 quantization and model pruning, and can run on 256KB memory devices to meet the offline prediction needs of scenarios without network coverage. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flow chart of the method of the present invention;
[0034] Figure 2 This is a structural diagram of the deep neural network for photovoltaic prediction in the present invention.
[0035] Figure 3 Design diagram for multi-column convolution in photovoltaic prediction method. DETAILED DESCRIPTION
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0037] refer to Figure 1 The present invention provides a photovoltaic prediction method, comprising the following steps:
[0038] Step S1, data preprocessing: First, given any input data I in ={P1,P2,P3,……,P m ,R n}, where P i ={d i1 ,d i2 ,…,d im |1≤i≤m,m∈N + ,d im ∈R},P iis the input photovoltaic data set of column i, m is the number of input photovoltaic feature columns, d im is the mth photovoltaic data value in the i-th column, R n This is a set of n influencing factors, supporting dynamic additions. These factors primarily include meteorological data and user behavior patterns. This involves aligning historical power generation data with meteorological data time series through time series regularization strategies, designing feature engineering, extracting the cumulative rate of change of irradiance, and constructing a device health index, which is then fed into the input layer.
[0039] It is necessary to use the time series regularization strategy to accurately find the optimal alignment path between the two series and complete the processing of missing values and outliers. The specific method is as follows:
[0040] First, the time series of historical power generation data and meteorological data are constructed into a two-dimensional cost matrix. The elements in the matrix represent the differences between corresponding points in the two series. Secondly, starting from the starting point of the matrix, the cumulative cost matrix is traversed step by step to calculate the cumulative cost matrix. The elements of the cumulative cost matrix represent the optimal path cost from the starting point to the current point. Starting from the end point of the cumulative cost matrix, backtracking along the optimal path is used to find the optimal alignment path between the two time series data. Then, missing values in the aligned data are checked and processed using interpolation filling. Outliers are identified and processed by removing the mean. The cumulative rate of change of irradiance is then calculated using the following formula:
[0041] , (1)
[0042] Where s is the size of the time window, which is set to 15 minutes, and V is the rate of change of irradiance. The cumulative rate of change of irradiance, which reflects the trend of irradiance change over time, is used to assist in analyzing its impact on power generation. When constructing the equipment health index (HI), irradiance, power generation, temperature, wind speed, wind direction, humidity, equipment operating time, maintenance records, and fault history records are selected as key features. These features are normalized to their maximum and minimum values so that they are on the same scale. The normalized feature set obtained after processing is set to Nor, and the health index is calculated using the following formula function:
[0043] (2)
[0044] where w i is the weight of feature i, which can be determined through expert experience or machine learning.
[0045] Step S2: Based on the preprocessing results of step S1, a spatiotemporal embedding layer is designed to perform spatiotemporal encoding on the data transmitted by the input layer and all meteorological data branches according to the time series characteristics and spatial series characteristics;
[0046] ,refer to Figure 3 Time series feature coding involves setting multi-column one-dimensional convolution operations with different kernel numbers according to the number of input features to capture local features, and dynamically setting the weights assigned to different time steps; spatial sequence feature coding mainly extracts the input features of spatial relationships as spatial sequence feature coding within the regular range of linear fitting based on the time resolution of the input data and the fluctuation law of sample sampling; then the outputs of the time coding layer and the space coding layer are structurally fused to form a unified spatiotemporal coding structure, and the information feature representation after spatiotemporal coding is output.
[0047] Step S3, define the neural network module, refer to Figure 2 In this module, a local-global attention layer is designed and an attention window parameter is added to capture the key information of the spatiotemporal features of the data in the spatiotemporal embedding layer. Then, a network layer is designed to achieve efficient feature transformation and enhancement, and meteorological condition constraints are injected through physical regularization terms to ensure that the model output is more consistent with physical laws.
[0048] In step S3, the custom neural network module is initialized and configured to accept two parameters: the model dimension d_model and the number of attention heads n_heads. A local-global attention layer is created within this module to implement the local-global attention mechanism. This layer adds an attention window parameter w_size = 24, indicating that the model performs attention calculations within a 24-hour time window. A class called the KAN_MLP network layer is also designed within the neural network module. This layer is the first in this field to combine the structure and features of both KANs (Kernel-Adaptive Networks) and Kolmogorov-Arnold networks, while also inheriting the structural advantages of MLPs (Multi-Layer Perceptrons). Multiple neural network sublayers are created within the network layer, forming a multi-layered subnetwork structure composed of multiple neural network sublayers stacked together. This design aims to gradually deepen feature representation and enhance the model's learning capabilities through a stacked approach. In each neural network sublayer, the input data first undergoes a linear transformation, expanding its dimension from d_model to d_model×4. Dynamically learnable radial basis functions are used as activation functions to introduce nonlinear characteristics and enhance physical interpretability. The activation function and weights are adaptively and dynamically adjusted to better match the distribution of the input data. After the activation function, the data undergoes another linear transformation, restoring its dimension from d_model×4 back to d_model, where d_model represents the input and output dimensions and d_model*4 represents the intermediate layer dimensions. The multi-layer structure of each neural network sublayer effectively transforms and enhances features. This ingenious design combines the characteristics of dual KANs with the multi-layer structure of MLPs, enabling the KAN_MLP layer to combine and approximate various functions, efficiently learning representations of complex functions and possessing stronger feature transformation capabilities while maintaining physical interpretability, thereby improving network performance.
[0049] After the model output, a physical constraint module is created to implement the physical regularization term. The output of the physical constraint module conforms to the laws of physics. Setting the parameter coef=0.1 indicates that the coefficient of the regularization term is 0.1, which is used to control the degree of influence of the regularization term on the model output. The output of the network layer is then added to the physical regularization term calculated by the physical constraint module, and meteorological condition constraints are injected to obtain the final output of the physical constraint module.
[0050] Step S4: Design a joint attention fusion layer, in which meteorological cross-attention calculation, gating weight generation, and attention result fusion are completed to effectively integrate key features from different sources and processed by the network layer (including spatiotemporal data, meteorological data, and device status data);
[0051] In step S4, a joint attention fusion layer is designed. Meteorological cross-attention calculations are performed in this fusion layer, using meteorological data as the key and historical power generation data as the query. The similarity between each power generation data point and all meteorological data points is compared to determine the degree of attention paid to different meteorological conditions. An attention weight matrix is then generated to represent the degree of attention paid by each power generation data point to each meteorological data point. Gated weight generation is then performed, and the meteorological cross-attention results obtained in the first step are input into the adaptive kernel network structure along with the local-global attention results from the previous step. By learning the characteristics of the input data, weights for the three attention types are generated. These weights represent the importance of the different attention results in the final fused output, and a weight vector containing the weights of the three attention types is output. Attention result fusion is then performed, and the weight vector generated in the previous step is used to perform a weighted sum of the three attention results. Specifically, each attention result is multiplied by its corresponding weight and then summed to obtain the fused output. This output comprehensively considers multiple information, such as meteorological cross-attention and local-global attention, and is used for subsequent photovoltaic power generation prediction.
[0052] Step S5, after step S4, designing the output layer, completing feature integration, linear transformation, setting activation function, physical regularization, and finally generating output of photovoltaic power prediction;
[0053] The features output by the joint attention fusion layer are further integrated through the output layer, and the input features are dimensionalized to ensure that they match the expected input of the output layer, so that the output layer can extract key information for the final prediction from them; at the same time, multiple fully connected layers are set, whose weights and bias parameters are learned during the training process to achieve nonlinear transformation of features and map the fused features to the target output space; a linear activation function is used in the fully connected layer, and the above-mentioned physical regularization term is called before the final prediction result is generated, and the output is additionally adjusted to ensure that the prediction result conforms to the physical laws.
[0054] During training, the entire prediction neural network uses the mean squared error (MSE) as the loss function to measure the difference between the predicted value and the true value. During training, the gradient of the loss function with respect to the model parameters is calculated, and the model parameters are updated using the backpropagation algorithm to optimize prediction performance.
[0055] Step S6, edge optimization, uses INT8 quantization and model pruning technology to compress this improved neural network model into a memory device to enable offline prediction. The specific method includes:
[0056] Step S61: First, evaluate the size of the current neural network model, including parameters such as weights and biases, and perform a benchmark test on the uncompressed model, recording performance indicators as a reference; calculate the absolute value of each weight and identify small, unimportant weights; then set a pruning threshold based on the model size and performance requirements, set weights less than the threshold to 0, remove corresponding connections, remove redundant weights and connections, and reconstruct the model structure; evaluate the importance of each layer by removing and testing the performance layer by layer, and then remove layers with less impact on performance based on the evaluation results, and gradually iterate the pruning to fine-tune the model;
[0057] Step S62: prepare a small calibration dataset for parameter adjustment during the quantization process; convert the model's weights from floating point numbers to INT8 representation; calculate the scaling factor and zero point for each weight layer to maintain numerical stability; convert the model's activation values from floating point numbers to INT8 representation; reconstruct the model based on the results of pruning and quantization to ensure the correct model structure; then use TensorRT, an inference engine optimized for INT8, and convert the quantized model to a format supported by the edge device, deploy the model to an edge device with 256KB of memory, and perform offline prediction tests on the edge device to verify the performance and accuracy of the model; finally, perform performance analysis and iterative optimization.
[0058] This paper proposes a photovoltaic forecasting method for power system scheduling, renewable energy management, and smart grid optimization. This method combines the core features of both the Kernel-Adaptive Network (KAN) and the Kolmogorov-Arnold Network (KAN), and incorporates the efficient feature extraction of multi-core Conv1Ds and an improved Transformer encoder structure. To avoid confusion, the Kernel-Adaptive Network structure will be referred to as KAN_1, and the Kolmogorov-Arnold Network as KAN_2. Considering that traditional photovoltaic prediction models mainly rely on linear statistical models, single deep learning architectures, or a combination of multiple frameworks, these existing prediction methods have certain limitations when dealing with complex spatiotemporal correlations and variable influencing factors. In this paper, a neural network called TransKANs-Fusion is designed. By considering the dual KAN structure and characteristics to capture the complex relationship between different influencing factors, and using the improved network encoder with multi-core one-dimensional convolutional connections to process historical time series photovoltaic characteristics and the coupled time dependence and spatial correlation in the influencing factors, more accurate and reliable long- and short-term photovoltaic output forecasts can be achieved.
[0059] 1) Improved Transformer encoders—TransKANs. While traditional Transformer encoders have achieved great success in natural language processing, they still have room for improvement in terms of complexity and effectiveness when processing time series data. To this end, this paper improves the Transformer encoder as follows:
[0060] (a) Position encoding: adding position information based on temporal and spatial sequences to the input elements and embedding meteorological features as position biases, enabling the model to understand the order and spatial temporal relationships in the sequence and enhance sensitivity to environmental changes.
[0061] (b) Local-global sparse attention: When processing long sequences of multidimensional data, a sliding window technique is used to limit the scope of local calculations, reduce the amount of computation, and improve the efficiency and accuracy of global attention allocation.
[0062] (c) Adaptive Dynamic Data Quantile Normalization: To address the non-Gaussian distribution of photovoltaic data and optimize data stability, we modified the traditional normalization layer into an element-by-element dynamic control layer. This core concept is the core concept of the reference layer. The output of this layer is defined as output = σ(x) ∗ tanh(x), where σ is the sigmoid function and x is the feature output from the previous layer. This design is particularly effective because layer normalization in the Transformer often produces a tanh-shaped input-output mapping.
[0063] (d) To effectively extract temporal and spatial correlation features from the photovoltaic output sequence, this paper employs a multi-kernel one-dimensional convolutional layer. This design not only captures local patterns in the input sequence but also enhances understanding of global characteristics. The multi-kernel strategy allows the model to analyze data at different scales, improving the comprehensiveness and accuracy of feature extraction.
[0064] (e) KAN_1 is used to adapt to the characteristics of non-stationary data. The present invention aims to capture complex patterns in input data by dynamically adjusting its kernel function. It is particularly suitable for dealing with rapidly changing influencing factors such as meteorological conditions. The network designed based on the KAN_2 super theorem can decompose multivariate continuous functions into single variable continuous functions and their combinations, thereby adapting to different input patterns more flexibly. It is mainly used to identify long-term trends and short-term fluctuations in historical photovoltaic output data. The traditional activation function is replaced by a dynamically learnable radial basis function. At the same time, by considering the characteristics of KAN_2, the learnable parameters are placed on the connection edges instead of on the nodes. For the KAN_2 operation of the Lth layer, its formula can be expressed as:
[0065] Output=f (W edge *x+b edge ) (3)
[0066] Where f is a nonlinear activation function, W edge and b edge These are the weights and biases on the edges, respectively, to adapt to the nonlinearity of varying meteorological conditions. A regularization term is injected into the photovoltaic equation to enforce the physical relationship of P = η⋅G⋅(1−kΔT). This activation function breaks through traditional black-box models and can reveal the impact of predictions under physical conditions.
[0067] (f) Reduce computational complexity and improve efficiency. The MatMul-free Channel Mixer layer is introduced into the neural network module of the present invention. In the network layer calculation, ternary weights (-1, 0, +1) are used instead of traditional floating-point weights, and matrix multiplication is converted into an efficient ternary accumulation operation, thereby greatly reducing the computational complexity and effectively improving the computational efficiency of the neural network model, making it more suitable for applications in resource-constrained environments.
[0068] The key module structure of the improved Transformer encoder is compared with the key modules of the traditional Transformer encoder. The advantages of the improvement are summarized in Table 1.
[0069] Table 1 Module Traditional Transformer TransKANs-Fusion Solution Improvements Advantages Attention Mechanism Global self-attention Local-Global Sparse Attention By limiting the local calculation range through sliding windows, the efficiency of long sequence processing is improved by 40%. Positional encoding Sine function encoding Learnable weather sensing position encoding Embed meteorological characteristics such as temperature and irradiance as position deviation Normalization layer Layer Normalization Adaptive Quantile Normalization for Dynamic Data Optimizing stability for non-Gaussian distribution of photovoltaic data
[0070] 2) A multi-attention joint fusion mechanism connected to TransKANs was designed to dynamically fuse three types of attention: spatiotemporal, meteorological, and device status. Using a specific fusion strategy (dynamic weight gating (weights generated by KANs)), the results of these different types of attention were effectively integrated to form a comprehensive feature representation that captures the long-term and short-term dependencies of the power generation sequence. The specific calculation method, fusion strategy, and function of the multi-attention joint fusion mechanism are shown in Table 2.
[0071] Table 2 Attention Type Calculation method Fusion Strategy Function Description Spatiotemporal attention Multi-head self-attention + temporal convolution Dynamic weight gating (weights generated by KAN) Capturing the long-term and short-term dependencies of power generation sequences Meteorological Cross-Attention Query-Key meteorological characteristics cross Attention scaling based on Pearson correlation coefficient Focus on high-impact factors (such as cloud cover) Device state attention Graph Attention Network (GAT) Adjacency matrix combined with device topology Modeling spatial correlation between photovoltaic arrays
[0072] The multi-attention joint fusion mechanism takes the output of TransKANBlock as input, and the output formula is as follows:
[0073] (4)
[0074] where w i is a learnable weight, and σ is a Sigmoid function to achieve dynamic weighting.
[0075] Finally, in the output layer, feature maps from different channels are spliced together to form a richer feature representation that contains information at multiple time scales. After passing through multiple fully connected layers, the fused features are mapped to the target output space to generate the final photovoltaic prediction results.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A photovoltaic prediction method, characterized in that: The steps include: Step S1, data preprocessing: mainly align the time series of historical power generation data with the meteorological data through the time series regularization strategy, design feature engineering, extract the cumulative change rate of irradiance, and construct the equipment health index, which is input into the input layer; Step S2: Based on the preprocessing results of step S1, a spatiotemporal embedding layer is designed to perform spatiotemporal encoding on the data transmitted by the input layer and all meteorological data branches according to the time series characteristics and spatial series characteristics; Step S3: Define a neural network module. Design a local-global attention layer in this module and add attention window parameters to capture key information about the spatiotemporal features of the data in the spatiotemporal embedding layer. Then design a network layer to achieve efficient feature transformation and enhancement. In addition, inject meteorological constraints through physical regularization terms to ensure that the model output is more consistent with physical laws. Step S4: Design a joint attention fusion layer, in which meteorological cross-attention calculation, gating weight generation, and attention result fusion are completed to effectively integrate key features from different sources and processed by the network layer; Step S5, after step S4, designing the output layer, completing feature integration, linear transformation, setting activation function, physical regularization, and finally generating output of photovoltaic power prediction; Step S6, edge optimization, compresses this improved neural network model to the memory device through INT8 quantization and model pruning technology, enabling offline prediction.
2. A photovoltaic prediction method according to claim 1, characterized in that: In step S1, firstly, any input processing data I is given in ={P1,P2,P3,……,P m ,R n }, where P i ={d i1 ,d i2 ,…,d im |1≤i≤m,m∈N + ,d im ∈ }, P i is the input photovoltaic data set of column i, m is the number of input photovoltaic feature columns, d im is the mth photovoltaic data value in the i-th column, R n It is a set of all influencing factors n, which supports dynamic addition and mainly includes meteorological data and user behavior patterns.
3. A photovoltaic prediction method according to claim 1, characterized in that: In step S1, it is necessary to use the time series regularization strategy to accurately find the optimal alignment path between the two sequences and complete the processing of missing values and outliers. The specific method is as follows: First, the time series of historical power generation data and meteorological data are constructed into a two-dimensional cost matrix. The elements in the matrix represent the differences between corresponding points in the two series. Secondly, starting from the starting point of the matrix, the cumulative cost matrix is traversed step by step to calculate the cumulative cost matrix. The elements of the cumulative cost matrix represent the optimal path cost from the starting point to the current point. Starting from the end point of the cumulative cost matrix, backtracking along the optimal path is used to find the optimal alignment path between the two time series data. Then, missing values in the aligned data are checked and processed using interpolation filling. Outliers are identified and processed by removing the mean. The cumulative rate of change of irradiance is then calculated using the following formula: , (1) Where s is the size of the time window, V is the rate of change of irradiance, The cumulative rate of change of irradiance, which reflects the trend of irradiance change over time, is used to assist in analyzing its impact on power generation; When constructing the equipment health index, irradiance, power generation, temperature, wind speed, wind direction, humidity, equipment operating time, maintenance records, and fault history records are selected as key features. These features are normalized to their maximum and minimum values so that they are on the same scale. The resulting normalized feature set is set as Nor, and the following formula function is used to calculate the health index: (2) where w i is the weight of feature i.
4. A photovoltaic prediction method according to claim 1, characterized in that: Step S2, time series feature encoding involves setting multiple columns of one-dimensional convolution operations with different kernel numbers according to the number of input features to capture local features, and dynamically setting the weights assigned to different time steps; spatial series feature encoding, mainly based on the time resolution of the input data and the fluctuation pattern of sample sampling, extracts the input features of the spatial relationship within the regular range of linear fitting as spatial series feature encoding; The outputs of the time coding layer and the space coding layer are then structurally fused to form a unified spatiotemporal coding structure, and the information feature representation after spatiotemporal coding is output.
5. A photovoltaic prediction method according to claim 3, characterized in that: In step S3, the custom neural network module is first initialized and set to accept two parameters, namely the model dimension d_model and the number of attention heads n_heads; a local-global attention layer is created in the module to implement the local-global attention mechanism. This layer adds the attention window parameter w_size=24, and multiple neural network sublayers are created in the network layer; in each neural network sublayer, the input data first undergoes a linear transformation to expand the dimension from d_model to d_model×4, which uses a dynamically learnable radial basis function as the activation function to introduce nonlinear characteristics and enhance physical interpretability; after the activation function, the data undergoes another linear transformation to restore the dimension from d_model×4 back to d_model.
6. A photovoltaic prediction method according to claim 5, characterized in that: Step S3: Create a physical constraint module after the model output to implement the physical regularization term. The output of the physical constraint module conforms to the laws of physics. Set the parameter coef=0.1 to control the influence of the regularization term on the model output. Then add the output of the network layer to the physical regularization term calculated by the physical constraint module, and inject meteorological condition constraints to obtain the final output of the physical constraint module.
7. A photovoltaic prediction method according to claim 1, characterized in that: In step S4, a joint attention fusion layer must first be designed. In this fusion layer, meteorological cross-attention calculation is first performed. Meteorological data is required as the key and historical power generation data as the query. By comparing the similarity between each power generation data point and all meteorological data points, the degree of attention to different meteorological conditions is determined; then an attention weight matrix is generated to represent the degree of attention of each power generation data point to each meteorological data point; secondly, gated weight generation is performed, and the meteorological cross-attention results obtained in the first step are input into the adaptive kernel network structure together with the local-global attention results of the previous step; by learning the characteristics of the input data, the weight distribution of the three types of attention is generated; then a weight vector is output, which contains the weight values of the three types of attention; followed by attention result fusion, and the weight vector generated in the previous step is used to perform weighted summation of the three types of attention results.
8. A photovoltaic prediction method according to claim 1, characterized in that: The key features in step S4 include spatiotemporal data, meteorological data, and equipment status data.
9. A photovoltaic prediction method according to claim 6, characterized in that: In step S5, the features output by the joint attention fusion layer are further integrated through the output layer, and the input features are dimensionally adjusted to ensure that they match the expected input of the output layer; at the same time, multiple fully connected layers are set, whose weights and bias parameters are learned during the training process to achieve nonlinear transformation of features and map the fused features to the target output space; a linear activation function is used in the fully connected layer, and the above-mentioned physical regularization term is called before the final prediction result is generated, and the output is additionally adjusted to ensure that the prediction result conforms to the physical laws.
10. The photovoltaic prediction method according to claim 1, wherein: In step S6, the edge optimization method includes: Step S61: First, evaluate the size of the current neural network model, including weights and biases, and perform a benchmark test on the uncompressed model, recording performance indicators as a reference; calculate the absolute value of each weight and identify small, unimportant weights; then set a pruning threshold based on the model size and performance requirements, set weights less than the threshold to 0, remove corresponding connections, remove redundant weights and connections, and reconstruct the model structure; evaluate the importance of each layer by removing and testing the performance layer by layer, then remove layers based on the evaluation results, and iterate pruning step by step to fine-tune the model; Step S62: prepare a small calibration dataset for parameter adjustment during the quantization process; convert the model's weights from floating point numbers to INT8 representation; calculate the scaling factor and zero point for each weight layer to maintain numerical stability; convert the model's activation values from floating point numbers to INT8 representation; reconstruct the model based on the results of pruning and quantization to ensure the correct model structure; then use TensorRT, an inference engine optimized for INT8, and convert the quantized model to a format supported by the edge device, deploy the model to an edge device with 256KB of memory, and perform offline prediction tests on the edge device to verify the performance and accuracy of the model; finally, perform performance analysis and iterative optimization.
Citation Information
Patent Citations
Short-term photovoltaic output prediction method based on multi-model fusion
CN114781723A
Photovoltaic forecasting method based on deep attention network and clear sky radiation prior fusion
CN114971058A
Multi-data fusion photovoltaic power prediction algorithm integrating meteorological factor data
CN115994605A
Photovoltaic power prediction method based on multi-scale space-time diagram attention convolutional network
CN117154704A
Hybrid photovoltaic power prediction method and system based on multi-source data fusion
US20220373984A1
Cited By
Buoy meteorological data restoration method based on space-time double-attention neural network
CN120832480A
Water supply network leakage node and time window joint positioning method based on hybrid neural network
CN121388739A
Photovoltaic power generation amount prediction method based on multi-model cooperation, time sequence feature correction and fusion
CN122757886A