A Photovoltaic Power Prediction Method and System Based on KAN
By adopting a KAN-based method in the photovoltaic power prediction model, using the learnable activation function and multi-head attention mechanism, the problem of insufficient accuracy when processing high-dimensional input features is solved, and higher prediction accuracy and training efficiency are achieved.
Patent Information
- Application Number
- CN202510059930.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-15
AI Technical Summary
When existing photovoltaic power prediction models deal with deep-level complex high-dimensional input features, it is difficult to capture sufficient detailed information, resulting in limited prediction accuracy.
The photovoltaic power prediction method based on KAN (Kolmogolov-Arnold Network) is adopted to improve the approximation ability of the model to high-dimensional complex functions using the learnable activation function, and spatial and timing features are extracted through the improved deep KAN and multi-head attention mechanism.
The nonlinear fitting effect of the model is significantly enhanced, allowing the model to capture deeper input features, improve the accuracy of photovoltaic power prediction, and improve training efficiency and stability.
Smart Images

Figure CN119482456B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wind power prediction, and particularly to a photovoltaic power prediction method and system based on KAN. Background Art
[0002] Deep learning methods have now become the research objects of many scholars in the field of photovoltaic prediction. Existing models such as Long Short-Term Memory (LSTM), Bidirectional Long Short-Term Memory (BiLSTM), Convolutional Neural Network (CNN), and their hybrid models CNN-LSTM and CNN-BiLSTM are often used to capture the complex mapping relationship between climate characteristics and actual photovoltaic output power, so as to achieve the purpose of prediction. However, the deficiencies of the above existing models are that they all place non-learnable activation functions on neurons, resulting in limited non-linear fitting ability of the models. When encountering deep, complex, and high-dimensional input features, the above models often have difficulty capturing sufficient detailed information, resulting in limited prediction accuracy. Summary of the Invention
[0003] The purpose of the present invention is to propose a photovoltaic power prediction method and system based on KAN. With KAN (Kolmogorov-Arnold network) as the basic architecture, a day-ahead prediction model for photovoltaic output is constructed. By using learnable activation functions, the approximation ability for high-dimensional complex functions is improved, the non-linear fitting effect of the model is significantly enhanced, enabling the model to capture deeper input features, and thus improving the accuracy of photovoltaic power prediction.
[0004] The present invention is achieved by at least one of the following technical solutions.
[0005] A photovoltaic power prediction method based on KAN includes the following steps:
[0006] 1) Collect the starting capacity value, weather forecast, and actual power generation data of the photovoltaic power station as the original data set, preprocess the data in the original data set, and divide it into a training set and a test set;
[0007] 2) Construct a day-ahead prediction model for photovoltaic output and train the day-ahead prediction model for photovoltaic output;
[0008] 3) Test and evaluate the trained day-ahead prediction model for photovoltaic output;
[0009] 4) Use the day-ahead prediction model for photovoltaic output to predict photovoltaic power.
[0010] Furthermore, in step 1), the preprocessing of the data in the original dataset includes the following steps:
[0011] a) Set the weather features corresponding to an output power of 0 to 0;
[0012] b) Replace the data points with missing or out-of-normal-range weather features and output power with the average value of the two adjacent time instants;
[0013] c) Perform normalization processing on the data of the same type according to the following formula:
[0014] ;
[0015] where is the u-th data in the original dataset and represents the data at one time instant. The original dataset includes all the installed capacity values, weather forecasts, and actual power generation data at a certain time instant. U is the number of data in the original dataset, and are the maximum and minimum values of the data of the same type in the original dataset respectively, and is the normalized data value;
[0016] d) Remove the data corresponding to the time instants with normalized power values greater than 1 or less than 0;
[0017] e) Take the data of the same day as a sample, divide the data processed through steps a) - d) into 600 - 1200 samples, and divide the samples into a training set and a test set according to a ratio.
[0018] Furthermore, in step 2), the day-ahead photovoltaic power prediction model includes a spatial feature extraction module, a temporal feature extraction module, a spatio-temporal feature fusion and predicted power output module; the spatial feature extraction module and the temporal feature extraction module extract features in parallel, and the extracted features are input into the spatio-temporal feature fusion and predicted power output module to output the photovoltaic power prediction result.
[0019] Furthermore, the spatial feature extraction module includes an improved deep KAN. The improved deep KAN includes multiple basic units connected in sequence, and each basic unit consists of two stacked KAN layers and a skip connection; the spatial feature extraction module finally outputs a result vector.
[0020] Furthermore, the temporal feature extraction module combines the multi-head attention mechanism with KAN, specifically including:
[0021] (1) Map the input features to the query, key, and value vector spaces respectively using three independent two-layer KANs to obtain the query matrix , key matrix , value matrix :
[0022] ;
[0023] where is the input feature matrix with dimension , and the dimensions of , , , , are all , , , , , , are all KANs with dimension , is the number of time points of the input sample, is the input feature dimension, is the number of heads of the multi-head attention mechanism;
[0024] (2) Block the query matrix , key matrix , value matrix by columns respectively. The block matrices with the same subscript belong to the same head of the attention mechanism, and each head calculates the attention independently:
[0025] ;
[0026] ;
[0027] where is the calculation result of the g-th head of the attention mechanism, , , , are the query matrix, key matrix, and value matrix of the g-th head of the attention mechanism respectively, and there are a total of G heads of the attention mechanism; is the activation function, is a matrix, means operating on each row vector of . Let a certain row vector be , then 's calculation is:
[0028] ;
[0029] Apply the activation function to the vector to adjust the sum of the elements of the vector to 1, is the -th The number of elements is \(J\), where \(J\) is the vector length, ;
[0030] (3) Concatenate the calculation results of the multi-head attention mechanism and output the time series feature extraction results through a fully connected layer:
[0031] ;
[0032] ;
[0033] where is the matrix after concatenating the calculation results of the multi-head attention mechanism, is the calculation result of the \(G\)-th head attention mechanism, , are the weight matrix and bias of the fully connected layer, is the output result of the time series feature extraction module.
[0034] Further, the spatio-temporal feature fusion and predicted power output module concatenates the extraction results of the spatial feature extraction module and the time series feature extraction module, and outputs the photovoltaic power prediction result through a fully connected layer:
[0035] ;
[0036] where , are the output results of the spatial feature extraction module and the time series feature extraction module respectively, , are the weight matrix and bias of the fully connected layer, is the predicted photovoltaic power.
[0037] Further, in step 2), the training of the photovoltaic output day-ahead prediction model includes:
[0038] Input the weather features of the training set into the photovoltaic output day-ahead prediction model. The photovoltaic output day-ahead prediction model outputs the predicted photovoltaic power. Calculate the loss function value according to the following formula, and aim to minimize the loss function. Iteratively update the model parameters repeatedly until the loss function no longer changes significantly:
[0039] ;
[0040] where, is the root mean square error of is the true output power sequence in the training set, is the p -th element of the true output power sequence, is the model output prediction sequence, is thep Element is the loss function value, and P is and the sequence length of
[0041] Furthermore, in step 3), the day-ahead prediction model test of photovoltaic output includes:
[0042] Input the weather characteristics of the test set into the trained day-ahead prediction model of photovoltaic output. The day-ahead prediction model of photovoltaic output outputs the predicted photovoltaic power, and evaluates the prediction accuracy of the day-ahead prediction model of photovoltaic output through evaluation indicators:
[0043]
[0044]
[0045] where is the monthly average root mean square error, is the actual power at time is the predicted photovoltaic power at time is the on-line capacity at time is the number of time points in the whole month, is the monthly average accuracy rate, is the monthly average qualified rate, is the predicted qualified judgment result at time equal to 1 means qualified, equal to 0 means unqualified.
[0046] Furthermore, in step 4), predicting the photovoltaic output of the day to be predicted specifically includes the following steps:
[0047] 41) Collect the real-time weather forecast data of the day to be predicted;
[0048] 42) Perform the above-mentioned preprocessing on the real-time weather forecast data;
[0049] 43) Input the preprocessed weather characteristic data into the day-ahead prediction model of photovoltaic output that meets the test requirements, and obtain the initial prediction through model calculation, where represents the prediction value at the o th time point of the day;
[0050] 44) According to the on-line capacity at each moment, perform anti-normalization calculation on the initial prediction output to obtain the final power prediction value. The anti-normalization calculation is:
[0051] ;
[0052] in For the o The final prediction result at a time point is for o Always on capacity.
[0053] A system for implementing the photovoltaic power prediction method based on KAN includes a photovoltaic output day-ahead prediction model, wherein the photovoltaic output day-ahead prediction model includes a spatial feature extraction module, a temporal feature extraction module, a spatiotemporal feature fusion and a power output prediction module;
[0054] The spatial feature extraction module is used to extract the spatial features of variables at the same time point in the sample data;
[0055] The time series feature extraction module is used to extract the time series features of variables at different time points in the sample data;
[0056] The function of the spatiotemporal feature fusion and predicted power output module is to fuse the spatial feature with the temporal feature and output the predicted photovoltaic power.
[0057] Compared with the prior art, the present invention has the following beneficial effects:
[0058] 1) Higher prediction accuracy: Compared with traditional LSTM, BiLSTM, CNN-LSTM, CNN-BiLSTM and other models, the present invention adopts KAN as the basic architecture to construct a photovoltaic output day-ahead prediction model. It uses learnable activation functions to improve the approximation ability of high-dimensional complex functions and significantly enhances the nonlinear fitting effect of the model, so that the model can capture deeper input features, thereby improving the accuracy of photovoltaic power prediction.
[0059] 2) Higher training efficiency and stability: Compared with the original deep KAN, the present invention effectively alleviates the gradient vanishing problem caused by the increase in the number of network layers by introducing a residual structure in KAN, thereby ensuring the training efficiency of the deep network and the stability of the model.
[0060] 3) Capability of extracting spatiotemporal features: This invention not only uses the improved deep KAN to extract the spatial features of the data, but also combines the attention mechanism with the powerful nonlinear fitting ability of KAN to extract the temporal features of the data, and integrates the spatiotemporal features to output the prediction results. In this way, the photovoltaic output day-ahead prediction model can simultaneously capture the interaction relationship of the data in the spatial dimension and the dynamic change law in the temporal dimension, showing higher prediction accuracy and stronger generalization ability.
[0061] 4) Wide applicability: The method of the present invention can not only be used for photovoltaic power prediction, but also can be extended to other time series prediction tasks that need to process high-dimensional complex inputs, such as wind power prediction and load demand prediction, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a schematic diagram of the shallow KAN structure of the embodiment;
[0063] Figure 2 It is a schematic diagram of the structure of a photovoltaic power prediction system based on KAN of the embodiment;
[0064] Figure 3 It is a comparison chart of the training effects before and after the improvement of the deep KAN;
[0065] Figure 4 It is a result chart of power prediction using Scenario 4 by each model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0067] A photovoltaic power prediction system based on KAN in this embodiment includes a photovoltaic output day-ahead prediction model. The photovoltaic output day-ahead prediction model is based on the Kolmogorov-Arnold Networks (KAN) as the basic architecture. The structural schematic diagram of the whole model is as Figure 2 shown. The whole photovoltaic output day-ahead prediction model is divided into three modules: a spatial feature extraction module, a temporal feature extraction module, and a spatio-temporal feature fusion and predicted power output module. The photovoltaic output day-ahead prediction model takes weather features as input and the photovoltaic power prediction value as output. The weather features include humidity, temperature, irradiance, pressure, wind speed, wind direction, and so on.
[0068] The Kolmogorov-Arnold Network (KAN) is a neural network model inspired by the Kolmogorov-Arnold Representation Theorem (KART). Different from neural network models such as LSTM, BiLSTM, CNN-LSTM, and CNN-BiLSTM, KAN places learnable activation functions on the weights, making its approximation ability for high-dimensional complex continuous functions stronger. Taking the multivariate continuous function on a bounded domain in Equation (1) as an example, according to KART, is expressed as a two-layer nested superposition of finite single-variable continuous functions:
[0069] ;
[0070] ;
[0071] ;
[0072] ;
[0073] where the independent variable of the function has a dimension of , ∈ , and the value range of the independent variable is , is the set of real numbers; the value ranges of both the independent variable and the dependent variable of the function are R , q is the function number; the value range of the independent variable of the function is [0,1], and the value range of the dependent variable is R, p is the number of the function.
[0074] Equation (1) is the calculation principle of the shallow KAN. Taking a two-dimensional input variable and a one-dimensional output variable as an example, the schematic diagram of the shallow KAN network structure is as shown in Figure 1 . is the learnable activation function , and its expression is:
[0075] ;
[0076] where is the B-spline curve, is the independent variable , is the number of the B-spline curve, The corresponding learnable parameter is , To control the overall size of the learnable activation function.
[0077] To generalize the network and enhance the non-linear mapping ability of KAN, a two-dimensional univariate function matrix is used to define the KAN layer, and it is proved that the KAN layers can be stacked to construct a deep KAN to fit more complex functions. The two-dimensional univariate function matrix is shown in Equation (3):
[0078] ;
[0079] where is the element in the th row and th column of the two-dimensional univariate function matrix corresponds to the activation function shown in Equation (2), is the input dimension of the KAN layer, , .
[0080] Figure 1 The KAN shown in is composed of two stacked KAN layers. The input of the first KAN layer is , and the output is . The input of the second KAN layer is .
[0081] The role of the spatial feature extraction module described in the present invention is to extract the variable features of the sample data at the same time point, and to learn the influence of the interaction between the input weather variables at the same time point on the photovoltaic power at this time point. Since the mapping relationship between the weather features and the output power time is very complex, the spatial feature extraction module needs to have a strong non-linear fitting ability to fully learn the key rules. Therefore, the present invention uses the improved deep KAN as the spatial feature extraction module. However, when the number of KAN layers increases, the training efficiency of the photovoltaic output day-ahead prediction model usually decreases. The reason is that when the number of KAN layers increases, the gradient becomes smaller and even disappears as the number of layers increases, resulting in very slow parameter updates and the model is difficult to converge. To solve this problem, the present invention adds a residual structure to the original deep KAN. By introducing skip connections (such as the skip connections shown in the spatial feature extraction module in Figure 2 ), the outputs between layers are directly added together, so that the gradient can be more directly passed back to the shallower layers during the backpropagation process, thereby alleviating the problem of gradient disappearance.
[0082] As an embodiment, the improved deep KAN in this embodiment takes two KAN layers (Kanlayer) and a skip connection as a basic unit, as shown in the basic unit in the spatial feature extraction module of Figure 2 , and the following number is the serial number. There are 2 n KAN layers, and a total of n basic units. This network structure is a residual structure. Let the final output result vector of the spatial feature extraction module be , and its dimension depends on the output dimension of the last KAN layer. By stacking multiple basic units, the improved deep KAN is constructed. The improved deep KAN solves the problem of difficult training and can well extract the spatial relationship between input features.
[0083] The function of the temporal feature extraction module is to extract the variable features at different time points in the sample data, mainly to learn the influence of the weather features at different time points on the output power at the current time point. Therefore, the output of the photovoltaic temporal features at each time point needs to use the weather features at all time points as inputs. By combining the multi-head attention mechanism with the improved deep KAN, the temporal feature extraction module is built. The specific process is as follows:
[0084] (1) Use three independent two-layer KANs to map the input features into the query, key, and value vector spaces respectively, and obtain the query matrix , the key matrix , and the value matrix :
[0085] ;
[0086] where is the input feature matrix with a dimension of , and the dimensions of , , are all , , , are all KANs with a dimension of , is the number of time points of the input sample, is the input feature dimension, is the number of heads of the multi-head attention mechanism;
[0087] (2) Block the query matrix , the key matrix , and the value matrix by columns respectively. The block matrices with the same subscript belong to the same head of the attention mechanism, and each head performs attention calculation independently:
[0088] ;
[0089] ;
[0090] where is the calculation result of the g-th head attention mechanism, , , are the query matrix, key matrix, and value matrix of the g-th head attention mechanism respectively, and there are a total of G head attention mechanisms; is the activation function. Let the vector , is calculated as:
[0091] ;
[0092] Apply the activation function to the vector to adjust the sum of each element of the vector to 1. is the -th element of the vector , J is the length of the vector, ;
[0093] (3) Concatenate the calculation results of the multi-head attention mechanism and output the time series feature extraction result through the fully connected layer:
[0094] ;
[0095] ;
[0096] where is the matrix after concatenating the calculation results of the multi-head attention mechanism, is the calculation result of the G-th head attention mechanism, , are the weight matrix and bias of the fully connected layer, is the output result of the time series feature extraction module.
[0097] The function of the spatio-temporal feature fusion and predicted power output module is to fuse the time series features extracted by the time series feature extraction module with the spatial features extracted by the spatial feature extraction module, and at the same time output the predicted photovoltaic power value.
[0098] In order to avoid further deepening the depth of the entire model, the present invention adopts a network model structure for parallel extraction of spatial features and time series features, and adopts a form of parallel extraction of time series features and spatial features. Compared with the serial extraction method of spatio-temporal features, parallel extraction is beneficial to reducing the network depth and the training difficulty of the network model. After extracting the spatio-temporal features, the extraction results are concatenated and the photovoltaic output prediction result is output through the fully connected layer. The process is as follows:
[0099] ;
[0100] wherein , are the output results of the spatial feature extraction module and the temporal feature extraction module respectively, , are the weight matrix and bias of the fully connected layer, is the predicted photovoltaic power.
[0101] To implement the photovoltaic power prediction method of the described photovoltaic power prediction system based on KAN, the entire photovoltaic power prediction process can be divided into historical data collection, historical data preprocessing, model offline training, model testing, and real-time data collection, real-time data preprocessing, and model real-time prediction phases. Before model offline training, model testing, and model real-time prediction, it is necessary to preprocess the input data. The data used during model offline training and model testing is historical data, while the data used during model real-time prediction is real-time data obtained from the photovoltaic power station. The model can be used for real-time prediction only after passing the test requirements after offline training. The photovoltaic power prediction method based on KAN in this embodiment specifically includes the following steps:
[0102] 1) Collect historical data;
[0103] Obtain historical data from the photovoltaic power station to be applied as the original data set, and these data need to include weather characteristics, actual photovoltaic output power, and actual installed capacity.
[0104] 2) Data preprocessing, and perform data preprocessing according to the following steps;
[0105] a) Set the weather characteristics corresponding to an output power of 0 to 0, because at this time the photovoltaic power generation unit has not started working yet. Setting the weather characteristics to zero can eliminate the influence of this human factor on the model training logic.
[0106] b) Replace the data points with missing or abnormal weather characteristics and output power with the average value of the adjacent two time instants. If the number of missing data time points in a day is too large (more than 20% of all data points in a day), then discard the sample. Abnormal data points are data points outside the normal range, and the normal range of data points is provided by the photovoltaic power station.
[0107] c) Different types of data have different units and different numerical magnitudes. To reduce the difficulty of the photovoltaic output day-ahead prediction model in learning data characteristics, perform normalization processing on the same type of data according to the following formula:
[0108] ;
[0109] wherein is the u-th data in the original dataset and represents the data at a certain moment. The original dataset includes all weather forecast values and actual power values at a certain moment. U is the number of data in the original dataset and and are the maximum and minimum values of the same type of data in the original dataset respectively, and is the data value after normalization.
[0110] For power data, take the installed capacity at this moment, and take 0.
[0111] d) Remove the data at the corresponding moments where the power value after normalization is greater than 1 or less than 0.
[0112] e) Among the data processed through steps a) - d), divide the data belonging to the same day into N samples, where H samples are the training set and m samples are the test set.
[0113] 3) Construct and train the day-ahead photovoltaic output prediction model. The data used for model training is from the training set, aiming to optimize the model parameters.
[0114] Input the weather features of the training set into the day-ahead photovoltaic output prediction model. The day-ahead photovoltaic output prediction model outputs the predicted photovoltaic power. Calculate the loss function value according to the following formula, aiming to minimize the loss function, and repeatedly iterate and update the model parameters until the loss function no longer changes significantly:
[0115] ;
[0116] where, is the root mean square error of is the true output power sequence in the training set, is the p -th element of the true output power sequence, is the model output prediction sequence, is the p -th element of the model output prediction sequence, is the loss function value, is and the sequence length of
[0117] 4) Model testing;
[0118] Use the data in the test set to test the day-ahead photovoltaic output prediction model, aiming to evaluate the performance of the day-ahead photovoltaic output prediction model, test the generalization ability of the day-ahead photovoltaic output prediction model on unknown data, and find out whether the day-ahead photovoltaic output prediction model is overfitted. Then, select and adjust the day-ahead photovoltaic output prediction model to ensure that the performance of the day-ahead photovoltaic output prediction model meets the requirements in actual applications. Essentially, it is a simulation experiment to simulate the actual application effect of the day-ahead photovoltaic output prediction model.
[0119] Specifically, input the weather characteristics of the test set into the trained day-ahead photovoltaic output prediction model. The day-ahead photovoltaic output prediction model outputs the predicted photovoltaic power, and evaluate the day-ahead photovoltaic output prediction model through the prediction accuracy. According to the national standard, use the monthly average accuracy (the accuracy defined by the national standard), the monthly average qualification rate to evaluate the prediction accuracy of the day-ahead photovoltaic output prediction model:
[0120]
[0121]
[0122] where is the monthly average root mean square error, is the actual power at time is the predicted photovoltaic power at time is the installed capacity at time is the number of time points in the whole month, is the predicted qualification judgment result at time equal to 1 means qualified, equal to 0 means unqualified.
[0123] According to the national standard, the prediction accuracy of the day-ahead photovoltaic output prediction model must reach a monthly average accuracy of more than 85%, and the monthly average qualification rate is more than 85%.
[0124] 5) Real-time data acquisition;
[0125] The day-ahead photovoltaic output prediction model that meets the test requirements can be used in actual projects. In actual applications, it is necessary to obtain the weather characteristics of numerical weather forecasts in real time as the input of the day-ahead photovoltaic output prediction model for photovoltaic output prediction.
[0126] 6) Data preprocessing;
[0127] Before inputting the weather characteristic data into the day-ahead photovoltaic output prediction model, it is necessary to perform the steps in 2) on the weather characteristic data Normalization operation, where 、 takes the same value as in step 2).
[0128] 7) Real-time prediction of photovoltaic output for the day-ahead;
[0129] Input the processed weather features into the day-ahead photovoltaic output prediction model for prediction, and calculate the initial prediction through the model and output it where represents the predicted value at the o th time point of the day, with a total of 96 time points. Subsequently, according to the starting capacity at each moment, inverse normalization calculation is performed on the initial prediction to obtain the final power prediction value. The inverse normalization calculation formula is as follows:
[0130] ;
[0131] where is the final prediction result at the o th time point, is o the starting capacity at the moment.
[0132] As a specific embodiment, this embodiment conducts a simulation experiment based on the above system and method. The original data for the simulation comes from a certain photovoltaic power prediction competition. The original data for the simulation contains 66,720 pieces of data from a certain photovoltaic power plant between 2016 and 2018. The time interval between each piece of data is 15 minutes, including six weather feature forecast values of irradiance, wind speed, wind direction, temperature, pressure, and humidity, as well as the numerical feature of actual power generation. After preprocessing the original data for the simulation, 66,720 pieces of data are obtained; among the 66,720 pieces of data, the data on the same day is used as a sample, and a total of 695 samples are divided. Among these 695 samples, 600 samples are the training set and 95 samples are the test set. Since this data set is desensitized data, the subsequent simulation experiment results shown are all based on the data values after min_max normalization. And since this data set lacks starting capacity data, the maximum output power value in the data is taken to approximate the starting capacity at all moments.
[0133] The present invention proposes an improved deep KAN for the problem that the original deep KAN is difficult to train. The improved deep KAN serves as a spatial feature extraction module in the day-ahead photovoltaic output prediction model of the present invention to extract the features of the input data. To highlight the improvement effect, now the six weather feature forecast values of irradiance, wind speed, wind direction, temperature, pressure, and humidity are used as the input of the improved deep KAN and the original deep KAN, and the predicted power is output. Record the change of the loss value of the improved deep KAN and the original deep KAN during the training process with the number of iterations, as Figure 3。
[0134] From Figure 3 As can be seen, during the first round of training, the improved Deep KAN has a larger loss value compared to the original Deep KAN. This may be caused by the unreasonable setting of the initial values of the parameters of the improved Deep KAN. However, it does not affect the final training effect. Starting from the second round of iteration, the improved Deep KAN can basically converge to the optimal value, while the loss value of the original Deep KAN always remains at a relatively high level and cannot be further reduced. This shows that the improvement of the original Deep KAN in the present invention alleviates the problem of difficult training, enabling the improved Deep KAN to fit the data more accurately and quickly.
[0135] In order to highlight the powerful feature extraction ability of the improved Deep KAN of the present invention by comparison, a feature ablation experiment is now conducted on the existing day-ahead prediction models (Long Short-Term Memory (LSTM), Bidirectional Long Short-Term Memory (BiLSTM), Convolutional Long Short-Term Memory Network (CNN-LSTM), Convolutional Bidirectional Long Short-Term Memory Network (CNN-BiLSTM)) and the day-ahead prediction model of photovoltaic power output built in the present invention.
[0136] The feature ablation experiment studies the influence of each input feature on the model prediction result by increasing or decreasing the input features, and then reflects the extraction ability of each model for different input features. The Pearson correlation coefficient and Spearman correlation coefficient between the output power and each input variable are calculated using the following formula to screen out effective input features, and further distinguish shallow-level features and deep-level features as different input schemes to evaluate the extraction ability of the model for input features at different levels. The calculation results are shown in Table 1. The Spearman correlation coefficient of each input feature is greater than or close to 0.8, indicating that each feature belongs to an effective input feature and can theoretically play a role in improving the prediction accuracy; while the Pearson correlation coefficients vary greatly, and the higher the Pearson correlation coefficient value, the higher the linear correlation between the input feature and the output power, the shallower the level, and the easier it is for the model to extract. Therefore, the input features are divided into four parts according to the level of depth with the Pearson correlation coefficient as the evaluation criterion: shallow-level input features (irradiance), ordinary input features (temperature, pressure), relatively deep-level input features (wind direction), and deep-level input features (wind speed, humidity). The corresponding four input feature schemes are as follows:
[0137] Scheme 1: Only use the shallow-level input features as the input variables of the model;
[0138] Solution 2: Use the shallow input features and ordinary input features as the model input variables;
[0139] Solution 3: Use the shallow input features, ordinary input features, and relatively deep input features as the model input variables;
[0140] Solution 4: Use the shallow input features, ordinary input features, relatively deep input features, and deep input features as the model input variables.
[0141] Calculate the Pearson correlation coefficient and Spearman correlation coefficient of each input variable using the following formula:
[0142] ;
[0143] where is the Pearson correlation coefficient between variables X and Y , and are the sample means of variables X and Y respectively, and are the X and Y th samples of variables
[0144] The Spearman correlation coefficient X between variable Y and variable is:
[0145] ;
[0146] where is the corresponding rank difference, is the rank of the X th sample in variable is the rank of the Y th sample in variable
[0147] Table 1 Calculation results of the correlation characteristics between weather features and output power
[0148]
[0149] Use the trained day-ahead photovoltaic output prediction model to make predictions on the test set. The calculation results of the evaluation indicators for the experimental results of the above four solutions are shown in Table 2. Among them The change rate refers to the The rate of change. Since the feature input scheme adds features from shallow to deep, the rate of change can reflect the change in the model's feature extraction ability.
[0150] Table 2 Results of Feature Ablation Experiments
[0151]
[0152] The bold characters in Table 2 represent the optimal indicators of all models in this scheme. Since the CNN-LSTM model, CNN-BiLSTM model, and the model of the present invention are all models for multi-input features, while Scheme 1 has only one input feature, so Scheme 1 only conducts experiments on the BiLSTM model and LSTM model.
[0153] 1) The monthly average root mean square error from Scheme 1 to Scheme 4 shows that the model of the present invention performs optimally in all cases, indicating that when the input features are the same, the model of the present invention has stronger feature extraction ability compared with the existing models (LSTM model, BiLSTM model, CNN-LSTM model, CNN-BiLSTM model), and can better learn the potential law between the input features and the output power, so as to achieve higher prediction accuracy.
[0154] 2) From the rate of change between adjacent schemes, when the model of the present invention adds new effective features, it is constantly optimized. When changing from Scheme 2 to Scheme 3, the of the LSTM model and CNN-BiLSTM model remains unchanged, and the deterioration of the BiLSTM model indicates that the three models cannot effectively extract the newly added input features; although the rate of change of the CNN-LSTM model is better than that of the model of the present invention, the index value is not as good as that of the model of the present invention. When changing from Scheme 3 to Scheme 4, the deepest-level features are added. The index of the LSTM model is not significantly optimized, and the deterioration of the CNN-LSTM model and CNN-BiLSTM model indicates that the three models have insufficient ability to extract this deep-level feature; while the optimization degree of the model of the present invention is the largest, indicating that the model of the present invention has strong ability to extract the newly added deepest-level input features. Generally speaking, the optimization amplitude of the model of the present invention is the largest, indicating that the model of the present invention has strong feature extraction ability, which is also one of the key factors for the highest prediction accuracy of the model designed by the present invention.
[0155] Figure 4It is the prediction result of the power values of each model, showing the power value prediction results of each model using Scenario 4 on a typical day in the test set, and the prediction effect can be seen more intuitively.
[0156] Taking Scenario 4 as the input feature to train the model of the present invention, and using the samples of a certain three months in 95 samples of the test set to test the trained model, the monthly average accuracy rate ( ), monthly average qualified rate ( ) of the model of the present invention are shown in Table 3. Obviously, the model of the present invention can meet the requirements of the national industry standard after training.
[0157] Table 3 Model Test Results
[0158]
[0159] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principle and practical application of the present invention, so that those skilled in the art can understand and utilize the present invention well.
Claims
1. A photovoltaic power prediction method based on KAN, characterized in that: The following steps are involved: 1) Collect historical data and preprocess it to form a data set; 2) Divide the data set into training set and test set; 3) Construct a photovoltaic output day-ahead prediction model and train the photovoltaic output day-ahead prediction model; 4) Test and evaluate the trained photovoltaic output day-ahead forecasting model; 5) Use the photovoltaic output day-ahead prediction model to predict photovoltaic power; The photovoltaic output day-ahead prediction model includes a spatial feature extraction module, a temporal feature extraction module, and a spatiotemporal feature fusion and power prediction output module; the spatial feature extraction module and the temporal feature extraction module extract features in parallel, and the extracted features are input into the spatiotemporal feature fusion and power prediction output module to output photovoltaic power prediction results; The spatial feature extraction module includes an improved deep KAN, which includes a plurality of basic units, the plurality of basic units are connected in sequence, and each basic unit is composed of two stacked KAN layers and a skip connection; The spatial feature extraction module finally outputs a vector.
2. A photovoltaic power prediction method based on KAN according to claim 1, characterized in that: Data preprocessing and data set division include the following steps: a) Set the weather characteristics corresponding to the output power of 0 to 0; b) Replace the data points where weather characteristics and output power are missing or outside the normal range with the average value of two adjacent moments; c) Perform min_max normalization on the same type of data according to the following formula: ; in is the original dataset data The u data, data =[ data 1,… data u … data U ] represents the data at a certain moment, the original data set data Includes all weather forecast values and actual power values at a certain moment, U is the number of data in the original data set, max( data ) and min( data ) are the original data sets data The maximum and minimum values of the same type of data, for Normalized data value; d) Remove the data corresponding to the time when the normalized power value is greater than 1 or less than 0; e) Taking the data of the same day as a sample, dividing the data processed by step a) to step d) into N samples, and dividing the N samples into a training set and a test set in proportion.
3. The photovoltaic power prediction method based on KAN according to claim 1, characterized in that: The temporal feature extraction module combines the multi-head attention mechanism with KAN, specifically including: (1) Use three independent two-layer KANs to map the input features to the query, key, and value vector spaces respectively, and obtain the query matrix Q , key matrix K , value matrix V : ; in X The dimension is [ n timestep , n input ] is the input feature matrix, Q , K , V The dimensions of n timestep , n heads ], , , The dimensions are [ n input , (2 n input +1), n heads ] of KAN, is the number of time points of the input sample, is the input feature dimension, is the number of heads of the multi-head attention mechanism; (2) Divide the query matrix Q, key matrix K, and value matrix V into blocks by column. Block matrices with the same subscript belong to the same head attention mechanism, and each head performs attention calculation independently: ; ; in is the calculation result of the g-th attention mechanism, , , They are the query matrix, key matrix, and value matrix of the g-th attention mechanism, and there are a total of G-head attention mechanisms; is the activation function, is a matrix, It means Operate on each row vector of A row vector is ,but The calculation is: ; Pair Vector Using activation functions , used to adjust the sum of each element of the vector to 1, For vector No. j elements, J is the length of the vector, j ∈(1,2,……,J); (3) Concatenate the calculation results of the multi-head attention mechanism and output the time series feature extraction results through the fully connected layer: ; ; in It is the concatenated matrix of the calculation results of the multi-head attention mechanism. is the calculation result of the G-th attention mechanism, time , time is the weight matrix and bias of the fully connected layer, It is the output result of the time series feature extraction module.
4. The photovoltaic power prediction method based on KAN according to claim 3 is characterized in that: The spatiotemporal feature fusion and predicted power output module combines the extraction results of the spatial feature extraction module and the temporal feature extraction module, and outputs the photovoltaic power prediction result through the fully connected layer: ; in , The output results of the spatial feature extraction module and the temporal feature extraction module are respectively, , is the weight matrix and bias of the fully connected layer, is the predicted photovoltaic power.
5. The photovoltaic power prediction method based on KAN according to claim 1, characterized in that: In step 3), the photovoltaic output day-ahead prediction model training includes the following contents: The weather characteristics of the training set are input into the photovoltaic output day-ahead prediction model. The photovoltaic output day-ahead prediction model outputs the predicted photovoltaic power. The loss function value is calculated according to the following formula. With the goal of minimizing the loss function, the model parameters are updated repeatedly until the loss function no longer changes significantly: ; in, for The root mean square error, is the true output power sequence in the training set, is the first p elements, Output a prediction sequence for the model, Output prediction sequence for the model p element, is the loss function value, P is and The length of the sequence.
6. The photovoltaic power prediction method based on KAN according to claim 1, characterized in that: In step 4), the photovoltaic output day-ahead prediction model test includes: The weather characteristics of the test set are input into the trained photovoltaic output day-ahead prediction model, and the photovoltaic output day-ahead prediction model outputs the predicted photovoltaic power. The prediction accuracy of the photovoltaic output day-ahead prediction model is evaluated by the evaluation index: in is the monthly average root mean square error, P pi for i Actual power at any moment, P Mi for i The predicted photovoltaic power at each moment, C i for i Always-on capacity, n is the time point of the whole month, C R is the monthly average accuracy, Q R is the monthly average pass rate, B i for i Predict the qualified judgment result at any time, B i Equal to 1 means qualified, B i A value of 0 means failure.
7. The photovoltaic power prediction method based on KAN according to claim 1, characterized in that: In step 5), predicting the photovoltaic output on the day to be predicted specifically includes the following steps: a) Collect real-time weather forecast data for the day to be predicted; b) performing the aforementioned preprocessing on the real-time weather forecast data; c) Input the pre-processed weather characteristic data into the photovoltaic output day-ahead prediction model that meets the test requirements, and obtain the initial prediction through model calculation ,in Represents the day o The predicted value at each time point; d) According to the startup capacity at each moment, the initial prediction output is denormalized to obtain the final power prediction value, denormalized calculation: ; in For the o The final prediction result at a time point is for o Always on capacity.
8. A system for implementing a photovoltaic power prediction method based on KAN according to any one of claims 1 to 7, characterized in that: It includes a photovoltaic output day-ahead prediction model, which includes a spatial feature extraction module, a temporal feature extraction module, a spatiotemporal feature fusion and a power output prediction module; The spatial feature extraction module is used to extract the spatial features of variables at the same time point in the sample data; The time series feature extraction module is used to extract the time series features of variables at different time points in the sample data; The function of the spatiotemporal feature fusion and predicted power output module is to fuse the spatial feature with the temporal feature and output the predicted photovoltaic power.
Citation Information
Patent Citations
Photovoltaic power ultra-short-term prediction method and system based on itransfomer
CN119312054A