Power system short-term net load prediction method based on RBF neural network-ridgelet Transformer

Through the short-term net load prediction method of power system based on RBF neural network-ribic wave Transformer, the combination of Transformer and RBF neural network is used to solve the problem of poor generalization ability of existing models, achieving higher prediction accuracy and accuracy.

CN120338532APending Publication Date: 2025-07-18STATE GRID SHANDONG ELECTRIC POWER CO JIMO POWER SUPPLY CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510221536.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing load prediction model cannot fully capture the complexity of net load changes in the power system, resulting in poor generalization capabilities of the model and low accuracy of the prediction results.

Method used

The short-term net load prediction method of power system based on RBF neural network-spine wave Transformer is adopted. By obtaining historical net load power and meteorological influencing factors data, the Transformer input layer, encoding layer and output layer are used for training, and combined with the RBF neural network and the multi-head self-attention mechanism, the spinal wave function is used for prediction.

Benefits of technology

It improves the generalization ability and prediction accuracy of the prediction model, reduces prediction errors, and meets the accuracy requirements of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338532A_ABST
    Figure CN120338532A_ABST
Patent Text Reader

Abstract

The invention discloses a power system short-term net load prediction method based on an RBF neural network-ridgelet Transformer, and the method comprises the following steps: obtaining historical net load power data and meteorological influence factor data, carrying out the data preprocessing, and dividing the preprocessed data into a training set and a prediction set; training the short-term net load prediction model of the power system based on the RBF neural network-ridgelet Transformer by using the data in the training set, and storing the trained model; the model comprises a Transform input layer, a Transform coding layer and a Transform output layer, and the Transform coding layer and the Transform output layer are connected with each other; and inputting data in the prediction set into the trained model, and outputting a net load power prediction value at a to-be-predicted moment. According to the method disclosed by the invention, the generalization ability and the prediction performance of the prediction model can be enhanced, and the prediction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power grid load forecasting, and particularly to a short-term net load forecasting method for power systems based on RBF neural network-ridge wave Transformer. Background Technique

[0002] Power enterprises are facing huge challenges in reducing operating costs and ensuring the safe and stable operation of power systems, and accurate forecasting of power demand has become a key factor among them.

[0003] The net load of a power system refers to the value obtained by subtracting the power generation output of renewable energy sources such as wind power and photovoltaic power from the electricity load. As the proportion of renewable energy sources such as wind power and photovoltaic power in the power system increases, the power generation output of renewable energy sources and the electricity load will be affected by meteorological factors, resulting in an increase in the volatility of the net load, thus posing higher requirements for the regulation ability of the power system. The research on power system net load forecasting has become an important part of the energy management system in a new type of power system, and the accuracy of its forecasting results directly affects the analysis results of subsequent safety checks of the power grid. Accurate short-term net load forecasting of power systems is of great significance for aspects such as power grid dynamic state estimation, load dispatching, and reducing generation costs. Only on the basis of accurately forecasting power demand can an efficient and economical power generation and supply plan be formulated, the unit output be reasonably arranged, and the safe and reliable power supply for the vast number of electricity customers be ensured.

[0004] In addition, the short-term net load forecasting of power systems also provides auxiliary decision-making basis for power supply enterprises to optimize the industrial structure and rationally allocate resources. In a market environment, power grid companies can implement load control measures targeted according to the short-term load forecasting results, improve economic benefits and market competitiveness. Through scientific and reasonable load forecasting, enterprises can better seize opportunities and reduce risks in the power market, providing strong support for the development of the power industry.

[0005] However, existing load forecasting models may not be able to fully capture the complexity of net load changes, especially there may be deficiencies in model parameter selection and model structure, resulting in poor generalization ability of the model and low accuracy of forecasting results. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a short-term net load forecasting method for power systems based on RBF neural network-ridge wave Transformer, aiming to enhance the generalization ability and forecasting performance of the forecasting model and improve the forecasting accuracy.

[0007] To achieve the above object, the technical solution of the present invention is as follows:

[0008] A short-term net load prediction method for power systems based on RBF neural network-ridge wave Transformer, comprising the following steps:

[0009] Step 1: Obtain historical net load power data and meteorological influence factor data, and perform data preprocessing. Divide the preprocessed data into a training set and a prediction set;

[0010] Step 2: Use the data in the training set to train a short-term net load prediction model for power systems based on RBF neural network-ridge wave Transformer, and save the trained model;

[0011] The model includes a Transformer input layer, a Transformer encoding layer, and a Transformer output layer;

[0012] The Transformer input layer includes an RBF neural network and an input fully connected layer, which are used to obtain high-dimensional features of the input historical net load power data and meteorological influence factor data, and use them as the input of the Transformer encoding layer;

[0013] The Transformer encoding layer includes multiple encoding units stacked in sequence. Each encoding unit includes a multi-head self-attention mechanism layer and a feed-forward fully connected layer. The input of each encoding unit first performs multi-head self-attention calculation through the multi-head self-attention mechanism layer, and then uses residual connection to sum the input of the encoding unit and the result of the multi-head self-attention calculation; the calculation result is normalized and then input into the feed-forward fully connected layer. After passing through the residual connection, it is summed with the output of the feed-forward fully connected layer and then normalized to obtain the output of the encoding unit. The output of the last encoding unit is the output of the Transformer encoding layer;

[0014] The Transformer output layer includes an output fully connected layer and a ridge wave function;

[0015] Step 3: Input the data in the prediction set into the trained model, and output the predicted value of the net load power at the moment to be predicted.

[0016] In the above solution, the calculation formula of the Transformer input layer is as follows:

[0017] X = WF RBF (x) + b

[0018] In the formula, x is the input vector composed of historical net load power and meteorological influence factor data; F RBF(·) is the RBF neural network calculation function; W and b are the parameter matrix and bias vector of the fully connected layer respectively; X is the high-dimensional net load power and meteorological data vector output by the input layer of the Transformer, which is used as the input of the encoding layer of the Transformer.

[0019] In the above solution, the RBF neural network includes an RBF neural network input layer, an RBF neural network hidden layer, and an RBF neural network output layer.

[0020] In a further technical solution, the RBF neural network input layer directly maps the input vector to the RBF neural network hidden layer, and there is no need for a weight connection between the two.

[0021] The RBF neural network hidden layer is the core of the RBF neural network. The Gaussian function is used as the activation function, and its relational expression is:

[0022]

[0023] In the formula, ‖·‖ represents the Euclidean distance, i represents the i-th neuron in the RBF neural network hidden layer, X = [x1, x2, …, x m T is the input matrix, x m represents the m-th component value of the input vector; μ i is the center of the radial basis function of the i-th RBF neural network hidden layer, m is the number of neurons in the RBF neural network hidden layer; σ i is the variance of the radial basis function of the i-th RBF neural network hidden layer.

[0024] In a further technical solution, the mapping from the RBF neural network hidden layer space to the RBF neural network output layer space is linear. The connection weight between the RBF neural network hidden layer and the RBF neural network output layer is adjustable. The output of the RBF neural network output layer is the linear weighted sum of the output of the RBF neural network hidden layer, that is:

[0025]

[0026] In the formula, w i is the connection weight corresponding to the RBF neural network output layer and the RBF neural network hidden layer; F(X, w i , μ i ) represents the output value of the RBF neural network output layer, m is the number of neurons in the hidden layer; G(‖X - μ i ‖) is the output value of the RBF neural network hidden layer.

[0027] In the above solution, the multi-head self-attention mechanism is based on the single-angle self-attention calculation and performs multiple groups of parallel self-attention calculations.​

[0028] In a further technical solution, the calculation formula for single-angle self-attention is as follows:

[0029] E·[W q ,W k ,W v = [Q, K, V]

[0030]

[0031] In the formula, W q , W k , W v are parameter matrices, which are respectively used to calculate the query vector Q, the key vector K, and the value vector V. d k is the dimension number of K, E is the historical net load power feature, and A(Q, K, V) is the self-attention expression result of the historical net load power feature.

[0032] In an even further technical solution, the calculation formula for the multi-head self-attention mechanism is as follows:

[0033] A h = A(Q h , K h , V h )

[0034] A MH (E) = W mh ·[A1, …, A H

[0035] In the formula, h is the self-attention head number; A h is the self-attention expression result of the h-th number; Q h , K h , V h are respectively the query vector, the key vector, and the value vector of the h-th self-attention head; H is the total number of self-attention heads; W mh is the multi-head self-attention parameter matrix, which is used to complete the summary calculation of multi-angle self-attention; A MH (E) is the multi-head self-attention calculation result of E.

[0036] In the above solution, the calculation formula for the output layer of Transformer is as follows:

[0037] Y = ψ γ (W o X S + b o )

[0038] In the formula, Y is the predicted result of the net load power output by the output layer of Transformer; X S ​is the output data of the Transformer encoding layer; W o , b o are the parameter matrix and bias vector in the fully connected layer of the Transformer output layer, and ψ γ is the ridgelet function.

[0039] In the above solution, the expression of the ridgelet function is as follows:

[0040]

[0041] In the formula, γ is the parameter space, that is:

[0042] γ = (a, u, g); a, g ∈ R; a > 0; u ∈ S d-1

[0043] In the formula, a represents the scale vector of the ridgelet function; u represents the direction vector of the ridgelet function, and ‖u‖ = 1; g represents the position vector of the ridgelet function; S d-1 represents the d - 1 dimensional space; n represents the input vector of the ridgelet function.

[0044] Through the above technical solution, the short - term net load prediction method for power systems based on RBF neural network - ridgelet Transformer provided by the present invention has the following beneficial effects:

[0045] The present invention constructs a short - term net load prediction model for power systems based on RBF neural network - ridgelet Transformer. The input layer of the Transformer uses an RBF neural network and a fully connected layer to replace the word embedding and position encoding links in the traditional Transformer model, obtains higher - dimensional features about load and meteorology, and uses the high - dimensional feature data with more detailed information as the input of the encoding layer.

[0046] The Transformer encoding layer analyzes the input net load power features through the multi-head self-attention mechanism. This process will loop N times in total to achieve deep learning of the input feature data. The main advantage of the Transformer multi-head self-attention mechanism is that it can obtain multiple different analysis conclusions through multi-angle analysis, improving the comprehensiveness of the model's problem analysis and the understanding ability of the net load power features. The other computational parts of the Transformer encoding layer are residual connection and normalization. The main function of the residual connection is to introduce a direct path across layers by direct summation, allowing gradients to propagate more easily to the upper network, thus promoting the training and optimization of the network, making the training of the model more stable and enabling it to explore the hierarchical structure of the network more deeply. The normalization step scales the data to the data range of 0 - 1, which can ensure that the inputs of each layer of the network have a similar distribution, thus avoiding the activation values of some neurons being too large or too small, and playing a role in accelerating convergence and improving the generalization ability of the model.

[0047] The Transformer output layer uses the ridgelet function. The ridgelet function can effectively represent the directional singular features in the signal, enhancing the generalization ability and prediction performance of the prediction model. Through the prediction simulation and testing of the actual power grid net load, it is confirmed that the proposed prediction model can effectively improve the prediction accuracy. Brief Description of the Drawings

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.

[0049] Figure 1 It is a schematic diagram of a short-term net load prediction model for a power system based on an RBF neural network - ridgelet Transformer disclosed in an embodiment of the present invention.

[0050] Figure 2 It is a schematic diagram of the RBF neural network structure. Detailed Embodiments

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention.

[0052] The present invention provides a short-term net load prediction method for a power system based on an RBF neural network - ridgelet Transformer, including the following steps:

[0053] Step 1: Obtain historical net load power data and meteorological influence factor data, and perform data preprocessing, and divide the preprocessed data into a training set and a prediction set.

[0054] Meteorological influencing factors include the daily maximum temperature, daily minimum temperature, humidity, weather type, etc.

[0055] Preprocessing includes data transformation such as handling missing values and outliers, and data standardization.

[0056] Step 2: Use the data in the training set to train the short-term net load prediction model of the power system based on the RBF neural network-ridge wave Transformer, and save the trained model.

[0057] As Figure 1 shown, the short-term net load prediction model of the power system based on the RBF neural network-ridge wave Transformer includes a Transformer input layer, a Transformer encoding layer, and a Transformer output layer.

[0058] 1. Transformer input layer

[0059] The Transformer input layer uses an RBF neural network and an input fully connected layer to replace the word embedding and position encoding links in the traditional Transformer model, obtains higher-dimensional features of historical net load and meteorological and other influencing factors, and uses the high-dimensional feature data with more detailed information as the input of the encoding layer. The calculation formula of the Transformer input layer is as follows:

[0060] X = WF RBF (x) + b

[0061] In the formula, x is the input vector composed of historical net load power and meteorological influencing factor data; F RBF (·) is the RBF neural network calculation function; W and b are the parameter matrix and bias vector of the fully connected layer respectively; X is the high-dimensional net load power and meteorological data vector output by the Transformer input layer, which is used as the input of the Transformer encoding layer.

[0062] As Figure 2 shown, the RBF neural network is a multi-layer forward network structure, including an RBF neural network input layer, an RBF neural network hidden layer, and an RBF neural network output layer.

[0063] The RBF neural network input layer directly maps the input vector to the RBF neural network hidden layer, and there is no need for a weight connection between them; the RBF neural network hidden layer is the core of the RBF neural network, and the Gaussian function is used as the activation function, and its relationship is:

[0064]

[0065] Where, ‖·‖ represents the Euclidean distance, i represents the i-th hidden layer neuron of the RBF neural network, X = [x1, x2, …, x m T is the input matrix, and x m represents the m-th component value of the input vector; μ i is the center of the radial basis function of the i-th hidden layer of the RBF neural network, and m is the number of hidden layer neurons of the RBF neural network; σ i is the variance of the radial basis function of the i-th hidden layer of the RBF neural network.

[0066] The mapping from the space of the hidden layer of the RBF neural network to the space of the output layer of the RBF neural network is linear. The connection weights between the hidden layer and the output layer of the RBF neural network are adjustable. The output of the output layer of the RBF neural network is the linear weighted sum of the output of the hidden layer of the RBF neural network, that is:

[0067]

[0068] Where, w i is the connection weight corresponding to the output layer and the hidden layer of the RBF neural network; F(X, w i , μ i ) represents the output value of the output layer of the RBF neural network, and m is the number of hidden layer neurons; G(‖X - μ i ‖) is the output value of the hidden layer of the RBF neural network.

[0069] 2. Transformer Encoder Layer

[0070] The Transformer encoder layer analyzes the input historical net load power features through the multi-head self-attention mechanism. This process will loop N times in total to achieve in-depth learning of the input feature data. The attention mechanism mimics the behavior of humans concentrating on important information. By assigning calculation weights to the input features, the model can recognize the feature information that needs to be focused on. The main advantage of the Transformer multi-head self-attention mechanism is that it can obtain multiple different analysis conclusions through multi-angle analysis, improving the comprehensiveness of the model's analysis of problems and the understanding ability of the net load power features.

[0071] ​The Transformer encoding layer consists of N encoding units stacked in sequence (N can take values from 2 to 6), and each encoding unit includes a multi-head self-attention mechanism layer and a feed-forward fully connected layer. The input of each encoding unit first undergoes multi-head self-attention calculation through the multi-head self-attention mechanism layer, and then a residual connection is adopted, that is, allowing the input of the encoding layer to be directly passed to a deeper layer of the network, and the input of the encoding unit and the result of the multi-head self-attention calculation are summed; the calculation result is normalized and then input into the feed-forward fully connected layer, and after another residual connection, it is summed with the output of the feed-forward fully connected layer and then normalized to obtain the output of the encoding unit. The output of the last encoding unit is the output of the Transformer encoding layer.

[0072] The self-attention calculation formula for a single angle is as follows:

[0073] E·[W q ,W k ,W v =[Q,K,V]

[0074]

[0075] In the formula, W q 、W k 、W v are parameter matrices, which are used to calculate the query vector Q, the key vector K, and the value vector V respectively. d k is the dimension number of K, E is the historical net load power feature, and A(Q,K,V) is the self-attention expression result of the historical net load power feature. Through the information retrieval mode, the dot product key-value pair Q·K T is calculated to reflect the correlation between different photovoltaic feature data. After softmax normalization, it is multiplied by V to complete the calculation of the attention expression A of the input net load power feature. d k is the dimension number of K, and the key-value pair needs to be divided by for scaling calculation to prevent the problem of gradient disappearance caused by the key-value pair having too large a value.

[0076] Based on the self-attention calculation of a single angle, the multi-head self-attention mechanism performs multiple groups of parallel self-attention calculations. The calculation formula of the multi-head self-attention mechanism is as follows:

[0077] A h =A(Q h ,K h ,V h )

[0078] A MH (E)=W mh ·[A1,…,A H

[0079] ​where h is the self-attention head number; A h is the self-attention expression result of the h-th; Q h , K h , V h are the query vector, key vector, and value vector of the h-th self-attention head respectively; H is the total number of self-attention heads; W mh is the multi-head self-attention parameter matrix, which is used to complete the summary calculation of multi-angle self-attention; A MH (E) is the multi-head self-attention calculation result of E.

[0080] The encoding layer analyzes the temporal features of the input data through the multi-head attention mechanism, highlights the key features that need to be concerned, and obtains the correlation vectors between different data.

[0081] In the encoding layer, the multi-head self-attention mechanism and the feed-forward fully connected layer are two main parts. By stacking these two parts multiple times, a deep model structure can be established, which helps the model better learn the complex features and high-level representations of the input data, and improves the analysis ability and generalization performance of the model. The other computational parts of the encoding layer are residual connection and normalization. The main role of the residual connection is to introduce a direct path across layers by direct summation, allowing the gradient to propagate more easily to the upper network, thereby promoting the training and optimization of the network, making the training of the model more stable and enabling it to explore the hierarchical structure of the network more deeply. The normalization step scales the data to the data range of 0-1, which can ensure that the inputs of each layer of the network have a similar distribution, thus avoiding the activation values of some neurons being too large or too small, and thus playing a role in accelerating convergence and improving the generalization ability of the model. After the above steps, the output of the Transformer encoding layer can be represented by X S .

[0082] 3. Transformer Output Layer

[0083] The Transformer output layer includes an output fully connected layer, and the ridgelet function is used as the activation function.

[0084] The calculation formula of the Transformer output layer is as follows:

[0085] Y = ψ γ (W o X S + b o )

[0086] where Y is the predicted result of the net load power output by the Transformer output layer; X S is the output data of the Transformer encoding layer; W o , b ois the parameter matrix and bias vector in the fully connected layer of the Transformer output layer, ψ γ is the ridgelet function.

[0087] The ridgelet transform theory is a multi-scale method that can effectively represent anisotropic singularities. The definition of the ridgelet function used therein can be described in detail as follows: Assume that the Fourier transform corresponding to the smoothing function ψ: R d →R meets the following admissibility condition:

[0088]

[0089] In the formula, ψ is the admissible function. Among them, σ is the independent variable and d is the spatial dimension. Then the ridge function ψ generated by the admissible function ψ that satisfies the above conditions γ is called the ridgelet function, and its expression is:

[0090]

[0091] In the formula, γ is the parameter space, that is:

[0092] γ = (a, u, g); a, g ∈ R; a > 0; u ∈ S d-1

[0093] In the formula, a represents the scale vector of the ridgelet function; u represents the direction vector of the ridgelet function, and ‖u‖ = 1; g represents the position vector of the ridgelet function; S d-1 represents the d - 1 dimensional space; n represents the input vector of the ridgelet function.

[0094] Step 3: Input the data in the prediction set into the trained model, and output the predicted value of the net load power at the moment to be predicted. In the actual application process, the obtained historical net load power data and meteorological influence factor data can be preprocessed and then input into the trained model, and the model automatically outputs the predicted value of the net load power at the moment to be predicted.

[0095] To verify the effectiveness of the short-term net load prediction model (Model I) for power systems based on RBF neural network - ridgelet Transformer proposed by the present invention, the short-term net load prediction model for power systems based on ridgelet Transformer (Model II), the short-term net load prediction model for power systems based on Transformer (Model III), and the short-term net load prediction model for power systems based on the conventional BP neural network (Model IV) are respectively used to predict the net load power on a certain day. The comparison of the prediction errors of the 4 models is shown in Table 1, where E MAPE is the mean absolute error, and E MAX is the maximum relative error.

[0096] Table 1 Comparison of Prediction Errors of Four Prediction Models

[0097]

[0098] As can be seen from Table 1, for the short-term net load prediction model of the power system using the conventional BP neural network, the mean absolute error is 9.83%, and the maximum relative error is 17.72%. For the short-term net load prediction model of the power system based on Transformer, the mean absolute error is reduced by 1.37% compared with the short-term net load prediction model of the power system using the conventional BP neural network, and the maximum relative error is reduced by 2.12%. For the short-term net load prediction model of the power system based on ridgelet Transformer, the mean absolute error is reduced by 1.87% compared with the short-term net load prediction model of the power system based on Transformer, and the maximum relative error is reduced by 3.07%. For the short-term net load prediction model of the power system based on RBF neural network-ridgelet Transformer proposed in the present invention, the mean absolute error is reduced by 4.56% compared with the short-term net load prediction model of the power system using the conventional BP neural network, and the maximum relative error is reduced by 8.90%; the mean absolute error is reduced by 3.19% compared with the short-term net load prediction model of the power system based on Transformer, and the maximum relative error is reduced by 6.78%; the mean absolute error is reduced by 1.32% compared with the short-term net load prediction model of the power system based on ridgelet Transformer, and the maximum relative error is reduced by 3.71%. The model of the present invention is the most ideal among all models, which indicates that the proposed short-term net load prediction model of the power system based on RBF neural network-ridgelet Transformer can effectively improve the prediction accuracy and meet the actual prediction requirements.

[0099] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A short-term net load forecasting method for power systems based on RBF neural network-ridge wave Transformer, characterized in that It includes the following steps: Step 1: Obtain historical net load power data and meteorological influence factor data, and perform data preprocessing. Divide the preprocessed data into a training set and a prediction set; Step 2: Use the data in the training set to train the short-term net load prediction model of the power system based on the RBF neural network - ridgelet Transformer, and save the trained model; The model includes a Transformer input layer, a Transformer encoding layer, and a Transformer output layer; The Transformer input layer includes an RBF neural network and an input fully connected layer, which are used to obtain the high-dimensional features of the input historical net load power data and meteorological influence factor data, and use them as the input of the Transformer encoding layer; The Transformer encoding layer includes multiple encoding units stacked in sequence. Each encoding unit includes a multi-head self-attention mechanism layer and a feed-forward fully connected layer. The input of each encoding unit first performs multi-head self-attention calculation through the multi-head self-attention mechanism layer, and then uses residual connection to sum the input of the encoding unit and the result of the multi-head self-attention calculation; The calculation result is normalized and then input into the feed-forward fully connected layer. After passing through the residual connection, it is summed with the output of the feed-forward fully connected layer and then normalized to obtain the output of the encoding unit. The output of the last encoding unit is the output of the Transformer encoding layer; The Transformer output layer includes an output fully connected layer and a ridgelet function; Step 3: Input the data in the prediction set into the trained model, and output the predicted value of the net load power at the moment to be predicted.

2. The short-term net load prediction method for power systems based on RBF neural network-ridge wave Transformer according to claim 1, characterized in that The calculation formula of the Transformer input layer is as follows: X = WF RBF (x) + b where x is the input vector composed of historical net load power and meteorological influence factor data; F RBF (·) is the RBF neural network calculation function; W and b are the parameter matrix and bias vector of the fully connected layer respectively; X is the high-dimensional net load power and meteorological data vector output by the Transformer input layer, which is used as the input of the Transformer encoding layer.

3. The short-term net load prediction method for power systems based on RBF neural network-ridge wave Transformer according to claim 1 or 2, characterized in that, The RBF neural network includes an RBF neural network input layer, an RBF neural network hidden layer, and an RBF neural network output layer.

4. The short-term net load prediction method for power systems based on RBF neural network-ridge wave Transformer according to claim 3, wherein, The RBF neural network input layer directly maps the input vector to the RBF neural network hidden layer, and there is no need for a weight connection between them; The RBF neural network hidden layer is the core of the RBF neural network. The Gaussian function is used as the activation function, and its relationship is: Where, ‖·‖ represents the Euclidean distance, i represents the i-th hidden layer neuron of the RBF neural network, X = [x1, x2, …, x m T is the input matrix, and x m represents the m-th component value of the input vector; μ i is the center of the radial basis function of the i-th hidden layer of the RBF neural network, and m is the number of neurons in the hidden layer of the RBF neural network; σ i is the variance of the radial basis function of the i-th hidden layer of the RBF neural network.​ 5. The short-term net load prediction method for power systems based on RBF neural network-ridge wave Transformer according to claim 3, characterized in that, The mapping from the RBF neural network hidden layer space to the RBF neural network output layer space is linear. The connection weights between the RBF neural network hidden layer and the RBF neural network output layer are adjustable. The output of the RBF neural network output layer is the linear weighted sum of the output of the RBF neural network hidden layer, that is: where, w i is the connection weight between the output layer and the hidden layer of the RBF neural network; F(X, w i , μ i ) represents the output value of the output layer of the RBF neural network, m is the number of neurons in the hidden layer; G(‖X - μ i ‖) is the output value of the hidden layer of the RBF neural network.

6. The short-term net load prediction method for power systems based on RBF neural network-ridge wave Transformer according to claim 1, characterized in that, The multi-head self-attention mechanism is based on the single-angle self-attention calculation and performs multiple groups of parallel self-attention calculations.

7. The short-term net load prediction method for power systems based on RBF neural network-ridge wave Transformer according to claim 6, characterized in that, The calculation formula of the single-angle self-attention is as follows: E·[W q ,W k ,W v = [Q, K, V] Where, W q , W k , W v are parameter matrices, used to calculate the query vector Q, the key vector K, and the value vector V respectively. d k is the dimension number of K, E is the historical net load power feature, and A(Q, K, V) is the self-attention expression result of the historical net load power feature.

8. The short-term net load prediction method for a power system based on an RBF neural network - ridge wave Transformer according to claim 7, wherein The calculation formula of the multi-head self-attention mechanism is as follows: A h = A(Q h , K h , V h ) A MH (E) = W mh ·[A1, …, A H ​ Where h is the self-attention head number; A h is the self-attention expression result of the h-th number; Q h , K h , V h are the query vector, key vector, and value vector of the h-th self-attention head respectively; H is the total number of self-attention heads; W mh is the multi-head self-attention parameter matrix, which is used to complete the aggregation calculation of multi-angle self-attention; A MH (E) is the multi-head self-attention calculation result of E.

9. The short-term net load prediction method for power systems based on RBF neural network-ridge wave Transformer according to claim 7, characterized in that, The calculation formula of the Transformer output layer is as follows: Y = ψ γ (W o X S + b o ) where Y is the predicted result of the net load power output by the output layer of the Transformer; X S is the output data of the encoding layer of the Transformer; W o , b o are the parameter matrix and bias vector in the fully connected layer of the output layer of the Transformer, and ψ γ is the ridge wave function.

10. The short-term net load prediction method for power systems based on RBF neural network-ridge wave Transformer according to claim 9, characterized in that, The expression of the ridgelet function is as follows: In the formula, γ is the parameter space, that is: γ = (a, u, g); a, g ∈ R; a > 0; u ∈ S d-1 Wherein, a represents the scale vector of the ridgelet function; u represents the direction vector of the ridgelet function and ‖u‖ = 1; g represents the position vector of the ridgelet function; S d-1 represents a (d - 1)-dimensional space; n represents the input vector of the ridgelet function.