Power load probability prediction method and system based on neural network quantile regression model and multiple linear regression
By combining a neural network quantile regression model and a multiple linear regression method with longitudinal data analysis and seasonal trend decomposition, the problem of power load forecasting with load uncertainty in new power systems was solved, and high-precision long-term power load probability forecasting was achieved.
Patent Information
- Application Number
- CN202511399731.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-02-06
AI Technical Summary
Existing power load forecasting methods suffer from insufficient accuracy when faced with the uncertainty and randomness of loads in new power systems. In particular, they are unable to effectively reflect meteorological information and temporal fluctuations in long-term power load probabilistic forecasting, and are difficult to process high-dimensional time series data.
Using a neural network quantile regression model and multiple linear regression, a factor analysis model is constructed through longitudinal data analysis, Pearson correlation analysis, and seasonal trend decomposition. Combined with nonparametric kernel density technology, this model is used to predict the probability of electricity load.
It improves the accuracy and computational efficiency of power load forecasting, accurately reflects the randomness and temporal correlation of load, provides hourly load forecasting results, and is suitable for long-term power load probabilistic forecasting.
Smart Images

Figure CN121481284A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system load forecasting technology, and more specifically, to a method and system for probabilistic power load forecasting based on a neural network quantile regression model and multiple linear regression. Background Technology
[0002] With the increasingly prominent contradiction between global energy supply and demand and the continuous advancement of energy conservation and emission reduction policies, the power industry is facing an urgent need for energy structure transformation. Against this backdrop, large-scale grid connection of renewable energy has become an inevitable choice for power system development. However, its inherent intermittent and volatile characteristics significantly reduce the controllability of the power supply side and significantly enhance system randomness. This change has led to the traditional power system gradually evolving from a purely demand-side stochastic system into a complex system with stochastic coupling on both the supply and demand sides. On the demand side of the power system, traditional industrial loads and residential loads inherently possess a certain degree of randomness due to factors such as climate change and natural disasters, macro-industrial restructuring, and changes in the energy market. In the future, with the continuous increase of new loads such as electric vehicles and high-speed rail loads, and the implementation of demand-side management policies enhancing the autonomy of load-side user behavior, demand-side randomness will inevitably increase accordingly.
[0003] Demand-side dynamic load confidence assessment methods refer to obtaining load forecast intervals at different confidence levels through probabilistic load forecasting. Compared with traditional power load forecasting, probabilistic load forecasting can better characterize future load changes. By setting confidence intervals and assigning upper and lower limits to the load for the forecast period, probabilistic load forecasting can comprehensively depict the load uncertainty under the new power system. Due to the integration of new energy sources and large-scale electric vehicles into the grid, the grid load is gradually changing from a traditional rigid load to a flexible load. Coupled with the application of demand response measures, grid uncertainty and volatility have further increased, and multiple uncertainties in grid load have gradually become a normalized characteristic of the new power system.
[0004] Based on different probabilistic modeling methods, power load prediction methods can be divided into statistical methods and artificial intelligence methods. Statistical methods perform statistical analysis on historical load data to obtain correlation information or probabilistic information on load state transitions, and then make predictions based on this. Common statistical methods include regression analysis and Bayesian network methods. Statistical methods can intuitively reflect the characteristics and patterns of the predicted object and its influencing factors. However, because load and its influencing factors contain many implicit characteristics and patterns, some of which cannot be intuitively reflected by existing feature analysis methods, the accuracy of load probabilistic prediction methods is limited. Currently, the application of statistical load probabilistic prediction methods is relatively limited.
[0005] Artificial intelligence (AI) methods are widely used in power load forecasting due to their outstanding ability to mine implicit data features and approximate complex functions. These methods include traditional AI network methods such as artificial neural networks and support vector machines, as well as deep learning algorithms. In load probability forecasting models based on AI networks, training samples require historical load probability distribution information. However, historical experience data cannot provide load probability data, resulting in a lack of training samples for probability forecasting. Therefore, it is impossible to directly train the AI network to calculate the load probability distribution of the output samples. To address this problem, AI networks typically use probability distribution estimation methods to generate the load probability distribution. These methods are divided into parametric and non-parametric methods. Parametric methods use known probability distribution functions to model the load probability distribution, outputting complete probability distribution information. They are computationally simple and have advantages in scenarios with few samples. These methods first assume that the load or load forecasting error follows a probability distribution function, and then estimate the parameters of the probability distribution function based on the point forecasting results. While parametric methods can easily adopt mature point forecasting techniques to improve prediction accuracy, they usually require pre-setting the probability distribution function. The pre-set probability distribution may deviate significantly from the actual probability distribution, thus affecting the accuracy of probability forecasting. Nonparametric methods can overcome the shortcomings of parametric methods. They do not require assumptions about the probability distribution function of the object being predicted. Instead, they use kernel density estimation or quantile regression methods to generate the load probability distribution. Summary of the Invention
[0006] To address the above problems, this invention proposes a power load probability prediction method based on a neural network quantile regression model and multiple linear regression, comprising:
[0007] Standardized data were obtained using longitudinal data analysis.
[0008] Based on Pearson correlation analysis, the key influencing factors of power load in the standardized data were identified, and a factor analysis model was constructed to quantify the influence weight of the key influencing factors on power load.
[0009] Based on seasonal trend decomposition, a neural network quantile regression model is constructed to fit key influencing factors with different influence weights and obtain quantile prediction results.
[0010] Based on the quantile prediction results, the continuous probability distribution curve of the common load factor is estimated using nonparametric kernel density techniques to obtain the interval prediction results.
[0011] Optionally, standardized data can be obtained using longitudinal data analysis, including:
[0012] Using hours as the time scale, the annual load data is represented as a 24-dimensional standardized electricity load time series vector P, corresponding to the standard load time series for each of the 24 hours of the day, as shown in the following formula:
[0013]
[0014] Where D is the total number of days in the sample observation period, and P i Let P be the time series of the daily variation of the standard load at time i. h,d The load standard value at time h on day d is derived from the original value. It is obtained through the following standardization:
[0015]
[0016] Where, μ h and S h represents the sample mean and standard deviation of the original load at time h, respectively.
[0017] Optionally, based on Pearson correlation analysis, key influencing factors of electricity load in the standardized data are identified, and a factor analysis model is constructed to quantify the influence weights of these key influencing factors on electricity load, including:
[0018] Calculate the covariance matrix S of the historical observation sample data;
[0019] Calculate the eigenvalues of the covariance matrix S and obtain the corresponding unit eigenvectors e1, e2, ..., e p ;
[0020] Extract the first r largest eigenvalues and their corresponding eigenvectors, and calculate the factor loading matrix. The calculation formula is as follows:
[0021]
[0022] Based on the factor loading matrix The sum of squares of the elements in the i-th row and the diagonal elements s of the covariance matrix S ii Calculate the special factor ε i variance Thus, we obtain D(ε):
[0023]
[0024] After obtaining the common factor vector, we replace the original p-dimensional variables with mutually independent r-dimensional common factor vectors.
[0025] Optionally, the hidden layer transformation function f of the neural network quantile regression model (h) Choose the hyperbolic tangent tanh function, and the output layer transformation function f.(o) By choosing the iso-function, the neural network quantile regression model exhibits the following nonlinear relationship:
[0026]
[0027] The estimation of the parameter vector in a neural network quantile regression model is to minimize an error function of the following form:
[0028]
[0029] Where T is the sample size, Y i This represents the value of the response variable Y for the i-th training sample. Let I(·) represent the conditional quantile of the i-th response variable, and let I(·) be the following indicator function:
[0030]
[0031] To avoid overfitting in the neural network quantile regression model, a penalty term is added to the objective function. In this case, the parameter estimation of the neural network quantile regression model is transformed into the following optimization problem:
[0032]
[0033] in, Here, ||.||2 is the hidden layer weight, ||.||2 is the L2 norm, and ρ is the penalty parameter;
[0034] Substituting the obtained optimal estimated parameters back into the neural network quantile regression model, the conditional quantile predicted value of the response variable Y can be calculated based on the input explanatory variable X. When the quantile τ takes continuous values in the interval (0,1), the conditional quantile curve... This is the conditional probability distribution curve of Y.
[0035] Optionally, a neural network quantile regression model is constructed based on seasonal trend decomposition to fit key influencing factors with different influence weights, and quantile prediction results are obtained, including:
[0036] Estimating the trend subsequence T t And remove it from the original data to obtain a non-trend time series:
[0037] D t =Y t -T t (k) (10)
[0038] Among them, T t (k) It is the trend subsequence calculated in the current iteration;
[0039] For the trend-removed sequence D t Locally weighted regression is used to smooth the data, thus obtaining the seasonal subsequence. Locally weighted regression is a nonparametric regression method that uses a sliding window to weight and fit the data. The specific formula is shown below:
[0040]
[0041] To update the trend subsequence, the de-seasonalized sequence needs to be low-pass filtered. The low-frequency signal obtained from the low-pass filtering corresponds to the trend subsequence T. t The high-frequency signal corresponds to the seasonal subsequence S. t The specific formula is as follows:
[0042]
[0043] Using the new trend subsequence T t (k+1) Recalculate the seasonal subsequence;
[0044]
[0045] Subtracting the updated seasonal subsequence from the original time series yields the de-seasonalized time series;
[0046]
[0047] Seasonless sequence Perform local weighted regression smoothing to obtain a new trend subsequence;
[0048]
[0049] If the changes in the trend subsequence and the seasonal subsequence satisfy the convergence condition, the algorithm terminates with convergence. If the convergence condition is not satisfied, the algorithm returns to step one and continues iterating. The convergence condition is as follows:
[0050] |T t (k+1) -T t (k) |<ε (16)
[0051] Where ε is the set convergence threshold;
[0052] Seasonal trend decomposition reveals that the load exhibits periodic changes over a 24-hour period. Therefore, a time penalty term is added to the error function of the neural network quantile regression model.
[0053]
[0054] In the formula, n m,tThe sample date t is the number of months from the month of the prediction month, with the furthest period having the smallest time weight and the most recent period having the smallest time weight. m,t =α 0 =1 is the maximum, α is the time weight parameter determined based on the prediction error, and the error function thus becomes:
[0055]
[0056] Optionally, based on the quantile prediction results, the continuous probability distribution curve of the common load factor is estimated using nonparametric kernel density techniques to obtain the interval prediction results, including:
[0057] The set of conditional quantiles obtained from the neural network quantile regression model can be regarded as a set of randomly sampled samples that follow the same distribution as Y. The probability density function of Y under the condition of input X can be calculated using nonparametric kernel density estimation.
[0058]
[0059] In the formula, h is the bandwidth, K(*) is the kernel function, n is the number of samples, and Y... i To obtain the n equally spaced quantiles τ of the predictor variable Y for the neural network quantile regression model. i (i = 1, ..., n) Conditional quantiles.
[0060] Alternatively, the following multiple linear regression model can be used:
[0061] Y = β0 + β1X1 + β2X2 + ... + β p X p (20)
[0062] Wherein, β1~β p It is the regression coefficient, X i Y is the influencing factor, and Y is the total load.
[0063] Optionally, the method also includes:
[0064] A multiple linear regression model is constructed to predict changes in load size, and the interval prediction results are adjusted accordingly.
[0065] Furthermore, this invention also proposes a power load probability prediction system based on a neural network quantile regression model and multiple linear regression, comprising:
[0066] Standardized units are used to obtain standardized data using longitudinal data analysis methods.
[0067] The identification unit is used to identify the key influencing factors of power load in the standardized data based on the Pearson correlation analysis method, and to construct a factor analysis model to quantify the influence weight of the key influencing factors on power load.
[0068] The fitting unit is used to fit key influencing factors with different influence weights to a neural network quantile regression model based on seasonal trend decomposition, and to obtain quantile prediction results.
[0069] The first prediction unit is used to estimate the continuous probability distribution curve of the common load factor based on the quantile prediction results and to obtain the interval prediction results by using nonparametric kernel density techniques.
[0070] Optionally, standardized data can be obtained using longitudinal data analysis, including:
[0071] Using hours as the time scale, the annual load data is represented as a 24-dimensional standardized electricity load time series vector P, corresponding to the standard load time series for each of the 24 hours of the day, as shown in the following formula:
[0072]
[0073] Where D is the total number of days in the sample observation period, and P i Let P be the time series of the daily variation of the standard load at time i. h,d The load standard value at time h on day d is derived from the original value. It is obtained through the following standardization:
[0074]
[0075] Where, μ h and S h represents the sample mean and standard deviation of the original load at time h, respectively.
[0076] Optionally, based on Pearson correlation analysis, key influencing factors of electricity load in the standardized data are identified, and a factor analysis model is constructed to quantify the influence weights of these key influencing factors on electricity load, including:
[0077] Calculate the covariance matrix S of the historical observation sample data;
[0078] Calculate the eigenvalues of the covariance matrix S and obtain the corresponding unit eigenvectors e1, e2, ..., e p ;
[0079] Extract the first r largest eigenvalues and their corresponding eigenvectors, and calculate the factor loading matrix. The calculation formula is as follows:
[0080]
[0081] Based on the factor loading matrix The sum of squares of the elements in the i-th row and the diagonal elements s of the covariance matrix S ii Calculate the special factor εi variance Thus, we obtain D(ε):
[0082]
[0083] After obtaining the common factor vector, we replace the original p-dimensional variables with mutually independent r-dimensional common factor vectors.
[0084] Optionally, the hidden layer transformation function f of the neural network quantile regression model (h) Choose the hyperbolic tangent tanh function, and the output layer transformation function f. (o) By choosing the iso-function, the neural network quantile regression model exhibits the following nonlinear relationship:
[0085]
[0086] The estimation of the parameter vector in a neural network quantile regression model is to minimize an error function of the following form:
[0087]
[0088] Where T is the sample size, Y i This represents the value of the response variable Y for the i-th training sample. Let I(·) represent the conditional quantile of the i-th response variable, and let I(·) be the following indicator function:
[0089]
[0090] To avoid overfitting in the neural network quantile regression model, a penalty term is added to the objective function. In this case, the parameter estimation of the neural network quantile regression model is transformed into the following optimization problem:
[0091]
[0092] in, Here, ||.||2 is the hidden layer weight, ||.||2 is the L2 norm, and ρ is the penalty parameter;
[0093] Substituting the obtained optimal estimated parameters back into the neural network quantile regression model, the conditional quantile predicted value of the response variable Y can be calculated based on the input explanatory variable X. When the quantile τ takes continuous values in the interval (0,1), the conditional quantile curve... This is the conditional probability distribution curve of Y.
[0094] Optionally, a neural network quantile regression model is constructed based on seasonal trend decomposition to fit key influencing factors with different influence weights, and quantile prediction results are obtained, including:
[0095] Estimating the trend subsequence T t And remove it from the original data to obtain a non-trend time series:
[0096] D t =Y t -T t (k) (10)
[0097] Among them, T t (k) It is the trend subsequence calculated in the current iteration;
[0098] For the trend-removed sequence D t Locally weighted regression is used to smooth the data, thus obtaining the seasonal subsequence. Locally weighted regression is a nonparametric regression method that uses a sliding window to weight and fit the data. The specific formula is shown below:
[0099]
[0100] To update the trend subsequence, the de-seasonalized sequence needs to be low-pass filtered. The low-frequency signal obtained from the low-pass filtering corresponds to the trend subsequence T. t The high-frequency signal corresponds to the seasonal subsequence S. t The specific formula is as follows:
[0101]
[0102] Using the new trend subsequence T t (k+1) Recalculate the seasonal subsequence;
[0103]
[0104] Subtracting the updated seasonal subsequence from the original time series yields the de-seasonalized time series;
[0105]
[0106] Seasonless sequence Perform local weighted regression smoothing to obtain a new trend subsequence;
[0107]
[0108] If the changes in the trend subsequence and the seasonal subsequence satisfy the convergence condition, the algorithm terminates with convergence. If the convergence condition is not satisfied, the algorithm returns to step one and continues iterating. The convergence condition is as follows:
[0109] |T t (k+1) -T t(k) |<ε (16)
[0110] Where ε is the set convergence threshold;
[0111] Seasonal trend decomposition reveals that the load exhibits periodic changes over a 24-hour period. Therefore, a time penalty term is added to the error function of the neural network quantile regression model.
[0112]
[0113] In the formula, n m,t The sample date t is the number of months from the month of the prediction month, with the furthest period having the smallest time weight and the most recent period having the smallest time weight. m,t =α 0 =1 is the maximum, α is the time weight parameter determined based on the prediction error, and the error function thus becomes:
[0114]
[0115] Optionally, based on the quantile prediction results, the continuous probability distribution curve of the common load factor is estimated using nonparametric kernel density techniques to obtain the interval prediction results, including:
[0116] The set of conditional quantiles obtained from the neural network quantile regression model can be regarded as a set of randomly sampled samples that follow the same distribution as Y. The probability density function of Y under the condition of input X can be calculated using nonparametric kernel density estimation.
[0117]
[0118] In the formula, h is the bandwidth, K(*) is the kernel function, n is the number of samples, and Y... i To obtain the n equally spaced quantiles τ of the predictor variable Y for the neural network quantile regression model. i (i = 1, ..., n) Conditional quantiles.
[0119] Alternatively, the following multiple linear regression model can be used:
[0120] Y = β0 + β1X1 + β2X2 + ... + β p X p (20)
[0121] Wherein, β1~β p It is the regression coefficient, X i Y is the influencing factor, and Y is the total load.
[0122] Optionally, the system may also include:
[0123] The second prediction unit is used to construct a multiple linear regression model to predict changes in load size and adjust the interval prediction results.
[0124] In another aspect, the present invention also provides a computing device, comprising: one or more processors;
[0125] A processor is used to execute one or more programs;
[0126] When the one or more programs are executed by the one or more processors, the method described above is implemented.
[0127] In another aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the method described above.
[0128] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0129] This invention provides a method for predicting the probability of power load based on a neural network quantile regression model and multiple linear regression. The method includes: obtaining standardized data using longitudinal data analysis; identifying key influencing factors of power load in the standardized data based on Pearson correlation analysis, and constructing a factor analysis model to quantify the influence weights of these key factors on the power load; fitting a neural network quantile regression model based on seasonal trend decomposition to key influencing factors with different influence weights to obtain quantile prediction results; and estimating the continuous probability distribution curve of common load factors using nonparametric kernel density techniques to obtain interval prediction results. This invention's method for predicting the probability of power load based on a neural network quantile regression model and multiple linear regression effectively improves prediction accuracy. The model's computational efficiency is significantly improved through the construction of Pearson correlation analysis and factor analysis models, fully demonstrating the practical value and potential impact of this invention. Attached Figure Description
[0130] Figure 1 This is a flowchart of the method of the present invention;
[0131] Figure 2 This is a flowchart illustrating the application of the factor analysis model during the implementation of this invention.
[0132] Figure 3 This is a flowchart illustrating the application of the load curve probability prediction method based on QRNN and multiple linear regression in the implementation process of this invention.
[0133] Figure 4 This is a graph showing the results of the seasonal trend decomposition of the load during the implementation of this invention.
[0134] Figure 5 This is a graph showing the prediction results of the multiple linear regression model during the implementation of this invention;
[0135] Figure 6 This is a graph showing the load prediction results for the 90% confidence interval during the implementation of this invention;
[0136] Figure 7 This is a graph showing the load prediction results for the 95% confidence interval during the implementation of this invention;
[0137] Figure 8 This is a graph showing the predicted results for the last week of 2024 during the implementation of this invention. Detailed Implementation
[0138] Exemplary embodiments of the invention will now be described with reference to the accompanying drawings. However, the invention may be embodied in many different forms and is not limited to the embodiments described herein. These embodiments are provided to fully and completely disclose the invention and to fully convey its scope to those skilled in the art. The terminology used in the exemplary embodiments illustrated in the drawings is not intended to limit the invention. In the drawings, the same units / elements are referred to by the same reference numerals.
[0139] Unless otherwise stated, the terms used herein (including technical terms) have their common meaning as understood by one of ordinary skill in the art. Furthermore, it is understood that terms defined in commonly used dictionaries should be understood to have a meaning consistent with the context of their relevant field, and not to be interpreted as having an idealized or overly formal meaning.
[0140] Example 1:
[0141] While some progress has been made in short-term and ultra-short-term load probability forecasting, research on long-term power load probability forecasting is significantly lacking and suffers from the following problems: First, meteorological forecast information closely related to load is not taken into account, resulting in low forecast accuracy; second, the forecast results are mostly monthly or daily average power, failing to provide hourly power time-series fluctuation information; and third, methods suitable for long-term power load probability forecasting are lacking, and as the forecast volume increases, long-term time scales lead to the "curse of dimensionality," making solutions difficult. To ensure that long-term power load probability forecasting results accurately reflect the randomness, volatility, and time-series correlation of actual load, it is necessary to study long-term power load curve probability forecasting methods that can take into account meteorological forecast information, socio-economic factors, and date types, and are convenient for handling high-dimensional time-series vectors.
[0142] To address the aforementioned problems, this invention proposes a power load probability prediction method based on a neural network quantile regression model and multiple linear regression, such as... Figure 1 As shown, it includes:
[0143] Step 1: Obtain standardized data using longitudinal data analysis.
[0144] Step 2: Identify the key influencing factors of power load in the standardized data based on Pearson correlation analysis, and construct a factor analysis model to quantify the influence weight of the key influencing factors on power load;
[0145] Step 3: Based on seasonal trend decomposition, construct a neural network quantile regression model to fit key influencing factors with different influence weights and obtain quantile prediction results;
[0146] Step 4: Based on the quantile prediction results, the continuous probability distribution curve of the common load factor is estimated using nonparametric kernel density techniques to obtain the interval prediction results.
[0147] The following examples will further illustrate steps 1-5 above:
[0148] The specific process is as follows: Figure 2-3 As shown, it includes:
[0149] (1) Standardized data were obtained using longitudinal data analysis.
[0150] The data used in this invention is based on an hourly time scale. Therefore, the original annual load data is represented as a 24-dimensional standardized power load time series vector P, corresponding to the standard load time series for each of the 24 hours of the day:
[0151]
[0152] In the formula, D is the total number of days in the sample observation period; P i Let P be the time series of the daily variation of the standard load at time i; h,d The load standard value at time h on day d is derived from the original value. It is obtained through the following standardization:
[0153]
[0154] In the formula, μ h and S h represents the sample mean and standard deviation of the original load at time h, respectively.
[0155] (2) Based on the Pearson correlation analysis method, the key influencing factors of the power load in the standardized data were identified, and a factor analysis model was constructed to quantify the influence weight of the key influencing factors on the power load.
[0156] ①Pearson correlation analysis is a statistical method used to measure the degree of linear correlation between two variables. This section uses the Pearson correlation coefficient calculation formula to quantitatively analyze the correlation between nine influencing factors—precipitation, east-west wind speed, north-south wind speed, dew point temperature, temperature, surface pressure, daily maximum temperature, daily minimum temperature, and daily average temperature—and electricity load. The correlation calculation formula is as follows:
[0157]
[0158] In the formula: cov(x,y) represents the covariance between x and y; σ x and σ y Then these represent the standard deviations of x and y, respectively.
[0159] The top five factors in terms of absolute value were selected as key factors for subsequent analysis: daily average temperature, daily minimum temperature, surface pressure, precipitation, and daily maximum temperature.
[0160] ② Factor analysis is a multivariate statistical method that simplifies multidimensional vectors through dimensionality reduction techniques. Its basic idea is to analyze the correlation between multivariate data to obtain the independent latent factors that describe this correlation, thereby achieving the purpose of data dimensionality reduction and explaining complex problems with a few variables.
[0161] Let the multidimensional population of variables be X = [X1, ..., X2]. i ,…,X p ] T The general model for factor analysis is as follows:
[0162] X = AF + ε (4)
[0163]
[0164] In the formula, F = [f1, f2, ..., f r ] T Let A be a common factor vector, representing r common influencing factors that are not directly observable but objectively exist. ij ) p×r Let a be the factor loading matrix, and let a be the matrix element. ij For variable X i In common factor f j The load on it reflects the variable X i For common factor f j The relative importance of factors. The factor loading matrix indirectly explains the correlation between different variables by reflecting the correlation between different original variables on the same common factor. ε=[ε1,ε2,…,ε p ] T This is a special factor vector, representing the portion of the multidimensional variable X that cannot be explained by common factors.
[0165] The steps and formulas for solving the factor analysis model are as follows:
[0166] Step 1: Calculate the covariance matrix S of the historical observation sample data;
[0167] Step 2: Calculate the eigenvalues of the covariance matrix S and obtain the corresponding unit eigenvectors e1, e2, ..., e p ;
[0168] Step 3: Extract the first r largest eigenvalues and their corresponding eigenvectors, and calculate the factor loading matrix A:
[0169]
[0170] Step 4: Based on the factor loading matrix The sum of squares of the elements in the i-th row and the diagonal elements s of the covariance matrix S ii Calculate the special factor ε i variance Thus, we obtain D(ε):
[0171]
[0172] After obtaining the common factor vector, we replace the original p-dimensional variables with mutually independent r-dimensional common factor vectors, thereby achieving the effect of dimensionality reduction and simplification.
[0173] (3) Based on seasonal trend decomposition, a neural network quantile regression model is constructed to fit key influencing factors with different influence weights, and the quantile prediction results are obtained:
[0174] ① Quantile regression uses the conditional quantiles of the explained variable Y to regress the input variable X, providing a detailed reflection of the influence of the input variable X on the location, distribution, and shape of the explained variable Y across different ranges. Consider an outcome Y influenced by K factors x1, x2, ..., x... K The influence of quantile regression model is shown below:
[0175] Q Y (τ|X)=β0(τ)+β1(τ)x1+β2(τ)x2+...+β K (τ)x K ≡X T ·β(τ) (9)
[0176] In the formula: Q Y (τ|X) represents the response variable Y in the explanatory variable X = [x1, x2, ..., x]. K ] TConditional τ quantile under given conditions; β(τ)=[β0(τ),β1(τ),β2(τ),...,β K (τ)] T This is the regression coefficient vector, which varies with the quantile τ.
[0177] Neural networks, due to their use of nonlinear kernel functions, are well-suited for handling complex nonlinear objects. Therefore, this invention combines neural networks with quantile regression models to obtain a quantile regression neural network (QRNN). This model utilizes the nonlinear kernel function of the neural network to analyze the complex nonlinear effects of explanatory variables on response variables. The QRNN model uses a three-layer backpropagation (BP) neural network, where the hidden layer contains J nodes. Its model representation is as follows:
[0178]
[0179] In the formula b represents the output layer weights. (o) (τ) represents the output layer offset; f (o) g is the output layer transformation function; j (τ) represents the hidden layer node data, satisfying:
[0180]
[0181] In the formula Hidden layer weights; For hidden layer offset; f (h) This is the hidden layer transition function.
[0182] The hidden layer transition function f of the QRNN model constructed in this invention (h) Choose the hyperbolic tangent tanh function, and the output layer transformation function f. (o) By choosing the equivalent function, the model as a whole exhibits the following nonlinear relationship:
[0183]
[0184] The parameter estimation formula for quantile regression models is similar; the estimation of the parameter vector in a QRNN model also aims to minimize an error function of the following form:
[0185]
[0186] In the formula, T is the sample size; Y i This represents the value of the response variable Y for the i-th training sample; Let I(·) represent the conditional quantile of the i-th response variable; I(·) is the following indicator function:
[0187]
[0188] To prevent the QRNN model from overfitting, a penalty term is added to the objective function. At this point, the parameter estimation of the QRNN model transforms into the following optimization problem:
[0189]
[0190] In the formula ρ represents the hidden layer weights, ||.||2 represents the L2 norm, and ρ represents the penalty parameter.
[0191] Substituting the obtained optimal estimated parameters back into the model, the conditional quantile predicted values of the response variable Y can be calculated based on the input explanatory variable X. When the quantile τ takes continuous values in the interval (0,1), the conditional quantile curve... This is the conditional probability distribution curve of Y.
[0192] ② The seasonal trend decomposition algorithm is a method for time series decomposition. It uses locally weighted regression to divide a time series into three main subsequences: trend, seasonal, and residual subsequences. This method is suitable for handling complex, non-stationary time series data. The trend sequence represents the overall trend of the time series; the seasonal sequence implies the periodicity of the time series; and the residual sequence represents noise signals that cannot be explained by the trend and seasonal sequences. Given the original time series Yt, the formula for seasonal trend decomposition is:
[0193] Y t =S t +T t +R t (16)
[0194] In the formula, S t For seasonal subsequences; T t For trend subsequences; R t This is a residual subsequence. The specific steps are as follows:
[0195] Step 1: Estimate the trend subsequence T t And remove it from the original data to obtain a non-trend time series:
[0196] D t =Y t -T t (k) (17)
[0197] In the formula, T t (k) It is the trend subsequence calculated in the current iteration.
[0198] Step 2: For the trend-removed sequence D tLocally weighted regression is performed to smooth the data, thus obtaining the seasonal subsequence. Locally weighted regression is a nonparametric regression method that uses a sliding window to weight and fit the data. The specific formula is shown below:
[0199]
[0200] Step 3: To update the trend subsequence, the de-seasonalized sequence needs to be low-pass filtered. The low-frequency signal obtained from the low-pass filtering corresponds to the trend subsequence T. t The high-frequency signal corresponds to the seasonal subsequence S. t The specific formula is as follows:
[0201]
[0202] Step 4: Use the new trend subsequence T t (k+1) Recalculate the seasonal subsequence.
[0203]
[0204] Step 5: Subtract the updated seasonal subsequence from the original time series to obtain the de-seasonalized time series.
[0205]
[0206] Step Six: De-seasonalize the sequence Local weighted regression smoothing is performed to obtain a new trend subsequence.
[0207]
[0208] Step 7: If the changes in the trend subsequence and the seasonal subsequence satisfy the convergence condition, the algorithm converges. If the convergence condition is not met, return to Step 1 and continue iterating. The convergence condition is as follows:
[0209] |T t (k+1) -T t (k) |<ε (23)
[0210] In the formula, ε is the set convergence threshold.
[0211] Seasonal trend decomposition reveals that the load exhibits periodic changes over a 24-hour period. Therefore, a time penalty term is added to the error function of the QRNN model.
[0212]
[0213] In the formula, n m,tThe sample date t is the number of months from the month of the prediction month, with the furthest period having the smallest time weight and the most recent period having the smallest time weight. m,t =α 0 =1 is the maximum. α is the time weight parameter determined based on the prediction error. The error function thus becomes:
[0214]
[0215] (4) Based on the quantile prediction results, the continuous probability distribution curve of the common load factor is estimated using nonparametric kernel density technology to obtain the interval prediction results.
[0216] The set of conditional quantiles obtained from the QRNN model can be viewed as a set of randomly sampled samples that follow the same distribution as Y. The probability density function of Y under the condition of input X can be calculated using nonparametric kernel density estimation.
[0217]
[0218] In the formula, h is the bandwidth, K(*) is the kernel function, n is the number of samples, and Y... i To obtain the n equally spaced quantiles τ of the predictor variable Y for the QRNN prediction model. i (i = 1, ..., n) Conditional quantiles.
[0219] (5) Construct a multiple linear regression model to predict changes in load size and adjust the interval prediction results.
[0220] This study uses a multiple linear regression model to explain the relationship between socioeconomic factors and load size, and to predict changes in load size. The multiple linear regression model is as follows:
[0221] Y = β0 + β1X1 + β2X2 + ... + β p X p (27)
[0222] In the formula, β1~β p It is the regression coefficient; X i Y represents the total load. This invention selects GDP, primary industry GDP, secondary industry GDP, tertiary industry GDP, total population, import and export situation, consumer price index, and per capita disposable income as relevant socioeconomic factors.
[0223] In the embodiments, its effects are as follows: Figure 3-8As shown. To more clearly demonstrate the purpose, technical solution, and advantages of this invention, this invention selects the Shandong power grid's electricity load dataset from 2021 to 2024, and obtains hourly meteorological data for relevant years from the ERA5 dataset. This invention uses the data from 2021 to 2023 as the training set and the data from 2024 as the test set, and uses a QRNN model to predict the load interval. Furthermore, a multiple linear regression model is fitted using relevant socioeconomic data and load data from 2010 to 2023, and the load scale for 2024 is predicted using the socioeconomic data from 2024. Experimental results show that the 70% confidence interval covers 75.5% of the actual values, the 90% confidence interval covers 91.8% of the actual values, and the 95% confidence interval covers 96.1% of the actual values, indicating that the predicted intervals basically cover the actual values; simultaneously, the average absolute percentage error of the predicted values is 7.67%, indicating a small prediction error and high model accuracy.
[0224] Example 2:
[0225] This invention also proposes a power load probability prediction system 200 based on a neural network quantile regression model and multiple linear regression, comprising:
[0226] Standardization unit 201 is used to obtain standardized data using longitudinal data analysis.
[0227] The identification unit 202 is used to identify the key influencing factors of power load in the standardized data based on the Pearson correlation analysis method, and to construct a factor analysis model to quantify the influence weight of the key influencing factors on power load.
[0228] Fitting unit 203 is used to fit key influencing factors with different influence weights to a neural network quantile regression model based on seasonal trend decomposition, and to obtain quantile prediction results.
[0229] The first prediction unit 204 is used to estimate the continuous probability distribution curve of the common load factor based on the quantile prediction results using nonparametric kernel density techniques to obtain the interval prediction results. The standardized data obtained using longitudinal data analysis includes:
[0230] Using hours as the time scale, the annual load data is represented as a 24-dimensional standardized electricity load time series vector P, corresponding to the standard load time series for each of the 24 hours of the day, as shown in the following formula:
[0231]
[0232] Where D is the total number of days in the sample observation period, and P i Let P be the time series of the daily variation of the standard load at time i.h,d The load standard value at time h on day d is derived from the original value. It is obtained through the following standardization:
[0233]
[0234] Where, μ h and S h represents the sample mean and standard deviation of the original load at time h, respectively.
[0235] Specifically, based on Pearson correlation analysis, key influencing factors of electricity load in the standardized data were identified, and a factor analysis model was constructed to quantify the influence weights of these key influencing factors on electricity load, including:
[0236] Calculate the covariance matrix S of the historical observation sample data;
[0237] Calculate the eigenvalues of the covariance matrix S and obtain the corresponding unit eigenvectors e1, e2, ..., e p ;
[0238] Extract the first r largest eigenvalues and their corresponding eigenvectors, and calculate the factor loading matrix. The calculation formula is as follows:
[0239]
[0240] Based on the factor loading matrix The sum of squares of the elements in the i-th row and the diagonal elements s of the covariance matrix S ii Calculate the special factor ε i variance Thus, we obtain D(ε):
[0241]
[0242] After obtaining the common factor vector, we replace the original p-dimensional variables with mutually independent r-dimensional common factor vectors.
[0243] Among them, the hidden layer transformation function f of the neural network quantile regression model (h) Choose the hyperbolic tangent tanh function, and the output layer transformation function f. (o) By choosing the iso-function, the neural network quantile regression model exhibits the following nonlinear relationship:
[0244]
[0245] The estimation of the parameter vector in a neural network quantile regression model is to minimize an error function of the following form:
[0246]
[0247] Where T is the sample size, Y i This represents the value of the response variable Y for the i-th training sample. Let I(·) represent the conditional quantile of the i-th response variable, and let I(·) be the following indicator function:
[0248]
[0249] To avoid overfitting in the neural network quantile regression model, a penalty term is added to the objective function. In this case, the parameter estimation of the neural network quantile regression model is transformed into the following optimization problem:
[0250]
[0251] in, Here, ||.||2 is the hidden layer weight, ||.||2 is the L2 norm, and ρ is the penalty parameter;
[0252] Substituting the obtained optimal estimated parameters back into the neural network quantile regression model, the conditional quantile predicted value of the response variable Y can be calculated based on the input explanatory variable X. When the quantile τ takes continuous values in the interval (0,1), the conditional quantile curve... This is the conditional probability distribution curve of Y.
[0253] Among them, a neural network quantile regression model based on seasonal trend decomposition is constructed to fit key influencing factors with different influence weights, and the quantile prediction results are obtained, including:
[0254] Estimating the trend subsequence T t And remove it from the original data to obtain a non-trend time series:
[0255] D t =Y t -T t (k) (10)
[0256] Among them, T t (k) It is the trend subsequence calculated in the current iteration;
[0257] For the trend-removed sequence D t Locally weighted regression is used to smooth the data, thus obtaining the seasonal subsequence. Locally weighted regression is a nonparametric regression method that uses a sliding window to weight and fit the data. The specific formula is shown below:
[0258]
[0259] To update the trend subsequence, the de-seasonalized sequence needs to be low-pass filtered. The low-frequency signal obtained from the low-pass filtering corresponds to the trend subsequence T. t The high-frequency signal corresponds to the seasonal subsequence S. t The specific formula is as follows:
[0260]
[0261] Using the new trend subsequence T t (k+1) Recalculate the seasonal subsequence;
[0262]
[0263] Subtracting the updated seasonal subsequence from the original time series yields the de-seasonalized time series;
[0264]
[0265] Seasonless sequence Perform local weighted regression smoothing to obtain a new trend subsequence;
[0266]
[0267] If the changes in the trend subsequence and the seasonal subsequence satisfy the convergence condition, the algorithm terminates with convergence. If the convergence condition is not satisfied, the algorithm returns to step one and continues iterating. The convergence condition is as follows:
[0268] |T t (k+1) -T t (k) |<ε (16)
[0269] Where ε is the set convergence threshold;
[0270] Seasonal trend decomposition reveals that the load exhibits periodic changes over a 24-hour period. Therefore, a time penalty term is added to the error function of the QRNN model.
[0271]
[0272] In the formula, n m,t The sample date t is the number of months from the month of the prediction month, with the furthest period having the smallest time weight and the most recent period having the smallest time weight. m,t =α 0 =1 is the maximum, α is the time weight parameter determined based on the prediction error, and the error function thus becomes:
[0273]
[0274] Among them, based on the quantile prediction results, the continuous probability distribution curve of the common load factor is estimated using nonparametric kernel density techniques to obtain the interval prediction results, including:
[0275] The set of conditional quantiles obtained from the neural network quantile regression model can be regarded as a set of randomly sampled samples that follow the same distribution as Y. The probability density function of Y under the condition of input X can be calculated using nonparametric kernel density estimation.
[0276]
[0277] In the formula, h is the bandwidth, K(*) is the kernel function, n is the number of samples, and Y... i To obtain the n equally spaced quantiles τ of the predictor variable Y for the neural network quantile regression model. i (i = 1, ..., n) Conditional quantiles.
[0278] The multiple linear regression model is as follows:
[0279] Y = β0 + β1X1 + β2X2 + ... + β p X p (20)
[0280] Wherein, β1~β p It is the regression coefficient, X i Y is the influencing factor, and Y is the total load.
[0281] The system also includes:
[0282] The second prediction unit 205 is used to construct a multiple linear regression model to predict changes in load size and adjust the interval prediction results.
[0283] The present invention provides a power load probability prediction method based on a neural network quantile regression model and multiple linear regression, which can effectively improve prediction accuracy. The model's computational efficiency can be effectively improved through the construction of Pearson correlation analysis and factor analysis models, fully demonstrating the practical value and potential impact of the invention.
[0284] Example 3:
[0285] Based on the same inventive concept, this invention also provides a computer device, which includes a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to implement corresponding method flows or corresponding functions, thereby implementing the steps of the methods in the above embodiments.
[0286] Example 4:
[0287] Based on the same inventive concept, this invention also provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the steps of the method in the above embodiments.
[0288] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0289] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0290] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0291] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0292] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0293] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for predicting the probability of electricity load based on a neural network quantile regression model and multiple linear regression, characterized in that, include: Standardized data were obtained using longitudinal data analysis. Based on Pearson correlation analysis, the key influencing factors of power load in the standardized data were identified, and a factor analysis model was constructed to quantify the influence weight of the key influencing factors on power load. Based on seasonal trend decomposition, a neural network quantile regression model is constructed to fit key influencing factors with different influence weights and obtain quantile prediction results. Based on the quantile prediction results, the continuous probability distribution curve of the common load factor is estimated using nonparametric kernel density techniques to obtain the interval prediction results.
2. The power load probability prediction method based on neural network quantile regression model and multiple linear regression according to claim 1, characterized in that, The standardized data obtained using longitudinal data analysis includes: Using hours as the time scale, the annual load data is represented as a 24-dimensional standardized electricity load time series vector P, corresponding to the standard load time series for each of the 24 hours of the day, as shown in the following formula: Where D is the total number of days in the sample observation period, and P i Let P be the time series of the daily variation of the standard load at time i. h,d The load standard value at time h on day d is derived from the original value. It is obtained through the following standardization: Where, μ h and S h represents the sample mean and standard deviation of the original load at time h, respectively.
3. The power load forecasting method based on neural network quantile regression model and multiple linear regression under different probabilities as described in claim 1, characterized in that, The method based on Pearson correlation analysis identifies key influencing factors of electricity load in the standardized data, and constructs a factor analysis model to quantify the influence weight of the key influencing factors on electricity load, including: Calculate the covariance matrix S of the historical observation sample data; Calculate the eigenvalues of the covariance matrix S and obtain the corresponding unit eigenvectors e1, e2, ..., e p ; Extract the first r largest eigenvalues and their corresponding eigenvectors, and calculate the factor loading matrix. The calculation formula is as follows: Based on the factor loading matrix The sum of squares of the elements in the i-th row and the diagonal elements s of the covariance matrix S ii Calculate the special factor ε i variance Thus, we obtain D(ε): After obtaining the common factor vector, we replace the original p-dimensional variables with mutually independent r-dimensional common factor vectors.
4. The power load probability prediction method based on neural network quantile regression model and multiple linear regression according to claim 1, characterized in that, The hidden layer transformation function f of the neural network quantile regression model (h) Choose the hyperbolic tangent tanh function, and the output layer transformation function f. (o) By choosing the iso-function, the neural network quantile regression model exhibits the following nonlinear relationship: The estimation of the parameter vector in a neural network quantile regression model is to minimize an error function of the following form: Where T is the sample size, Y i This represents the value of the response variable Y for the i-th training sample. Let I(·) represent the conditional quantile of the i-th response variable, and let I(·) be the following indicator function: To avoid overfitting in the neural network quantile regression model, a penalty term is added to the objective function. In this case, the parameter estimation of the neural network quantile regression model is transformed into the following optimization problem: in, Here, ||.||2 is the hidden layer weight, ||.||2 is the L2 norm, and ρ is the penalty parameter; Substituting the obtained optimal estimated parameters back into the neural network quantile regression model, the conditional quantile predicted value of the response variable Y can be calculated based on the input explanatory variable X. When the quantile τ takes continuous values in the interval (0,1), the conditional quantile curve... This is the conditional probability distribution curve of Y.
5. The power load probability prediction method based on neural network quantile regression model and multiple linear regression according to claim 1, characterized in that, The neural network quantile regression model constructed based on seasonal trend decomposition fits key influencing factors with different influence weights to obtain quantile prediction results, including: Estimating the trend subsequence T t And remove it from the original data to obtain a non-trend time series: D t =Y t -T t (k) (10) Among them, T t (k) It is the trend subsequence calculated in the current iteration; For the trend-removed sequence D t Locally weighted regression is used to smooth the data, thus obtaining the seasonal subsequence. Locally weighted regression is a nonparametric regression method that uses a sliding window to weight and fit the data. The specific formula is shown below: To update the trend subsequence, the de-seasonalized sequence needs to be low-pass filtered. The low-frequency signal obtained from the low-pass filtering corresponds to the trend subsequence T. t The high-frequency signal corresponds to the seasonal subsequence S. t The specific formula is shown below: Using the new trend subsequence T t (k+1) Recalculate the seasonal subsequence; Subtracting the updated seasonal subsequence from the original time series yields the de-seasonalized time series; Seasonless sequence Perform local weighted regression smoothing to obtain a new trend subsequence; If the changes in the trend subsequence and the seasonal subsequence satisfy the convergence condition, the algorithm converges and terminates. If the convergence condition is not satisfied, return to step one and continue iterating. The convergence condition is as follows: |T t (k+1) -T t (k) |<ε (16) Where ε is the set convergence threshold; Seasonal trend decomposition reveals that the load exhibits periodic changes over a 24-hour period. Therefore, a time penalty term is added to the error function of the neural network quantile regression model. In the formula, n m,t The sample date t is the number of months from the month of the prediction month, with the furthest period having the smallest time weight and the most recent period having the smallest time weight. m,t =α 0 =1 is the maximum, α is the time weight parameter determined based on the prediction error, and the error function thus becomes:
6. The power load probability prediction method based on neural network quantile regression model and multiple linear regression according to claim 1, characterized in that, The interval prediction results are obtained by estimating the continuous probability distribution curve of the common load factor based on the quantile prediction results using nonparametric kernel density techniques, including: The set of conditional quantiles obtained from the neural network quantile regression model can be regarded as a set of randomly sampled samples that follow the same distribution as Y. The probability density function of Y under the condition of input X can be calculated using nonparametric kernel density estimation. In the formula, h is the bandwidth, K(*) is the kernel function, n is the number of samples, and Y... i To obtain the n equally spaced quantiles τ of the predictor variable Y for the neural network quantile regression model. i (i = 1, ..., n) Conditional quantiles.
7. The power load probability prediction method based on neural network quantile regression model and multiple linear regression according to claim 1, characterized in that, The multiple linear regression model is as follows: Y=β0+β1X1+β2X2+...+β p X p (20) Wherein, β1~β p It is the regression coefficient, X i Y is the influencing factor, and Y is the total load.
8. The power load probability prediction method based on neural network quantile regression model and multiple linear regression according to claim 1, characterized in that, The method further includes: A multiple linear regression model is constructed to predict changes in load size, and the interval prediction results are adjusted accordingly.
9. A power load probability prediction system based on a neural network quantile regression model and multiple linear regression, characterized in that, include: Standardized units are used to obtain standardized data using longitudinal data analysis methods. The identification unit is used to identify the key influencing factors of power load in the standardized data based on the Pearson correlation analysis method, and to construct a factor analysis model to quantify the influence weight of the key influencing factors on power load. The fitting unit is used to fit key influencing factors with different influence weights to a neural network quantile regression model based on seasonal trend decomposition, and to obtain quantile prediction results. The first prediction unit is used to estimate the continuous probability distribution curve of the common load factor based on the quantile prediction results and to obtain the interval prediction results by using nonparametric kernel density techniques.
10. The power load probability prediction system based on neural network quantile regression model and multiple linear regression according to claim 9, characterized in that, The standardized data obtained using longitudinal data analysis includes: Using hours as the time scale, the annual load data is represented as a 24-dimensional standardized electricity load time series vector P, corresponding to the standard load time series for each of the 24 hours of the day, as shown in the following formula: Where D is the total number of days in the sample observation period, and P i Let P be the time series of the daily variation of the standard load at time i. h,d The load standard value at time h on day d is derived from the original value. It is obtained through the following standardization: Where, μ h and S h represents the sample mean and standard deviation of the original load at time h, respectively.
11. The power load probability prediction system based on neural network quantile regression model and multiple linear regression according to claim 9, characterized in that, The method based on Pearson correlation analysis identifies key influencing factors of electricity load in the standardized data, and constructs a factor analysis model to quantify the influence weight of the key influencing factors on electricity load, including: Calculate the covariance matrix S of the historical observation sample data; Calculate the eigenvalues of the covariance matrix S and obtain the corresponding unit eigenvectors e1, e2, ..., e p ; Extract the first r largest eigenvalues and their corresponding eigenvectors, and calculate the factor loading matrix. The calculation formula is as follows: Based on the factor loading matrix The sum of squares of the elements in the i-th row and the diagonal elements s of the covariance matrix S ii Calculate the special factor ε i variance Thus, we obtain D(ε): After obtaining the common factor vector, we replace the original p-dimensional variables with mutually independent r-dimensional common factor vectors.
12. The power load probability prediction system based on neural network quantile regression model and multiple linear regression according to claim 9, characterized in that, The hidden layer transformation function f of the neural network quantile regression model (h) Choose the hyperbolic tangent tanh function, and the output layer transformation function f. (o) By choosing the iso-function, the neural network quantile regression model exhibits the following nonlinear relationship: The estimation of the parameter vector in a neural network quantile regression model is to minimize an error function of the following form: Where T is the sample size, Y i This represents the value of the response variable Y for the i-th training sample. Let I(·) represent the conditional quantile of the i-th response variable, and let I(·) be the following indicator function: To avoid overfitting in the neural network quantile regression model, a penalty term is added to the objective function. In this case, the parameter estimation of the neural network quantile regression model is transformed into the following optimization problem: in, Here, ||.||2 is the hidden layer weight, ||.||2 is the L2 norm, and ρ is the penalty parameter; Substituting the obtained optimal estimated parameters back into the neural network quantile regression model, the conditional quantile predicted value of the response variable Y can be calculated based on the input explanatory variable X. When the quantile τ takes continuous values in the interval (0,1), the conditional quantile curve... This is the conditional probability distribution curve of Y.
13. The power load probability prediction system based on neural network quantile regression model and multiple linear regression according to claim 9, characterized in that, The neural network quantile regression model constructed based on seasonal trend decomposition fits key influencing factors with different influence weights to obtain quantile prediction results, including: Estimating the trend subsequence T t And remove it from the original data to obtain a non-trend time series: D t =Y t -T t (k) (10) Among them, T t (k) It is the trend subsequence calculated in the current iteration; For the trend-removed sequence D t Locally weighted regression is used to smooth the data, thus obtaining the seasonal subsequence. Locally weighted regression is a nonparametric regression method that uses a sliding window to weight and fit the data. The specific formula is shown below: To update the trend subsequence, the de-seasonalized sequence needs to be low-pass filtered. The low-frequency signal obtained from the low-pass filtering corresponds to the trend subsequence T. t The high-frequency signal corresponds to the seasonal subsequence S. t The specific formula is shown below: Using the new trend subsequence T t (k+1) Recalculate the seasonal subsequence; Subtracting the updated seasonal subsequence from the original time series yields the de-seasonalized time series; Seasonless sequence Perform local weighted regression smoothing to obtain a new trend subsequence; If the changes in the trend subsequence and the seasonal subsequence satisfy the convergence condition, the algorithm converges and terminates. If the convergence condition is not satisfied, return to step one and continue iterating. The convergence condition is as follows: Where ε is the set convergence threshold; Seasonal trend decomposition reveals that the load exhibits periodic changes over a 24-hour period. Therefore, a time penalty term is added to the error function of the neural network quantile regression model. In the formula, n m,t The sample date t is the number of months from the month of the prediction month, with the furthest period having the smallest time weight and the most recent period having the smallest time weight. m,t =α 0 =1 is the maximum, α is the time weight parameter determined based on the prediction error, and the error function thus becomes:
14. The power load probability prediction system based on neural network quantile regression model and multiple linear regression according to claim 9, characterized in that, The interval prediction results are obtained by estimating the continuous probability distribution curve of the common load factor based on the quantile prediction results using nonparametric kernel density techniques, including: The set of conditional quantiles obtained from the neural network quantile regression model can be regarded as a set of randomly sampled samples that follow the same distribution as Y. The probability density function of Y under the condition of input X can be calculated using nonparametric kernel density estimation. In the formula, h is the bandwidth, K(*) is the kernel function, n is the number of samples, and Y... i To obtain the n equally spaced quantiles τ of the predictor variable Y for the neural network quantile regression model. i (i = 1, ..., n) Conditional quantiles.
15. The power load probability prediction system based on neural network quantile regression model and multiple linear regression according to claim 9, characterized in that, The multiple linear regression model is as follows: Y=β0+β1X1+β2X2+...+β p X p (20) Wherein, β1~β p It is the regression coefficient, X i Y is the influencing factor, and Y is the total load.
16. The power load probability prediction system based on neural network quantile regression model and multiple linear regression according to claim 9, characterized in that, The system also includes: The second prediction unit is used to construct a multiple linear regression model to predict changes in load size and adjust the interval prediction results.
17. A computer device, characterized in that, include: One or more processors; A processor is used to execute one or more programs; When the one or more programs are executed by the one or more processors, the method described in any one of claims 1-8 is implemented.
18. A computer-readable storage medium, characterized in that, It contains a computer program, which, when executed, implements the method as described in any one of claims 1-8.
Citation Information
Cited By
Pruning lightweight-based SPR connection quality tracing method and system
CN121903644A