Adaptive bayesian identification method for long-term prediction of industrial chain power load
By employing an adaptive Bayesian identification method and utilizing frequency domain analysis and sparse regression models, the problem of the inability to comprehensively consider the power load characteristics of upstream and downstream industrial chains in existing technologies has been solved, achieving low-cost, high-precision, and highly interpretable power load forecasting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies struggle to coordinate the power load characteristics of upstream and downstream industrial chains within industrial enterprises, cannot accurately capture multi-scale periodic structures, and suffer from high computational costs and low physical interpretability.
An adaptive Bayesian identification method is adopted, and a regularized sparse regression model is constructed through frequency domain analysis and an adaptive basis function library to screen out key periodic components and perform long-term predictions in conjunction with time indexing.
It enables accurate capture of multi-scale periodic characteristics of power load with low computational cost, improves physical interpretability and prediction stability, and adapts to the needs of different industries.
Smart Images

Figure CN121581330B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of data analysis, and particularly relates to an adaptive Bayesian identification method for long-term prediction of power load of an industrial chain. BACKGROUND
[0002] Industrial enterprises, as the main load side of the power grid, have power consumption data that are characterized by high-frequency collection, cross-regional, cross-industry and long-period, and these data have become a key resource for describing the operation rhythm and energy utilization structure of the industrial chain.
[0003] In the upstream and downstream industrial chain, the power load characteristics of different industries differ significantly and have multi-scale time-dependent properties. Upstream industries (such as steel, chemical industry, building materials) usually have high load levels, long production cycles and relatively stable shift systems, and their intraday fluctuations are relatively smooth, but they show significant periodicity at the weekly, monthly and seasonal scales. Downstream industries (such as equipment manufacturing, electronic products) are greatly affected by market demand, order fluctuations and environmental factors, and show stronger seasonality and uncertainty. This "multi-industry, multi-scale, multi-mode" characteristic poses extremely high requirements for unified load modeling.
[0004] In the prior art, methods for such medium and long-term power load prediction are mainly divided into two categories: traditional statistical time series analysis methods and artificial intelligence methods based on deep learning.
[0005] The first category is represented by the autoregressive integrated moving average model and its seasonal extension model. This type of method has certain advantages in dealing with approximately linear, weakly non-stationary load sequences by parameterizing trend and seasonal components, and the model structure is relatively simple and easy to interpret. However, when faced with strong nonlinear characteristics, multiple periodic superimpositions (such as daily, weekly and annual cycles existing at the same time) and sudden structural changes in industrial load, traditional statistical models often fail to capture complex cross-scale patterns, resulting in a significant decrease in prediction accuracy.
[0006] The second category is the deep learning-based method that has emerged in recent years, such as Long-Short Term Memory (LSTM), Gate Recurrent Unit (GRU) and Transformer architecture. These models can automatically extract high-dimensional features from the original sequence through deep network structures, effectively alleviating the nonlinearity and multi-scale problems, and have made significant progress in prediction accuracy. However, deep learning models usually have the following significant defects:
[0007] Lack of interpretability (black box model): the internal parameters of the model are numerous and the structure is complex, which is difficult to directly explain the physical meaning behind the prediction results (such as specific periodic components or trend driving forces), which is a great challenge for power system planning and industrial analysis that requires high transparency and credibility.
[0008] High training cost and dependence on data volume: such models are very sensitive to sample size and hyperparameter adjustment, and the training process consumes a lot of computing resources, and is prone to overfitting in industrial scenarios with sparse data or a large amount of noise.
[0009] It is difficult to balance the differences between multiple industries: existing models often have difficulty adapting to both high-inertia upstream industries and highly volatile downstream industries in a unified framework, and there is a lack of a general modeling method that can maintain structural simplicity while adapting to different industrial characteristics. SUMMARY
[0010] In view of the problems in the prior art, the present application provides an adaptive Bayesian identification method for long-term prediction of industrial chain electricity load, which at least partially solves the problems of inability to coordinate upstream and downstream industrial chain characteristics, inability to accurately capture multi-scale periodic structure, high computational cost and low physical interpretability in the prior art.
[0011] The present disclosure provides an adaptive Bayesian identification method for long-term prediction of industrial chain electricity load, comprising:
[0012] S1: Obtain daily high-frequency active power data of target industrial users, and aggregate the active power data into daily frozen power time series;
[0013] S2: Discrete Fourier transform is performed on the daily frozen power time series to obtain amplitude spectrum, and the spectral complexity index is calculated to determine the number of dominant frequencies and specific frequency components;
[0014] S3: According to the screened dominant frequency, an adaptive basis function library containing sine and cosine functions is constructed;
[0015] S4: A sparse regression model containing regularization is established, the daily frozen power time series is represented as a linear combination of basis functions in the adaptive basis function library, and the sparse coefficient vector is obtained by solving an optimization problem, so as to screen out key periodic components;
[0016] S5: Using the identified sparse coefficient vector and adaptive basis function library, the daily frozen power prediction value of the future time step is calculated through time index extrapolation.
[0017] Optionally, in step S1, the daily high-frequency active power data of the target industrial user is acquired, and the active power data is aggregated into a daily frozen power time series, including constructing a formula of daily frozen power consumption, which is:
[0018] ;
[0019] wherein is the pre-processed 15-minute average active power.
[0020] Optionally, in step S2, the calculation formula of the spectral complexity index is:
[0021] ,
[0022] wherein, and both represent the amplitude spectrum values arranged in descending order, is the maximum amplitude; the spectral complexity index is used to quantify the complexity of the periodic structure, and the spectral complexity index is also used to assist in determining the number of dominant frequencies .
[0023] Optionally, in step S3, the construction method of the adaptive basis function library includes:
[0024] selecting dominant frequencies based on the amplitude spectrum, denoted as , the normalized frequency of which is defined as ;
[0025] For each normalized frequency , two real-valued basis functions are constructed:
[0026] ;
[0027] All basis functions are collected into a design matrix ;
[0028] wherein, is the total length of the daily frozen power time series.
[0029] Optionally, the daily frozen power time series is represented as a linear combination of the linear combination of the basis functions in the adaptive basis function library by establishing a sparse regression model containing regularization, and the formula is:
[0030] ,
[0031] wherein is the coefficient related to the th basis function, is the th basis function in the input the function value at the step.
[0032] Optionally, in step S4, the optimization objective function of the sparse regression model is:
[0033] ,
[0034] where, is the Euclidean norm, is the norm, is a regularization parameter controlling the sparsity level, is the coefficient vector, is the daily frozen vector, is the real number set.
[0035] Optionally, in step S5, using the identified sparse coefficient vector and the adaptive basis function library, the daily frozen power prediction value at the future time step is calculated by time index extrapolation, including:
[0036] For any future date index , the prediction value of the day is given by:
[0037] where is the basis function evaluation vector at the , is the estimated value of the coefficient related to the th basis function, is the prediction value, is the estimated value of the coefficient vector.
[0038] Optionally, it also includes the step of multi-industry joint modeling, which includes:
[0039] For industry , a multivariate model is established,
[0040] The multivariate model is rewritten as where is the coefficient matrix;
[0041] The sparse coefficient matrix is identified by solving the optimization problem with regularization ;
[0042] The industry-level load prediction is calculated by ;
[0043] where is the Frobenius norm, denotes the sum of the absolute values of all entries, is a regularization parameter, is an estimate of the coefficient matrix, is a real matrix.
[0044] Optionally, the method further comprises a step of introducing auxiliary features, the introducing auxiliary features comprising:
[0045] collecting the sequence of auxiliary features to form a feature matrix;
[0046] constructing an augmented design matrix based on the feature matrix;
[0047] replacing the sparse regression problem with where an augmented coefficient vector comprising both periodic basis coefficients and auxiliary variable coefficients, is the augmented design matrix, is a regularization parameter.
[0048] Optionally, the multi-industry joint modeling comprises:
[0049] concatenating the industry-level matrices of a selected group of industries along columns into a joint matrix;
[0050] establishing a joint model based on the joint matrix;
[0051] performing joint sparse identification by solving where is the joint model, is a coefficient matrix, is a regularization parameter;
[0052] generating a prediction for all industries participating in the joint modeling based on the solution.
[0053] The adaptive Bayesian identification method for long-term prediction of power load of an industry chain provided by the application can accurately capture complex multi-scale periodic characteristics in power load through frequency domain analysis and an adaptive basis function library, improves the robustness of the method through regularized sparse regression, and avoids overfitting through sparsity, so that the method is stable on unseen data and can be reliably used for medium and long term prediction. High frequency data is aggregated into a time series, the amplitude spectrum is obtained through discrete Fourier transform, and a basis function library is constructed, improving the physical interpretability. Aggregating data into a time series reduces the dimensionality, the discrete Fourier transform is efficiently calculated, the basis function library is constructed and the sparse regression problem is solved, reducing the calculation cost. Thus, the characteristics of upstream and downstream industry chains can be coordinated, the multi-scale periodic structure can be accurately captured, and the method has low calculation cost and high physical interpretability. BRIEF DESCRIPTION OF DRAWINGS
[0054] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which like reference characters refer to like parts throughout the figures, and in which:
[0055] Figure 1 A flowchart of an adaptive Bayesian identification method for long-term prediction of industrial chain electricity load provided by an embodiment of the present disclosure;
[0056] Figure 2 A comparison diagram of time series prediction of ARIMA and SBL in the steel industry provided by an embodiment of the present disclosure;
[0057] Figure 3 A comparison diagram of time series prediction of SINDy and LSTM in the steel industry provided by an embodiment of the present disclosure;
[0058] Figure 4 A comparison diagram of time series prediction of GRU and ResNet in the photovoltaic industry provided by an embodiment of the present disclosure;
[0059] Figure 5 A comparison diagram of time series prediction of ARIMA and SBL in the photovoltaic industry provided by an embodiment of the present disclosure;
[0060] Figure 6 A comparison diagram of time series prediction of SINDy and LSTM in the photovoltaic industry provided by an embodiment of the present disclosure;
[0061] Figure 7 A comparison diagram of time series prediction of GRU and ResNet in the photovoltaic industry provided by an embodiment of the present disclosure;
[0062] Figure 8 A comparison diagram of time series prediction of ARIMA and SBL in the chemical industry provided by an embodiment of the present disclosure;
[0063] Figure 9 A comparison diagram of time series prediction of SINDy and LSTM in the chemical industry provided by an embodiment of the present disclosure;
[0064] Figure 10 A comparison diagram of time series prediction of GRU and ResNet in the chemical industry provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0065] The embodiments of the present disclosure will be described in detail below with reference to the drawings.
[0066] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0067] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The illustrations only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0068] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0069] like Figure 1 As shown, this embodiment discloses an adaptive Bayesian identification method for long-term forecasting of electricity load in an industrial chain, specifically an adaptive Bayesian system identification method for long-term forecasting of electricity load in upstream and downstream industries; the method includes the following steps:
[0070] Step S1: Data collection and daily frozen value construction. Obtain daily high-frequency active power data of target industry users and aggregate them into daily frozen power time series. The daily frozen power time series reflects the long-term trend and periodic information of users or industries.
[0071] Step S2: Frequency domain feature analysis and dominant frequency selection. Perform discrete Fourier transform on the daily frozen power time series to obtain the amplitude spectrum and calculate the spectral complexity index to determine the number of dominant frequencies and specific frequency components.
[0072] Step S3: Construction of the adaptive sine basis function library. Based on the selected dominant frequencies, an adaptive basis function library containing sine and cosine functions is constructed to form the design matrix.
[0073] Step S4: Sparse coefficient identification, establishing a system containing... The regularized sparse regression model represents the daily frozen electricity time series as a linear combination of basis functions in the basis function library, and obtains the sparse coefficient vector by solving the optimization problem, thereby screening out the key periodic components.
[0074] Step S5: Long-term load extrapolation forecast. Using the identified sparse coefficients and basis function library, the daily frozen power forecast value for future time steps is calculated through extrapolation of the time index.
[0075] The formula for constructing daily frozen electricity consumption is defined as follows:
[0076] ;
[0077] in It is the preprocessed 15-minute average active power. Equation (1) compresses the intraday fine-grained fluctuations into a scalar that retains the total consumption, thus providing a simple and intuitive representation for the identification of medium- and long-term trends and dominant cycles.
[0078] Based on this, multi-scale representations can be constructed as needed. At the weekly scale, daily frozen values for a user within the same week can be aggregated to form the weekly frozen value. The daily frozen value is defined as the sum of all daily frozen values within the week, thus characterizing the weekly load level. Similarly, a monthly frozen value can be defined, or a sliding window smoothing can be applied to the daily series to highlight seasonality and long-term trends while mitigating short-term random disturbances. Furthermore, the daily frozen series can be decomposed into a long-term baseline component and a short-term fluctuation component, thereby distinguishing between structural and transient changes and supporting forecasting and scenario analysis at different time scales.
[0079] At the industry level, for the industry It can be seen from Select A representative user group was identified, and an industry-level load vector was constructed.
[0080] ,
[0081] Thus, multivariate time series are obtained. In this way, the evolution of industrial-level loads and the heterogeneity among representative firms can be jointly captured.
[0082] Based on the above data, the load forecasting problem for upstream and downstream industries can be formalized into a standard time series forecasting task. For a single user... The daily freezing sequences observed during the study period are as follows:
[0083] ,
[0084] During the model training process, the first day frozen value is used as the historical observation to build the prediction model. The next prediction task is to estimate the power consumption of the future days based on this historical window, i.e.
[0085] ,
[0086] This results in a daily load prediction trajectory for the user in the prediction period.
[0087] At the industry level, for the industry load vector sequence , the prediction task is to estimate the future day industry-level load vector given a historical window of length :
[0088] ,
[0089] where each is a dimensional vector representing the predicted power consumption of a representative enterprise in the industry .
[0090] Considering the upstream and downstream structure of the industrial chain, a joint multi-industry prediction task can also be defined. Let and represent the upstream and downstream industry sets, respectively. The joint time series:
[0091] ,
[0092] can be used to build a model to capture the dynamic changes of multiple industries and identify the periodic structure they share. On this basis, it can be analyzed how common dominant frequencies and long-term trends are reflected in different industries, and how to use these patterns in a joint modeling setting to improve prediction performance. This modeling provides a quantitative basis for medium and long-term planning and capacity allocation decisions of the power system.
[0093] The formula for calculating the spectral complexity index is:
[0094] ,
[0095] where represents the amplitude spectrum value arranged in descending order, is the maximum amplitude; this index is used to quantify the complexity of the periodic structure, and when Approaching 1 indicates that the spectrum is dominated by a small number of frequencies, and the periodic structure is relatively simple. This index is used to assist in determining the number of dominant frequencies .
[0096] The construction method of the adaptive basis function library includes:
[0097] Based on the amplitude spectrum, a set of dominant frequencies is selected. Suppose that frequency indices are selected, denoted as , corresponding to the angular frequencies . The normalized frequency of the th component is defined as:
[0098] ,
[0099] For each selected normalized frequency , two real-valued basis functions are constructed:
[0100] ,
[0101] where is the time index, ; this results in a basis function library containing sinusoidal functions, each corresponding to a dominant oscillatory mode of the sequence. These basis functions have interpretability, as each pair of functions captures a periodic pattern characterized by a specific period and phase offset. Collecting all the basis functions into a matrix yields:
[0102] ,
[0103] This matrix serves as the design matrix for the subsequent sparse identification step. The number of basis functions can be adaptively determined: starting from a moderate upper limit and adjusting according to the spectral complexity index and cross-validated performance.
[0104] For sparse coefficient identification, given the basis matrix and the daily frozen vector , the load sequence can be approximated as a linear combination of basis functions:
[0105] ,
[0106] where is the coefficient associated with the th basis function. In matrix form, the model can be written as:
[0107] ,
[0108] where is the coefficient vector, is the residual term.
[0109] To avoid overfitting and enhance interpretability, the model promotes sparsity in by introducing norm regularization. The optimization objective function of the proposed sparse regression model is:
[0110] ,
[0111] where is the Euclidean norm, is the norm, is the regularization parameter that controls the degree of sparsity. This step is solved using the LASSO algorithm, which shrinks most coefficients to zero, leaving only the dominant periodic components.
[0112] From the perspective of system identification, the non-zero elements reveal which periodic components are essential for reconstructing the observed load series. Their amplitudes determine the contribution of each component to the overall load curve, while the corresponding frequencies provide interpretable information about daily, weekly, or seasonal periodicity.
[0113] For long-term load extrapolation forecasting, the sparse coefficient vector is obtained by solving the above equation, which can be used to generate long-term predictions. For any future date index , the basis functions in equation (9) can be evaluated at by simple substitution. The predicted value for the th day is given by:
[0114] ,
[0115] where is the basis function evaluation vector at . Since are sinusoidal functions with fixed frequencies, the prediction naturally extrapolates the periodic patterns identified from historical data to the future.
[0116] The method further includes an extension step for multi-industry joint modeling:
[0117] Assuming industry has representative enterprises, with daily frozen loads where and For each day define the industry-level load vector:
[0118] ,
[0119] Stacking these vectors over time yields a matrix:
[0120] ,
[0121] If all firms within an industry use the same basis matrix in equation (10), the multivariate model can be written as:
[0122] ,
[0123] where is a coefficient matrix whose columns correspond to different firms.
[0124] To promote sparsity across the entire industry and reveal shared periodic structures, group or element-wise regularization can be introduced. By penalizing the element-wise norm:
[0125] ,
[0126] where is the Frobenius norm, denotes the sum of the absolute values of all entries in . This approach allows different firms within the same industry to share a library of basis functions while retaining their own characteristic coefficients, enabling joint forecasting across upstream and downstream industry chains.
[0127] Industry-level forecasts are obtained in a similar fashion. For future days evaluate the basis vectors and compute the industry-level load forecast as:
[0128] ,
[0129] The approach also includes a step to introduce auxiliary features:
[0130] In practical applications, auxiliary features, environmental variables, and aggregated production indicators, among other auxiliary features, can also be available and informative. These features can be incorporated into the model in a simple and transparent manner by augmenting the basis matrix .
[0131] Let each day There is an auxiliary feature vector where is the number of auxiliary variables. The sequence can be collected into a feature matrix:
[0132] ,
[0133] The augmented design matrix can be constructed as:
[0134] ,
[0135] The sparse regression problem (13) is replaced by:
[0136] ,
[0137] where includes both periodic basis coefficients and auxiliary variable coefficients. The penalty term continues to promote sparsity, so only auxiliary variables that are meaningful for load sequence reconstruction and prediction will be retained.
[0138] For industry-level modeling, the same extension idea applies, only replace with in equation (18). This design enables upstream and downstream industries to share a common periodic basis while capturing specific industry sensitivities to calendar, environmental, or production-related features through additional sparse coefficients.
[0139] Joint modeling of upstream and downstream industries includes:
[0140] To extend the sparse identification adaptive basis function method from a single industry to joint modeling of multiple upstream and downstream industries, the industry-level matrix organization introduced earlier can be organized into a block structure. Suppose a set of industries are selected for joint analysis, which includes both upstream and downstream industries of interest. For each , suppose representative firms have been selected, and their daily frozen sequences have been arranged into a matrix .
[0141] A direct joint representation can be obtained by concatenating the industry matrices along the column dimension:
[0142] ,
[0143] where , enumerates industries, and:
[0144] ,
[0145] is the total number of representative firms in all considered industries.
[0146] Using the same basis matrix (if auxiliary features are included, then ), the joint multivariate model can be written as:
[0147] ,
[0148] where is the coefficient matrix to be identified. The sparse optimization problem is:
[0149] ,
[0150] or, if auxiliary features are included, as a corresponding variant using .
[0151] This joint modeling form allows the upstream and downstream industries to share a common periodic basis on a daily scale, reflecting global periodicities such as the workweek and seasonality, while maintaining unique coefficient features for each firm and industry. By inspecting the resulting coefficient matrix, one can identify periodic components that are common to many industries, as well as components that are specific to certain industries or firms. Furthermore, predictions for all industries can be generated in a unified manner by applying the extrapolation formula (19) to each column of .
[0152] When sparse identification of adaptive basis functions is applied to multiple industries simultaneously, a consistent evaluation protocol is needed to evaluate performance at different aggregation levels. At the firm level, standard error indicators such as Mean Absolute Error (MAE) or Mean Absolute Percentage Error (MAPE) can be computed individually for each time series. At the industry level, these errors can be aggregated across firms within the same industry to form industry-level performance indicators, such as the mean or median MAPE. For joint modeling, it is often useful to report the error for a specific industry and the global error for all firms in . Let denote the industry-level error indicator for industry , and let denote the global error indicator computed over .
[0153] Experimental description and pre-processing, in particular:
[0154] To build the power consumption prediction model for power utility customers, legitimate data was collected from relevant industries. The dataset contains power consumption data of several users in three different power consumption fields: steel, photovoltaics (PV), and chemical industry, at 15-minute intervals. The data for each field is stored in a separate table file, with the structure shown in Table 1.
[0155] Table 1: Example of data storage format
[0156]
[0157] Each row in the dataset records the power consumption of a user in a certain industry at 96 observation points in that day. The power consumption records of most industrial users span one year, usually from July 1, 2023 to June 30, 2024.
[0158] Before building the model, missing values in the data must be handled. Due to sensor failure, transmission interruption, human error, or external interference, industrial user power consumption time series data inevitably contains missing values. If not handled properly, these missing values can undermine the stability and accuracy of the model. Simply deleting records containing missing values not only results in the loss of valuable historical data, but also disrupts the time continuity of the sequence, which may cause the model prediction to deviate. Therefore, proper imputation of missing values is crucial to improve the prediction accuracy and robustness of the model. Considering the characteristics of industrial data, this paper adopts a hybrid imputation method. First, the K-Nearest Neighbors (KNN) mean imputation method using 8 adjacent points is applied to fill in the missing values. If the data point is still Not a Number (NaN) after KNN imputation, a forward fill method is used for secondary imputation. If the element is the first value in the sequence and is still NaN after the previous step, a global mean is finally applied for imputation.
[0159] Daily frozen values represent the total daily power consumption, which is calculated by summing the 96 observation points from 00:00 to 24:00. In this study, columnar summation is performed using the pandas library in Python. The result is added as a new column named "Daily Frozen" to the dataset, thus recording the total daily power consumption of each industrial user.
[0160] Model input and evaluation indicators, specifically:
[0161] To prevent data leakage, the dataset will be split into a training set and a test set. The training set will be used to build and train the predictive model, while the test set will be used to evaluate the performance of the model. This division ensures that the model's performance on unseen data can be effectively validated, thereby guaranteeing the model's generalization ability and prediction accuracy.
[0162] The "daily freeze value" is used as a one-dimensional feature of the time series to build a univariate time series prediction model. This task selects the daily freeze value data with continuous one-year records (July 1, 2023 to June 30, 2024). The daily freeze value data of the first 11 months is used as the training set input, while the data of the last month is used as the test set for the final model performance evaluation.
[0163] This experiment uses MAPE as the main evaluation indicator. The formula definition of MAPE is:
[0164] ,
[0165] where is the actual value, is the predicted value, is the number of observations. MAPE is particularly suitable for industrial and utility power consumption prediction because it can scale the prediction error to the demand level, similar to MAE.
[0166] Compared with MAE and other scale-dependent indicators, MAPE is particularly suitable for power consumption prediction in industrial and utility scenarios. The main value of MAPE in industry lies in its scale independence and interpretability. The power consumption of different industrial sectors (such as cement and photovoltaic) and individuals can differ by several orders of magnitude. MAE reports errors in the original units (such as kWh), which makes it impossible to directly compare the prediction performance between large steel plants and smaller new energy facilities. On the contrary, MAPE expresses errors in percentage form, providing a normalized accuracy measure that allows for fair comparison of model performance across heterogeneous groups, which is crucial for benchmarking in the utility sector. In addition, power utility companies and system operators often manage risk and assess performance based on relative errors. For a plant that consumes 100,000 kWh per day, an absolute error of 1,000 kWh may be insignificant (1% error), but for a user who consumes only 5,000 kWh per day, this is very significant (20% error). MAPE directly translates into percentage, providing immediate industrial context, which helps decision-makers such as grid planners and energy traders understand the relative impact of prediction errors, thereby better aligning with key performance indicators in the industry.
[0167] The prediction results and comparative analysis include:
[0168] The proposed SIABF model obtains power consumption predictions and its performance is compared with several established benchmark models. To comprehensively assess the robustness and prediction accuracy of the SIABF framework, the following models are chosen as comparison benchmarks. These models represent the most advanced and theoretically sound benchmark methods widely adopted in the current industrial power forecasting field. The comparison set includes statistical models like Autoregressive Integrated Moving Average (ARIMA), which is the baseline benchmark for linear time series analysis. It also incorporates machine learning and system identification models, such as Sparse Bayesian Learning (SBL) and Sparse Identification of Nonlinear Dynamics (SINDy), which are valued for their potential in providing sparse and interpretable representations of system dynamics. Finally, advanced deep learning models are included, specifically Long Short-Term Memory (LSTM) networks, Gated Recurrent Unit (GRU) networks, and Residual Network (ResNet), in order to benchmark the capabilities of SIABF against those methods renowned for capturing complex, long-term nonlinear temporal dependencies in modern energy consumption time series.
[0169] The performance of the SIABF model is systematically compared with the benchmark models in all three industrial sectors using the MAPE metric on the specified test set (data from the last month). The comparison results are summarized in Table 2, demonstrating the superior ability of the SIABF model in capturing the complex, potentially physically informative temporal dependencies in energy consumption data.
[0170] Table 2: Representative MAPE comparison in three industries:
[0171] Model Steel Photovoltaic Chemical SIABF 7.029% 4.502% 9.022% ARIMA 9.709% 22.947% 8.420% SBL 11.483% 15.303% 14.143% SINDy 55.953% 33.908% 64.315% LSTM 9.734% 18.988% 16.742% GRU 9.919% 15.132% 18.570% ResNet 10.640% 27.771% 19.064%
[0172] To gain a deeper understanding of the model's behavior, representative prediction results are specifically selected to highlight characteristic performance differences. This includes high-load industries (e.g., steel industry) to assess the model's robustness under high variance and inertia, as well as volatile users (e.g., photovoltaic or new energy industry) to evaluate the model's ability to handle irregular and rapidly changing consumption patterns. Visualizations depicting daily frozen value curves are presented, comparing the predicted daily frozen values with the actual values for selected cases, aiming to qualitatively demonstrate the superior accuracy and stability of SIABF predictions compared to leading benchmark models.
[0173] As Figure 2 shown by the preliminary comparison between SIABF and SBL, SIABF highlights the enhanced stability of the prediction and a tighter fit. While SBL exhibits significant bias, especially around the consumption peaks, SIABF maintains a more reliable trajectory. Moreover, ARIMA and SBL models fail to accurately learn the overall trend, struggling particularly in coping with data volatility, exhibiting significant phase lags and amplitude errors.
[0174] As Figure 3 shown by the comparison between system identification method SINDy and LSTM, the challenges faced by both methods are evident. SINDy exhibits high variance and significant lags, failing to effectively capture the time series structure. In contrast, while LSTM exhibits better responsiveness, it still struggles in accurately predicting the sharp rises and falls in the industrial load curve.
[0175] As Figure 4 shown, SIABF effectively tracks the true consumption curve, even successfully predicting the observed sharp drop and rebound around June 22. In contrast, traditional statistical and sparse learning methods struggle significantly. ARIMA and SBL models fail to accurately learn the overall trend and struggle particularly in coping with data volatility, exhibiting significant phase lags and amplitude errors. This result confirms that by incorporating implicit structural information, SIABF is highly effective in minimizing prediction errors and provides the most accurate and reliable prediction for the complex operational demands of the steel industry.
[0176] Due to high volatility and dependence on external factors such as weather, photovoltaic (PV) industry predictions exhibit unique characteristics. The following detailedly showcases the comparative prediction results for a representative PV user.
[0177] As Figure 5 shown, the first set of comparisons, the proposed SIABF model and the sparse benchmark SBL model both struggle to cope with the typical sharp and rapid fluctuations of the PV load curve. However, SIABF exhibits a significant advantage, with its physical information structure enabling it to maintain a trajectory closer to the actual consumption, while SBL displays significant, lagging bias, failing to quickly adapt to sudden changes in power generation. ARIMA completely fails, exhibiting a smooth, almost static straight line, missing all temporal dependencies, resulting in a massive MAPE of 22.947%. SINDy performs even worse, with a MAPE of 33.908%, producing unstable and severely misaligned predictions.
[0178] The performance comparison between SINDy and the deep learning model LSTM is shown in Figure 6SINDy struggles to cope with the high-frequency noise inherent in photovoltaic data, leading to very inaccurate and unstable predictions. LSTM, with its recurrent structure, can better capture overall trends, but still suffers from the common limitation of over-smoothing extreme peaks and troughs when dealing with non-stationary renewable energy data.
[0179] The prediction performance is crucial for the photovoltaic industry, which is highly sensitive to external factors, leading to highly volatile load curves. As shown in Figure 7 , the superior ability of the SIABF model in this extremely challenging environment is confirmed. Visually, the SIABF model consistently maintains excellent fit with the true electricity consumption curve, successfully capturing both the overall trend and the dramatic daily fluctuations in photovoltaic operation characteristics. In sharp contrast, the traditional sparse model performs abysmally.
[0180] The electricity load curve of the chemical industry typically has high inertia and strong seasonality, reflecting its continuous production operations. As shown in the figure, a representative chemical industry's comparative prediction results. Analysis Figure 8 in the preliminary comparison between the SIABF model and the SBL sparse learning baseline. Although the electricity consumption curve of the chemical industry is generally smoother than that of the steel or photovoltaic industry, SBL still shows significant deviation from the actual data, especially during the transition phase. In contrast, the SIABF model utilizes its underlying structure to provide predictions that closely match the actual values, demonstrating excellent tracking ability. In addition, the numerical results in the table show that the ARIMA model has the lowest MAPE of 8.420%, slightly better than the SIABF. Specifically, the two key sharp electricity consumption drops (abrupt changes) shown in the figure are completely missed by the ARIMA prediction. This indicates that although ARIMA is slightly better in average fit due to linear stability, it is very unsuitable for predicting key nonlinear operation events. On the contrary, both SBL and SINDy models show significant deviation from the actual data, confirming their limited suitability in dealing with the nonlinear dynamics of industrial processes.
[0181] As shown in Figure 9 , SINDy and LSTM. The SINDy framework faces great difficulty in modeling the complex, unobserved dependencies in the chemical process, leading to highly amplified errors and poor temporal alignment. LSTM, benefiting from its architecture designed specifically for sequence data, produces more reasonable predictions, effectively capturing the overall low-frequency trend in electricity consumption patterns.
[0182] The electricity consumption data of the chemical industry represents a relatively stable, high-inertia, and strongly potentially periodic load. As shown in Figure 10As shown, the SIABF model reveals the intricate performance of the field. It is clear that the SIABF model provides an excellent fit to the true curve, accurately tracking the trajectory of electricity consumption throughout the test period.
[0183] Empirical results on representative enterprises from the steel, photovoltaic, and chemical industries show that the proposed sparse system identification framework (instantiated as the SIABF model) can provide competitive or superior prediction accuracy compared to classical statistical models and deep learning benchmarks. In high volatility scenarios (such as photovoltaic users), whose load curves exhibit strong meteorological and operational driven fluctuations, the adaptive periodic basis functions and the Bayesian state-space formulation enable the model to capture both fast-changing and long-term seasonal patterns, achieving significantly lower MAPE than competing methods.
[0184] The disclosed sparse system identification framework with adaptive periodic basis functions is used for long-term load forecasting of upstream and downstream industrial chains. For daily frozen power time series, the proposed SIABF method combines frequency domain analysis, spectral complexity index for adaptive basis selection, and Regularized sparse regression to construct interpretable models that explicitly decompose industrial loads into a small number of dominant daily, weekly, and seasonal components. The framework naturally extends from single-enterprise sequences to multivariate industry-level models and joint multi-industry formulations, thus providing a unified and scalable approach to modeling heterogeneous industrial users in new power systems.
[0185] The above describes the basic principles of the disclosure in combination with specific embodiments, but it should be noted that the advantages, advantages, effects, etc. mentioned in the disclosure are only examples and are not limiting, and these advantages, advantages, effects, etc. cannot be considered as each embodiment of the disclosure must have. In addition, the above disclosed specific details are only for the purpose of example and for the purpose of understanding, and the above details do not limit the disclosure to the above specific details.
[0186] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0187] The above description has been given for the purpose of illustration and description. Furthermore, this description does not intend to limit embodiments of the disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those of skill in the art will recognize certain modifications, permutations, additions, sub-combinations, and sub-combinations thereof.
Claims
1. An adaptive Bayesian identification method for long-term forecasting of electricity load in an industrial chain, characterized in that, include: S1: Obtain daily high-frequency active power data of target industry users and aggregate the active power data into a daily frozen power time series; S2: Perform a discrete Fourier transform on the daily frozen power time series to obtain the amplitude spectrum and calculate the spectral complexity index to determine the number of dominant frequencies and the specific frequency components. S3: Based on the selected dominant frequencies, construct an adaptive basis function library containing sine and cosine functions; S4: Establish a sparse regression model with regularization, represent the daily frozen power time series as a linear combination of basis functions in the adaptive basis function library, and obtain the sparse coefficient vector by solving the optimization problem, thereby screening out the key periodic components. S5: Using the identified sparse coefficient vector and adaptive basis function library, calculate the daily frozen power forecast for future time steps through extrapolation of time index; In step S2, the formula for calculating the spectral complexity index is: , in, and All represent amplitude spectrum values arranged in descending order. This is the maximum amplitude; the spectral complexity index is used to quantify the complexity of periodic structures, and it is also used to help determine the number of dominant frequencies. ; In step S3, the method for constructing the adaptive basis function library includes: Based on amplitude spectrum selection The dominant frequency, denoted as Its normalized frequency is defined as ; For each normalized frequency Construct two real-valued basis functions: ; Collect all basis functions into the design matrix middle; in, The total length of the daily frozen power time series; The formula for establishing a sparse regression model with regularization, which represents the daily frozen electricity time series as a linear combination of basis functions in the adaptive basis function library, is as follows: , in Is with the first The coefficients related to the basis functions It is the first Each basis function in the input The function value at that location; In step S4, the objective function of the sparse regression model is: , in, It is the Euclidean norm. yes Norm, It is a regularization parameter that controls the degree of sparsity. It is a coefficient vector. The vector was frozen on that day. It is the set of real numbers.
2. The adaptive Bayesian identification method for long-term forecasting of electricity load in the industrial chain according to claim 1, characterized in that, In step S1, daily high-frequency active power data of the target industry users is obtained, and the active power data is aggregated into a daily frozen power time series, including constructing a formula for daily frozen power consumption, wherein the formula is: ; in It is the average active power after 15 minutes of preprocessing.
3. The adaptive Bayesian identification method for long-term forecasting of electricity load in the industrial chain according to claim 1, characterized in that, In step S5, using the identified sparse coefficient vector and adaptive basis function library, the daily frozen power forecast value for future time steps is calculated through extrapolation of the time index, including: For any future date index , No. The predicted value for the day is given by the following formula: ,in Is The basis function evaluation vector at that point. Is with the first Estimates of the coefficients related to the basis functions. It is a predicted value. It is an estimate of the coefficient vector.
4. The adaptive Bayesian identification method for long-term forecasting of electricity load in the industrial chain according to claim 3, characterized in that, It also includes a multi-industry joint modeling step, which includes: For the industry Establish a multivariate model , Rewrite the multivariate model as follows: ,in It is a coefficient matrix; By solving the problem with Regularization optimization problem To identify sparse coefficient matrices; Industrial load forecasting calculate; in It is the Frobenius norm. express The sum of the absolute values of all entries in the list. It is a regularization parameter. It is an estimate of the coefficient matrix. It is a real matrix.
5. The adaptive Bayesian identification method for long-term forecasting of electricity load in the industrial chain according to claim 4, characterized in that, It also includes the step of introducing auxiliary features, which include: Collect auxiliary feature sequences to form a feature matrix; Constructing augmented design matrices based on eigenma matrices; Replace the sparse regression problem with ,in An augmented coefficient vector that includes both periodic base coefficients and auxiliary variable coefficients. To augment the design matrix, It is the regularization parameter.
6. The adaptive Bayesian identification method for long-term forecasting of electricity load in the industrial chain according to claim 5, characterized in that, The multi-industry joint modeling includes: The industry-level matrices of the selected set of industries are concatenated column-wise to form a joint matrix; Establish a joint model based on the joint matrix; By solving Joint sparsity discrimination is performed, where, For the joint model, It is a coefficient matrix. It is a regularization parameter; Based on the solution results, a unified prediction is generated for all industries participating in the joint modeling.
Citation Information
Patent Citations
A power load prediction method based on a Bayesian regularization neural network
CN109409614A
Single-module follow-up control method and system
CN118818992A