Electric quantity prediction method based on small sample data
Through the combination of four interval dynamic discrete binning and multi-ridge regression model, combined with the dynamic parameter adjustment mechanism, the overfitting and multi-step prediction error accumulation problems in small sample data prediction are solved, high-precision and stable power prediction are achieved, and dynamic adaptability is achieved.
Patent Information
- Application Number
- CN202510507085.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In power systems, the existing small sample data prediction methods have problems such as overfitting, multi-step prediction error accumulation, poor model adaptability, and failure to effectively combine the periodicity and volatility of power data, making it difficult to achieve high-precision and stable multi-step prediction under limited data.
The power data is preprocessed by four interval dynamic discrete binning method, and the data characteristic characterization method is adaptively adjusted based on the periodicity and volatility of the data. Based on the ridge regression model, multiple sub-models with differentiated hyperparameter configuration are constructed, the weight coefficient is optimized through the convex optimization algorithm with constraints, and a dynamic parameter adjustment mechanism is introduced to dynamically adjust the hyperparameters of the model through the feedback relationship between the average absolute percentage error and the threshold.
It significantly improves the accuracy and stability of small sample data prediction, can achieve multi-step prediction under limited data, has dynamic adaptability, and is suitable for complex scenarios of power systems.
Smart Images

Figure CN120031264A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of load-side data prediction in power systems, and specifically relates to a method for predicting power based on small sample data, especially for scenarios with limited historical data (data volume ≤ 100). Background Art
[0002] In power system planning and dispatching, power forecasting is one of the core technologies to ensure the safe operation of the power grid and optimize energy distribution. Traditional forecasting methods (ARIMA, LSTM, Transformer, etc.) rely on a large amount of historical data (hundreds to tens of thousands of samples) to capture the complex characteristics of time series. However, in practical applications, we often face small sample data scenarios: emerging power markets or newly built power generation facilities lack long-term historical data; sudden events such as extreme weather and equipment failures lead to interruptions in data continuity; daily data predicts insufficient coverage of the future time of monthly data, etc.
[0003] Existing small sample data prediction scenarios have the following limitations: traditional machine learning models such as neural networks are prone to overfitting problems when the amount of data is insufficient, resulting in inaccurate prediction results; recursive prediction methods such as the ARIMA model have exponentially amplified errors when the prediction step increases, resulting in the accumulation of multi-step prediction errors, making it difficult to meet the prediction accuracy requirements of more than 6 steps; static model parameters such as regularization coefficients use fixed hyperparameters and cannot dynamically adjust parameters according to the latest data, resulting in poor model adaptability and insufficient generalization ability.
[0004] In recent years, researchers have tried to improve the prediction effect of small samples through methods such as transfer learning and data enhancement, but the following problems still exist: dependence on external data. Transfer learning requires data support from similar fields, which is difficult to obtain in power scenarios; in terms of data generation, such as the generative adversarial network (GAN) model, the data generated under the condition of a small amount of data is of poor quality, and a large number of examples are required, and the real-time performance is poor; the existing methods do not effectively combine the periodicity and volatility of power data.
[0005] Therefore, there is an urgent need for a small sample prediction method designed for the characteristics of the power system, which can achieve high-precision and stable multi-step prediction with limited data and has dynamic adaptability. Summary of the invention
[0006] The present invention provides a method for predicting electricity based on small sample data. First, the historical electricity data is preprocessed and the data is dynamically discretized into four intervals for binning. Combining the periodicity and volatility of the electricity data, the data feature representation method is adaptively adjusted to improve the learning ability of the prediction model. Secondly, based on the ridge regression basic model, a A ridge regression prediction sub-model with differentiated hyperparameter configurations is used to optimize the The weight coefficients of the sub-models are optimized and solved, and the loss function of the multi-ridge regression model is reset. The optimized weight coefficients are put into the model as the preliminary prediction model, and the The sub-models are re-set to minimize the prediction mean square error. Finally, a dynamic parameter adjustment mechanism is added to the model. Based on the mean absolute percentage error ( ) and the set threshold to dynamically adjust the model’s hyperparameters, improve the model’s generalization and robustness, and ensure that the prediction error meets the accuracy requirements.
[0007] A power prediction method based on small sample data includes:
[0008] Collect historical electricity data with the smallest time unit of month to construct the initial data set, where the total sample size is . The raw data is preprocessed, including missing value interpolation, outlier correction and normalization operations, to eliminate the interference of dimensional differences on model training. Multi-dimensional time series analysis is carried out on the preprocessed electricity data, including: overall trend analysis, periodicity detection and peak and valley feature extraction. Overall trend analysis, extract long-term trend items through time series decomposition (STL decomposition), draw trend change graphs, and effectively identify the growth or decay rules of electricity data; periodicity detection, calculate the spectrum energy based on Fourier transform (FFT), combined with autocorrelation function (ACF) analysis, to determine whether the data has significant periodicity, the overall period length of the data and the local characteristic period length.
[0009] Based on the above analysis, the present invention proposes a four-interval dynamic discretization binning method. It breaks through the limitations of traditional equal-width and equal-frequency divisions and adaptively determines the initial interval boundaries based on the data distribution characteristics. Different from the existing methods that mechanically divide intervals, the present invention combines the statistical characteristics of the data and adds 25% quantiles each through cumulative distribution calculations to dynamically generate four intervals of low power, medium-low power, medium-high power, and high power, thereby more accurately reflecting the essential distribution of the data. The rules are as follows: Define the rule function for dynamically demarcating the initial interval boundaries: Where: Indicates historical power data Quantile; Indicates the interval boundary; Interval bounding function: Where: Indicates low battery interval; Indicates the medium-low power range and the medium-high power range; Indicates high power interval; is the minimum value in the overall data; is the maximum value in the overall data.
[0010] The dynamic definition rules of interval boundaries of the present invention are designed for the periodicity and volatility unique to power data, and aim to improve the feature characterization capabilities in small sample scenarios through intelligent binning strategies. This method uses quantile binning as the basic framework to adapt to the time-varying characteristics of electricity. For data with significant differences in peak and valley periods or drastic seasonal changes, which can easily lead to loss of feature information, the present invention introduces a dynamic adjustment mechanism: in response to changes in load curves caused by summer cooling and winter heating, a seasonal adaptation strategy is adopted to dynamically adjust the binning rules according to the peak and valley periods and seasonal characteristics of the data, and the new boundaries are gradually transitioned according to the weighted average to achieve dynamic adjustment of the interval. The dynamic adjustment formula is specifically: Where: is the historical electricity quantile of the new dynamic range; is the historical electricity quantile initially set; The current historical electricity data quantile; is the smoothing coefficient.
[0011] The original data is mapped to the discretization interval according to the above rules. After the four-interval dynamic discretization binning, the discretization label with the timestamp is generated. The categorical features of the data are used as input features for model training to enhance the ability to characterize the power fluctuation pattern.
[0012] The power prediction method based on small sample data is constructed The sub-models are ridge regression models with different regularization parameters. The different characteristic patterns of the data are captured through model diversity, and the constrained convex optimization algorithm is combined to calculate the optimal weights and adaptive mechanism to improve the stability and accuracy of the prediction.
[0013] The power prediction method based on small sample data is aimed at Different regularization parameters are used to set the loss function of the multi-ridge regression model. The specific formula is: Where: is the set loss function; is the sample size; is the number of features; It is The data Features represents the intercept term; It is The regression coefficient of each feature; is the regularization parameter of different models; For the training set The true value of the data.
[0014] By adjusting the regularization parameter , to balance model complexity and generalization ability. Different regularization parameters The model can capture the fluctuation characteristics of different scales and directions in the data, and show complementarity in prediction: low regularization parameter Models such as , the model has strong fitting ability and can capture short-term fluctuations in data, but it is easy to produce overfitting problems in the model; high regularization parameters Models such as It can improve the generalization ability of the model and smooth out noise interference, but ignores detailed features, resulting in poor prediction accuracy. In order to cover the diverse characteristics of the data and ensure that regularization strengths of different orders of magnitude are covered, The regularization parameter of the model The values are evenly distributed symmetrically in a specific interval, and the logarithmic interval coefficient is set. The specific rules are as follows: Where: and According to the number of models Determined, such as ,but , ; is the total number of models. To avoid computational redundancy, the value is limited to 5-8; for The first model in the is the logarithmic spacing coefficient.
[0015] In the method of electricity prediction based on small sample data, the weight allocation of the model is a weight allocation mechanism based on constraint optimization to achieve dynamic allocation through data modeling and numerical optimization. The specific goals include: minimizing the prediction error. By combining the optimal weights, the mean square error after multi-model fusion is reduced; a new weight normalization constraint is set, that is, the sum of the weights of all models is 1. , avoid excessive bias towards a single model; set weight boundary constraints , prevent extreme weight allocation from causing prediction fluctuations and improve model stability; combine gradient optimization and numerical solution methods to improve the convergence speed of the optimization process and ensure good computational efficiency under small sample data conditions; The objective function of the constrained optimization algorithm is to minimize the prediction mean square error using numerical optimization tools. The constrained optimization objective function is: The constraints are: Where: is the weight vector of the model; For the The weight of each model; For the Model The predicted value of the step; For the The true value of the step; is the total number of models; is the prediction step length; The sequential least squares programming algorithm is used to solve the optimal weight. The initial weight is defined as: Where: is the initial weight; is the total number of models; After iteratively solving the optimal solution under the current conditions, calculate the objective function gradient: . Set a fixed learning rate , weight adjustment . Update the weight vector according to the objective function gradient calculation result , so that the weights gradually approach the optimal solution that satisfies the constraints. Indicates passing The weight after iterations; Represents the optimal solution under current conditions; according to The basis model is reset to minimize the prediction mean square error, and its function is: Where: To minimize the prediction mean square error; For the The weight of each model; is the prediction step length; is the total number of models; For the The model is The predicted value of the step; For the The true value of the step.
[0016] The present invention proposes an adaptive hyperparameter adjustment mechanism based on prediction error, which dynamically adjusts the regularization parameter by real-time monitoring of the prediction error. , balancing the model's fitting ability and generalization performance. The mean absolute percentage error is used as the adjustment trigger indicator, and the mean absolute percentage error calculation formula is: Where: is the mean absolute percentage error; is the prediction step length; is the actual value; is the predicted value.
[0017] Set the error threshold to ,when , triggers parameter adjustment. The parameter adjustment rule is: when the error exceeds the threshold, the regularization parameter is proportionally attenuated, and the attenuation ratio range is 0.8 to 0.95; the regularization attenuation rule is: , Where: is the attenuation ratio if and only if , trigger the effect; is the new regularization parameter after solving; is the original initial regularization parameter; Dynamic adjustment based on the error margin , the specific adjustment formula is: Add a border protection mechanism. If the preset range is exceeded, the boundaries of the interval are reset.
[0018] The electricity forecasting method based on small sample data further generates rolling statistical features in the training set, including the mean and standard deviation of the data in the window period, to reflect the short-term fluctuation law of the time series; the window length matches the seasonal cycle of the electricity data to strengthen the embedded relationship between the time series and the electricity data; the standard deviation is used to quantify the data fluctuation amplitude and distinguish normal fluctuations from abnormal noise. The rolling mean and standard deviation calculation formula is: Where: is the window length; For training set In the window Monthly rolling average; For training set In the window Monthly rolling standard deviation; For the training set The true value of the data.
[0019] Add dynamic adjustment mechanism, if the training is concentrated Standard deviation within window If the sudden increase or decrease exceeds 3 times of the previous data and the next data, the window length is adjusted, the main period of the data is identified through Fourier transform or autocorrelation analysis, and the window length is automatically calibrated. , The value range is The data standard deviation, as a subsidiary condition of the adaptive mechanism, is fed back to the regularization parameter dynamic adjustment mechanism and participates in the dynamic adjustment of the regularization parameter.
[0020] The power prediction method based on small sample data, the preset threshold is , and the prediction error of the next M data points is evaluated by the mean absolute percentage. Greater than the preset threshold , the regularization adjustment mechanism, weight adjustment mechanism and dynamic window adjustment mechanism are triggered simultaneously, and the three mechanisms coordinately adjust the model parameters until the conditions are met.
[0021] Compared with the prior art, the present invention has the following advantages: Traditional prediction models, such as LSTM and ARIMA models, rely on a large amount of historical data to ensure prediction accuracy. The present invention effectively alleviates the problem of small samples ( ) scenario, it can still maintain high-precision prediction capabilities when samples are severely insufficient, solving the high dependence of existing technologies on data scale.
[0022] Existing recursive prediction methods are prone to step-by-step error amplification in multi-step predictions. The present invention significantly suppresses the error accumulation effect through the collaborative design of rolling statistical feature extraction and four-interval dynamic discretization binning, combined with a dynamic weight constraint mechanism, to ensure the consistency and reliability of future multi-step predictions.
[0023] Traditional models mostly use fixed hyperparameters, which are difficult to adapt to the non-stationary characteristics of power data. The present invention introduces an adaptive parameter adjustment mechanism based on error feedback. By real-time monitoring of errors and dynamically attenuating regularization parameters, the model can quickly respond to complex scenarios such as load mutations and seasonal fluctuations, and has stronger generalization than static models.
[0024] Through the closed-loop linkage of four-interval dynamic discretization binning, multi-model fusion, weight optimization, adaptive adjustment and other links, a multi-level collaborative optimization system is formed, breaking through the limitations of single technology improvements, and achieving systematic improvement in prediction accuracy, stability, real-time and other dimensions. It has stronger comprehensive advantages than the isolated solutions of existing technologies.
[0025] These innovative designs enable the present invention to demonstrate greater robustness and industrial application value in small-sample electricity forecasting scenarios, providing new technical support for key applications such as power grid dispatching and new energy management. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The details and advantages of the present invention are further described in detail through specific implementation cases. The drawings in the specification are used to intuitively present the key technical points of the present invention and to assist in explaining its superiority. For ordinary technicians in this field, other relevant diagrams can be derived based on these drawings without creative work to further understand the present invention.
[0027] Attached Figure 1 , basic flow chart;
[0028] Attached Figure 2 , Four-interval dynamic discretization binning flow chart;
[0029] Attached Figure 3 , Flowchart of dynamic tuning of regularization parameters;
[0030] Attached Figure 4 , Extraction diagram of the change trend of the original data through STL;
[0031] Attached Figure 5 , Fourier transform spectrum energy diagram of original data;
[0032] Attached Figure 6 , the autocorrelation function diagram of the original data;
[0033] Attached Figure 7 , distribution diagram of peak and valley values within the period;
[0034] Attached Figure 8 , single original ridge regression true value predicted value map;
[0035] Attached Fig. 9 , ARIMA model true value predicted value chart;
[0036] Attached Fig.10 , the true value predicted value diagram of the model of the present invention. DETAILED DESCRIPTION
[0037] The invention is further described in detail through the accompanying drawings. The implementation cases are only used to more fully illustrate the specific technical content of the invention. The invention is not limited to the implementation cases shared, but can be applied to other fields.
[0038] like Figures 1 to 10 As shown, a method for power prediction based on small sample data includes the following steps:
[0039] The electricity consumption data with the smallest unit of month is collected, and the data volume is less than 100. The data is the monthly electricity consumption data of the whole society in a certain province (unit: 10,000 kWh), and the time range is from 2017 to 2024, with a total of 96 data. No additional features such as temperature, wind speed, and air pressure are added. Only the single variable of electricity consumption is used to analyze and predict the data. First, the data is preprocessed, the data is tested for outliers and missing values, and the data is normalized. The judgment of abnormal data is whether the fluctuation amount of the original data and the data on both sides exceeds , as the basis for judgment. If the number of outliers is less than 3, the outliers will not be processed; if the number of data is greater than 3, the outliers will be replaced by the mean of the left and right sides of the data. The same method is used for missing data. If the missing values are less than 3, they will not be processed. If the missing values are greater than three, the average of the left and right sides of the data will be used to fill them.
[0040] After the data preprocessing is completed, the present invention adopts a multi-stage analysis method to realize the periodic feature mining and dynamic binning of the electricity data: first, the data structure is decomposed into trend items, seasonal items and residual items using time series decomposition, and the long-term change trend and overall trend of the data are analyzed; on this basis, the fast Fourier transform is used to perform frequency domain analysis on the trended data, focusing on identifying local periodic fluctuations outside the annual cycle, and the power spectrum density detection finds that the data has significant 4-month local periodic fluctuations. In order to verify the reliability of the cycle length, the autocorrelation function analysis is further introduced, and the autocorrelation coefficients of different lag orders are calculated. It is confirmed that there is a significant peak at the lag of 4 months, and the valley sequence is regularly spaced, thereby confirming the significance of the local cycle length.
[0041] Based on the results of the above analysis, the data is dynamically discretized into four intervals and binned. The initial interval boundaries are calculated using quantiles (25%, 50%, 75%): According to these quantiles, four intervals are initially divided. For low battery range, For low to medium power range, For medium to high power range, This is the high power range.
[0042] The maximum and minimum values in the local cycle are used to determine whether the interval is reasonable, focusing on the data in summer and winter, which have strong seasonality. Determine whether the high power interval and low power interval can fully cover the power data of the two quarters. If the seasonal data exceeds the current interval boundary, the boundary of the interval is adjusted by weight, according to the formula: . The initial value is 0.7, and the boundaries of the final four intervals are determined by dynamically adjusting the interval boundaries.
[0043] Map the data to the discretization intervals according to the above rules, and generate discretization labels with timestamps after dynamic discretization of four intervals. The categorical characteristics of . Specifically expressed as:
[0044] All data are labeled and classified by the change of dynamic intervals. This method is designed for the label determination of each interval according to the maximum and minimum values of each interval under the condition of small sample data. The limited power data is mapped to a limited interval, short-term random fluctuations are filtered, the interference of single abnormal data on model feature learning is reduced or reduced, and the risk of model overfitting is reduced.
[0045] After completing the four-interval dynamic discretization binning process, the present invention further generates rolling statistical features to enhance the ability to extract periodic features. The specific implementation is as follows: Based on the inherent periodicity of power data, the sliding window size is set to , by calculating the training set In the window Average power data within the window of the month With standard deviation , quantifies the intensity of local load fluctuations. Reflects the average load level within the cycle, standard deviation Characterize the discrete degree of load, and the two work together to capture the stability and abnormality characteristics in periodic fluctuations. Add a dynamic adjustment mechanism. If the data standard deviation If the sudden increase or decrease exceeds 3 times of the previous data and the next data, the window length is adjusted, the main period of the data is identified through Fourier transform or autocorrelation analysis, and the window length is automatically calibrated. , The value range is 2-24. It should be noted that the parameters involved in the above formula Indicates the test set Step index of step prediction and training set index Strictly distinguish between application scenarios. This design significantly improves the model's ability to learn periodic evolution patterns through dynamic calculation of windowed statistics.
[0046] Based on the ridge regression model, a multi-ridge regression model is constructed, in which , that is, construct 5 ridge regression models. In order to cover the diversity of data characteristics and ensure that regularization strengths of different orders of magnitude are covered, The regularization parameter of the model The values are evenly distributed symmetrically within a specific interval, based on the logarithmic spacing coefficient , Calculate the baseline value of the regularization parameter for each model .
[0047] The constrained optimization algorithm is The ridge regression model assigns weights and the initial weights are obtained according to the initial weight formula , using the least squares planning algorithm, under the constraints of: minimizing the mean square error of the fused prediction results in the test set, and satisfying the weight sum of 1 and weight boundary constraints: each weight is between 0 and 1. Iteratively solve the weights. Set boundary condition checks in the algorithm to ensure that the results after each iteration are within the boundary conditions. After iteratively solving the optimal solution under the current conditions, calculate the objective function gradient: . Set a fixed learning rate , weight adjustment . Update the weight vector according to the objective function gradient calculation result , so that the weights gradually approach the optimal solution that satisfies the constraints. The weights after optimization are .
[0048] Based on the previous analysis, the data is divided into training set and test set. , 90 data in the data set are used as training sets, and the remaining 6 data are used as test sets. Discretized labels, rolling mean, and rolling mean square error are input into the model as auxiliary features. And set the number of lags. The number of lags set in this example is Based on the latest weights for the future months of data and set the threshold to The MAPE error is calculated for the prediction results. If the error result MAPE is greater than the set threshold, the regularization parameter attenuation mechanism is triggered. The regularization parameter attenuation formula is: , . Dynamically adjust according to the error limit , the specific adjustment formula is: . A boundary protection mechanism is also added to the program. If If it exceeds the preset range, the regularization parameters are reset. If the MAPE of the predicted data still exceeds the threshold under the optimal regularization parameters, the error information is fed back to the weight optimization module through the feedback mechanism to re-solve the optimal weight.
[0049] In the example verification process, the ARIMA model and the original ridge regression model are added for comparison. Only the power data is used. The training set data is 90, and the test set data is 6. When predicting the ARIMA model, no additional temperature, weather and other features are added; when training the original ridge regression model, only the original data is used, and the data after the four-interval dynamic discretization bins and data labels are not added. Only the original power data after normalization is used as the input feature to train the data. The training set of the data is 90 data, and the test set data is 6. When the original ridge regression model and the model of the present invention are compared and predicted, the parameters are all adopted , , ,in It represents the lag number; It represents the length of the rolling window. is the number of partitions. The other parameters of the original ridge regression model are all unadjusted data. The data prediction results are shown in the following table: Table 1 MSE and MAPE data of each model prediction Model MSE MAPE Single Ridge Regression Model 4.30 6.78% ARIMA Model 6.36 9.51% Model of the present invention 1.42 2.18%
[0050] From Table 1, it can be intuitively seen that the model of the present invention is optimal in both MSE and MAPE, and the model of the present invention improves the MAPE of the single ridge regression model by 67.85% under the condition of small sample data.
Claims
1. A method for predicting power consumption based on small sample data, characterized in that: The steps include: Firstly, the historical electricity data is preprocessed, and the preprocessed data is dynamically discretized into four intervals. In combination with the periodicity and volatility of electricity data, the data feature representation method is adaptively adjusted to improve the learning ability of the prediction model. Secondly, based on the ridge regression basic model, a A ridge regression prediction sub-model with differentiated hyperparameter configurations is used to optimize the The weight coefficients of the sub-models are optimized and solved, and the loss function of the multi-ridge regression model is reset. The optimized weight coefficients are put into the model as the preliminary prediction model, and the The sub-models are re-set to minimize the prediction mean square error; finally, a dynamic parameter adjustment mechanism is added to the model: based on the mean absolute percentage error ( ) and the set threshold to dynamically adjust the model's hyperparameters, improve the model's generalization and robustness, and ensure that the prediction error meets the accuracy requirements; through the ridge regression model under the optimal weight condition, the future ( ) months’ forecast values to get the electricity consumption forecast results.
2. The method for predicting power consumption based on small sample data according to claim 1, characterized in that: The four-interval dynamic discretization binning process adaptively determines the initial interval boundary based on the data distribution characteristics; combined with the statistical characteristics of the data, 25% quantiles are added through cumulative distribution calculation to dynamically generate four intervals: low power, medium-low power, medium-high power, and high power. The rule function for dynamic programming of the initial boundary is: Where: Indicates historical power data Quantile; Indicates the interval boundary; Based on the initial interval boundaries, a dynamic adjustment mechanism is introduced: in view of the load curve changes caused by summer cooling and winter heating, a seasonal adaptation strategy is adopted to dynamically adjust the binning rules according to the peak and valley periods and seasonal characteristics of the data. The new boundaries are gradually transitioned according to the weighted average to achieve dynamic adjustment of the interval; the dynamic adjustment formula is as follows: Where: is the historical electricity quantile of the new dynamic range; is the historical electricity quantile initially set; The current historical electricity data quantile; is the smoothing coefficient.
3. The method for predicting power consumption based on small sample data according to claim 1, characterized in that: The base models are ridge regression models with different regularization parameters. The loss function of the multi-ridge regression model is reset. The specific formula is: Where: is the set loss function; is the sample size; is the number of features; It is The data Features represents the intercept term; It is The regression coefficient of each feature; is the regularization parameter of different models; The training set The true value of the data.
4. The method for predicting power consumption based on small sample data according to claim 1, characterized in that: Convex optimization algorithm with constraints The weight coefficients of each sub-model are optimized and solved, and a weight allocation mechanism based on constrained optimization is realized through data modeling and numerical optimization. The constrained optimization objective function is: The constraints are: Where: is the weight vector of the model; For the The weight of each model; For the Model The predicted value of the step; For the The true value of the step; is the total number of models; is the prediction step length; The initial weights are set to: Where: is the initial weight; After iteratively solving the optimal solution under the current conditions, calculate the objective function gradient: ; Set a fixed learning rate , weight adjustment ; Update the weight vector according to the objective function gradient calculation result , so that the weight gradually approaches the optimal solution that satisfies the constraints; Indicates passing The weight after iterations; It represents the optimal solution under the current conditions.
5. The method for predicting power consumption based on small sample data according to claim 1, characterized in that: according to The basis model is reset to minimize the prediction mean square error, and its function is: Where: To minimize the prediction mean square error; For the The weight of each model; is the prediction step length; is the total number of models; For the The model is The predicted value of the step; For the The actual value of the step.
6. The method for predicting power consumption based on small sample data according to claim 1, characterized in that: The specific method of dynamically adjusting the model hyperparameters is: using the mean absolute percentage error ( ) as the adjustment trigger indicator, setting the error threshold to ,when When , the parameter adjustment is triggered, and the regularization parameter is decayed proportionally, with the decay ratio ranging from 0.8 to 0.95; the regularization decay rule is: , Where: is the attenuation ratio if and only if , triggers the effect; is the new regularization parameter after solving; is the original initial regularization parameter.
7. The method for predicting power consumption based on small sample data according to claim 1, characterized in that: Rolling statistical features are further generated in the training set, including the mean and standard deviation of the data within the window period. The length of the window is dynamically adjusted based on whether the rolling mean and standard deviation have changed suddenly. The window length matches the seasonal cycle of the power data and strengthens the embedded relationship between the time series and the power data. The rolling mean and standard deviation calculation formula is: Where: is the window length; For training set In the window Monthly rolling average; For training set In the window Monthly rolling standard deviation; The training set The true value of the data.
8. The method for predicting power consumption based on small sample data according to claim 1, characterized in that: The preset threshold is , and the prediction error of the next M data points is evaluated by the mean absolute percentage; when the model prediction results exceed 3 Greater than the preset threshold , the regularization adjustment mechanism, weight adjustment mechanism and dynamic window adjustment mechanism are triggered simultaneously, and the three mechanisms coordinately adjust the model parameters until the conditions are met.
Citation Information
Patent Citations
Electric quantity prediction system and prediction method thereof
CN113205223A
Power grid load prediction method and device based on XGBoost-LSTM
CN114862032A
Medium and long term power load prediction method and system based on improved seasonal ARIMA model
CN116995668A
Electric quantity prediction method and system based on combination of multiple electric quantity prediction models
CN117933457A
Short-term load prediction method and system based on LSTM-DNN hybrid neural network
CN118970924A
Cited By
Multivariable time delay nonlinear industrial process modeling method cooperatively driven by data and knowledge
CN120874581A