A power load rolling prediction method and system based on large model time series modeling
Patent Information
- Application Number
- CN202611008099.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-08
AI Technical Summary
[0007]本申请通过提供一种基于大模型时序建模的电力负荷滚动预测方法及系统,旨在解决现有电力负荷预测中窗口僵化无法响应电网结构突变、大模型全量更新算力开销大无法支撑高频滚动、以及协变量物理语义缺失导致极端工况失准的技术问题
本申请通过提供一种基于大模型时序建模的电力负荷滚动预测方法及系统,建立负荷波动和电网运行工况结合的双因子联合决策机制,在电网结构发生变化时,强制切换至高波动模式并缩短输入窗口,避免了预测模型因过度依赖陈旧历史数据而产生的误判,确保在电网拓扑改变、新能源出力剧变等极端场景下,预测结果仍能准确反映当前的物理运行特性。
Smart Images

Figure CN122532904B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power load forecasting technology, and in particular to a method and system for rolling power load forecasting based on large-scale time-series modeling. Background Technology
[0002] Electricity load forecasting is a core basis for grid dispatching, unit allocation, and electricity market transactions. With the construction of new power systems, the high proportion of renewable energy integration, frequent extreme weather events, and the widespread application of power electronic equipment, load characteristics have become highly random, volatile, and time-varying nonlinear, posing a serious challenge to traditional electricity load forecasting methods.
[0003] While existing power load forecasting methods based on large models (such as CN120184939A and CN121529526A) have made some progress, the following technical problems still urgently need to be solved:
[0004] Rigid window mechanism: Fixed-length time windows are commonly used. When the power grid experiences abnormal outages, important equipment is under maintenance, or there are sudden weather changes, the outdated historical data within the fixed window will dilute recent key features, resulting in a lag in model response and an inability to capture structural changes in load characteristics.
[0005] Addressing the bottleneck of computing power: General-purpose time series models (such as TimesFM and TimeGPT) typically have hundreds of millions of parameters; if existing technologies are to achieve rolling forecasts, they need to frequently perform full fine-tuning or incremental training on the models, resulting in huge computational overhead and making it difficult to meet the low-latency requirements of real-time grid dispatch.
[0006] Insufficient data quality and feature utilization: Data preprocessing often uses simple threshold filtering, which can easily lead to the accidental deletion of real power grid faults; covariates (meteorological and economic data) often use simple numerical embedding, ignoring the inherent mechanisms of meteorological physical quantities and socio-economic activities, resulting in distorted predictions under extreme conditions. Summary of the Invention
[0007] This application provides a rolling forecasting method and system for power load based on large-scale time-series modeling. It aims to address the technical problems in existing power load forecasting, such as rigid windows that cannot respond to sudden changes in grid structure, high computational overhead from full-scale model updates that cannot support high-frequency rolling, and inaccuracies due to missing physical semantics of covariates under extreme conditions. The method utilizes a dual-channel architecture to achieve lightweight incremental learning and introduces a physically guided encoding structure to enhance the physical consistency of feature representation, thereby achieving low-latency and highly robust rolling forecasting while ensuring prediction accuracy.
[0008] This application provides a rolling forecasting method for power load based on large-scale time-series modeling, including: Step S10: Obtain historical power load data and multidimensional covariate data of the target area, and preprocess the historical power load data and multidimensional covariate data to construct an initial time series dataset containing load characteristics, meteorological characteristics, calendar characteristics and socio-economic characteristics. Step S20: Based on the dynamic window mechanism, traverse the initial time series dataset, adaptively extract the target data segment corresponding to the current prediction time domain according to the load fluctuation degree and power grid operating conditions, and determine the rolling duration; Step S30: Apply the pre-trained time series large model to construct a dual-channel incremental update architecture, input the load features in the target data segment into the architecture, extract deep load time series features, and keep the static base channel parameters in the architecture frozen, only update the parameters of the dynamic residual channel in the architecture. Step S40: Input the meteorological features, calendar features, and socioeconomic features in the target data segment into the physical-guided coding structure to generate a covariate feature vector with physical constraints. Step S50: The deep load time series features and the covariate feature vector are spatiotemporally aligned and deeply fused, and the fused feature tensor is input into the updated time series large model to output the rolling prediction results of power load for the future preset period and the corresponding confidence interval.
[0009] Further, in step S10, the historical power load data and multidimensional covariate data are preprocessed, specifically including: Step S11: Perform multi-level anomaly detection on historical power load data. First, invalid data exceeding the operating limits of electrical equipment is removed based on hard thresholds of physical quantities. Second, density-based spatial clustering algorithm is used to identify outliers in the load curve and mark them as missing values. Consistency verification is performed on multi-dimensional covariate data to detect abrupt noise and logical conflicts in meteorological data. Step S12: For historical power load data marked with missing values, multiple interpolation method is used to complete the data: linear interpolation is performed using adjacent time period data first. If the missing span exceeds a preset threshold, a similar day matching algorithm is introduced to retrieve load curves of the same date type from the historical database for weighted completion. Step S13: Extract load features, meteorological features, calendar features, and socioeconomic features from the historical power load data and multidimensional covariate data after filling in missing values, and align and stitch them according to timestamps to form an initial time-series dataset. The load characteristics include active power, reactive power, voltage amplitude, and current amplitude extracted from historical power load data; the meteorological characteristics include temperature, humidity, wind speed, air pressure, irradiance, and precipitation probability extracted from multidimensional covariate data; the calendar characteristics include date type, holiday identifier, season identifier, and time identifier extracted from multidimensional covariate data; and the socioeconomic characteristics include industry type, operating rate, GDP growth rate, and electricity price index extracted from multidimensional covariate data.
[0010] Furthermore, step S20 includes the following detailed steps: Step S21, set the baseline window length Baseline rolling duration First fluctuation threshold Second fluctuation threshold To obtain the current operating conditions of the power grid, including changes in the power grid topology, the maintenance status of key equipment, and fluctuations in the output of new energy sources; Step S22: Extract and analyze the load characteristics of historical power load data in the initial time series dataset, and calculate the load fluctuation rate within a preset historical period before the current time. ; Step S23, load fluctuation rate With the first fluctuation threshold Second fluctuation threshold By comparing the data and considering the current power grid operating conditions, a comprehensive judgment strategy is used to adaptively extract target data segments with a set window length from the initial time series dataset, and the rolling duration is dynamically set.
[0011] The comprehensive judgment strategy includes: when If an abnormal power grid outage, critical equipment maintenance, or fluctuations in renewable energy output exceeding a preset ratio are detected, the system is classified as a high-fluctuation mode, and a window of length is extracted from the initial time-series dataset. The target data segment, where the weight coefficients And set the scrolling duration to The weighting coefficient ; when Furthermore, if the power grid operates smoothly, meaning no abnormal power outages, major equipment maintenance, or fluctuations in new energy output exceeding a preset ratio are detected, the mode is determined to be a medium fluctuation mode, and the window length is [not specified]. The target data segment maintains the baseline scrolling duration. ; when Furthermore, when the power grid is in a fully connected and fully protected normal operating state, it is determined to be in low fluctuation mode, and the window length is [value missing]. The target data segment, where the weight coefficients And set the scrolling duration to The weighting coefficient .
[0012] Further, in step S30, the dual-channel incremental update architecture includes: Static base channel: Loads converged general temporal pattern parameters from the pre-trained temporal large model; during the training of the dual-channel incremental update architecture, this static base channel is frozen and only performs forward inference, without participating in gradient updates; Dynamic residual channel: As the only learnable subnetwork, it achieves nonlinear transformation of features through low-rank matrix factorization and is inserted after the query, key, and value projection matrix of the self-attention layer of the Transformer encoder in the pre-trained temporal large model, as well as after the up-dimensional linear layer of the feedforward network. During the incremental update phase, only the gradient of the dynamic residual channel is calculated and updated, so that the model can adapt to the latest load change pattern without forgetting the general time series pattern.
[0013] Further, in step S40, the encoding structure of the physical guide includes: Meteorological physical quantity coding submodule: Based on thermodynamic principles, it converts the raw meteorological characteristic data in the target data segment into physical equivalent indicators that reflect human comfort or equipment heat dissipation efficiency. Socioeconomic graph coding submodule: used to construct a heterogeneous graph network containing socioeconomic features of the target data fragment, and extract load transmission features between different industries through graph attention mechanism; Semantic enhancement submodule: Used to jointly represent the calendar features in the target data fragment through sine and cosine position encoding and holiday semantic embedding vectors to generate time semantic features.
[0014] This application also provides a rolling forecasting system for power load based on large-scale time-series modeling, including: The time-series dataset construction module is used to acquire historical power load data and multidimensional covariate data of the target area, and to preprocess the historical power load data and multidimensional covariate data to construct an initial time-series dataset containing load characteristics, meteorological characteristics, calendar characteristics and socio-economic characteristics. Dynamic window decision module: used to traverse the initial time series dataset based on the dynamic window mechanism, and adaptively extract the target data segment corresponding to the current prediction time domain according to the load fluctuation degree and power grid operating conditions; Large Model Incremental Training Module: Used to construct a dual-channel incremental update architecture using a pre-trained time series large model. The load features in the target data segment are input into the architecture to extract deep load time series features. At the same time, the static base channel parameters in the architecture are kept frozen, and the parameters of the dynamic residual channel in the architecture are updated only. Physically constrained coding module: used to input meteorological features, calendar features and socio-economic features in the target data segment into a physically guided coding structure to generate a covariate feature vector with physical constraints; Inference and prediction module: It is used to perform spatiotemporal alignment and deep fusion of the deep load time series features and the covariate feature vector, and input the fused feature tensor into the time series large model after parameter update, and output the rolling prediction results of power load for the future preset period and the corresponding confidence interval.
[0015] This application discloses the following technical effects: This application provides a rolling forecasting method and system for power load based on large-scale time-series modeling. It establishes a two-factor joint decision-making mechanism that combines load fluctuations and grid operating conditions. When the grid structure changes, it forces a switch to a high-fluctuation mode and shortens the input window, avoiding misjudgments caused by the forecasting model's over-reliance on outdated historical data. This ensures that the forecasting results can still accurately reflect the current physical operating characteristics under extreme scenarios such as changes in grid topology and dramatic changes in renewable energy output.
[0016] Secondly, to address the problem that general time series large models have a huge number of parameters and full fine-tuning is difficult to meet the timeliness of rolling prediction, this invention constructs a dual-channel incremental update architecture. During the training process, the parameters of the static base channel are kept constant, and only the gradient of the dynamic residual channel is updated. This enables high-frequency rolling prediction under ordinary computing power environment and breaks through the computing power bottleneck of large models for real-time control of power grid.
[0017] Existing methods often use meteorological and economic data as simple numerical inputs, lacking constraints from physical mechanisms, leading to prediction distortion when there are no similar historical samples. This invention introduces a physics-guided coding structure to calculate physical equivalent indicators with clear thermodynamic meanings, and combines this with socio-economic mapping to capture the transmission effects of the industrial chain. This deep integration of physical mechanisms and data-driven approaches endows the model with extremely strong extrapolation capabilities. Especially under extreme conditions such as public health emergencies and the introduction of new policies where there are "no similar historical days," it can still provide high-confidence predictions with physically consistent logic, effectively avoiding precipitous drops or inflated predictions. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a rolling forecasting method for power load based on large-scale time-series modeling, provided in an embodiment of this application.
[0019] Figure 2 This is a schematic diagram of the structure of a rolling forecasting system for power load based on large-scale time-series modeling, provided in an embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application will be provided in conjunction with the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] Example 1: This application provides a rolling forecasting method for power load based on large-scale time-series modeling. The method is applied to the day-ahead and intraday rolling dispatching scenario of a provincial power grid, aiming to solve the challenges of rolling power load forecasting under conditions of high-proportion renewable energy integration, frequent extreme weather events, and changes in grid structure. Figure 1 As shown, the method includes: Step S10: Obtain historical power load data and multidimensional covariate data of the target area, and preprocess the historical power load data and multidimensional covariate data to construct an initial time series dataset containing load characteristics, meteorological characteristics, calendar characteristics and socio-economic characteristics.
[0022] In this embodiment, SCADA measurement data (historical power load data) and multidimensional covariate data of the target area from July 1 to July 31 are obtained through the data acquisition module, with a sampling frequency of 15 minutes per sampling.
[0023] First, multi-level anomaly detection is performed on historical power load data, including: Physical boundary detection: Set upper and lower threshold values for power load. ,in The minimum load record for the region is set at 100MW. The rated capacity of the main transformer is set at 6000MW. If the load value at a certain sampling point is not within the upper or lower limit threshold, it is judged as a physical fault and marked. For example, if a 500kV main transformer trips at 14:05 on July 10, resulting in a local load sampling value of 0, the system determines... Marked as a suspected bad pixel.
[0024] Density-based spatial clustering (DBSCAN) algorithm for identifying outliers: defining the feature vector for each sample point. for ,in for Active power at any given time express The change in load at a given time relative to the previous time. express The change in load at time 1 relative to the previous two time points. and They are respectively The mean and standard deviation of the load one hour prior to the time point (4 sampling points); setting the neighborhood radius of the DBSCAN algorithm. for and minimum sample size The value is 4 (corresponding to the number of sampling points within 1 hour), of which This represents the standard deviation of the daily load curves over the past 30 days. The clustering process of the DBSCAN algorithm is as follows: Randomly select an unvisited sampling point and obtain its feature vector. ;calculate of Points within the neighborhood That is, all that satisfy sampling points Quantity, Indicates the calculation of the L2 norm; if Then create a new cluster C, and transfer the current cluster to the new cluster. At each sampling point, a new cluster is added to C, and all density-reachable points in C are recursively added to the cluster; if Then the current The sampling points at any given time are marked as noise points (outliers). For example, the load surge of 5800MW at 12:00 on July 15th, although not exceeding the physical boundary, has a significantly different density compared to surrounding points in the characteristic space. <4, which was identified as an outlier by DBSCAN and marked as missing.
[0025] Contextual correlation detection: The Isolation Forest algorithm is used for logical verification. Temperature and relative humidity data are extracted from multidimensional covariate data, and the feature space is defined as follows. ,in for Temperature at any moment for Relative humidity at any time; training the isolated forest model, setting the number of trees. Subsampling rate Calculate the power load sample Path length and abnormal scores ,in For sample size The average path length over time. It is a sample The expected value of the path length in all trees; if and If a value of 0 is detected as a logical contradiction ("high temperature but abnormally low load"), it is confirmed as a data acquisition malfunction. For example, regarding the value of 0 at 14:05 on July 10, the system detected that the temperature at that time was 38℃ and the humidity was 60%. The logical contradiction confirms that the value of 0 is a data acquisition malfunction rather than a true requirement to return to zero, and it is therefore removed.
[0026] Secondly, consistency checks are performed on the multi-dimensional covariate data. By setting threshold values for temperature, humidity, and wind speed, abrupt noise in each meteorological parameter is identified. Logical conflicts in the meteorological data are detected by establishing the physical correlation between temperature and humidity. For the missing segments of historical power load data identified in the multi-level anomaly detection, intelligent repair is achieved using a similar-day matching algorithm combined with weighted completion, including: Define date and Similarity measurement function for:
[0027] in, The range of values is The weighting coefficients satisfy + In this embodiment, the values are set to 0.4, 0.3, and 0.3 respectively. This indicates the date type encoding: weekdays = 1, weekends = 0.5, and public holidays = 0. and These are the daily average temperature and daily average relative humidity, respectively. For example, when searching for dates with a similarity of less than 0.2 in the historical database, July 3rd (also a Wednesday with a temperature of 37°C) is matched as the best similar day. Based on the best similar day found, a linear weighted method is used to complete the missing values in the power load data. The completion formula is expressed as follows: ,in for The load value after real-time completion Best similarity day Load value at any given time The missing segment is adjacent to the previous and next segments. Load values at each moment; The weight for similar days is set to 0.7 in this embodiment; The number of adjacent time periods is set to 3 in this embodiment.
[0028] Finally, four feature vectors were extracted and constructed from the preprocessed historical power load data and multidimensional covariate data to form an initial time-series dataset; this dataset includes load features, meteorological features, calendar features, and socioeconomic features. In this embodiment, load characteristics Depend on Active power, reactive power, three-phase voltage, and three-phase current at any given time, and Load mean and standard deviation over the past hour; meteorological characteristics Depend on The calendar features are composed of temperature, relative humidity, wind speed, irradiance, and atmospheric pressure at any given time; the calendar features are encoded using periodic coding, and the encoded calendar features... Represented as: ,in It represents the number of hours in a day, is a positive integer, and ranges from 0 to 23; Represents the day of the week, with values ranging from positive integers 0 to 6; Indicates holiday markers, 1 for holidays and 0 for non-holidays; socio-economic characteristics. Include Real-time industrial operating rates, semiconductor industry load share, and real-time electricity prices.
[0029] Step S20: Based on the dynamic window mechanism, traverse the initial time series dataset, adaptively extract the target data segment corresponding to the current prediction time domain according to the load fluctuation degree and power grid operating conditions, and determine the rolling duration.
[0030] Traditional methods adjust the input window solely based on statistical fluctuations in load data, neglecting the structural impact of changes in grid topology on load characteristics. This leads to models continuing to use outdated historical data when grid operating conditions change abruptly, resulting in "inertial misjudgments." This invention constructs a two-factor joint decision-making mechanism combining load fluctuations and grid operating conditions. By quantifying the physical state of the grid, it adaptively adjusts the input window length and rolling duration. The implementation process is as follows: Load volatility is quantified based on the following formula: in, This represents the load volatility; a larger value indicates more severe volatility. The sliding window length is set to 96 in this embodiment (corresponding to a sampling frequency of 24 hours and 15 minutes). For the first Active power at each sampling point For the past The average load of each sampling point.
[0031] The operating conditions of the power grid are quantified, and a power grid operating condition vector is defined. ,in This is the N-1 interruption flag, which is 1 when an N-1 interruption fault occurs in the power grid, and 0 otherwise. This is the equipment maintenance indicator. It is 1 when there is a planned maintenance for important equipment (such as main transformers or lines), and 0 otherwise. The power change rate of the tie line, and , for Constant communication line power, Rated capacity of the tie line.
[0032] like (This embodiment sets a first preset threshold) (0.15) or If the power grid operating conditions or load change drastically, indicating a high fluctuation mode, then the window length for capturing the target data segment should be set. (In this embodiment) Reference window length (168 hours), rolling duration (In this embodiment) Baseline rolling duration (for 1 hour) like (This embodiment sets a second preset threshold) (0.05) and When load changes are in a low-fluctuation mode, the window length is extended to 336 hours, and the rolling duration is extended to 2 hours. For example, at 09:00 on July 15th, the power grid plans to maintain a 500kV tie line; at this time... It belongs to a low volatility pattern, but The system was forced to switch to high volatility mode, and the window was shortened from 168h to 84h.
[0033] Step S30: Apply the pre-trained time series large model to construct a dual-channel incremental update architecture, input the load features in the target data segment into the architecture, extract deep load time series features, and keep the static base channel parameters in the architecture frozen, only updating the parameters of the dynamic residual channel in the architecture.
[0034] General-purpose time-series large-scale models typically have hundreds of millions of parameters (e.g., the TimesFM-256M has approximately 250 million parameters). Using full fine-tuning not only incurs enormous computational overhead and high memory consumption, but also easily leads to catastrophic forgetting of the general-purpose time-series knowledge learned during pre-training. This results in the model being unable to adapt to the high-frequency rolling prediction requirements of the power grid every 15 minutes. This invention proposes a dual-channel architecture of static basis freezing and dynamic residual channel updating. This architecture decouples model parameters into a general knowledge carrier (static) and a scenario adaptation plugin (dynamic), updating only a very small percentage of parameters (<1%). While maintaining prediction accuracy, it achieves minute-level incremental updates. Implementation details are as follows: In this embodiment, a pre-trained TimesFM-256M time series model is loaded. This model has been trained on massive amounts of publicly available time series data and has a powerful ability to extract the period (daily, weekly, quarterly) of general time series data. All parameters of the static base channel are traversed. The list of parameters that need to be traversed and frozen, along with their positions and functions in the model, is shown in Table 1. Table 1. Parameters to be frozen in the large time series model TimesFM-256M
[0035] Freeze all parameters in the static base channel by setting the attribute of the above parameter to "param.requires_grad =False" in the PyTorch code.
[0036] The target data segment extracted from the power load data through a dynamic windowing mechanism is input into the pre-trained time series large model, and the primary load feature representation is output through forward propagation. ,in For training batch size, In each layer of the Transformer of the pre-trained temporal large model, a low-rank adaptation module (LoRA) is inserted as a dynamic residual channel; for the original weight parameters... Introduce low-rank decomposition: ,in The weight matrix after low-rank decomposition and It is a low-rank matrix that is updated during model training. Initialize using He. Initialize as an all-zero matrix; This is a scaling factor, set to 32, used to balance the gradient scale in the early stages of training; The rank of the matrix is set to 8, which is much smaller than the original dimension (256), greatly reducing the number of parameters. In this embodiment, LoRA is inserted into the projection matrices of the query (Q), key (K), and value (V) in the self-attention mechanism; and LoRA is inserted into the linear layer that maps features from 256 dimensions to 1024 dimensions in the feedforward network.
[0037] The pre-trained temporal model has approximately 256M parameters. Now, only the newly added LoRA parameters need to be trained, which only number about 0.5M, representing only 0.2% of the total parameters. This significantly reduces computational overhead and memory usage. Incremental training of the pre-trained temporal model only involves training the parameters in the dynamic residual channels. and The matrix is set to `requires_grad = True`, and mean squared error loss combined with L2 regularization is used to prevent overfitting. During backpropagation and parameter update optimization, the AdamW optimizer is used with a learning rate set to 0.0003, and only updates... and The weight.
[0038] The deep load time series features extracted by the final model Represented as: ,in This refers to the incremental features output by the dynamic residual channel, used to correct the output of the static basis. To address the issue of redundant calculations in rolling prediction, a caching mechanism is introduced: if the overlap rate between the input window of the current time step and the previous time step exceeds 90% (i.e., only one new sampling point has been added), the static basis channel is not recalculated; instead, the previous time step's data is directly read from the cache. The system performs forward inference only on the newly added sampling point and concatenates it with cached features. This reduces the time for a single inference cycle from 28 seconds to less than 5 seconds, significantly improving system response speed.
[0039] Step S40: Input the meteorological features, calendar features, and socioeconomic features in the target data segment into the physical-guided coding structure to generate a covariate feature vector with physical constraints.
[0040] Existing methods often directly input temperature and humidity as independent values into the MLP (Mean Level Programming), lacking physical mechanism constraints, leading to prediction distortion under extreme weather conditions. This invention introduces physical equivalent indicators and socio-economic transmission maps, injecting physical mechanisms into the feature encoding process. The specific implementation flow is as follows: In this embodiment, the equivalent temperature and humidity index is calculated based on the meteorological features extracted from the initial time-series dataset. : ,in Temperature in Celsius Relative humidity, equivalent temperature and humidity index The higher the value, the hotter it feels, and the greater the electrical load required (e.g., for air conditioning). Taking into account the combined effects of temperature and humidity on human comfort and the energy efficiency of air conditioning equipment, it has more physical significance than temperature alone.
[0041] Based on the socioeconomic features extracted from the initial time-series dataset, socioeconomic graph encoding is performed to construct an industry heterogeneity graph, whose nodes... Representing industries (such as silicon, wafers, packaging, and manufacturing), edge Representing upstream and downstream supply relationships, the node characteristic is the operating rate; based on the constructed industry heterogeneity graph, the load transmission characteristics between different industries are extracted through a graph attention mechanism. For example, historical power load data for the same period may differ in time patterns due to certain reasons. The model no longer relies on historical load values, but instead uses socio-economic maps to infer the physical transmission relationship that the load increases by 300MW for every 10% increase in the resumption rate, thus avoiding a precipitous drop in the predicted value.
[0042] Calendar features extracted from the initial time-series dataset are transformed into vectors easily understood by the model. These vectors are then encoded using sine and cosine coding to represent the hours. Transform it into a periodic vector, for example, =23 and =0 is very close in the vector space, which conforms to the cyclical nature of time; construct a learnable holiday embedding matrix, for example, encode Spring Festival, National Day, and weekdays into specific vectors, so that the model can learn the semantic information that the load is low during Spring Festival and high during weekdays; concatenate the sine and cosine encoded vectors with the holiday embedding matrix to generate temporal semantic features. .
[0043] Step S50: The deep load time series features and the covariate feature vector are spatiotemporally aligned and deeply fused, and the fused feature tensor is input into the updated time series large model to output the rolling prediction results of power load for the future preset period and the corresponding confidence interval.
[0044] Existing methods often output a single predicted value, failing to differentiate the reliability of the prediction results, and the model is susceptible to oscillations caused by acquisition noise, increasing scheduling risks. This invention employs a cross-attention mechanism for deep fusion and integrates Monte Carlo Dropout to directly output the confidence interval of the predicted value during the inference phase, quantifying uncertainty. The specific implementation process is as follows: In this embodiment, the deep load timing characteristics are based on Unix timestamps. With covariate eigenvectors (Including equivalent temperature and humidity index) Load transmission characteristics and temporal semantic features Spatiotemporal alignment is performed; a cross-attention mechanism is employed based on load characteristics. For querying, based on covariate features Using the key value, the two features are deeply fused. Based on the calculated cross-attention weights, the temperature feature is assigned a high weight during high-temperature periods, and the operating rate feature is automatically assigned a high weight during night shifts.
[0045] The fused feature tensor is input into the updated temporal model, keeping the Dropout layer active (unlike traditional inference where Dropout is disabled), and 100 random sampling inferences are performed; the mean of the 100 predicted outputs is then calculated. and variance The final output of the model, the rolling forecast of electricity load, is expressed as: (95% confidence interval). For example, the peak power load forecast for the evening of July 15th is 5200MW ± 15MW. The dispatcher sees that the confidence interval is very narrow and judges the forecast to be reliable. Based on this, the dispatcher accurately arranges the power generation plan, reducing the spinning reserve by 500MW and saving dispatching costs. If the interval is too wide (such as ± 100MW), the dispatcher is prompted to make conservative arrangements for the reserve.
[0046] Example 2: The power load rolling forecasting system based on large-model time series modeling provided in this embodiment of the invention can execute the power load rolling forecasting method based on large-model time series modeling provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method, such as... Figure 2 As shown, it includes: The time-series dataset construction module is used to acquire historical power load data and multidimensional covariate data of the target area, and to preprocess the historical power load data and multidimensional covariate data to construct an initial time-series dataset containing load characteristics, meteorological characteristics, calendar characteristics and socio-economic characteristics. Dynamic window decision module: used to traverse the initial time series dataset based on the dynamic window mechanism, and adaptively extract the target data segment corresponding to the current prediction time domain according to the load fluctuation degree and power grid operating conditions; Large Model Incremental Training Module: Used to construct a dual-channel incremental update architecture using a pre-trained time series large model. The load features in the target data segment are input into the architecture to extract deep load time series features. At the same time, the static base channel parameters in the architecture are kept frozen, and the parameters of the dynamic residual channel in the architecture are updated only. Physically constrained coding module: used to input meteorological features, calendar features and socio-economic features in the target data segment into a physically guided coding structure to generate a covariate feature vector with physical constraints; Inference and prediction module: It is used to perform spatiotemporal alignment and deep fusion of the deep load time series features and the covariate feature vector, and input the fused feature tensor into the time series large model after parameter update, and output the rolling prediction results of power load for the future preset period and the corresponding confidence interval.
[0047] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.
[0048] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A rolling forecasting method for power load based on large-scale time series modeling, characterized in that, The method includes: Step S10: Obtain historical power load data and multidimensional covariate data of the target area, and preprocess the historical power load data and multidimensional covariate data to construct an initial time series dataset containing load characteristics, meteorological characteristics, calendar characteristics and socio-economic characteristics. Step S20: Based on the dynamic window mechanism, traverse the initial time-series dataset and adaptively extract the target data segment corresponding to the current prediction time domain according to the load fluctuation degree and power grid operating conditions. when If an abnormal power grid outage, critical equipment maintenance, or fluctuations in renewable energy output exceeding a preset ratio are detected, the system is classified as a high-fluctuation mode, and a window of length is extracted from the initial time-series dataset. The target data segment, in which For load fluctuation rate, The first fluctuation threshold, The baseline window length and weighting coefficients are given. And set the scrolling duration to ,in, The baseline rolling duration, weighting coefficient ; when Furthermore, if the power grid operates smoothly, meaning no abnormal power outages, major equipment maintenance, or fluctuations in new energy output exceeding a preset ratio are detected, the mode is determined to be a medium fluctuation mode, and the window length is [value missing]. The target data segment maintains the baseline scrolling duration. ,in This is the second fluctuation threshold; when Furthermore, when the power grid is in a fully connected and fully protected normal operating state, it is determined to be in low fluctuation mode, and the window length is [value missing]. The target data segment, where the weight coefficients And set the scrolling duration to The weighting coefficient ; Step S30: Apply the pre-trained time series large model to construct a dual-channel incremental update architecture, input the load features in the target data segment into the architecture, extract deep load time series features, and keep the static base channel parameters in the architecture frozen, only update the parameters of the dynamic residual channel in the architecture. The dual-channel incremental update architecture includes: Static base channel: Loads converged general temporal pattern parameters from the pre-trained temporal large model; during the training of the dual-channel incremental update architecture, this static base channel is frozen and only performs forward inference, without participating in gradient updates; Dynamic residual channel: As the only learnable subnetwork, it achieves nonlinear transformation of features through low-rank matrix factorization and is inserted after the query, key, and value projection matrix of the self-attention layer of the Transformer encoder in the pre-trained temporal large model, as well as after the up-dimensional linear layer of the feedforward network. During the incremental update phase, only the gradient of the dynamic residual channel is calculated and updated, so that the model can adapt to the latest load change pattern without forgetting the general time series pattern. Step S40: Input the meteorological features, calendar features, and socioeconomic features in the target data segment into the physical-guided coding structure to generate a covariate feature vector with physical constraints. The coding structure for physical guidance includes: Meteorological physical quantity coding submodule: Based on thermodynamic principles, it converts the raw meteorological characteristic data in the target data segment into physical equivalent indicators that reflect human comfort or equipment heat dissipation efficiency. Socioeconomic graph coding submodule: used to construct a heterogeneous graph network containing socioeconomic features of the target data fragment, and extract load transmission features between different industries through graph attention mechanism; The semantic enhancement submodule is used to jointly represent the calendar features in the target data fragment using sine and cosine positional encoding and holiday semantic embedding vectors to generate time semantic features. Step S50: The deep load time series features and the covariate feature vector are spatiotemporally aligned and deeply fused, and the fused feature tensor is input into the time series large model after parameter update, and the rolling prediction results of power load for the future preset period and the corresponding confidence interval are output.
2. The rolling forecasting method for power load based on large-scale time series modeling as described in claim 1, characterized in that, In step S10, the historical power load data and multidimensional covariate data are preprocessed, specifically including: Step S11: Perform multi-level anomaly detection on historical power load data. First, invalid data exceeding the operating limits of electrical equipment is removed based on hard thresholds of physical quantities. Second, density-based spatial clustering algorithm is used to identify outliers in the load curve and mark them as missing values. Consistency verification is performed on multi-dimensional covariate data to detect abrupt noise and logical conflicts in meteorological data. Step S12: For historical power load data marked with missing values, multiple interpolation method is used to complete the data: linear interpolation is performed using adjacent time period data first. If the missing span exceeds a preset threshold, a similar day matching algorithm is introduced to retrieve load curves of the same date type from the historical database for weighted completion. Step S13: Extract load features, meteorological features, calendar features, and socioeconomic features from the historical power load data and multidimensional covariate data after filling in missing values, and align and stitch them according to timestamps to form an initial time series dataset.
3. The rolling forecasting method for power load based on large-scale time series modeling as described in claim 2, characterized in that, The load characteristics include active power, reactive power, voltage amplitude, and current amplitude extracted from historical power load data. The meteorological features include temperature, humidity, wind speed, air pressure, irradiance, and precipitation probability extracted from multidimensional covariate data; The calendar features include date type, holiday identifier, season identifier, and time identifier extracted from multidimensional covariate data; The socioeconomic characteristics include industry type, operating rate, GDP growth rate, and electricity price index extracted from multidimensional covariate data.
4. The rolling forecasting method for power load based on large-scale time series modeling as described in claim 1, characterized in that, Step S20 includes the following detailed steps: Step S21, set the baseline window length Baseline rolling duration First fluctuation threshold Second fluctuation threshold To obtain the current operating conditions of the power grid, including changes in the power grid topology, the maintenance status of key equipment, and fluctuations in the output of new energy sources; Step S22: Extract and analyze the load characteristics of historical power load data in the initial time series dataset, and calculate the load fluctuation rate within a preset historical period before the current time. ; Step S23, load fluctuation rate With the first fluctuation threshold Second fluctuation threshold By comparing the data and considering the current power grid operating conditions, a comprehensive judgment strategy is used to adaptively extract target data segments with a set window length from the initial time series dataset, and the rolling duration is dynamically set.
5. The rolling forecasting method for power load based on large-scale time series modeling as described in claim 1, characterized in that, The load characteristics in the target data segment are input into the dual-channel incremental update architecture to extract deep load time-series characteristics, specifically including: The load features from the target data segment are input into the static basis channel, and feature extraction is performed by a multi-layer Transformer encoder to obtain a primary load feature representation containing long-period dependencies. ; Representing the primary load characteristics The input is fed into the dynamic residual channel, where the feature space is dynamically transformed and enhanced through low-rank matrix decomposition, outputting deep load time-series features. ; And represent the primary load characteristics of the static base channel output. The features are stored in the cache. In the next rolling prediction, if the window length has not changed, the cached features are directly called to reduce redundant calculations.
6. A rolling forecasting system for power load based on large-scale time-series modeling, characterized in that, The system is used to implement the rolling forecasting method for power load based on large-model time series modeling as described in any one of claims 1-5, and the system comprises: The time-series dataset construction module is used to acquire historical power load data and multidimensional covariate data of the target area, and to preprocess the historical power load data and multidimensional covariate data to construct an initial time-series dataset containing load characteristics, meteorological characteristics, calendar characteristics and socio-economic characteristics. Dynamic window decision module: used to traverse the initial time series dataset based on the dynamic window mechanism, and adaptively extract the target data segment corresponding to the current prediction time domain according to the load fluctuation degree and power grid operating conditions; Large Model Incremental Training Module: Used to construct a dual-channel incremental update architecture using a pre-trained time series large model. The load features in the target data segment are input into the architecture to extract deep load time series features. At the same time, the static base channel parameters in the architecture are kept frozen, and the parameters of the dynamic residual channel in the architecture are updated only. Physically constrained coding module: used to input meteorological features, calendar features and socio-economic features in the target data segment into a physically guided coding structure to generate a covariate feature vector with physical constraints; Inference and prediction module: It is used to perform spatiotemporal alignment and deep fusion of the deep load time series features and the covariate feature vector, and input the fused feature tensor into the time series large model after parameter update, and output the rolling prediction results of power load for the future preset period and the corresponding confidence interval.
Citation Information
Patent Citations
Electric power short-term load prediction method and system based on time sequence large model
CN120184939A
Electric power total-factor unified load prediction method and system based on large time sequence model
CN121529526A
Power load prediction method and system based on end-cloud collaboration
CN122315625A