Machine learning based power load short term forecasting system

CN122532911BActive Publication Date: 2026-09-08HANGZHOU KAICHANG ELECTRIC POWER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611031321.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-08
Estimated Expiration
2046-07-13

AI Technical Summary

Technical Problem

[0007]第一、对输入数据的质量缺乏系统的评估机制

Benefits of technology

[0030]In this machine learning-based short-term power load forecasting system, this solution employs a full-link data quality assurance mechanism to internalize anomaly detection and quality assessment as core components. It utilizes coupled embedding modeling of load and meteorological data to detect hidden anomalies and outputs a global quality score, suppressing data disturbances from the source and mitigating their impact on forecasting. Simultaneously, a cross-modal dynamic gating fusion unit uses the quality score as a continuously modulated signal to generate dynamic fusion weights, achieving soft decision-making by enhancing high-quality features and suppressing low-quality features, ensuring the robustness of multi-source fusion. Finally, a three-dimensional parallel architecture is constructed using frequency domain adaptive decomposition, causal graph attention, and multi-scale temporal feature extraction. This automatically matches decomposition parameters and constrains attention to be calculated only on causal paths, eliminating spurious correlations. Multi-scenario adaptive forecasting and an online incremental update mechanism based on residual trend detection enable the system to autonomously adapt to differentiated power consumption scenarios and continuously evolve even as performance degrades, significantly improving forecast accuracy and long-term operational reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122532911B_ABST
    Figure CN122532911B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of power load prediction, in particular to a power load short-term prediction system based on machine learning, which comprises a cross-modal dynamic gating fusion unit, a multi-scene adaptive prediction unit and a model online adaptive updating unit. The present application uses a full-link data quality guarantee mechanism to internalize anomaly detection and quality evaluation into the core link of the system, and uses the coupling of load and weather to embed modeling to detect hidden anomalies. At the same time, the cross-modal dynamic gating fusion unit uses the quality score as a continuous modulation signal to generate a dynamic fusion weight, realizes soft decision of high-quality feature enhancement and low-quality feature suppression, and adopts frequency domain adaptive decomposition, causal graph attention and multi-scale time series feature extraction to form a three-dimensional parallel architecture, which automatically matches the decomposition parameters and restricts the attention calculation only on the causal path. Finally, the multi-scene adaptive prediction and the online incremental updating mechanism based on residual trend detection significantly improve the prediction accuracy and long-term operation reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power load forecasting technology, and more specifically, to a short-term power load forecasting system based on machine learning. Background Technology

[0002] Short-term load forecasting typically refers to predicting the trend and specific values ​​of power load changes over the next few hours to days. This work is a core component of power grid dispatching and operation, and its accuracy directly affects the rationality of generation plans, the efficiency of spinning reserve capacity allocation, and the quality of decision-making in the electricity market. Accurate short-term load forecasting enables grid dispatching departments to plan unit start-ups and shutdowns in advance, optimize fuel procurement, and reduce generation costs; it ensures frequency stability and voltage safety of the system under sudden disturbances; and it provides a reliable load baseline for demand response resource allocation, supporting electricity retailers and load aggregators in formulating precise pricing strategies. Therefore, improving the accuracy of short-term load forecasting has always been a key focus in the field of power system engineering and academic research.

[0003] Existing short-term load forecasting methods can be summarized into two main technical approaches:

[0004] The first category is statistical methods based on traditional time series analysis. These methods treat load as a time series signal, using linear or generalized linear mathematical models to characterize the historical autocorrelation structure and trend characteristics of the load. Representative methods include autoregressive integral moving average models and their seasonal extensions, exponential smoothing, and linear regression models. These methods have a mature theoretical foundation and strong parameter interpretability, and have achieved relatively reliable application results during periods of relatively stable load behavior. However, the formation mechanism of electricity load is highly complex, influenced by nonlinear meteorological factors such as temperature, humidity, and sunshine conditions, and by the combined modulation of human behavioral factors such as social production activities, holiday arrangements, and electricity price signals. Statistical methods, limited by their linear modeling framework, have limited ability to fit these nonlinear relationships. When load behavior undergoes structural changes or encounters extreme weather conditions, prediction errors often increase significantly.

[0005] The second category is intelligent prediction methods based on machine learning. These methods automatically extract the mapping relationship between load and various influencing factors from a large amount of historical data by building data-driven learning models.

[0006] Despite the progress made by the aforementioned machine learning methods, the following shortcomings still exist in practical engineering applications:

[0007] First, there is a lack of a systematic evaluation mechanism for the quality of input data. Existing forecasting systems typically assume that the input data is complete, accurate, and consistent, but the quality of data in real-world engineering environments is often difficult to guarantee. Load data may contain missing values ​​and abrupt changes due to acquisition terminal failures, communication interruptions, or storage anomalies; meteorological data may be subject to biases and noise due to sensor drift, transmission delays, or data source switching; user electricity consumption behavior tag data and calendar information data also face the risk of entry errors and update delays.

[0008] Second, the dynamic relationships between multi-source heterogeneous data have not been fully explored. Load forecasting relies on the collaborative support of various heterogeneous data sources, including the time dependence of the load series itself, the cross-correlation between load and meteorological variables such as temperature and humidity, and the contextual constraints of date type and user behavior tags.

[0009] Third, the differentiated characteristics of different regions, seasons, and electricity consumption patterns are difficult to effectively cover with a single prediction model. The statistical characteristics and behavioral patterns of electricity load exhibit significant spatiotemporal heterogeneity: industrial load shows a stable periodic pattern, residential load exhibits greater random fluctuations, and commercial load is significantly affected by holiday effects; the load curve shape of the same region varies greatly in different seasons; and there are even more fundamental differences in load behavior characteristics between regions with different climate zones and different levels of economic development.

[0010] The three shortcomings mentioned above are not isolated but interconnected, forming a negative synergistic effect: data quality issues exacerbate the difficulty of multimodal fusion, and the degradation of fusion results further amplifies the impact of scene differences on prediction accuracy; conversely, changes in data distribution during scene switching may make data quality problems more subtle and diverse. Current technologies have not yet provided a systematic solution to these three shortcomings.

[0011] Therefore, there is an urgent need for a short-term power load forecasting system that can systematically assess the quality of multi-source input data, dynamically adjust the multimodal feature fusion weights based on data quality, and automatically adapt to the differentiated characteristics of different power consumption scenarios. Summary of the Invention

[0012] The purpose of this invention is to provide a short-term power load forecasting system based on machine learning, thereby solving the problems mentioned in the background art.

[0013] To achieve the above objectives, the machine learning-based short-term power load forecasting system includes a multi-source data acquisition unit, a time-series-meteorological coupling anomaly detection unit, a multi-modal data quality assessment unit, a frequency domain adaptive decomposition unit, a causal graph attention feature extraction unit, and a multi-scale time-series feature extraction unit.

[0014] The multi-source data acquisition unit is used to acquire historical load data, historical meteorological data, user electricity consumption behavior tag data, and calendar information data from the power dispatch automation system, meteorological data service interface, power marketing business system, and calendar database, respectively, and perform time alignment processing on the four types of data to form a unified structured dataset indexed by timestamp.

[0015] The time-series-meteorological coupled anomaly detection unit is used to map load time-series data into load embedding vectors through a load embedding network, and to map multivariate meteorological data into meteorological embedding vectors through a meteorological embedding network. The load embedding vectors and meteorological embedding vectors are concatenated to form a coupling vector. A pre-built coupling probability density model is used to calculate the probability value of this coupling vector, and an anomaly detection threshold is set. Perform anomaly detection;

[0016] When the probability value is not less than the preset anomaly detection threshold Do not mark when;

[0017] When the probability value is less than the preset anomaly detection threshold An exception marker is generated at the appropriate time;

[0018] The multimodal data quality assessment unit is used to generate a multidimensional quality score vector for each input data set consisting of load data, meteorological data, user electricity consumption behavior tag data, and calendar information data, based on the data missing rate, mutation statistics, and physical consistency index. ;

[0019] The frequency domain adaptive decomposition unit is used to respond to the load data quality score. state;

[0020] When load data quality score Less than the preset quality threshold At that time, the autocorrelation function is calculated for the historical load sequence. Based on the local maxima of the autocorrelation function, the basic period parameters of the trend component, periodic component and fluctuation component are determined. Based on the basic period parameters, the load sequence is subjected to adaptive frequency domain decomposition, and the trend component sequence, periodic component sequence and fluctuation component sequence are output.

[0021] When load data quality score Not less than the preset quality threshold At that time, skip the frequency domain decomposition step;

[0022] The causal graph attention feature extraction unit is used to learn a directed acyclic causal graph from the variables of the structured dataset through conditional independence test, and to perform graph attention calculation with causal constraints based on the directed acyclic causal graph. The attention weight is calculated only between node pairs connected by directed edges in the causal graph. The attention weight between node pairs not connected by directed edges is forcibly set to zero, and the causal attention feature representation is output.

[0023] The multi-scale temporal feature extraction unit employs a multi-branch parallel temporal convolutional network and a bidirectional long short-term memory network to perform multi-scale temporal feature extraction on the load subsequence or the original load sequence. Each branch is configured with causal dilated convolutions with different dilation coefficients to cover dependencies at different time scales, and outputs a multi-scale temporal feature representation.

[0024] It also includes a cross-modal dynamic gating fusion unit, a multi-scenario adaptive prediction unit, and a model online adaptive update unit;

[0025] The cross-modal dynamic gating fusion unit is communicatively connected to the causal graph attention feature extraction unit, the multi-scale temporal feature extraction unit, and the multi-modal data quality assessment unit, and is used to dynamically calculate the temporal feature gating weights based on the multi-dimensional quality scoring vector. and causal feature gating weights ,by and The multi-scale temporal feature representation and the causal attention feature representation are weighted and summed respectively to generate a fused feature representation;

[0026] The multi-scene adaptive prediction unit includes a scene classification subnetwork and multiple scene-specific prediction subnetworks;

[0027] The scenario classification subnetwork is used to determine the load behavior scenario category to which the current input data belongs based on the fusion feature representation, and selects the corresponding scenario-specific prediction subnetwork to perform forward calculation on the fusion feature representation according to the scenario category, and outputs the load prediction value sequence at each time point within the prediction period.

[0028] The model online adaptive update unit is used to maintain the prediction residual cache queue, perform trend testing on the residual sequence, and when a statistically significant positive trend is detected in the residual sequence, it performs online updates on the system parameters using the most recent incremental sample set.

[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0030] In this machine learning-based short-term power load forecasting system, this solution employs a full-link data quality assurance mechanism to internalize anomaly detection and quality assessment as core components. It utilizes coupled embedding modeling of load and meteorological data to detect hidden anomalies and outputs a global quality score, suppressing data disturbances from the source and mitigating their impact on forecasting. Simultaneously, a cross-modal dynamic gating fusion unit uses the quality score as a continuously modulated signal to generate dynamic fusion weights, achieving soft decision-making by enhancing high-quality features and suppressing low-quality features, ensuring the robustness of multi-source fusion. Finally, a three-dimensional parallel architecture is constructed using frequency domain adaptive decomposition, causal graph attention, and multi-scale temporal feature extraction. This automatically matches decomposition parameters and constrains attention to be calculated only on causal paths, eliminating spurious correlations. Multi-scenario adaptive forecasting and an online incremental update mechanism based on residual trend detection enable the system to autonomously adapt to differentiated power consumption scenarios and continuously evolve even as performance degrades, significantly improving forecast accuracy and long-term operational reliability. Attached Figure Description

[0031] Figure 1 This is a flowchart illustrating the overall method of the present invention; Detailed Implementation

[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0034] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0035] Please see Figure 1As shown, a short-term power load forecasting system based on machine learning is provided, including a multi-source data acquisition unit, a time-series-meteorological coupling anomaly detection unit, a multi-modal data quality assessment unit, a frequency domain adaptive decomposition unit, a causal graph attention feature extraction unit, and a multi-scale time-series feature extraction unit.

[0036] The multi-source data acquisition unit is used to acquire historical load data, historical meteorological data, user electricity consumption behavior tag data, and calendar information data from the power dispatch automation system, meteorological data service interface, power marketing business system, and calendar database, respectively, and performs time alignment processing on the four types of data to form a unified structured dataset indexed by timestamp.

[0037] The time-series-meteorological coupling anomaly detection unit maps load time-series data into load embedding vectors via a load embedding network, and meteorological multivariate data into meteorological embedding vectors via a meteorological embedding network. The load embedding vector and the meteorological embedding vector are concatenated to form a coupling vector. A pre-built coupling probability density model is used to calculate the probability value of this coupling vector, and an anomaly detection threshold is set. Perform anomaly detection;

[0038] When the probability value is not less than the preset anomaly detection threshold Do not mark when;

[0039] When the probability value is less than the preset anomaly detection threshold An exception marker is generated at the appropriate time;

[0040] The multimodal data quality assessment unit generates a multidimensional quality score vector for each input data set consisting of load data, meteorological data, user electricity consumption behavior tag data, and calendar information data, based on the missing data rate, mutation statistics, and physical consistency indicators. ;

[0041] Frequency domain adaptive decomposition unit is used for response load data quality scoring state;

[0042] When load data quality score Less than the preset quality threshold At that time, the autocorrelation function is calculated for the historical load sequence. Based on the local maxima of the autocorrelation function, the basic period parameters of the trend component, periodic component and fluctuation component are determined. Based on the basic period parameters, the load sequence is subjected to adaptive frequency domain decomposition, and the trend component sequence, periodic component sequence and fluctuation component sequence are output.

[0043] The causal graph attention feature extraction unit is used to learn a directed acyclic causal graph from the variables of a structured dataset through a conditional independence test, and to perform graph attention computation with causal constraints based on the directed acyclic causal graph. The attention weight is calculated only between node pairs connected by directed edges in the causal graph, and the attention weight between node pairs not connected by directed edges is forced to zero. The unit outputs a causal attention feature representation.

[0044] The multi-scale temporal feature extraction unit employs a multi-branch parallel temporal convolutional network and a bidirectional long short-term memory network to perform multi-scale temporal feature extraction on the load subsequence or the original load sequence. Each branch is configured with causal dilated convolutions with different dilation coefficients to cover dependencies at different time scales and outputs multi-scale temporal feature representations.

[0045] It also includes a cross-modal dynamic gating fusion unit, a multi-scenario adaptive prediction unit, and a model online adaptive update unit;

[0046] The cross-modal dynamic gating fusion unit communicates with the causal graph attention feature extraction unit, the multi-scale temporal feature extraction unit, and the multi-modal data quality assessment unit to dynamically calculate the temporal feature gating weights based on the multi-dimensional quality score vector. and causal feature gating weights ,by and The multi-scale temporal feature representation and the causal attention feature representation are weighted and summed respectively to generate a fused feature representation;

[0047] The multi-scene adaptive prediction unit includes a scene classification subnetwork and multiple scene-specific prediction subnetworks;

[0048] The scenario classification subnetwork is used to determine the load behavior scenario category to which the current input data belongs based on the fusion feature representation, and selects the corresponding scenario-specific prediction subnetwork to perform forward calculation on the fusion feature representation according to the scenario category, and outputs the load prediction value sequence at each time point within the prediction period.

[0049] The online adaptive update unit is used to maintain the prediction residual cache queue, perform trend tests on the residual sequence, and when a statistically significant positive trend is detected in the residual sequence, it performs online updates to the system parameters using the most recent incremental sample set.

[0050] The specific plan is as follows:

[0051] First, in order to ensure the subsequent forecasting work, it is necessary to collect load data of the current power area in advance. This solution obtains historical load data, historical meteorological data, user electricity consumption behavior tag data and calendar information data from the power dispatch automation system, meteorological data service interface, power marketing business system and calendar database through a multi-source data acquisition unit.

[0052] The load data includes a sequence of power load values ​​sampled at fixed time intervals.

[0053] Meteorological data includes multidimensional meteorological variable sequences such as temperature, humidity, wind speed, air pressure, precipitation, cloud cover, and solar radiation intensity.

[0054] User electricity consumption behavior tagging data includes user category tags, time-of-use electricity pattern tags, and industry type tags.

[0055] Calendar information data includes date type, season identifier, and holiday identifier.

[0056] After completing the data collection, in order to ensure the consistency of data time, the four types of data were time aligned to form a unified structured dataset indexed by timestamp.

[0057] In addition, in order to realize real-time detection of collected power load data and respond to abnormal data, this scheme constructs a time-series-meteorological coupled anomaly detection unit to map load time-series data into load embedding vectors through a load embedding network, and to map meteorological multivariate data into meteorological embedding vectors through a meteorological embedding network.

[0058] Because there is an inherent physical correlation between electricity load and meteorological conditions—for example, high temperatures drive up air conditioning load, and low cloud cover and high radiation affect photovoltaic power generation output, thus altering the net load pattern—when a set of load data and a set of meteorological data deviate from the statistical norm of this inherent correlation in the embedding space, it indicates that at least one type of data may be anomaly. Based on this, this scheme concatenates the load embedding vector and the meteorological embedding vector to form a coupling vector, and uses a pre-built coupling probability density model to calculate the probability value of this coupling vector. The specific method is as follows:

[0059] The load embedding network considers the temporal characteristics of load data. Since load sequences are typical time-series signals, their local patterns, such as rising edges, falling edges, peaks, valleys, and fluctuation frequencies, contain key information for judging data quality. Therefore, this unit uses a temporal convolutional sub-network composed of three stacked one-dimensional convolutional layers as the load embedding network. The kernel sizes of the three convolutional layers are configured as 7, 5, and 3, respectively, with a stride of 1 for each. The first layer uses a larger kernel size of 7, which can capture the trend change patterns over a longer span in the load sequence, such as a load ramp-up or ramp-down process lasting several hours. The kernel size of the second layer is reduced to 5, focusing on periodic features over a medium time span, such as the alternating rhythm of morning and evening peaks. The kernel size of the third layer is further reduced to 3, extracting short-term local fluctuations and detailed texture information.

[0060] By stacking three convolutional layers, the embedded network abstracts multi-scale temporal feature representations of the load sequence layer by layer, from local details to global trends. Simultaneously, the activation functions of each layer employ... The function, with its slight slope on the negative half-axis, ensures the complete flow of information across all dimensions in the embedded vector, avoiding the traditional... The problem of loss of embedding space information caused by the function completely truncating the gradient when the input is negative.

[0061] Correspondingly, the design of the meteorological embedding network considers the multi-vector characteristics of meteorological data. Meteorological data consists of multiple variables such as temperature, humidity, wind speed, air pressure, precipitation, cloud cover, and solar radiation intensity, forming a multi-dimensional vector at the same time point. Furthermore, there are physical coupling relationships between these meteorological variables; for example, high temperatures are often accompanied by high radiation, and low air pressure is often accompanied by high humidity and precipitation. Therefore, in this scheme, the meteorological embedding network uses a fully connected feedforward sub-network with two hidden layers. The number of neurons in the two hidden layers is configured as 128 and 64, respectively. This forces the network to learn the nonlinear and high-order interaction features between meteorological variables in the 128-dimensional intermediate representation, and further extracts the essential meteorological features most relevant to load behavior in the 64-dimensional space, ultimately mapping them to a meteorological embedding vector of the same dimension as the load embedding vector. The activation function also adopts... function.

[0062] Dimensions of load embedding vector and meteorological embedding vector The values ​​are uniform, ranging from 32 to 128. During the training phase, a joint training strategy is employed, where a decoding and reconstruction module is added after the load embedding network and the meteorological embedding network. This module takes the concatenated vectors of the load embedding vector and the meteorological embedding vector as input to reconstruct the original load sequence and the original meteorological vector, respectively. The training loss function consists of a weighted sum of the load reconstruction error and the meteorological reconstruction error.

[0063] After the training phase, each pair of load sequences and meteorological vectors from all normal historical operational data are input into the corresponding embedding network to obtain load embedding vectors and meteorological embedding vectors. Each pair of embedding vectors is then concatenated along the dimensional direction to form a vector with dimension [missing information]. The coupling vectors, all Group training samples co-generated A coupling vector, using Each coupling vector is fitted with a dimension of... Determine the multivariate Gaussian distribution and calculate the parameters of the multivariate Gaussian distribution, specifically including the mean vector. and covariance matrix , where the mean vector It is calculated by the dimension-wise arithmetic mean of all coupled vectors, and the corresponding algorithm formula is as follows:

[0064] ;

[0065] Let be the coupling vector of the i-th training sample.

[0066] The calculated mean vector and covariance matrix Persistent storage is used as a coupled probability density model in the online anomaly detection phase.

[0067] In the specific anomaly detection process, for each prediction period, the input data pair to be detected includes a set of load time series data and a corresponding set of meteorological multivariate data. The load time series data and the meteorological multivariate data are mapped to load embedding vectors and meteorological embedding vectors, and the two are concatenated along the dimensional direction to form the current coupling vector. , coupling vector Substitute into the probability density function of the multivariate Gaussian distribution and calculate the corresponding probability value. The specific function expression is as follows:

[0068] ;

[0069] in The determinant of the covariance matrix. This represents the inverse of the covariance matrix.

[0070] And set an anomaly detection threshold. Perform anomaly detection and calculate all. The probability values ​​of the coupling vectors of each training sample under this multivariate Gaussian distribution will be... The probability values ​​are sorted in ascending order, and the 5th percentile is taken as the anomaly detection threshold. The initial setting value;

[0071] When the probability value ≥ Anomaly detection threshold When the current input data pair's coupled embedding representation falls within the distribution range of normal data, and the combination pattern of load and weather is consistent with historical normal behavior, the system determines that the input data pair is normal and does not generate an anomaly flag.

[0072] When the probability value <Preset anomaly detection threshold When the current input data pair deviates from the distribution area of ​​normal data in the coupled embedding space, it is determined that the input data pair is abnormal, an anomaly marker is generated, the timestamp and probability value of the abnormal event are recorded, and the anomaly marker is transmitted to the multimodal data quality assessment unit.

[0073] After completing the anomaly labeling, the multimodal data quality assessment unit receives data from the multi-source data acquisition unit and anomaly labels from the time-series-meteorological coupled anomaly detection unit, and generates a multidimensional quality score vector for each input data set consisting of load data, meteorological data, user electricity consumption behavior tag data, and calendar information data. ;

[0074] in The specific algorithm formula for scoring the quality of workload data is as follows:

[0075] ;

[0076] in This indicates the missing rate of workload data, and ,in This indicates the number of missing load sampling points in the current input time series. This indicates the total number of sampling points in the sequence. This represents the mutation detection statistic, and ,in , representing the load difference between the i-th sampling point and the previous sampling point. and Let these represent the mean and standard deviation of the historical normal load difference series, respectively. This indicates the preset mutation threshold. This indicates an indicator function; it takes the value 1 if the condition within the parentheses is true, and 0 otherwise. and This represents the penalty coefficient for the missing rate and the penalty coefficient for the mutation.

[0077] in, The specific algorithm formula for representing the quality score of meteorological data is as follows:

[0078] ;

[0079] in Indicates the rate of missing meteorological data, and , This represents the total number of missing meteorological variable-time point pairs. This represents the total number of meteorological variables (temperature, humidity, wind speed, air pressure, precipitation, cloud cover, solar radiation intensity, etc.). The number of time steps contained in the sequence. This represents the physical consistency test value, and , as well as These represent the penalty coefficients for meteorological missing rate and physical consistency, respectively.

[0080] The algorithm formula for representing the quality score of user electricity consumption behavior tag data is as follows:

[0081] ;

[0082] in , representing the missing label data rate. This indicates the number of user records with missing tags. Indicates the total number of user records. This represents the label consistency detection value, and , indicating user The tags exhibit logical conflicts. Logical conflicts refer to contradictions in the tags of the same user at different times or in different fields (such as inconsistencies between the industry type tag and the electricity consumption mode tag). This indicator is detected based on a predefined conflict rule base. and These represent the penalty coefficient for missing labels and the penalty coefficient for label consistency, respectively.

[0083] The algorithm for calculating the quality score of calendar information data is as follows:

[0084] ;

[0085] This represents the calendar data integrity check value. This indicates the number of missing or invalid calendar fields (such as date, day of the week, holiday identifier, etc., where null or illegal values ​​are present). Indicates the total number of calendar fields required. This represents the calendar integrity penalty coefficient in the above algorithms. Indicates that the operation is truncated to The interval is 0, which means the data is completely reliable, and 1 means it is completely unreliable.

[0086] Furthermore, the frequency domain adaptive decomposition unit is used for response load data quality scoring. state;

[0087] When load data quality score Less than the preset quality threshold At that time, the autocorrelation function is calculated for the historical load series. The specific calculation method is as follows:

[0088] First, a historical load sequence is given. Its length is The sequence at different lag orders The autocorrelation coefficient is calculated using the following formula:

[0089] ;

[0090] in Indicates the lag order as The autocorrelation coefficient at time , the range of values ​​is , Indicates the historical load sequence at time [time]. The load value, Representation of time interval Load values ​​for each step size, It is expressed as the arithmetic mean of the entire historical load series. This represents the preset maximum lag order, which can be taken as half the length of the historical sequence. The autocorrelation function R consists of all lag orders and their corresponding autocorrelation coefficients. .

[0091] Extract the lag orders corresponding to the three largest local maxima in the autocorrelation function, and sort them in ascending order. as well as ,Right now > > The basic periodic parameters of the trend component, periodic component, and fluctuation component are determined based on the local maxima of the autocorrelation function. The basic cycle as a trend component The fundamental period of the periodic component. Using the fundamental period as the basis for the fluctuation components, and performing adaptive frequency domain decomposition on the load sequence based on the fundamental period parameter, the trend component sequence, periodic component sequence, and fluctuation component sequence are output, as shown in the following expressions:

[0092] ;

[0093] in This represents a trend component sequence, corresponding to long-term trend changes. This represents a periodic component sequence, corresponding to periodic patterns such as daily cycles and weekly cycles. This represents a sequence of fluctuation components, corresponding to short-term random fluctuations. This represents the residual term, which typically has a very small or zero magnitude.

[0094] When load data quality score Not less than the preset quality threshold At that time, the frequency domain decomposition step is skipped, and the original load data is directly transferred to the multi-scale time series feature extraction unit.

[0095] Furthermore, in order to obtain the relationship between various data and power load, the causal graph attention feature extraction unit receives the structured dataset output by the multi-source data acquisition unit, extracts load variables, meteorological variables and calendar variables from the dataset as a set of node variables, learns the causal relationship graph structure between the node variables, and performs graph attention calculation of causal constraints based on the causal relationship graph structure.

[0096] The learning method for causal relationship graph structures is as follows:

[0097] For each pair of variables in the set of node variables and Given all other nodes, perform a conditional independence test:

[0098] If the test result is and If statistical dependencies still exist given all other nodes, then in a causal graph, this is: and Add an undirected edge between them;

[0099] Then, the undirected edges are oriented using time sequence constraints, that is, variables that occur earlier in time are directed to variables that occur later in time, thus transforming the undirected edges into directed edges, ultimately resulting in a directed acyclic causal graph. , where nodes represent variables and directed edges represent direct causal relationships between variables.

[0100] The specific process for calculating the graph attention corresponding to the causal constraints is as follows:

[0101] For cause-effect graphs For each node in the algorithm, a learnable linear transformation is used to map the node features into a query vector, a key vector, and a value vector.

[0102] compute nodes With nodes When considering attention weights between nodes, first check if there exists a node in the causal graph. Pointing to node Or from the node Pointing to node Directed edges:

[0103] If directed edges exist, the attention score is calculated using the conventional attention mechanism. ;

[0104] If no directed edge exists, the attention score is forcibly set to negative infinity. The attention weights are zero after normalization.

[0105] The formula for calculating attention weights is:

[0106] When node With nodes When there are directed edges in a causal graph: ,in and They are nodes and nodes The input feature vector has a dimension of , To query the weight matrix, the dimension is , The key weight matrix has dimensions of . , This represents the vector dot product operation. To determine the dimensions of the query vector and key vector, i.e., the dimensions of the attention head, This represents the scaling factor, used to prevent the dot product value from becoming too large as the vector dimension increases;

[0107] When node With nodes When there are no directed edges in the cause-effect graph: ;

[0108] right application The function obtains normalized attention weights. Pay attention weight AND value vector Weighted summation yields the nodes The output representation vector ,Right now ,in Represents a node In causality diagram The causal neighbor set, Represents the value weight matrix, Represents a node The input feature vector.

[0109] Specifically, in order to extract time-series features, this scheme uses a multi-scale time-series feature extraction unit to receive the load subsequence output by the frequency domain adaptive decomposition unit and the structured dataset output by the multi-source data acquisition unit, and performs multi-scale time-series feature extraction. The specific scheme is as follows:

[0110] First, multi-scale temporal feature extraction employs a multi-branch parallel temporal convolutional network structure, with the following specific configuration:

[0111] The first branch is configured with a causal dilated convolution with a kernel size of 3 and a dilation coefficient of 1, and the receptive field covers short-term dependencies.

[0112] The second branch configures a causal dilated convolution with a kernel size of 3 and a dilation coefficient of 2, which covers the dependencies at medium time scales in the receptive field.

[0113] The third branch configures a causal dilated convolution with a kernel size of 3 and a dilation coefficient of 4, and the receptive field covers long-term dependencies.

[0114] After each of the three branches undergoes a convolutional operation, a bidirectional long short-term memory network is used to perform temporal modeling on the outputs of each branch. The outputs of the three branches are concatenated along the feature dimension and mapped to a unified feature dimension through a fully connected projection layer, resulting in a multi-scale temporal feature representation. When the frequency domain adaptive decomposition unit outputs a trend component sequence, a periodic component sequence, and a fluctuation component sequence, the three branches respectively process the three sub-sequences: the first branch processes the fluctuation component sequence. The second branch processes the periodic component sequence. The third branch processes the trend component sequence. This is to achieve a correspondence between subsequence features and receptive field scales. The specific calculation method is as follows:

[0115] First, assume that the output state of the frequency domain adaptive decomposition unit is determined by an indicator variable. mark:

[0116] Frequency domain decomposition has been performed, and the output consists of three subsequences—trend component sequences. Periodic component sequences and fluctuation component sequences ;

[0117] Frequency domain decomposition was not performed; the output is the original load sequence L.

[0118] The input sequence of the b-th branch The corresponding output rules are as follows:

[0119] ;

[0120] ;

[0121] ; .

[0122] Secondly, regarding branches Calculate the output scalar of the causal dilated convolution at time t. ,Right now:

[0123] ;

[0124] in, This is the kernel size; all three branches use the same configuration. For branches The expansion coefficient, i.e., the first branch , Indicates branch The learnable weight parameters at position s, Indicates branch The input sequence at time... elements, Indicates branch The bias term of the convolutional layer performs the above convolution operation at each time t in the sequence to obtain the branch. Convolutional output sequence ;

[0125] The convolution output sequence is then processed. Input a bidirectional long short-term memory network and perform temporal encoding simultaneously from both the forward and reverse directions.

[0126] Hidden state of a positive long short-term memory network at time t ;

[0127] Hidden state of the reverse long short-term memory network at time t ;

[0128] The bidirectional output at time t is obtained by concatenating the forward and reverse hidden states, i.e. For the entire sequence Perform the above operation at each time point, and take the last time point. bidirectional output as a branch The temporal characteristics representation, i.e. ,in Representing a dimension as The vector, For branches The dimensions of hidden units in Long Short-Term Memory (LSTM) networks.

[0129] Finally, the temporal feature representations of the three branches are concatenated along the feature dimensions to form a fused multi-scale feature vector. ,Right now and multi-scale feature vectors Mapped to a unified feature dimension through a fully connected projection layer. To obtain the final multi-scale temporal feature representation ,Right now ,in This represents the bias vector of the fully connected projection layer. This represents the learnable weight matrix of the fully connected projection layer. Represents a nonlinear activation function, using function.

[0130] In addition, the cross-modal dynamic gating fusion unit receives the causal attention feature representation output by the causal graph attention feature extraction unit, the multi-scale temporal feature representation output by the multi-scale temporal feature extraction unit, and the multi-dimensional quality score vector output by the multi-modal data quality assessment unit. The specific functional implementation process is as follows:

[0131] First, the gating weights are calculated, and the multidimensional quality score vector is... middle as well as The average value is used as a quality index of load-related modes. ,Will as well as The average value is used as a quality index for auxiliary modes. ,according to and Gating weights for calculating time series features Gating weights for causal features The specific algorithm formula is as follows:

[0132] ;

[0133] ;

[0134] in Represents the time-series characteristic temperature coefficient, controlling Quality indicators for load-related modes Sensitivity Represents the causal characteristic temperature coefficient, controlling Quality indicators for auxiliary modes Sensitivity It is a natural exponential function, and .

[0135] Subsequently, multi-scale temporal feature representation Representation of causal attention features The calculated gating weights are weighted and summed to generate a fused feature representation. :

[0136] ;

[0137] Finally, the cross-modal dynamic gating fusion unit fuses the feature representations. The data is sent to a multi-scene adaptive prediction unit, which includes a scene classification subnetwork and multiple scene-specific prediction subnetworks.

[0138] The scene classification subnetwork receives the fused feature representation output by the cross-modal dynamic gating fusion unit. Based on this fusion feature representation, the system determines the load behavior scenario to which the current input data belongs. Scenario categories include, but are not limited to, weekday scenarios, weekend scenarios, and holiday scenarios. The scenario classification subnetwork consists of a two-layer fully connected subnetwork and a... The classification layer consists of two fully connected subnetworks with hidden layer neurons configured to have 64 and 32 neurons respectively. The activation function used is... function, The classification layer outputs the probability of belonging to each scene category, and the specific calculation method is as follows:

[0139] First, input the fused feature representation. ;

[0140] The corresponding first fully connected layer (64 neurons) Function activation): ;

[0141] The corresponding second fully connected layer (32 neurons) Function activation): ;

[0142] Third-level classification output ( (Probability of belonging to 3 types of scenarios) ;

[0143] in , This represents the weight matrix of the first fully connected layer. This represents the first-level bias vector. This represents the output hidden vector of the first layer. This represents the weight matrix of the second fully connected layer. This represents the second-layer bias vector. This represents the output hidden vector of the second layer. Represents the classification layer weight matrix. This represents the bias vector for the classification layer. This represents the probability vector of scene affiliation.

[0144] The multi-scene adaptive prediction unit is configured with three scene-specific prediction subnetworks, corresponding to weekday, weekend, and holiday scenarios, respectively. Each scene-specific prediction subnetwork consists of a three-layer fully connected subnetwork, with the number of hidden layer neurons configured to be 128, 64, and 32, respectively. The activation function is […]. The function outputs a sequence of predicted load values ​​for each time point within the prediction period. During prediction execution, based on the assignment probability output by the scene classification sub-network, the scene-specific prediction sub-network corresponding to the scene category with the highest assignment probability is selected to fuse the feature representation. A forward calculation is performed to generate a sequence of load forecast values. The specific calculation method is as follows:

[0145] The system is configured with three identical but parameter-independent scene-specific prediction sub-networks, corresponding to weekday scenes (s=1), weekend scenes (s=2), and holiday scenes (s=3), respectively. Each network performs forward computation in the following layers, fusing feature representations. Input to multi-scene adaptive prediction unit:

[0146] The first fully connected layer (128 neurons) Function activation): ;

[0147] The second fully connected layer (64 neurons) Function activation): ;

[0148] The third fully connected layer (32 neurons) Function activation): ;

[0149] Output layer (predicted load value sequence, no activation function or linear output): ;

[0150] in, Indicates scene index, These correspond to weekdays, weekends, and public holidays, respectively. Indicates the first The first layer weight matrix of the network for each scenario, Indicates the first The bias vector of the first layer of the network for each scene Indicates the first The first layer hidden vector of the network for each scene, Indicates the first The second-layer weight matrix of the network for each scenario Indicates the first The second layer bias vector of the network for each scene No. The second-layer hidden vector of the network for each scene Indicates the first The third layer weight matrix of the network for each scenario Indicates the first The bias vector of the third layer of the network for each scenario Indicates the first The third layer hidden vector of the network for each scene Indicates the first Each scenario's network output layer weight matrix, Indicates the first Each scene's network output layer bias vector Indicates the first The load value sequence of the network in each scenario.

[0151] Finally, the attribution probability vector output by the scene classification subnetwork is used as the basis for classification. The scene category with the highest probability is selected as the scene determination result for the current input data.

[0152] Furthermore, the online adaptive update unit communicates with the multi-scenario adaptive prediction unit, the cross-modal dynamic gating fusion unit, and the multimodal data quality assessment unit to maintain the prediction residual cache queue, perform trend checks on the residual sequence, and when a statistically significant positive trend is detected in the residual sequence, the system parameters are updated online using the most recent incremental sample set. The specific implementation method is as follows:

[0153] First, after each complete forecast cycle, i.e., at the end of the forecast period, the actual load value sequence within that period is obtained. This sequence is then compared point by point with the load forecast value sequence previously output by the system. The root mean square error of the forecast residual for that forecast cycle is calculated. Let the forecast period contain a total of [number missing] days. The predicted value sequence for the nth prediction period at each time point is: The corresponding actual load value sequence is The corresponding root mean square error of the prediction residual The system maintains a fixed-capacity first-in-first-out prediction residual cache queue. The queue capacity is configured as follows: This corresponds to the number of residual records for each hour within a week. Each queue element is a tuple containing the timestamp for that prediction period. and the corresponding root mean square error of the residuals ,Right now When a new record arrives, if the queue is full, the oldest record is removed from the front of the queue, and the new record is inserted at the rear of the queue, ensuring that the queue always keeps the most recent record. The residual information for each period. This first-in-first-out mechanism ensures that residual analysis is always based on the latest and limited historical window, avoiding interference from outdated data in trend determination and ensuring the statistical validity of the test sample size.

[0154] At the same time, the system every interval Residual trend detection is automatically triggered once per prediction period, and a trend detection time queue is set. The residual sequence in is , , Arranged chronologically by timestamp, the Mann-Kendall trend test was used to detect the statistical significance of monotonicity in the sequence. The Mann-Kendall test statistic S was calculated as follows:

[0155] ;

[0156] in It is a symbolic function, and When the statistic S > 0, it indicates that the sequence has an upward trend; when S < 0, it indicates that the sequence has a downward trend. The larger the absolute value of S, the more obvious the trend.

[0157] corresponding Statistics from Normalization yields, i.e. When the sample size When large enough ( >10), the statistic S approximately follows a normal distribution with a mean of zero and a variance given by the following formula:

[0158] ;

[0159] in This indicates the number of groups with the same value in the sequence. Indicates the first The number of identical values ​​in a group, and the final standardized test statistic. Represented as:

[0160] ;

[0161] Furthermore, the update trigger condition requires that two conditions be met simultaneously: the residual sequence exhibits a positive trend, and the statistical significance of the trend reaches a preset threshold. The specific judgment logic is as follows:

[0162] Trigger Update and ,in This indicates the preset significance threshold, with a default value of 0.05;

[0163] The physical meaning of the above dual conditions is as follows: The residual direction was confirmed to be increasing rather than decreasing or remaining flat, ruling out the possibility that the model performance improved or fluctuated but did not deteriorate. It is necessary to ensure that the upward trend is not a random phenomenon caused by fluctuations in random sampling, but a statistically significant systematic degradation.

[0164] When no positive trend is detected in the residual sequence, or the positive trend is not statistically significant, the system remains in normal operation and does not trigger the update process. The predicted residual cache queue continues to record new residuals and evict old residuals, waiting for evaluation in the next detection cycle.

[0165] Once the online adaptive update process of the model is triggered, the system automatically performs the following sequence of operations:

[0166] Constructing an incremental training sample set: Extract data from the system data storage for the most recent W=72 prediction periods prior to the trigger time as incremental training samples. For each prediction period, the sample consists of two parts—the input part is all the input data used in the prediction for that period, including historical load sequences, meteorological data, user electricity consumption behavior tag data, and calendar information data; the tag part is the sequence of actual load values ​​obtained afterward within that period. A total of W sets of input-tag samples constitute the incremental training sample set. .

[0167] Configure incremental training hyperparameters: learning rate for incremental training Set as the learning rate during the initial training phase of the system. 0.1 times.

[0168] Parameter update: Utilize incremental training sample set With learning rate Iterative optimization is performed on all trainable parameters in the system. The range of trainable parameters covers all learning units throughout the system, specifically including: the weight parameters of the load embedding network and meteorological embedding network in the time-series-meteorological coupled anomaly detection unit; the parameters of the query weight matrix, key weight matrix, and value weight matrix in the causal graph attention feature extraction unit; the kernel weights and biases of the causal dilation convolutions in each branch of the multi-scale time-series feature extraction unit; all weights and biases of the bidirectional long short-term memory network; the weights and biases of the fully connected projection layer; and the temperature coefficient of the cross-modal dynamic gating fusion unit. and The system includes all weights and biases of the scene classification subnetwork and three scene-specific prediction subnetworks in the multi-scene adaptive prediction unit. The number of iteration rounds is set to 5 to 10, with a default of 5 rounds. Mini-batch stochastic gradient descent or its adaptive variant optimizer is used for parameter updates, traversing the incremental sample set once per round.

[0169] Post-update processing: After the incremental parameter update is complete, the system clears the prediction residual cache queue. All records are processed to prevent residuals from the old model before the update and those from the new model after the update from being mixed in the queue, thus avoiding interference from mixed residual data across model versions in subsequent trend detection. The system then resumes the normal online prediction and residual recording process, using the updated parameters to perform subsequent prediction tasks. The online model update mechanism enables the system to automatically adapt to the slow changes in the operating environment and data distribution during continuous operation, maintaining high prediction accuracy without the need for manual retraining.

[0170] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A short-term power load forecasting system based on machine learning, comprising a multi-source data acquisition unit, a time-series-meteorological coupling anomaly detection unit, a multi-modal data quality assessment unit, a frequency domain adaptive decomposition unit, a causal graph attention feature extraction unit, and a multi-scale time-series feature extraction unit; The multi-source data acquisition unit is used to acquire historical load data, historical meteorological data, user electricity consumption behavior tag data, and calendar information data from the power dispatch automation system, meteorological data service interface, power marketing business system, and calendar database, respectively, and perform time alignment processing on the four types of data to form a unified structured dataset indexed by timestamp. The time-series-meteorological coupled anomaly detection unit is used to map load time-series data into load embedding vectors through a load embedding network, and to map multivariate meteorological data into meteorological embedding vectors through a meteorological embedding network. The load embedding vectors and meteorological embedding vectors are concatenated to form a coupling vector. A pre-built coupling probability density model is used to calculate the probability value of this coupling vector, and an anomaly detection threshold is set. Perform anomaly detection; When the probability value is not less than the preset anomaly detection threshold Do not mark when; When the probability value is less than the preset anomaly detection threshold An exception marker is generated at the appropriate time; The multimodal data quality assessment unit is used to generate a multidimensional quality score vector for each input data set consisting of load data, meteorological data, user electricity consumption behavior tag data, and calendar information data, based on the data missing rate, mutation statistics, and physical consistency index. ; The frequency domain adaptive decomposition unit is used to respond to the load data quality score. state; When load data quality score Less than the preset quality threshold At that time, the autocorrelation function is calculated for the historical load sequence. Based on the local maxima of the autocorrelation function, the basic period parameters of the trend component, periodic component and fluctuation component are determined. Based on the basic period parameters, the load sequence is subjected to adaptive frequency domain decomposition, and the trend component sequence, periodic component sequence and fluctuation component sequence are output. When load data quality score Not less than the preset quality threshold At that time, skip the frequency domain decomposition step; The causal graph attention feature extraction unit is used to learn a directed acyclic causal graph from the variables of the structured dataset through conditional independence test, and to perform graph attention calculation with causal constraints based on the directed acyclic causal graph. The attention weight is calculated only between node pairs connected by directed edges in the causal graph. The attention weight between node pairs not connected by directed edges is forcibly set to zero, and the causal attention feature representation is output. The multi-scale temporal feature extraction unit employs a multi-branch parallel temporal convolutional network and a bidirectional long short-term memory network to perform multi-scale temporal feature extraction on the load subsequence or the original load sequence. Each branch is configured with causal dilated convolutions with different dilation coefficients to cover dependencies at different time scales, and outputs a multi-scale temporal feature representation. Its features include: a cross-modal dynamic gating fusion unit, a multi-scene adaptive prediction unit, and a model online adaptive update unit; The cross-modal dynamic gating fusion unit is communicatively connected to the causal graph attention feature extraction unit, the multi-scale temporal feature extraction unit, and the multi-modal data quality assessment unit, and is used to dynamically calculate the temporal feature gating weights based on the multi-dimensional quality score vector. and causal feature gating weights ,by and The multi-scale temporal feature representation and the causal attention feature representation are weighted and summed respectively to generate a fused feature representation; The multi-scene adaptive prediction unit includes a scene classification subnetwork and multiple scene-specific prediction subnetworks; The scenario classification subnetwork is used to determine the load behavior scenario category to which the current input data belongs based on the fusion feature representation, and selects the corresponding scenario-specific prediction subnetwork to perform forward calculation on the fusion feature representation according to the scenario category, and outputs the load prediction value sequence at each time point within the prediction period. The model online adaptive update unit is used to maintain the prediction residual cache queue, perform trend testing on the residual sequence, and when a statistically significant positive trend is detected in the residual sequence, it performs online updates on the system parameters using the most recent incremental sample set.

2. The short-term power load forecasting system based on machine learning according to claim 1, characterized in that, The method for anomaly determination in the time-series-meteorological coupled anomaly detection unit includes the following steps: S10. For each forecast period, the input data pair to be detected includes a set of load time series data and a set of corresponding meteorological multivariate data. S11. Map the load time-series data and meteorological multivariate data into load embedding vectors and meteorological embedding vectors, and concatenate the two along the dimensional direction to form the current coupling vector. ; S12, Coupling Vector Substitute into the probability density function of the multivariate Gaussian distribution and calculate the corresponding probability value. ; S13. Set the anomaly detection threshold. Perform anomaly detection; When the probability value ≥ Anomaly detection threshold When the current input data pair's coupled embedding representation falls within the distribution range of normal data, and the combination pattern of load and weather is consistent with historical normal behavior, the system determines that the input data pair is normal and does not generate an anomaly flag. When the probability value <Preset anomaly detection threshold When this occurs, it indicates that the current input data pair deviates from the distribution area of ​​normal data in the coupled embedding space, and the input data pair is determined to be abnormal, and an abnormality marker is generated.

3. The short-term power load forecasting system based on machine learning according to claim 1, characterized in that, The method for generating a multidimensional quality score vector in the multimodal data quality assessment unit includes the following steps: S20. Calculate the load data quality score. The specific algorithm formula is as follows: ; in This indicates the missing rate of workload data, and ,in This indicates the number of missing load sampling points in the current input time series. This indicates the total number of sampling points in the sequence. This represents the mutation detection statistic, and ,in , representing the load difference between the i-th sampling point and the previous sampling point. and Let these represent the mean and standard deviation of the historical normal load difference series, respectively. This indicates the preset mutation threshold. Indicates an indicator function, and This represents the penalty coefficient for deletion rate and the penalty coefficient for mutation rate; S21. Calculate the meteorological data quality score. The specific algorithm formula is as follows: ; in Indicates the rate of missing meteorological data, and , This represents the total number of missing meteorological variable-time point pairs. This represents the total number of meteorological variables. The number of time steps contained in the sequence. This represents the physical consistency test value, and , as well as These represent the penalty coefficient for meteorological missing rate and the penalty coefficient for physical consistency, respectively. S22. Calculate the quality score of user electricity consumption behavior tag data. The specific algorithm formula is as follows: ; in This indicates the missing label data rate. This indicates the number of user records with missing tags. Indicates the total number of user records. This represents the label consistency detection value, and , indicating user The tags exhibit logical conflicts. Logical conflicts refer to contradictions in the tags of the same user at different times or in different fields (such as inconsistencies between the industry type tag and the electricity consumption mode tag). This indicator is detected based on a predefined conflict rule base. and These represent the penalty coefficient for missing labels and the penalty coefficient for label consistency, respectively. S23. Calculate the quality score of calendar information data. The specific algorithm formula is as follows: ; This represents the calendar data integrity check value. Indicates the number of missing or invalid calendar fields. Indicates the total number of calendar fields required. This represents the calendar integrity penalty coefficient in the above algorithms. Indicates operation truncation to The interval is defined as follows: 0 indicates that the data is completely reliable, and 1 indicates that it is completely unreliable. S24. Summary Load Data Quality Score Meteorological data quality score User electricity consumption behavior tag data quality score and calendar information data quality score Construct a multidimensional quality score vector .

4. The short-term power load forecasting system based on machine learning according to claim 1, characterized in that, The steps for calculating the autocorrelation function of the historical load sequence in the frequency domain adaptive decomposition unit are as follows: S30, Given historical load sequence Its length is ; S31. Calculate historical load sequences autocorrelation coefficient The specific algorithm is as follows: ; in Indicates the lag order as The autocorrelation coefficient at time , the range of values ​​is , Indicates the historical load sequence at time [time]. The load value, Representation of time interval Load values ​​for each step size, It is expressed as the arithmetic mean of the entire historical load series. This indicates the preset maximum lag order; S32. Based on the autocorrelation coefficient The autocorrelation function R is obtained as follows: 。 5. The short-term power load forecasting system based on machine learning according to claim 1, characterized in that, The method for obtaining the directed acyclic causal graph in the causal graph attention feature extraction unit is as follows: S40. For each pair of variables in the set of node variables... and ; S41. For each pair of variables and Independence of execution conditions test: If the test result is and If statistical dependencies still exist given all other nodes, then in a causal graph, this is: and Add an undirected edge between them; S42. Using time sequence constraints, orient the undirected edges to transform them into directed edges, resulting in a directed acyclic causal graph. , where nodes represent variables and directed edges represent direct causal relationships between variables.

6. The short-term power load forecasting system based on machine learning according to claim 5, characterized in that, The method for calculating graph attention that performs causal constraints in the causal graph attention feature extraction unit is as follows: S401, Regarding cause-effect diagrams For each node in the algorithm, a learnable linear transformation is used to map the node features into a query vector, a key vector, and a value vector. S402, Computation Node With nodes When considering the attention weights between points, determine the existence of directed edges: If directed edges exist, the attention score is calculated using the conventional attention mechanism. The specific algorithm formula is as follows: ,in and They are nodes and nodes The input feature vector has a dimension of , To query the weight matrix, the dimension is , The key weight matrix has dimensions of . , This represents the vector dot product operation. To query the dimensions of the vector and key vector, Indicates the scaling factor; If no directed edge exists, the attention score is forcibly set to negative infinity, i.e. ; S403, to application The function obtains normalized attention weights. Pay attention weight AND value vector Weighted summation yields the nodes The output representation vector ,Right now ,in Represents a node In causality diagram The causal neighbor set, Represents the value weight matrix. Represents a node The input feature vector.

7. The short-term power load forecasting system based on machine learning according to claim 1, characterized in that, The method for outputting multi-scale temporal feature representations in the multi-scale temporal feature extraction unit includes the following steps: S50 employs a multi-branch parallel temporal convolutional network structure, with the following specific configuration: The first branch is configured with a causal dilated convolution with a kernel size of 3 and a dilation coefficient of 1, and the receptive field covers short-term dependencies. The second branch configures a causal dilated convolution with a kernel size of 3 and a dilation coefficient of 2, which covers the dependencies at medium time scales in the receptive field. The third branch is configured with a causal dilated convolution with a kernel size of 3 and a dilation coefficient of 4, and the receptive field covers long-term dependencies. S51. Assume the output state of the frequency domain adaptive decomposition unit is determined by an indicator variable. mark: Frequency domain decomposition has been performed, and the output consists of three subsequences—trend component sequences. Periodic component sequences and fluctuation component sequences ; Frequency domain decomposition was not performed; the output is the original load sequence L, where the input sequence of the b-th branch is... The corresponding output rules are as follows: ; ; ; S52, Regarding branches Calculate the output scalar of the causal dilated convolution at time t. ,Right now: ; in, This is the kernel size; all three branches use the same configuration. For branches The expansion coefficient, i.e., the first branch , Indicates branch The learnable weight parameters at position s, Indicates branch The input sequence at time... elements, Indicates branch Bias terms of convolutional layers; S53. Perform the above convolution operation on each time step t in the sequence to obtain the branch. Convolutional output sequence ; S54, then the convolutional output sequence Input a bidirectional long short-term memory network and perform temporal encoding simultaneously from both the forward and reverse directions to obtain the hidden state of the forward long short-term memory network at time t. And the hidden state of the inverse long short-term memory network at time t. ; S55, Bidirectional output at time t ; S56, For the entire sequence At each time point, execute steps S51-S55 to retrieve the last time point. bidirectional output as a branch The temporal characteristics representation, i.e. ; S57. Concatenate the temporal feature representations of the three branches along the feature dimension to form a fused multi-scale feature vector. and multi-scale feature vectors Mapped to a unified feature dimension through a fully connected projection layer. To obtain the final multi-scale temporal feature representation ,in This represents the bias vector of the fully connected projection layer. This represents the learnable weight matrix of the fully connected projection layer. This represents a non-linear activation function.

8. The short-term power load forecasting system based on machine learning according to claim 1, characterized in that, The method for generating fused feature representations in the cross-modal dynamic gating fusion unit includes the following steps: S60, Multidimensional quality scoring vector middle as well as The average value is used as a quality index of load-related modes. ,Will as well as The average value is used as a quality index for auxiliary modes. ,according to and Gating weights for calculating time series features Gating weights for causal features ; S61. Representing multi-scale temporal features Representation of causal attention features The calculated gating weights are weighted and summed to generate a fused feature representation. .

9. The short-term power load forecasting system based on machine learning according to claim 8, characterized in that, The method for generating the load forecast value sequence in the multi-scenario adaptive prediction unit includes the following steps: S70, The scene classification subnetwork receives the fused feature representation output by the cross-modal dynamic gating fusion unit. And based on this fusion feature representation, determine the load behavior scenario to which the current input data belongs; S71. Use an activation function to output the probability of each scene category, and use a hierarchical output method to calculate the scene category probability vector; S72. Use a scenario-specific prediction subnetwork to match the corresponding load behavior scenario, and configure a scenario-specific prediction subnetwork with the same structure and independent parameters as the number of load behavior scenarios. S73. Each network performs forward computation in the following layers to obtain the hidden vectors corresponding to different stages; S74. Calculate the load value sequence of the scene network based on the hidden vectors corresponding to different stages; S75. Based on the attribution probability vector output by the scene classification sub-network, select the scene category with the highest probability as the scene determination result of the current input data.

Citation Information

Patent Citations

  • Substation inspection double-layer inspection data model based on multi-source data fusion and application thereof

    CN118941268A

  • Slope three-dimensional deformation prediction method

    CN119295688A