Cross-dimension multi-scale fusion load prediction method based on multi-user load space-time correlation
Through the cross-dimensional and multi-scale fusion load forecasting method of multi-user load spatiotemporal correlation, combined with time and space characteristics, the nonlinearity and lack of spatial correlation of load forecasting in traditional methods are solved, and higher-precision load forecasting is achieved.
Patent Information
- Application Number
- CN202511339792.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Traditional power load forecasting methods have difficulty in effectively capturing the complex nonlinear characteristics and spatial correlations in load data, especially when dealing with sudden load changes and cross-regional load propagation effects, and the forecasting accuracy is insufficient.
A cross-dimensional and multi-scale fusion load forecasting method based on the spatiotemporal correlation of multi-user loads is adopted. Through the time scale embedding module, channel attention module, multi-scale spatiotemporal fusion module and linear projection prediction output module, the time and spatial dimension characteristics of user load are combined to perform multi-scale feature extraction and fusion, and the attention mechanism and residual connection mechanism are introduced to optimize the model performance.
It significantly improves the accuracy and stability of load forecasting, enhances the ability to identify short-term fluctuations and long-term trends, and improves the robustness and forecast accuracy of the model in complex environments.
Smart Images

Figure CN120832644A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of power load prediction, and particularly relates to a cross-dimension multi-scale fusion load prediction method based on multi-user load space-time correlation. BACKGROUND
[0002] Power load prediction is a key technology in energy management systems, and its accuracy directly affects the efficiency of power grid dispatching and power supply reliability. With the popularity of distributed energy, electric vehicles and diversified electrical equipment, user-side load data presents characteristics of high nonlinearity, multi-scale and strong randomness, and traditional prediction methods face significant challenges. Current mainstream prediction models mainly include time series analysis methods based on statistics (such as ARIMA and exponential smoothing) and end-to-end models based on deep learning (such as LNN, TCN and Transformer). Statistical methods are computationally efficient, but they are difficult to capture complex nonlinear features in load data. Deep learning models have advantages in feature extraction, but they still have limitations in multi-scale feature fusion. For example, LSTM networks can effectively model temporal dependencies, but they are not responsive to sudden load changes. The Transformer architecture is good at capturing long-term dependencies, but it is less sensitive to high-frequency fluctuations and has high computational complexity. In addition, existing methods only rely on time dimension modeling, ignoring the inherent spatial correlation in power systems. This single-dimensional modeling approach limits the model's generalization ability in complex scenarios, especially in industrial electricity scenarios, making it difficult to capture cross-regional load propagation effects, such as power grid cascading fluctuations caused by the start and stop of a production line, further increasing the difficulty of prediction. SUMMARY
[0003] The application aims to provide a cross-dimension multi-scale fusion load prediction method based on multi-user load space-time correlation to solve the technical problem of insufficient response to sudden load and gradual changes in traditional methods.
[0004] To solve the above technical problems, the specific technical solutions of the cross-dimension multi-scale fusion load prediction method based on multi-user load space-time correlation are as follows: A cross-dimension multi-scale fusion load prediction method based on multi-user load space-time correlation includes the following steps: Step 1: Obtain user load statistical data set; Step 2: Preprocess the obtained data and divide the data set into training set, validation set and test set; Step 3: Build a load prediction model based on cross-dimension multi-scale fusion; Step 4: Use the training set and validation set obtained in step 2 to train the cross-dimensional multi-scale fusion network, select the model parameters with the smallest validation loss on the validation set as the optimal model parameters, and use the test set to evaluate the performance of the optimal model; Step 5: Input the historical load monitoring data of the user to be predicted into the trained load forecasting model, and output the load forecast for a certain time granularity in the future.
[0005] Furthermore, the data described in step 2 is a user load statistics data set, which contains electricity load statistics of multiple users within a unified time range. The data has clear timestamp annotations and constitutes a multi-user, multi-time point time series data structure. All data are uniformly preprocessed by the system; then, a user-based division strategy is adopted to divide the data set into a training set, a validation set, and a test set.
[0006] Furthermore, the load forecasting model based on cross-dimensional multi-scale fusion described in step 3 includes a time scale embedding module, a channel attention module, a multi-scale spatiotemporal fusion module, and a linear projection prediction output module, each of which is used to perform the following steps: Step 3.1: Extract the temporal dependency and periodic variation characteristics of load data at different time scales through the time scale embedding module; Step 3.2: Adaptively capture key feature channels through the channel attention module to enhance the model's ability to focus on different load feature dimensions; Step 3.3: Interactively fuse features at different temporal and spatial resolutions through the multi-scale spatiotemporal fusion module to further integrate information at different scales to enhance the overall representation capability; Step 3.4: Input the fused deep features into the linear projection prediction output module to achieve high-precision prediction of future load change trends.
[0007] Furthermore, the step 3.1 includes the following steps: Step 3.1.1: For each data variable, the time series X i Perform normalization to normalize each element to the range of [-1,1]; Step 3.1.2: Project the normalized sequence onto dimensional space, and obtain the basic feature map; Step 3.1.3: Superimpose position embedding to encode sequence temporal position information; Step 3.1.4: Introduce a learnable global timestamp embedding to capture periodic temporal features; Step 3.1.5: Use the unified input representation to generate embeddings.
[0008] Further, the step 3.2 comprises the following steps: Step 3.2.1: Firstly, a global average pooling operation is performed on the input data to obtain the overall state of the time series in the entire life cycle, obtaining a global mean vector of each variable ; at the same time, a global standard deviation operation is used to measure the fluctuation amplitude of each variable in the time dimension, thereby obtaining a standard deviation vector reflecting the degree of dynamic change ; Step 3.2.2: the global mean vector and the standard deviation vector are respectively input into two independent 1x1 convolution layers, and high-order channel expressions are extracted through further convolution operations, thereby forming channel embedding features representing the importance of each variable and , then splicing and using convolution operation to extract high-level features ; Step 3.2.3: the high-level features are normalized, and the attention weight coefficients of each variable are generated through Sigmoid activation function ; the weight coefficients are multiplied with the original embedding vector element by element to generate the weighted variable feature representation tensor , realizing the weighted adjustment in the variable dimension; Finally, by merging the variable dimension and the channel dimension, the weighted variable feature matrix is converted into a shape of .
[0009] Further, the step 3.3 comprises the following steps: Step 3.3.1: in the multi-scale recognition layer of the multi-scale spatio-temporal fusion module, firstly, the embedded time series representation is converted from time domain to frequency domain by using fast Fourier transform, and different frequency components are decomposed, and the amplitude corresponding to each frequency point is calculated; Step 3.3.2: the first k frequency components with the largest amplitude value are selected as the main periodic signal source, so that the model focuses on the periodic components with the most significant influence on the sequence fluctuation, filters the noise frequency, and through the inverse relationship between the frequency and the sequence length L , the corresponding time scale is obtained; Step 3.3.3: the sequence is extended in the time dimension to ensure that the length of the subsequence is aligned with the time scale, and through the dimension reshaping operation, the completed is converted into a multi-scale feature vector suitable for graph convolution ; Step 3.3.4: In the adaptive graph convolution layer of the multi-scale spatiotemporal fusion module, receive the above reconstructed subsequence , through a learnable linear transformation matrix The first i The tensor corresponding to the scale Projecting back to the variable space yields , generate a trainable parameter matrix ; Step 3.3.5: Transform these two parameter matrices Multiply to get the original correlation matrix, filter the negative correlation through the activation function, and then normalize the weights between different nodes to get the adaptive adjacency matrix adapted to the current scale ; Step 3.3.6: Use Mixhop strategy to capture high-order correlation. First, define the power set of the adjacency matrix and apply it to the adaptive adjacency matrix. Perform multi-order power operations to obtain the j Order correlation information, and then each order matrix is combined with the input feature Perform graph convolution operations and output the results by splicing feature dimensions, enhance nonlinear expression capabilities through activation functions, and finally pass MLP The multi-layer perceptron performs dimensional projection and nonlinear mapping on the fusion features, and outputs the graph convolution. Project back to 3D tensor ; Step 3.3.7: In the multi-head attention mechanism module of the multi-scale spatiotemporal fusion module, focus on the feature interaction of the time scale dimension, first for the input tensor Adjust the structure by reshaping the dimensions, and then based on the query Q ,key K ,value V The triple mechanism uses a learnable matrix , generate attention input; Step 3.3.8: Apply attention weighting to the embedding sequence of each time scale to learn the relative importance of the time scale dimension. Combine the normalized amplitude of each frequency component to generate scale aggregation weights to achieve dynamic integration and fusion of multi-scale features. Step 3.3.9: Amplitudes corresponding to each scale calculated based on fast Fourier transform , generating scale aggregation weights , ensure that the weight sum is 1, then combine the attention output and scale weight, and obtain the output result of fusion multi-scale features through normalization and weighted aggregation .
[0010] Furthermore, the step 3.4 includes the following steps: Step 3.4.1: Introduce the residual connection mechanism, and perform element-wise addition on the original input features and the output features after multi-scale processing; Step 3.4.2: Convert the fused features into prediction results through a bilinear transformation, first define the prediction granularity function, and output the corresponding time granularity m for the first prediction task; Step 3.4.3: Use a learnable weight matrix to map the time dimension of the fused features from L to the prediction length T to obtain intermediate features , realizing cross-scale conversion from the history sequence length to the future prediction length while keeping the feature dimension unchanged, and then through a weight matrix , map the feature dimension of from to the variable number N to obtain variable space features , realizing prediction conversion from abstract features to specific variables; Step 3.4.4: Combine the bias vector to generate the prediction results of the future time steps .
[0011] Further, the step 4 includes the following steps: Use the training set to train the built prediction model, and use the trained model to infer the validation set, verify the model performance through the root mean square error RMSE and MAE two evaluation indexes, select the best model as the best model on the validation set, and use the test set to evaluate the performance of the best model.
[0012] The multi-dimensional multi-scale fusion load prediction method based on multi-user load space-time correlation has the following advantages: (1) The present application fuses cross-dimensional information and multi-scale time series features on the basis of the backbone network, significantly improves the understanding ability of the model to user load behavior through the design of multi-path fusion structure, and effectively enhances the accuracy and stability of the prediction results.
[0013] (2) The present application combines the multi-dimensional information of the spatial dimension and the time dimension of the user load time series data, enhances the features, and significantly improves the accuracy of load prediction.
[0014] (3) The present invention designs a multi-scale feature extraction module and adopts a sliding window processing strategy to obtain the load sequence characteristics at multiple time granularities, thereby improving the model's ability to perceive both short-term fluctuations and long-term trends, and enhancing the ability to recognize load periodicity and burstiness.
[0015] (4) The present invention introduces an attention mechanism to adjust the weights of important features in the fusion stage, so that the model focuses on key influencing factors, thereby suppressing the interference of redundant information and improving the robustness of load forecasting in complex environments.
[0016] (5) The present invention proposes a deep fusion module that gradually integrates cross-dimensional and multi-scale features in multiple stages, guiding the model to construct load forecast representation from coarse to fine, and then generating a load curve with greater spatiotemporal consistency and prediction accuracy.
[0017] (6) The present invention introduces a module cascade mechanism based on residual connection to deeply fuse features of each dimension while maintaining the integrity of information transmission, and performs end-to-end training through a joint loss function to optimize the overall performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of the cross-dimensional and multi-scale fusion load forecasting method based on spatiotemporal correlation of multi-user loads of the present invention; Figure 2 It is the structure diagram of the user load forecasting model; Figure 3 It is a diagram of the time scale embedded module structure; Figure 4 This is the structure diagram of the channel attention module; Figure 5 This is the structure diagram of the multi-scale spatiotemporal fusion module; Figure 6 It is the structure diagram of the adaptive graph convolution layer; Figure 7 This is the module structure diagram of the multi-head attention mechanism. DETAILED DESCRIPTION
[0019] In order to better understand the purpose, structure and function of the present invention, the following is a further detailed description of a cross-dimensional and multi-scale fusion load prediction method based on spatiotemporal correlation of multi-user loads in conjunction with the accompanying drawings.
[0020] like Figure 1 As shown, the present invention provides a cross-dimensional and multi-scale fusion load forecasting method based on the spatiotemporal correlation of multi-user loads, comprising the following steps: Step 1: Obtain user load statistics data set; Step 2: Preprocess the acquired data and divide the dataset into training set, validation set and test set; The data described is a user load statistics dataset, comprising the following time-series monitoring data: Regarding electricity load, electricity load statistics were collected from 320 users within a unified time range. The data is clearly timestamped, forming a typical multi-user, multi-time point time-series data structure. All data undergoes a systematic preprocessing process: first, data quality control is performed, including format unification and missing value handling. Feature normalization is then implemented to improve the stability and efficiency of model training. Finally, time-series sample construction methods such as sliding windows are used to generate time-correlated model input samples. This dataset is characterized by its clear data structure and consistent temporal granularity, making it suitable for multi-user load forecasting tasks. Regarding data partitioning, the dataset is divided into training, validation, and test sets based on user data, ensuring a more objective assessment of the model's generalization ability when faced with unknown user data. This data organization effectively supports the needs of cross-dimensional, multi-scale fusion load forecasting modeling based on the spatiotemporal correlation of multi-user loads.
[0021] The data sample a of the time series monitoring data is ,in m is the number of variables, i The time series of the variables is ,in is the time step.
[0022] Step 3: Build a load forecasting model based on cross-dimensional and multi-scale fusion; like Figure 2 As shown in Figure 1, the load forecasting model based on cross-dimensional multi-scale fusion includes a time scale embedding module, a channel attention module, a multi-scale spatiotemporal fusion module, and a linear projection prediction output module. Each module is used to perform the following steps: Step 3.1: If Figure 3 As shown in Figure 2, the time scale embedding module is used to extract the temporal dependence and periodic variation characteristics of load data at different time scales.
[0023] Step 3.1.1: For each data variable, the time series X i Perform normalization to normalize each element to the range of [-1,1]. The specific expression is as follows:
[0024] in, L represents the length of the time series, t Indicates the initial position of the prediction window.
[0025] Step 3.1.2: Project the normalized sequence onto dimensional space to obtain the basic feature map.
[0026] Step 3.1.3: Superimposed position embedding , encoding sequence time position information, specifically expressed as: where, is a mathematical symbol, representing the real number set, represents the position, i.e., the number of time steps in the time sequence.
[0027] Step 3.1.4: Introducing a learnable global timestamp embedding , capturing periodic time features such as minutes / hours, etc.
[0028] Step 3.1.5: Adopting a unified input representation to generate embeddings where, s , are variable dimensions and channel dimensions, respectively, and are expressed as follows:
[0029] where, is an adjustable hyperparameter used to control the contribution of the basic feature mapping, 1x1 convolution operation.
[0030] Step 3.2: As shown in Figure 4 , the channel attention module adaptively captures key feature channels, enhancing the model's attention ability to different load feature dimensions.
[0031] Step 3.2.1: First, perform global average pooling on the input data to obtain the overall state of the time series throughout its life cycle, obtaining the global mean vector for each variable; at the same time, use the global standard deviation operation to measure the fluctuation amplitude of each variable in the time dimension, thereby obtaining the standard deviation vector reflecting the degree of dynamic change. These two global statistical features represent sequence characteristics from different angles, and when combined, they not only retain overall trends but also enhance the response ability to local changes in time. The specific expression is as follows:
[0032] where, , are the average pooling features and standard deviation features, respectively, and are the global average pooling and global standard deviation operations, respectively.
[0033] Step 3.2.2: Combine the global mean vector and the standard deviation vector The two independent 1x1 convolutional layers are input respectively, and high-order channel expressions are extracted through further convolution operations to form channel embedding features representing the importance of each variable and Then, the high-level features are spliced and high-level features are extracted using convolution operations The specific expression is as follows:
[0034] wherein, is an activation function, is a splicing operation, and the feature variable is obtained after splicing.
[0035] Step 3.2.3: The high-level features are normalized, and the attention weight coefficients of each variable are generated through Sigmoid activation functions ; the weight coefficients are multiplied element by element with the original embedding vectors to generate a weighted variable feature representation tensor , which realizes the weighted adjustment in the variable dimension, thereby significantly enhancing the focusing ability of the model on key variables. The specific expression is as follows:
[0036] wherein, is an activation function, represents a batch normalization operation, is an attention weight.
[0037] Finally, by merging the variable dimension and the channel dimension, the weighted variable feature matrix is converted into a shape of .
[0038] Step 3.3: As shown in Figure 5 , the multi-scale spatio-temporal fusion module interacts and fuses the features at different time and spatial resolutions, further integrating information at different scales to enhance the overall representation ability.
[0039] The proposed multi-scale spatio-temporal fusion module is used to perform the following steps: Step 3.3.1: In the multi-scale identification layer of the multi-scale spatio-temporal fusion module, the embedded time series representation is first converted from the time domain to the frequency domain using the fast Fourier transform, the different frequency components are decomposed, and the amplitude corresponding to each frequency point is calculated. The specific expression is as follows:
[0040] wherein, represents performing fast Fourier transform on the input time series data, Indicates the calculation of each frequency component in the frequency domain The corresponding amplitude, represents the index of the channel dimension, and Respectively represent the extraction The real and imaginary parts of the channels are calculated by the function Perform the average operation on the channel dimension, vector Represents the calculated amplitude for each frequency. It is k frequency components, j Represents an imaginary unit .
[0041] Step 3.3.2: Select the front k frequency components As the main source of periodic signals, the model focuses on the periodic components that have the most significant impact on sequence fluctuations and filters out noise frequencies. and time series length L The inverse relationship is converted to obtain the corresponding time scale , the specific expressions are as follows:
[0042] in, Indicates that the largest amplitude is selected k The frequency index. is the frequency Convert to the corresponding time scale.
[0043] Step 3.3.3: To make the subsequences of different time scales be uniformly input into the graph convolution module, Perform time dimension expansion to ensure that the subsequence length is aligned with the time scale to avoid dimensional misalignment during convolution. Convert to multi-scale feature vectors suitable for graph convolution , the specific expressions are as follows:
[0044] in, for An integer multiple of . It is used to zero-extend the time series in the time dimension so that compatible, Representation based on time scale The reconstructed time series.
[0045] Step 3.3.4: If Figure 6 As shown, in the adaptive graph convolution layer of the multi-scale spatiotemporal fusion module, the above reconstructed subsequence is received , through a learnable linear transformation matrix The first i The tensor corresponding to the scale Projecting back to the variable space yields , generate a trainable parameter matrix ,in N is the number of variables, h To hide the dimension, the specific expression is as follows:
[0046] in, Indicates the i The variable space projection of the time scale, is a learnable weight matrix.
[0047] Step 3.3.5: Transform these two parameter matrices Multiply to get the original correlation matrix, filter the negative correlation through the activation function, and then normalize the weights between different nodes to get the adaptive adjacency matrix adapted to the current scale , the specific expressions are as follows:
[0048] in, ReLU is the activation function, ensuring that all negative values become zero, SoftMax The function is used to normalize the weights of different nodes.
[0049] Step 3.3.6: Use Mixhop strategy to capture high-order correlation. First, define the power set of the adjacency matrix and apply it to the adaptive adjacency matrix. Perform multi-order power operations to obtain the j Then, each order matrix is combined with the input feature Perform graph convolution operations and output the results by splicing feature dimensions, and enhance the nonlinear expression ability through activation functions. MLP The multi-layer perceptron performs dimensional projection and nonlinear mapping on the fusion features, and outputs the graph convolution. Project back to 3D tensor , which improves the model's ability to capture complex dependencies between variables. The specific expression is as follows:
[0050] Among them, the hyperparameters p is a power set of adjacency matrices, Represents the learned adjacency matrix Multiply by itself j Second-rate, Indicates the concatenation of neighbor features of each order by column, is the activation function, Represents dimensional projections and nonlinear mappings.
[0051] Step 3.3.7: If Figure 7 As shown, in the multi-head attention mechanism module of the multi-scale spatiotemporal fusion module, the feature interaction of the time scale dimension is focused. First, for the input tensor Adjust the structure by reshaping the dimensions, and then based on the query Q ,key K ,value V The triple mechanism uses a learnable matrix , generate attention input, the specific expression is as follows:
[0052] in, B is the batch size, represents the reshaping function, is the mapping matrix.
[0053] Step 3.3.8: Apply attention weighting to the embedding sequence of each time scale to learn the relative importance of the time scale dimension, and further combine the normalized amplitude of each frequency component to generate scale aggregation weights to achieve dynamic integration and fusion of multi-scale features, and improve the model's comprehensive extraction ability of periodic and asynchronous features. Reshape into , the specific expressions are as follows:
[0054] in, SoftMax Function is used to calculate attention weights and aggregate features. is the key vector in the multi-head attention k dimensionality to avoid gradient vanishing when the dimension is high. It is a multi-head self-attention function, which concatenates the multi-head attention outputs and projects them back to the original feature dimension. c is the number of attention heads, is the output projection matrix.
[0055] Step 3.3.9: Amplitudes corresponding to each scale calculated based on fast Fourier transform , generating scale aggregation weights , ensure that the weight sum is 1, then combine the attention output and scale weight, and obtain the output result of fusion multi-scale features through normalization and weighted aggregation , the specific expressions are as follows:
[0056] in, is the amplitude normalization value corresponding to the first i represents an amplitude normalization operation for weighted fusion of features of different time scales. represents normalizing the sample feature layer.
[0057] Step 3.4: input the fused deep features into the linear projection prediction output module to realize high-precision prediction of future load change trend.
[0058] The proposed linear projection prediction output module is used to perform the following steps: Step 3.4.1: a residual connection mechanism is introduced, which effectively alleviates the gradient vanishing problem in deep networks and improves the stability and expression ability of model training by element-wise addition of the original input features and the output features after multi-scale processing. The specific expression is as follows:
[0059] wherein, l is the network layer index.
[0060] Step 3.4.2: convert the fused features into prediction results through bilinear transformation, first define the prediction granularity function, which outputs the time granularity m for the first prediction task, and the specific expression is as follows:
[0061] wherein, is used to support multi-time granularity output, is the first prediction granularity (the prediction granularity includes hourly level such as 1 hour, 3 hours). m Step 3.4.3: use the learnable weight matrix
[0062] to map the time dimension of the fused features from to the prediction length L to obtain the intermediate features T , realizing cross-scale conversion from historical sequence length to future prediction length while keeping the feature dimension unchanged. Then through the weight matrix , the feature dimension of is mapped from to the variable number to obtain the variable space feature N , realizing prediction conversion from abstract features to specific variables, and the specific expression is as follows:
[0063] Step 3.4.4: Combine bias vector to generate the prediction results of future time steps , which is expressed as follows:
[0064] wherein, is the bias vector, is the load prediction value of future time steps.
[0065] Step 4: Train the cross-dimension multi-scale fusion network using the training set and the validation set obtained in step 2, select the model parameters with the minimum validation loss on the validation set as the best model parameters, and perform performance evaluation on the best model using the test set; In step 4, the training set is used to train the prediction model built, and the trained model is used to infer the validation set, and the model performance is verified by two evaluation indexes of root mean square error RMSE and MAE , the best model with the best performance on the validation set is selected, and the performance of the best model is evaluated using the test set, which is expressed as follows:
[0066] wherein, n is the number of test samples, and are the true value and the predicted value of the i th test sample, respectively.
[0067] Step 5: input the load history monitoring data of the user to be predicted into the user load prediction model completed by training, and output the load prediction of future time granularity.
[0068] It can be understood that the present application is described by some embodiments, and those skilled in the art know that various changes or equivalent replacements can be made to these features and embodiments without departing from the spirit and scope of the present application. In addition, under the guidance of the present application, these features and embodiments can be modified to adapt to specific conditions and materials without departing from the spirit and scope of the present application. Therefore, the present application is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application are within the scope of the present application.
Claims
1. A multi-dimensional multi-scale fusion load prediction method based on multi-user load space-time correlation, characterized in that, Comprising the following steps: Step 1: Obtain a user load statistics dataset, and pre-process the obtained data; Step 2: Build a load prediction model based on cross-dimension multi-scale fusion; the load prediction model comprises a time scale embedding module, a channel attention module, a multi-scale space-time fusion module, and a linear projection prediction output module, each module being configured to perform the following steps: Step 2.1: Extract the time series dependence and periodic variation characteristics of the load data at different time scales through the time scale embedding module; Step 2.2: Adaptively capture key feature channels through the channel attention module to enhance the model's attention ability to different load feature dimensions; Step 2.3: Interactively fuse features at different time and spatial resolutions through the multi-scale space-time fusion module to further integrate information at different scales to enhance overall representation ability; Step 2.4: Input the fused deep features into the linear projection prediction output module to achieve high-precision prediction of future load trends; Step 3: Input the load history monitoring data of the user to be predicted into the trained load prediction model to output the load prediction at a certain time granularity in the future.
2. The multi-user load spatio-temporal correlation based cross-dimension multi-scale fusion load prediction method according to claim 1, characterized in that, The data of step 1 is a user load statistics dataset, which contains the electricity load statistics data of multiple users within a unified time range, and the data has clear timestamp annotations, forming a multi-user and multi-time point time series data structure. All data are uniformly pre-processed by the system; then the data set is divided into training set, validation set and test set according to the user division strategy.
3. The multi-user load spatio-temporal correlation based cross-dimension multi-scale fusion load forecasting method according to claim 1, characterized in that, The step 2.1 comprises the following steps: Step 2.1.1: Time series for each data variable X i Normalization is performed to normalize each element to the range [-1, 1]; Step 2.1.2: Project the normalized sequence into a feature map by a 1D convolutional layer a 256-dimensional feature space. Step 2.1.3: Superimpose position embedding to encode sequence time position information; Step 2.1.4: Introduce a learnable global timestamp embedding to capture periodic time features; Step 2.1.5: Use a unified input representation to generate embedding.
4. The multi-user load spatio-temporal correlation based cross-dimension multi-scale fusion load forecasting method according to claim 1, characterized in that, The step 2.2 comprises the following steps: Step 2.2.1: First, a global average pooling operation is performed on the input data to obtain the overall state of the time series throughout its life cycle, resulting in a global mean vector for each variable ; At the same time, a global standard deviation operation is used to measure the fluctuation amplitude of each variable in the time dimension, resulting in a standard deviation vector reflecting the degree of dynamic change ; Step 2.2.2: The global mean vector and the standard deviation vector are input into two independent 1x1 convolutional layers, respectively, to extract high-order channel representations through further convolutional operations, thereby forming channel embedding features representing the importance of each variable and , which are then spliced and used for convolutional operations to extract high-level features ; Step 2.2.3: High-level features Normalized and processed by Sigmoid The activation function generates the attention weight coefficient for each variable ; The weight coefficient is multiplied element-by-element with the original embedding vector to generate the weighted variable feature representation tensor , to achieve weighted adjustment on variable dimensions; Finally, by merging the variable dimension with the channel dimension, the weighted variable feature matrix is transformed into a shape .
5. The multi-user load spatio-temporal correlation based cross-dimension multi-scale fusion load forecasting method according to claim 1, characterized in that, The step 2.3 comprises the following steps: Step 2.3.1: In the multi-scale recognition layer of the multi-scale space-time fusion module, first convert the embedded time series representation from the time domain to the frequency domain using fast Fourier transform, decompose the different frequency components, and calculate the amplitude corresponding to each frequency point; Step 2.3.2: Select the top k frequency components with the largest amplitude values As the main periodic signal source, let the model focus on the periodic component that has the most significant impact on the sequence fluctuations, filter out the noise frequencies, and convert the frequency to the corresponding time scale by the inverse relationship between the frequency L and the length of the sequence ; Step 2.3.3: On the sequence Perform temporal dimension expansion, ensuring that the sub-sequence length is aligned with the time scale, complete the Convert to multi-scale feature vector adapted for graph convolution ; Step 2.3.4: In the adaptive graph convolution layer of the multi-scale spatiotemporal fusion module, receive the above reconstructed subsequence , through a learnable linear transformation matrix The first i The tensor corresponding to the scale Projecting back to the variable space yields , generate a trainable parameter matrix ; Step 2.3.5: Multiply these two parameter matrices to get the original association matrix, filter the negative association through the activation function, and normalize the weights between different nodes to get the adaptive adjacency matrix that adapts to the current scale .
6. The multi-user load spatio-temporal correlation based cross-dimension multi-scale fusion load prediction method according to claim 5, characterized in that, The step 2.3 further comprises the following steps: Step 2.3.6: Use Mixhop strategy to capture high-order correlation. First, define the power set of the adjacency matrix and apply it to the adaptive adjacency matrix. Perform multi-order power operations to obtain the j Order correlation information, and then each order matrix is combined with the input feature Perform graph convolution operations and output the results by splicing feature dimensions, enhance nonlinear expression capabilities through activation functions, and finally pass MLP The multi-layer perceptron performs dimensional projection and nonlinear mapping on the fusion features, and outputs the graph convolution. Project back to 3D tensor ; Step 2.3.7: In the multi-head attention mechanism module of the multi-scale spatio-temporal fusion module, focus on the feature interaction of the time scale dimension, first for the input tensor Adjust the structure by dimension reshaping, and then generate attention input based on the triple mechanism of query Q , key K , value V , and learnable matrix . Step 2.3.8: Apply attention weighting to the embedded sequence of each time scale to learn the relative importance in the time scale dimension, combine the normalized amplitudes of each frequency component to generate scale aggregation weights, and realize dynamic integration and fusion of multi-scale features; Step 2.3.9: Calculate the amplitude of each scale corresponding to the fast Fourier transform , generate scale aggregation weight , ensure the weight sum is 1, and combine the attention output with the scale weight to get the output result of the fused multi-scale feature by normalized weighted aggregation .
7. The multi-user load space-time correlation based cross-dimensional multi-scale fusion load prediction method according to claim 1, characterized in that, The step 2.4 comprises the following steps: Step 2.4.1: Introduce residual connection mechanism by element-wise adding the original input features with the multi-scale processed output features Step 2.4.2: Convert the fused features into prediction results by a bilinear transformation, first define a prediction granularity function, for the i-th prediction task output corresponding time granularity m . ; Step 2.4.3: Utilizing the learnable weight matrix , the time dimension of the fused feature is mapped from L to the prediction length T to obtain the intermediate feature , realizing the cross-scale conversion from the history sequence length to the future prediction length while keeping the feature dimension unchanged, and then the feature dimension of is mapped from to the variable number N by the weight matrix to obtain the variable space feature , realizing the prediction conversion from the abstract feature to the specific variable; Step 2.4.4: Incorporate bias vector , generate prediction for future time steps .
8. The multi-user load space-time correlation based cross-dimension multiscale fusion load forecasting method according to claim 1, characterized in that, The step 3 comprises the following steps: The built prediction model is trained using the training set, and the trained model is used to infer the validation set, and the model performance is verified by the root mean square error RMSE and MAE Two evaluation indexes, and the best model is selected as the best model, and the performance of the best model is evaluated using the test set.
Citation Information
Patent Citations
Urban electricity consumption prediction method, system and device fusing time-frequency characteristics and storage medium
CN119227899A
Wind power prediction method and system based on multiple time scales
CN119646636A
Dynamic multivariate time series-oriented infrastructure load prediction method
CN120124789A
Electrical load prediction method and device based on spatial-temporal correlation
WO2025092993A1
Cited By
Power consumption dynamic prediction method based on physical-semantic dual-channel modulation
CN121303468A
Power system load prediction method and system based on hybrid deep learning architecture
CN121457847A
Non-intrusive non-stationary load identification method for statistically driving multi-scale fusion attention
CN121659039A
A statistical driving multi-scale fusion attention non-stationary load non-intrusive identification method
CN121659039B
Fixed-wing unmanned aerial vehicle timing sequence state prediction method and system fusing multi-scale embedding and grouping channel attention, and medium
CN121808718A