Power load prediction method based on multi-scale hypergraph neural network and large language model alignment
Through the multi-scale hypergraph neural network and large language model alignment method, the problem of insufficient alignment of natural language and power load sequence feature representation is solved, and more accurate and robust power load prediction is achieved, and the model's prediction ability in different scenarios is improved.
Patent Information
- Application Number
- CN202510546395.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The difficulty in effectively aligning natural language and multi-scale feature representations of power load sequences results in insufficient accuracy in power load prediction, especially in case of few samples or zero samples.
The method of alignment of multi-scale hypergraph neural network and large language model is adopted. Through the multi-scale feature extraction module, hyper-edge feature extraction module and cross-modal alignment module, combined with the multi-scale hybrid prompt mechanism, the large language model's understanding of the multi-scale mode of power load sequence is enhanced.
It improves the accuracy and robustness of power load prediction, can capture key features more accurately in complex and variable power load scenarios, and improves the universality and practicality of the model in different application scenarios.
Smart Images

Figure CN120433183A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power load prediction, and specifically relates to a power load prediction method based on multi-scale hypergraph neural network and large language model alignment. Background Art
[0002] With the intelligent transformation and upgrading of modern industries, the global energy consumption structure is undergoing profound changes. Driven by the dual forces of new urbanization and the digital development of manufacturing, the increasing urban power load has become a critical issue that needs to be addressed. Accurately predicting future power load changes not only enables precise coordination among power generation, transmission, and distribution, but also provides dynamic decision-making support for the resilient optimization of smart grids. Therefore, power load forecasting has become one of the key technological breakthroughs in the digital evolution of modern energy systems.
[0003] Real-world power load forecasting is a challenging task. First, due to the influence of extreme events and new construction, power load series often contain few or even no samples, posing a significant challenge to accurate load forecasting. Second, due to the influence of human activities, power load series exhibit varying patterns at different time scales. For example, household electricity consumption typically decreases in the morning and increases in the evening on a daily scale, and decreases during the week and increases on weekends on a weekly scale. Considering the multi-scale patterns of power load series generally yields more accurate forecasts than considering only a single scale.
[0004] To accurately predict power load, many classic network structures have emerged. Convolutional Neural Networks (CNNs), with their powerful feature extraction capabilities, excel in capturing local features. Graph Neural Networks (GNNs) excel at processing graph-structured data and can effectively mine the relationships between data. Recurrent Neural Networks (RNNs) have a natural advantage in processing sequential data and can capture temporal dependencies in the data. Transformers, with their self-attention mechanism, demonstrate excellent performance in processing long sequence data and parallel computing. However, these methods are proprietary models trained for specific scenarios, making them difficult to effectively migrate to different scenarios and struggling with situations with few or even zero samples.
[0005] In recent years, with the rise of large language models (LLMs) in natural language processing and computer vision, attempts have been made to apply them to power load forecasting. However, existing LLM-based methods have significant shortcomings. First, they ignore the differences between natural language and power load sequences in multi-scale semantic space. Natural language focuses on semantic understanding and logical reasoning, while power load sequences emphasize numerical changes and pattern recognition in the temporal dimension. Second, they lack specialized prompts designed for power load forecasting, hindering the full potential of LLMs in processing complex data and tasks. These two key issues pose challenges to the successful application of LLMs to power load forecasting and urgently require in-depth research and effective solutions. Summary of the Invention
[0006] In view of the above, the purpose of the present invention is to provide a power load forecasting method based on the alignment of a multi-scale hypergraph neural network and a large language model, so as to align the multi-scale feature representations of natural language and power load sequences, and by introducing a multi-scale hybrid prompt mechanism, use a variety of different prompts to enrich contextual information, thereby enhancing the large language model's ability to understand the multi-scale patterns of power load sequences. It has broad application prospects in the fields of power system operation, energy planning, and energy efficiency management.
[0007] To achieve the above-mentioned purpose, the present invention provides the following technical solutions:
[0008] An embodiment of the present invention provides a method for power load prediction based on alignment of a multi-scale hypergraph neural network and a large language model, comprising the following steps:
[0009] Preprocess the power load sequence and construct training samples;
[0010] In the multi-scale feature extraction module, the training samples and word embeddings based on the pre-trained large language model are mapped into multi-scale temporal feature representations and multi-scale text prototype representations respectively;
[0011] In the hyperedge feature extraction module, the hyperedge mechanism based on the multi-scale hypergraph neural network aggregates the multi-scale temporal feature representation into a multi-scale hyperedge feature representation;
[0012] In the cross-modal alignment module, the multi-scale text prototype representation and the multi-scale hyperedge feature representation are aligned based on cross-attention to obtain the aligned multi-scale feature representation, and a multi-scale hybrid prompt is constructed;
[0013] In the large model prediction module, the multi-scale mixed prompts and the aligned multi-scale feature representation are concatenated and input into the pre-trained large language model for power load prediction. Based on the prediction loss, the power load prediction model including the multi-scale feature extraction module, the hyperedge feature extraction module, the cross-modal alignment module and the large model prediction module is trained.
[0014] The new power load sequence is input into the trained power load prediction model to obtain the prediction result.
[0015] Preferably, in the multi-scale feature extraction module, the mapping of the training samples and the word embedding based on the pre-trained large language model into a multi-scale temporal feature representation and a multi-scale text prototype representation respectively includes:
[0016] Perform instance normalization on the training samples to obtain the instance normalized power load sequence T is the time window size. Based on the subsequence of the s-1th scale, the subsequence of the sth scale is generated through different aggregation windows. The subsequence of the sth scale is expressed as in is the power load sequence representation of the s-th scale time step t, D is the feature dimension, is the sequence length of the sth scale, N s-1 is the subsequence length of the s-1th scale, l s-1 is the aggregation window size of the s-1th scale, and the calculation formula of the aggregation process is as follows:
[0017]
[0018] Among them, Agg(·) is the aggregation function, θ s-1 is the learnable parameter of the aggregation function at the s-1th scale, is the input training sample after instance normalization;
[0019] Word embedding based on pre-trained large language model Where V is the vocabulary size, P is the hidden layer dimension of the pre-trained large language model, and the word embedding U is first transformed into the initial text prototype representation through linear mapping Where V′<<V, and then the multi-scale text prototype representation is obtained through linear mapping. The calculation formula of the linear mapping process is as follows:
[0020]
[0021] Among them, Linear(·) is the linear mapping function, U s is the text prototype representation of the s-th scale, U s-1 is the text prototype representation of the s-1th scale, λ s-1is the learnable parameter of the linear mapping function at the s-1th scale, V s is the number of prototype representations of the s-th scale text;
[0022] Finally, the S-scale temporal feature representation {X 1 ,…,X s ,…,X S} and text prototype representation {U 1 ,…,U s ,…,U S}, where X s and U s They are the temporal feature representation and text prototype representation of the s-th (1≤s≤S) scale respectively.
[0023] Preferably, in the hyperedge feature extraction module, the hyperedge mechanism based on the multi-scale hypergraph neural network aggregates the multi-scale temporal feature representation into a multi-scale hyperedge feature representation, including:
[0024] The multi-scale temporal feature {X 1 ,…,X s ,…,X S} is considered as a node and two parameters are initialized, namely the hyperedge embedding and node embedding Among them, X s is the temporal feature representation of the s-th (1≤s≤S) scale, M s is the number of hyperedges at the s-th scale, D is the feature dimension, and the point-edge correlation matrix H at the s-th scale is obtained by similarity calculation s , which is calculated as follows:
[0025]
[0026] in, and is a learnable parameter, the tanh(·) activation function is used to perform nonlinear transformation, the ReLU(·) activation function is used to eliminate weak connections, Linear(·) is a linear mapping function, the superscript T is a transpose operation, and finally the point-edge association matrix {H 1 ,…,H s ,…,H S};
[0027] The i-th hyperedge at scale s Hyperedge features Based on the point-edge correlation matrix H s The information is aggregated and the calculation formula is as follows:
[0028]
[0029] Among them, Avg(·) is the averaging operation, is the point-edge correlation matrix H under scale s s The hyperedge shown in Connected neighbor nodes, is the jth node under scale s The time feature representation of the scale s is the set of all hyperedge features. s , and finally get the hyperedge feature representation of S scales {ε 1 ,…,ε s ,…,ε S}.
[0030] Preferably, a sparsification strategy is used to calculate the point-edge correlation matrix of the sth scale, and the calculation formula is as follows:
[0031]
[0032] Among them, η∈[0,M s ] is the threshold of the TopK(·) function, which indicates the maximum number of neighbor hyperedges connected to the node, n is the node index, m is the hyperedge index, For all hyperedges connecting the nth node, is the s-th scale sparse point-edge correlation matrix, based on Get the final S-scale point-edge correlation matrix {H 1 ,…,H s ,…,H S}.
[0033] Preferably, in the cross-modal alignment module, the cross-attention-based alignment of the multi-scale text prototype representation and the multi-scale hyperedge feature representation to obtain the aligned multi-scale feature representation includes:
[0034] Hyperedge feature representation ε at scale s based on multi-scale hyperedge feature representation s and the text prototype representation U at scale s in the multi-scale text prototype representation s , first map it to the corresponding query key Sum in is the index of the number of heads in attention, is the total number of heads, and is the mapping matrix that can be learned under scale s, D is the feature dimension, P is the hidden layer dimension of the pre-trained large language model, Then, the cross attention is calculated to align the hyperedge feature representation and the text prototype representation. The calculation formula is as follows:
[0035]
[0036] Among them, Attn(·) is the cross attention mechanism, softmax(·) is the activation function, and the superscript T is the transposition operation. For scale s The output of each attention head is aggregated to obtain the aligned feature representation Z of the multi-head attention output at scale s (1≤s≤S) s , and finally obtain the aligned S scale multi-scale feature representation {Z 1 ,…,Z s ,…,Z S}.
[0037] Preferably, the multi-scale hybrid prompt includes:
[0038] Learning tips Expressed as in, is the learnable hint at the sth scale and is embedded by the learnable initialization, 1≤s≤S, S is the number of scales, L s is the length of the learnable cue at scale s, and D is the feature dimension;
[0039] Data-related tips These include data description prompts π that provide basic background information of the input power load sequence to the pre-trained large language model, task description prompts τ that guide the pre-trained large language model to understand and perform specific power load prediction-related tasks, and data statistics prompts μ that include statistical features of the power load sequence, including the input power load sequence and subsequences at different scales.
[0040] Ability enhancement tips Including logical thinking prompts σ that guide pre-trained large language models to gradually solve problems, and emotional manipulation prompts that simulate the impact of emotions on human decision-making and the inference-related class hints μ that guide the available methods of pre-training large language models for power load forecasting.
[0041] Preferably, in the large model prediction module, the step of concatenating the multi-scale mixed prompts with the aligned multi-scale feature representations and inputting the concatenated information into the large language model for power load prediction includes:
[0042] After obtaining the multi-scale mixed prompts, firstly transform the learnable prompts {P 1 ,...,P s ,...,P S} and the aligned multi-scale feature representation {Z 1 ,…,Z s ,…,Z S} is spliced, and then further spliced with data-related prompts and ability enhancement prompts, and the spliced results are input into the pre-trained large language model to obtain the output result. The calculation formula is as follows:
[0043]
[0044] Among them, [·,·] is the splicing operation, Z s is the aligned feature representation at the sth (1≤s≤S) scale, is the output of the pre-trained large language model LLMs(·).
[0045] Preferably, the main parameters of the pre-trained large language model are kept frozen during the training process.
[0046] Preferably, the output of the large language model is input into the linear layer for linear transformation, and the result after linear transformation is reverse instance normalized to obtain the power load prediction value for the next H steps. As a result of power load forecasting.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] (1) The present invention designs a hyperedge feature extraction module, which extracts multi-scale hyperedge feature representation from the multi-scale time feature representation of the multi-scale power load sequence based on the hyperedge mechanism, thereby enhancing the multi-scale semantic information in the semantic space of the power load sequence, enabling the model to more comprehensively and deeply understand the changing laws and internal connections of the power load data at different time scales. Therefore, when facing complex and changeable power load scenarios, the model can more accurately capture key features, provide richer and more accurate information support for subsequent prediction tasks, and effectively improve the accuracy and reliability of power load prediction.
[0049] (2) The present invention designs a cross-modal alignment module that can align the multi-scale feature representations of natural language and power load sequences, and introduces a multi-scale mixed prompt mechanism in the alignment process to generate diversified and targeted prompt information, guiding the large language model to more deeply explore the semantic connotations and potential laws behind the multi-scale patterns of power load sequences, thereby enhancing the large language model's ability to understand the multi-scale patterns of power load sequences. When dealing with complex prediction scenarios such as few samples or even zero samples, the model can rely on its enhanced understanding ability to better transfer existing knowledge and achieve more accurate and robust power load prediction, significantly improving the versatility and practicality of power load prediction technology in different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0051] Figure 1 This is an overall flow chart of a method for power load prediction based on multi-scale hypergraph neural network and large language model alignment provided by an embodiment of the present invention;
[0052] Figure 2 This is a general framework diagram of a power load forecasting method based on multi-scale hypergraph neural network and large language model alignment provided by an embodiment of the present invention;
[0053] Figure 3 Detailed schematic diagram of a multi-scale feature extraction module provided by an embodiment of the present invention;
[0054] Figure 4 Detailed schematic diagram of the hyperedge mechanism provided by an embodiment of the present invention;
[0055] Figure 5 This is a detailed schematic diagram of the multi-scale hybrid prompting mechanism provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0057] The inventive concept of the present invention is as follows: to address the technical problem in the prior art that when a large language model is introduced for power load prediction, the multi-scale feature representations of natural language and power load sequences cannot be effectively aligned, thereby affecting the prediction accuracy. The embodiment of the present invention provides a power load prediction method based on the alignment of a multi-scale hypergraph neural network and a large language model. First, the power load data is preprocessed and a training dataset is constructed. Second, a multi-scale feature extraction module is introduced to map the input power load sequence and the word embedding based on the pre-trained large language model into a multi-scale temporal feature representation and a multi-scale text prototype representation, respectively. Then, a hyperedge mechanism is introduced to aggregate the multi-scale temporal feature representation into a multi-scale hyperedge feature representation, thereby enhancing the semantic information in the multi-scale semantic space of the power load sequence. Then, a cross-modal alignment module is introduced to align the multi-scale feature representations of natural language and power load sequences, and a multi-scale hybrid prompt mechanism is introduced during the alignment process to provide multi-scale contextual information and enhance the large language model's ability to understand the multi-scale patterns of the power load sequence. Finally, the multi-scale hybrid prompt and the aligned multi-scale feature representation are fused and fed into the frozen main body of the pre-trained large language model to achieve prediction of the power load sequence.
[0058] Figure 1 This is an overall flow chart of a power load prediction method based on multi-scale hypergraph neural network and large language model alignment provided by an embodiment of the present invention. Figure 2 This is a general framework diagram of a power load prediction method based on multi-scale hypergraph neural network and large language model alignment provided by an embodiment of the present invention. Figure 1 and Figure 2 As shown, the embodiment provides a method for power load prediction based on multi-scale hypergraph neural network and large language model alignment, including the following steps:
[0059] The power load forecasting task is defined as: given a station or region, the power load sequence observation value T times before Predict the power load value at the next H moments
[0060] Step 1: Eliminate outliers and missing values from the power load, and divide the processed data into a training data set by sliding windows in chronological order.
[0061] The outliers and missing values in the given power load series are eliminated. Then, the time window size T is set manually based on experience, and the normalized data is divided using a fixed-length sliding step to obtain the training data set.
[0062] Step 2: Divide the training dataset into batches according to a fixed batch size, with a total number of batches B.
[0063] The training data set is divided into batches based on the batch size P set by experience. The total number of batches is B, which is calculated as follows:
[0064]
[0065] Among them, N Samples is the total number of samples in the training dataset.
[0066] Step 3: Sequentially select a batch of training samples with index b from the training dataset, where b∈{1,2,…,B}. Repeat steps 4-11 for each training sample in the batch.
[0067] Step 4: Perform instance normalization on the training samples so that they have zero mean and unit variance in the feature dimension.
[0068] The power load series in the training sample is subjected to instance normalization processing to convert the data into a standard normal distribution with zero mean and unit variance in the feature dimension. The calculation formula is as follows:
[0069]
[0070] Where x′ t,d is the value of the dth (1≤d≤D)th dimension time step t in the power load sequence after eliminating outliers, μ d is the average value of the dth dimension in the power load sequence, σ d is the variance of the dth dimension in the power load sequence, is the power load value of the dth dimension at time step t after normalization of the power load sequence instance.
[0071] Step 5: Normalize the power load sequence of the instance And the word embedding U based on the pre-trained large language model is input into the multi-scale feature extraction module to extract the multi-scale temporal feature representation {X 1 ,…,X s ,…,X S} and multi-scale text prototype representation {U 1 ,…,U s ,…,U S}. Where X s and U s They represent the temporal feature representation and text prototype representation of the s-th (1≤s≤S) scale respectively.
[0072] like Figure 3 As shown in (a), the normalized power load sequence of a given instance The multi-scale temporal feature extraction module generates S-scale subsequences through different aggregation windows. The scale here refers to the time scale, such as hours, days, weeks, etc. The subsequence of the s-th scale is expressed as in is the power load sequence representation of the s-th scale time step t, D is the feature dimension, is the sequence length of the sth scale, N s-1 is the subsequence length of the s-1th scale, l s-1 is the aggregation window size of the s-1th scale. The calculation formula of the specific aggregation process is as follows:
[0073]
[0074] Among them, Agg(·) is an aggregation function, such as convolution or pooling, θ s-1 is the learnable parameter of the aggregation function at the s-1th scale, is the input sample sequence after instance normalization.
[0075] like Figure 3 As shown in (b), given the word embedding based on the pre-trained large language model Where V is the vocabulary size and P is the hidden layer dimension of the pre-trained large language model. First, the word embedding U is transformed into the initial text prototype representation through linear mapping Where V′<<V. Then the multi-scale text prototype extraction module obtains the multi-scale text prototype representation through linear mapping. The calculation formula of the specific mapping process is as follows:
[0076]
[0077] Among them, Linear(·) is the linear mapping function, U s is the text prototype representation of the s-th scale, U s-1 is the text prototype representation of the s-1th scale. The scale here refers to the text scale corresponding to the time scale, such as words, sentences, paragraphs, etc., λ s-1 is the learnable parameter of the linear mapping function at the s-1th scale, V s is the number of prototype representations of the s-th scale text.
[0078] Finally, we get the temporal feature representation {X 1 ,…,X s ,…,X S} and text prototype representation {U 1 ,…,U s ,…,U S}, where X s and U sThey represent the temporal feature representation and text prototype representation of the s-th (1≤s≤S) scale respectively.
[0079] Step 6: In the hyperedge feature extraction module, the multi-scale temporal feature representation is aggregated into a multi-scale hyperedge feature representation through the hyperedge mechanism.
[0080] like Figure 4 As shown, firstly, the multi-scale time feature is represented as {X 1 ,…,X s ,…,X S} is considered as a node and two parameters are initialized, namely the hyperedge embedding and node embedding Among them, M s is a hyperparameter, indicating the number of hyperedges at the sth scale. Then, the point-edge correlation matrix H of the sth scale is obtained by similarity calculation. s , which is calculated as follows:
[0081]
[0082] in, and is a learnable parameter, the tanh(·) activation function is used to perform nonlinear transformations, the ReLU(·) activation function is used to eliminate weak connections, Linear(·) is a linear mapping function, and the superscript T is a transpose operation. In order to enhance the robustness of the model and reduce noise interference in subsequent calculations, a sparseness strategy is designed. Its calculation formula is as follows:
[0083]
[0084] Among them, η∈[0,M s ] is the threshold of the TopK(·) function, which indicates the maximum number of neighbor hyperedges connected to the node, n is the node index, m is the hyperedge index, For all hyperedges connecting the nth node, is the s-th scale sparse point-edge correlation matrix. After the sparse strategy, we finally get the S-scale point-edge correlation matrix {H 1 ,…,H s ,…,H S}.
[0085] The i-th hyperedge at scale s Hyperedge features Based on the point-edge correlation matrix H s The information is aggregated and the calculation formula is as follows:
[0086]
[0087] Among them, Avg(·) is the averaging operation, is the point-edge correlation matrix H under scale s s The hyperedge shown in Connected neighbor nodes, is the jth node under scale s The temporal feature representation of , the set of all hyperedge features under the fusion scale s is ε s , and finally get the hyperedge feature representation of S scales {ε 1 ,…,ε s ,…,ε S}.
[0088] Step 7: Represent the multi-scale hyperedge feature {ε 1 ,…,ε s ,…,ε S and multi-scale text prototype representation {U 1 ,…,U s ,…,U S} is sent to the cross-modal alignment module to obtain the aligned multi-scale feature representation {Z 1 ,…,Z s ,…,Z S}. Among them Z s It represents the feature representation of the sth (1≤s≤S) scale after alignment.
[0089] Hyperedge feature representation ε at a given scale s s and text prototype representation U s , first map it to the corresponding query key Sum in is the index of the number of heads in attention, The total number of heads. and is the learnable mapping matrix under scale s, Then, the hyperedge feature representation and text prototype representation are aligned by calculating the cross-attention. The calculation formula is as follows:
[0090]
[0091] Among them, Attn(·) is the cross attention mechanism, softmax(·) is the activation function, and the superscript T is the transposition operation. For scale s The output of each attention head can be aggregated to obtain the aligned feature representation Z of the multi-head attention output at scale s. s Finally, the feature representation after alignment at different scales is expressed as {Z1 ,…,Z s ,…,Z S}.
[0092] Step 8: Execute the multi-scale hybrid prompt mechanism by constructing a learnable prompt Data-related tips And ability enhancement tips Generating multi-scale hybrid cues.
[0093] like Figure 5 As shown, the multi-scale hybrid prompt mechanism uses different types of prompts (i.e., learning prompts Data-related tips And ability enhancement tips ) to enrich the input context information, thereby enhancing the large language model's ability to understand the multi-scale patterns of power load sequences.
[0094] Learnable Hints: We construct learnable hints to enhance the large language model’s ability to understand multi-scale temporal patterns in power load sequences. Multi-scale learnable hints can be expressed as in, is the learnable hint at the sth scale, which is embedded by the learnable initialization, L s is the length of the learnable cue at scale s. Learning is done through a supervised loss between the model’s output and the true labels.
[0095] Data related tips: Figure 5 As shown in (a), data-related prompts are constructed by introducing three components These are data description prompts π, task description prompts τ, and data statistics prompts μ. Data description prompts provide the large language model with basic background information about the input power load sequence; task description prompts guide the large language model to understand and perform specific power load prediction-related tasks; and data statistics prompts provide statistical features of the power load sequence, including the input power load sequence and subsequences at different scales. The calculation formula is as follows:
[0096]
[0097] Among them, tokenizer(·) is the word segmenter.
[0098] Ability enhancement tips: Figure 5 As shown in (b), the ability to enhance prompts is enhanced by introducing three components That is, logical thinking prompts σ, emotional manipulation prompts And reasoning-related prompts μ. Logical thinking prompts guide the large language model to solve problems step by step, thereby improving its multi-step reasoning ability in power load sequences; emotional manipulation prompts simulate the impact of emotions on human decision-making, using "emotional blackmail" to make the model more focused on the current task; and reasoning-related prompts provide a usable method to guide the large language model in power load prediction. The final ability-enhancing prompt The calculation formula is as follows:
[0099]
[0100] Among them, tokenizer(·) is the word segmenter.
[0101] In step 9, the multi-scale mixed prompt and the aligned multi-scale feature representation are concatenated, and the concatenated multi-scale feature representation is fed into the frozen main body of the pre-trained large language model to generate the output result of the pre-trained large language model.
[0102] After obtaining the multi-scale mixed prompts, firstly transform the learnable prompts P at different scales into s and the aligned multi-scale feature representation Z s Then combine it with the data related prompts and ability enhancement tips The concatenation is performed and the concatenated result is fed into the frozen main body of the pre-trained large language model to obtain the output result of the pre-trained large language model. The calculation formula is as follows:
[0103]
[0104] Among them, [·,·] is the splicing operation, is the output of the pre-trained large language model LLMs(·).
[0105] Step 10: Input the output of the pre-trained large language model into the linear layer for linear transformation, and perform reverse instance normalization on the result after linear transformation to obtain the power load prediction value for the next H steps.
[0106] After obtaining the output of the pre-trained large language model, it is input into the linear layer for linear transformation, and the result after linear transformation is reversed instance normalized to obtain the final H-step power load prediction value. The calculation formula for reverse instance normalization is as follows:
[0107]
[0108] in, is the power load prediction value of the dth dimension with time step t after reverse instance normalization of the power load data, o′ t,d It is the result of linear transformation of the pre-trained large language model output with the d-th dimension time step being t.
[0109] Step 11: Calculate the prediction loss That is, the predicted value of the training sample and the corresponding true label The error between .
[0110] The present invention uses the square error as the prediction loss to calculate the predicted value of the training sample and the corresponding true label The error between them is calculated as follows:
[0111]
[0112] Step 12, based on the loss of all samples in the batch Adjust the learnable network parameters, node embeddings, and scale embeddings in the entire power load forecasting model.
[0113] The loss of all samples in the batch The calculation formula is as follows:
[0114]
[0115] in, is the loss of the bth sample in the batch, and B is the number of samples in each batch. The learnable network parameters, node embeddings, and scale embeddings (all learnable parameters are denoted as θ) in the entire power load forecasting model are adjusted, and the update formula is as follows:
[0116]
[0117] Among them, γ is the learning rate.
[0118] Step 13: Repeat steps 3-12 until all batches of the training dataset participate in the training of the power load prediction model.
[0119] Step 14: Repeat steps 3-13 until the specified number of iterations is reached.
[0120] Step 15: Input the power load sequence to be predicted into the trained power load prediction model to obtain a prediction result.
[0121] After preprocessing the load sequence to be predicted, it is fed into a trained load prediction model. The load prediction model predicts the load values for the next H moments based on the load data within the historical window and uses the prediction results for subsequent analysis and decision-making.
[0122] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A power load forecasting method based on multi-scale hypergraph neural network and large language model alignment, characterized in that: The following steps are involved: Preprocess the power load sequence and construct training samples; In the multi-scale feature extraction module, the training samples and word embeddings based on the pre-trained large language model are mapped into multi-scale temporal feature representations and multi-scale text prototype representations respectively; In the hyperedge feature extraction module, the hyperedge mechanism based on the multi-scale hypergraph neural network aggregates the multi-scale temporal feature representation into a multi-scale hyperedge feature representation; In the cross-modal alignment module, the multi-scale text prototype representation and the multi-scale hyperedge feature representation are aligned based on cross-attention to obtain the aligned multi-scale feature representation, and a multi-scale hybrid prompt is constructed; In the large model prediction module, the multi-scale mixed prompts and the aligned multi-scale feature representation are concatenated and input into the pre-trained large language model for power load prediction. Based on the prediction loss, the power load prediction model including the multi-scale feature extraction module, the hyperedge feature extraction module, the cross-modal alignment module and the large model prediction module is trained. The new power load sequence is input into the trained power load prediction model to obtain the prediction result.
2. The power load forecasting method based on multi-scale hypergraph neural network and large language model alignment according to claim 1 is characterized in that: In the multi-scale feature extraction module, the training samples and the word embedding based on the pre-trained large language model are mapped into multi-scale temporal feature representation and multi-scale text prototype representation respectively, including: Perform instance normalization on the training samples to obtain the instance normalized power load sequence T is the time window size. Based on the subsequence of the s-1th scale, the subsequence of the sth scale is generated through different aggregation windows. The subsequence of the sth scale is expressed as in is the power load sequence representation of the s-th scale time step t, D is the feature dimension, is the sequence length of the s-th scale, N s-1 is the subsequence length of the s-1th scale, l s-1 is the aggregation window size of the s-1th scale, and the calculation formula of the aggregation process is as follows: Among them, Agg(·) is the aggregation function, θ s-1 is the learnable parameter of the aggregation function at the s-1th scale, is the input training sample after instance normalization; Word Embedding Based on Pre-trained Large Language Model Where V is the vocabulary size and P is the hidden layer dimension of the pre-trained large language model. First, the word embedding U is transformed into an initial text prototype representation through a linear mapping Where V ′ <<V, and then a multi-scale text prototype representation is obtained through a linear mapping. The calculation formula of the linear mapping process is as follows: Among them, Linear(·) is the linear mapping function, U s is the text prototype representation of the s-th scale, U s-1 is the text prototype representation of the s-1th scale, λ s-1 is the learnable parameter of the linear mapping function at the s-1th scale, V s is the number of prototype representations of the s-th scale text; Finally, the S-scale temporal feature representation {X 1 ,…,X s ,…,X S } and text prototype representation {U 1 ,…,U s ,…,U S }, where X s and U s They are the temporal feature representation and text prototype representation of the s-th (1≤s≤S) scale respectively.
3. The power load forecasting method based on multi-scale hypergraph neural network and large language model alignment according to claim 1 is characterized in that: In the hyperedge feature extraction module, the hyperedge mechanism based on the multi-scale hypergraph neural network aggregates the multi-scale temporal feature representation into a multi-scale hyperedge feature representation, including: The multi-scale temporal feature {X 1 ,…,X s ,…,X S } is considered as a node and two parameters are initialized, namely the hyperedge embedding and node embedding Among them, X s is the temporal feature representation of the s-th (1≤s≤S) scale, M s is the number of hyperedges at the s-th scale, D is the feature dimension, and the point-edge correlation matrix H at the s-th scale is obtained by similarity calculation s , which is calculated as follows: in, and is a learnable parameter, the tanh(·) activation function is used to perform nonlinear transformation, the ReLU(·) activation function is used to eliminate weak connections, Linear(·) is a linear mapping function, the superscript T is a transpose operation, and finally the point-edge association matrix {H 1 ,…,H s ,…,H S }; The i-th hyperedge at scale s Hyperedge features Based on the point-edge correlation matrix H s The information is aggregated and the calculation formula is as follows: Among them, Avg(·) is the averaging operation, is the point-edge correlation matrix H under scale s s The hyperedge shown in Connected neighbor nodes, is the jth node under scale s The time feature representation of the scale s is the set of all hyperedge features. s , the final hyperedge features of S scales are expressed as {ε 1 ,…,ε s ,…,ε S }.
4. The power load forecasting method based on multi-scale hypergraph neural network and large language model alignment according to claim 3 is characterized in that: The sparsification strategy is used to calculate the point-edge correlation matrix of the sth scale, and the calculation formula is as follows: Among them, η∈[0,M s ] is the threshold of the TopK(·) function, which indicates the maximum number of neighbor hyperedges connected to the node, n is the node index, m is the hyperedge index, For all hyperedges connecting the nth node, is the s-th scale sparse point-edge correlation matrix, based on Get the final S-scale point-edge correlation matrix {H 1 ,…,H s ,…,H S }.
5. The power load forecasting method based on multi-scale hypergraph neural network and large language model alignment according to claim 1 is characterized in that: In the cross-modal alignment module, the aligned multi-scale feature representation obtained by aligning the multi-scale text prototype representation and the multi-scale hyperedge feature representation based on cross attention includes: Hyperedge feature representation ε at scale s based on multi-scale hyperedge feature representation s and the text prototype representation U at scale s in the multi-scale text prototype representation s , first map it to the corresponding query key Sum in is the index of the number of heads in attention, is the total number of heads, and is the mapping matrix that can be learned under scale s, D is the feature dimension, P is the hidden layer dimension of the pre-trained large language model, Then, the cross attention is calculated to align the hyperedge feature representation and the text prototype representation. The calculation formula is as follows: Among them, Attn(·) is the cross attention mechanism, softmax(·) is the activation function, and the superscript T is the transposition operation. For scale s The output of each attention head is aggregated to obtain the aligned feature representation Z of the multi-head attention output at scale s (1≤s≤S) s , and finally obtain the aligned S scale multi-scale feature representation {Z 1 ,…,Z s ,…,Z S }.
6. The power load forecasting method based on multi-scale hypergraph neural network and large language model alignment according to claim 1 is characterized in that: Multi-scale hybrid cues include: Learning tips Expressed as in, is the learnable hint at the sth scale and is embedded by the learnable initialization, 1≤s≤S, S is the number of scales, L s is the length of the learnable cue at scale s, and D is the feature dimension; Data-related tips These include data description prompts π that provide basic background information of the input power load sequence to the pre-trained large language model, task description prompts τ that guide the pre-trained large language model to understand and perform specific power load prediction-related tasks, and data statistics prompts μ that include statistical features of the power load sequence, including the input power load sequence and subsequences at different scales. Ability enhancement tips Including logical thinking prompts σ that guide pre-trained large language models to gradually solve problems, and emotional manipulation prompts that simulate the impact of emotions on human decision-making and the inference-related class hints μ that guide the available methods of pre-training large language models for power load forecasting.
7. The power load forecasting method based on multi-scale hypergraph neural network and large language model alignment according to claim 6 is characterized in that: In the large model prediction module, the multi-scale mixed prompts and the aligned multi-scale feature representations are spliced and then input into the large language model for power load prediction, including: After obtaining the multi-scale mixed prompts, firstly transform the learnable prompts {P 1 ,...,P s ,...,P S } and the aligned multi-scale feature representation {Z 1 ,…,Z s ,…,Z S } is spliced, and then further spliced with data-related prompts and ability enhancement prompts, and the spliced results are input into the pre-trained large language model to obtain the output result. The calculation formula is as follows: Among them, [·,·] is the splicing operation, Z s is the aligned feature representation at the sth (1≤s≤S) scale, The output of the pre-trained large language model LLMs(1).
8. The power load forecasting method based on multi-scale hypergraph neural network and large language model alignment according to claim 1 is characterized in that: Keep the main parameters of the pre-trained large language model frozen during training.
9. The power load forecasting method based on multi-scale hypergraph neural network and large language model alignment according to claim 2 is characterized in that: The output of the large language model is input into the linear layer for linear transformation, and the result after linear transformation is reverse instance normalized to obtain the power load prediction value for the next H steps. As a result of power load forecasting.
Citation Information
Patent Citations
Pyramid-type recurrent neural network-based multi-scale power load prediction method
CN117077074A
Power risk prediction method based on lama2 big language model
CN118313657A
Non-Intrusive Load Decomposition Method Based on Informer Model Coding Structure
US20220397874A1