Power system source load prediction method and related device
By using an improved Decoder-Only architecture for large language models and a Shapley additive interpretation method, the problem of low source-load prediction accuracy in power systems is solved, achieving high-precision source-load prediction and model transparency, thereby improving the stability of power systems.
Patent Information
- Application Number
- CN202510995081.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
In power systems with a high proportion of renewable energy integration, existing technologies have low source-load prediction accuracy, making it difficult to meet the requirements for stable operation.
Based on a large language model with a Decoder-Only architecture, this paper improves the source-load prediction accuracy by replacing the input and output layers, fine-tuning the root mean square normalization module of the decoder layer during pre-training, freezing the attention module and the multilayer perceptron module, designing an adaptive output layer, and combining the Shapley additive interpretation method.
It effectively improves the accuracy of power system source-load forecasting, provides model transparency and interpretability, and enhances dispatchers' confidence in the forecast results.
Smart Images

Figure CN120875150A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power forecasting and relates to a power system source-load forecasting method and related apparatus. Background Technology
[0002] With the integration of a high proportion of renewable energy sources, the time-varying nature of power systems is gradually increasing. Business scenarios such as development planning, source-load forecasting, and power generation forecasting urgently require high-precision source-load forecasting capabilities to provide a foundation for understanding the complex characteristics of new power systems. However, the integration of a high proportion of renewable energy sources and power electronic equipment results in a high degree of uncertainty in the source-load characteristics of the power system, posing a severe challenge to the stable operation of the power system.
[0003] Artificial intelligence models can uncover hidden patterns from massive amounts of historical data, enabling accurate predictions of renewable energy output and load demand. However, when applying traditional AI models to power system source-load forecasting, the prediction accuracy is often low due to limitations such as small model size and limited training data. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a power system source load prediction method and related apparatus.
[0005] To achieve the above objectives, the present invention employs the following technical solution:
[0006] In a first aspect, the present invention provides a power system source-load prediction method, comprising: acquiring power system source-load influencing variables; inputting the power system source-load influencing variables into a pre-trained source-load prediction large model to obtain the power system source-load prediction result; wherein the pre-trained source-load prediction large model is obtained by: using a large language model with a Decoder-Only architecture as the base model, replacing the input layer of the base model with a patch layer, a normalization layer, and a first fully connected layer connected in sequence, and replacing the output layer of the base model with a second fully connected layer to obtain the source-load prediction large model and performing pre-training to obtain the pre-trained source-load prediction large model; wherein, during the pre-training process, the root mean square layer normalization module of the decoder layer of the source-load prediction large model is fine-tuned, and the attention module and multilayer perceptron module of the decoder layer are frozen; the output dimension of the first fully connected layer is the same as the hidden layer dimension of the base model; the output dimension of the second fully connected layer is the same as the time scale of the source-load prediction result.
[0007] Optionally, the normalization layer may be a reversible instance normalization layer.
[0008] Optionally, it also includes: using the Shapley additive interpretation method to obtain the average marginal gain of each variable in the power system source load influence variables on the power system source load prediction results; and obtaining the key variables of the power system source load prediction results based on the average marginal gain of each variable on the power system source load prediction results.
[0009] Optionally, the step of using the Shapley additive interpretation method to obtain the average marginal gain of each variable in the power system source load influence variables on the power system source load prediction results includes: obtaining the average marginal gain φ of each variable on the power system source load prediction results using the following formula. i :
[0010]
[0011] in, For the set of power system source-load influence variables; x i Let x be the i-th variable in the power system source-load influence variables; S is the set of power system source-load influence variables that does not include x. i The feature subset; m is the number of variables; f(x) is the pre-trained source load prediction large model based on The source load forecast results of the power system; f S (x) represents the source load prediction results of the pre-trained source load prediction large model based on the S power system.
[0012] In a second aspect, the present invention provides a power system source-load prediction system, comprising: a data acquisition module for acquiring power system source-load influencing variables; and a source-load prediction module for inputting the power system source-load influencing variables into a pre-trained source-load prediction large model to obtain the power system source-load prediction result; wherein the pre-trained source-load prediction large model is obtained by: using a Decoder-Only architecture large language model as the base model, replacing the input layer of the base model with a patch layer, a normalization layer, and a first fully connected layer connected in sequence, and replacing the output layer of the base model with a second fully connected layer, thereby obtaining the source-load prediction large model and performing pre-training to obtain the pre-trained source-load prediction large model; wherein, during the pre-training process, the root mean square layer normalization module of the decoder layer of the source-load prediction large model is fine-tuned, and the attention module and multilayer perceptron module of the decoder layer are frozen; the output dimension of the first fully connected layer is the same as the hidden layer dimension of the base model; and the output dimension of the second fully connected layer is the same as the time scale of the source-load prediction result.
[0013] Optionally, the normalization layer may be a reversible instance normalization layer.
[0014] Optionally, a model interpretation module is also included; the model interpretation module is used to: obtain the average marginal gain of each variable in the power system source load influence variables on the power system source load prediction results using the Shapley additive interpretation method; and obtain the key variables of the power system source load prediction results based on the average marginal gain of each variable on the power system source load prediction results.
[0015] Optionally, the model interpretation module is specifically used to: obtain the average marginal gain φ of each variable on the source-load prediction results of the power system using the following formula. i :
[0016]
[0017] in, For the set of power system source-load influence variables; x i Let x be the i-th variable in the power system source-load influence variables; S is the set of power system source-load influence variables that does not include x. i The feature subset; m is the number of variables; f(x) is the pre-trained source load prediction large model based on The source load forecast results of the power system; f S (x) represents the source load prediction results of the pre-trained source load prediction large model based on the S power system.
[0018] In a third aspect, the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the power system source-load prediction method described above.
[0019] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the power system source-load prediction method described above.
[0020] Compared with the prior art, the present invention has the following beneficial effects:
[0021] This invention presents a power system source-load prediction method based on a pre-trained large-scale source-load prediction model. It achieves power system source-load prediction based on power system source-load influencing variables. The pre-trained large-scale source-load prediction model is obtained by optimizing a Decoder-Only architecture large language model as the base model. It fully leverages the advantages of the large language model's large model structure and abundant training data. This is achieved by replacing the input layer of the base model with a sequentially connected patch layer, normalization layer, and first fully connected layer, and by replacing the output layer of the base model with a second fully connected layer. Furthermore, the output dimension of the first fully connected layer is designed to be the same as the hidden layer dimension of the base model, and the output dimension of the second fully connected layer is designed to be the same as the time scale of the source-load prediction result. This enables the adaptive application of the base model in source-load prediction, transferring the powerful text generation capabilities of the large oracle model to the source-load prediction scenario. Simultaneously, during pre-training, the root mean square normalization module of the decoder layer of the large-scale source-load prediction model is fine-tuned, and the attention module and multilayer perceptron module of the decoder layer are frozen. This effectively preserves the prior knowledge of the large oracle model, providing a knowledge foundation for source-load prediction and ultimately improving the accuracy of power system source-load prediction. Attached Figure Description
[0022] Figure 1 This is a flowchart of the power system source-load prediction method according to an embodiment of the present invention.
[0023] Figure 2 This is a block diagram of the power system source-load prediction system according to an embodiment of the present invention. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] The present invention will now be described in further detail with reference to the accompanying drawings:
[0027] See Figure 1 In one embodiment of the present invention, a power system source load prediction method is provided, which realizes high-precision source load prediction of the power system through a pre-trained large source load prediction model.
[0028] Specifically, the power system source-load prediction method of the present invention includes the following steps:
[0029] S1: Obtain the source-load influence variables of the power system.
[0030] S2: Input the power system source load influence variables into the pre-trained source load prediction large model to obtain the power system source load prediction results.
[0031] The pre-trained source-charge prediction model is obtained as follows: Using a Decoder-Only architecture-based large language model as the base model, the input layer of the base model is replaced with a patch layer, a normalization layer, and a first fully connected layer connected in sequence. The output layer of the base model is replaced with a second fully connected layer. This results in a pre-trained source-charge prediction model. During pre-training, the root mean square (RMS) normalization module of the decoder layer of the source-charge prediction model is fine-tuned, and the attention module and multilayer perceptron module of the decoder layer are frozen. The output dimension of the first fully connected layer is the same as the hidden layer dimension of the base model; the output dimension of the second fully connected layer is the same as the time scale of the source-charge prediction result.
[0032] This invention presents a power system source-load prediction method based on a pre-trained large-scale source-load prediction model. It achieves power system source-load prediction based on power system source-load influencing variables. The pre-trained large-scale source-load prediction model is obtained by optimizing a Decoder-Only architecture large language model as the base model. It fully leverages the advantages of the large language model's large model structure and abundant training data. This is achieved by replacing the input layer of the base model with a sequentially connected patch layer, normalization layer, and first fully connected layer, and by replacing the output layer of the base model with a second fully connected layer. Furthermore, the output dimension of the first fully connected layer is designed to be the same as the hidden layer dimension of the base model, and the output dimension of the second fully connected layer is designed to be the same as the time scale of the source-load prediction result. This enables the adaptive application of the base model in source-load prediction, transferring the powerful text generation capabilities of the large oracle model to the source-load prediction scenario. Simultaneously, during pre-training, the root mean square normalization module of the decoder layer of the large-scale source-load prediction model is fine-tuned, and the attention module and multilayer perceptron module of the decoder layer are frozen. This effectively preserves the prior knowledge of the large oracle model, providing a knowledge foundation for source-load prediction and ultimately improving the accuracy of power system source-load prediction.
[0033] Interpretive power system source-load influence variables are generally time-series variables used to represent the dynamic characteristics of the system over time. These typically include key time-series variables on both the power generation and load sides. Key time-series variables on the power generation side include variables related to renewable energy generation, traditional energy generation, and grid operation. Key time-series variables on the load side include weather-related variables, social activity variables (such as weekdays / holidays), and user behavior variables (such as electric vehicle charging schedules).
[0034] Explained, in mainstream large language model architectures, the Decoder-Only architecture generates sequences through autoregression, meaning each step predicts the next time point based on historical information, perfectly consistent with the recursive nature of time series prediction. Therefore, the Decoder-Only architecture is selected as the base model. For example, the Qwen large language model with a Decoder-Only architecture is chosen.
[0035] The Decoder-Only architecture is a neural network design based on a self-attention mechanism. It only includes the decoder part of the Transformer model, omitting the encoder. This architecture is widely used in natural language processing and sequence generation tasks due to its simplicity, efficiency, and powerful generative capabilities. The Qwen large-scale model is a general-purpose language model independently developed by Alibaba Cloud based on the Decoder-Only architecture.
[0036] Interpretively, since source load prediction is deterministic point prediction, and the input numerical sequence differs significantly from the language sequence used by the large language model in terms of physical meaning and data dimensionality, the input layer must be redesigned to project the time series data to the required dimensions of the base model. Specifically, the input layer of the base model is replaced with a patch layer, a normalization layer, and a first fully connected layer connected in sequence.
[0037] The patch layer aggregates adjacent time steps through patch operations to form subsequence blocks, enhancing local temporal information. Each input sample is first divided into multiple subsequences. The length of a subsequence is denoted as P, and the interval between two consecutive subsequences is denoted as S. When the interval S < 0, it means that adjacent subsequences overlap by S time points. When the interval S > 0, it means that adjacent subsequences are separated by S time points. For a sample with a sequence length of L, the patching process will generate a set of subsequences. The number of subsequences N is:
[0038] The normalization layer normalizes the input samples, accelerating model convergence and improving training stability.
[0039] The newly constructed input layer uses a first fully connected layer as the input embedding layer. The output dimension of the first fully connected layer is consistent with the hidden layer dimension of the base model, ensuring that the content input from the input layer into the base model matches the dimension d calculated by the attention mechanism within the base model.
[0040] X embed =X input ·W+b
[0041] Among them, X input Let X be an N×D dimensional input vector. embed The input vector sequence is embedded, and W and b are the parameters of the first fully connected layer, which maps the D-dimensional input vector at each time step to a d-dimensional vector.
[0042] Interpretive, during pre-training, the root mean square layer normalization module of the decoder layer of the large model predicting source loads is fine-tuned, and the attention module and multilayer perceptron module of the decoder layer are frozen.
[0043] Specifically, since the decoder layer of a large source-load prediction model generally uses multiple stacked layers, this embodiment will use a single-layer decoder layer as an example for explanation.
[0044] For the first root mean square normalization (RMSNorm) module of the Decoder layer, the input representation X NEntering the Decoder layer in a hidden state, the first step is root mean square normalization, calculated using the following formula:
[0045]
[0046] Here, γ1 is a learnable scaling parameter, and β1 is a learnable offset parameter. Without fine-tuning this module, the normalized distribution of the time-series data, after passing through this layer, will be influenced by γ1 and β1, aligning with the red curve representing the semantic data distribution. Due to the inherent distributional differences between the two types of data, this process will lead to instability during training. Therefore, the parameters of this module will be continuously fine-tuned during pre-training.
[0047] For the attention module of the Decoder layer, the normalized hidden states are entered into the attention module for computation. During attention computation, positional encoding is performed on the input sequence to enable the model to recognize the sequential relationships between tokens. Rotated Position Encoding (RoPE) is used to represent the sequential relationships between tokens.
[0048] Let the function f(·) denote the location encoding, then the query vector q and key vector k integrating location information are expressed as: Where m and n are the token positions in the sequence, respectively; x m x n These are the normalized hidden states X. norm The word vectors corresponding to the m-th and n-th tokens.
[0049] Based on the geometric properties of spatial vectors and the properties of complex numbers, the positional encoding of q and k is as follows:
[0050]
[0051]
[0052] Where, Θ={θ i =10000 -2(i-1) / d ,i∈[1,2,...,d / 2]}, It is a d-dimensional rotation matrix, W {q,k} θ is the parameter matrix for q and k. d / 2-1 A fixed set of angle parameters is used to construct the rotation matrix.
[0053] Applying RoPE to attention calculation yields the output Attn. out :
[0054]
[0055] Among them, W V M is the parameter matrix of the value matrix V. maskIt is the mask matrix at the corresponding position.
[0056] The attention output, after passing through the residual connection, will yield a new hidden state X. h,1 :
[0057] X h,1 =X norm +Attn out
[0058] As can be seen from the above process, since most of the knowledge learned by the basic model is stored in the attention module, and RoPE position encoding is an absolute position encoding without learnable parameters, the attention module is frozen during the pre-training of the large source-load prediction model to provide a knowledge base for subsequent source-load prediction.
[0059] For the second RMSNorm module in the Decoder layer, the output X of the attention module is... h,1 It is fed into the second RMSNorm module and normalized to X. h,rms The process is similar to that described above, so this module needs to be fine-tuned.
[0060] For the Multilayer Perceptron (MLP) module in the Decoder layer, X h,rms The input is fed into the MLP module for feature transformation, resulting in the output X of the Decoder layer. h,2 :
[0061] X h,2 =X h,1 +MLP(X h,rms )
[0062] =W down ·(SwiGLU(W gate X h,rms +b gate ))
[0063] ⊙(W up X h,rms +b up )+b down +X h,1
[0064] Where ⊙ represents the Hadamarda unit, W down W gate W up These represent the parameter matrices of the three linear layers in the MLP module, b down b gate b up The subscript represents the corresponding bias term, and SwiGLU(·) represents the SwiGLU activation function.
[0065] As can be seen, the MLP module performs non-linear mapping on the features extracted by the attention module, enhancing the model's expressive power, enabling cross-dimensional interaction of features, and uncovering complex relationships between features in the input time series. Therefore, this module is frozen to preserve the prior knowledge of the basic model.
[0066] Interpretive large language models have their outputs mapped to vectors of the same dimension as the semantic model's vocabulary. Therefore, for large source-load prediction models, a new output layer needs to be designed to replace the original output layer of the base model and retrained. In this implementation, the output layer of the base model is replaced with a second fully connected layer, and the output dimension of the second fully connected layer is designed to have the same time scale as the source-load prediction results.
[0067] In one possible implementation, the normalization layer employs a reversible instance normalization layer.
[0068] Interpretively, given that the input for time series prediction may be a multivariate sequence, and the sequence features of each variable are relatively independent, compared to the RMSNorm and Layer Norm commonly used in natural language tasks for normalization from the sample dimension, this implementation uses Reversible Instance Normalization (RevIN layer) as the normalization layer. The RevIN layer, based on the RevIN operation, relies only on the mean and variance to independently normalize each variable channel, preserving the integrity of variable features and improving training stability and convergence.
[0069] The RevIN layer normalization process is as follows:
[0070]
[0071] Where μ is the mean, σ is the variance, γ and β are learnable scaling and translation transformation parameters, and ε is a very small constant to avoid the case where the variance is 0. The sequence y predicted by the model will be inversely normalized using the same parameter values as the RevIN layer, calculated as follows:
[0072] In one possible implementation, the power system source load prediction method further includes: using the Shapley additive interpretation method to obtain the average marginal gain of each variable in the power system source load influencing variables on the power system source load prediction result; and obtaining the key variables of the power system source load prediction result based on the average marginal gain of each variable on the power system source load prediction result.
[0073] Interpretive Shapley Additive Interpretation (SHAP) is a machine learning model interpretation method based on Shapley values in game theory. It helps to understand the model's decision-making process and improves the model's transparency and credibility by quantifying the contribution of each feature to the model's output.
[0074] The SHAP method originates from the concept of Shapley value in game theory. Let the input feature set be... For a given input sample The model output is denoted as f(x), and the goal is to express this output as the sum of the linearly weighted contributions of each input feature: Where φ0 represents the reference output (such as the average output of the entire sample), φ i Indicates feature x i The average marginal gain of the current prediction results.
[0075] In one possible implementation, obtaining the average marginal gain of each variable in the power system source load influence variables on the power system source load prediction results using the Shapley additive interpretation method includes: obtaining the average marginal gain φ of each variable on the power system source load prediction results using the following formula. i :
[0076]
[0077] in, For the set of power system source-load influence variables; x i Let x be the i-th variable in the power system source-load influence variables; S is the set of power system source-load influence variables that does not include x. i The feature subset; m is the number of variables; f(x) is the pre-trained source load prediction large model based on The source load forecast results of the power system; f S (x) represents the source load prediction results of the pre-trained source load prediction large model based on the S power system.
[0078] For example, f S (x) can be estimated using Monte Carlo sampling or fitting approximation methods.
[0079] This formula reflects the characteristic x i The expected contribution of the marginal effect of the model output to each possible cooperative subset ensures the consistency and fairness of the interpretation results across the feature dimensions.
[0080] Based on the average marginal gain of each variable on the power system's source load forecast results, the key variables of the power system's source load forecast results can be analyzed. For example, variables with higher average marginal gains can be used as the key variables of the power system's source load forecast results, thereby providing an explanation for the power system's source load forecast results and helping decision-makers understand the reasons for the source load forecast results and fluctuations.
[0081] Meanwhile, based on the average marginal gain of each variable on the power system's source-load prediction results, it is also possible to optimize the structure of the large source-load prediction model, such as by eliminating unimportant variables to optimize the model's input layer structure.
[0082] In summary, the power system source-load prediction method of this invention uses the SHAP method to quantify the contribution of variables to the source-load prediction results, providing a decision basis for adjusting the model input layer structure and optimizing weights, providing transparency and interpretability to the source-load prediction model, and enhancing dispatchers' confidence in the prediction results.
[0083] The following are embodiments of the apparatus of the present invention, which can be used to execute embodiments of the method of the present invention. For details not disclosed in the apparatus embodiments, please refer to the embodiments of the method of the present invention.
[0084] See Figure 2 In another embodiment of the present invention, a power system source load prediction system is provided, which can be used to implement the above-mentioned power system source load prediction method. Specifically, the power system source load prediction system includes a data acquisition module and a source load prediction module.
[0085] The data acquisition module is used to acquire the power system source-load influence variables; the source-load prediction module is used to input the power system source-load influence variables into a pre-trained source-load prediction model to obtain the power system source-load prediction results; the pre-trained source-load prediction model is obtained in the following way: using a Decoder-Only architecture large language model as the base model, the input layer of the base model is replaced with a patch layer, a normalization layer, and a first fully connected layer connected in sequence, and the output layer of the base model is replaced with a second fully connected layer, thus obtaining the source-load prediction model, which is then pre-trained; during the pre-training process, the root mean square normalization module of the decoder layer of the source-load prediction model is fine-tuned, and the attention module and multilayer perceptron module of the decoder layer are frozen; the output dimension of the first fully connected layer is the same as the hidden layer dimension of the base model; the output dimension of the second fully connected layer is the same as the time scale of the source-load prediction results.
[0086] In one possible implementation, the normalization layer employs a reversible instance normalization layer.
[0087] In one possible implementation, the power system source-load prediction system further includes a model interpretation module; the model interpretation module is used to: use the Shapley additive interpretation method to obtain the average marginal gain of each variable in the power system source-load influence variables on the power system source-load prediction results; and obtain the key variables of the power system source-load prediction results based on the average marginal gain of each variable on the power system source-load prediction results.
[0088] In one possible implementation, the model interpretation module is specifically used to: obtain the average marginal gain φ of each variable on the source-load prediction results of the power system using the following formula. i :
[0089]
[0090] in, For the set of power system source-load influence variables; x i Let x be the i-th variable in the power system source-load influence variables; S is the set of power system source-load influence variables that does not include x. i The feature subset; m is the number of variables; f(x) is the pre-trained source load prediction large model based on The source load forecast results of the power system; f S (x) represents the source load prediction results of the pre-trained source load prediction large model based on the S power system.
[0091] All relevant content of each step involved in the aforementioned embodiments of the power system source load prediction method can be referenced to the functional description of the corresponding functional module of the power system source load prediction system in the embodiments of the present invention, and will not be repeated here.
[0092] The module division in this embodiment of the invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the invention can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0093] In another embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of a power system source-load prediction method.
[0094] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the power system source-load prediction method in the above embodiments.
[0095] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0096] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0097] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0098] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A power system source-load prediction method, characterized in that, include: Obtain the source-load influence variables of the power system; The power system source load influencing variables are input into a pre-trained large-scale source load prediction model to obtain the power system source load prediction results. The pre-trained source-charge prediction model is obtained as follows: Using a Decoder-Only architecture-based large language model as the base model, the input layer of the base model is replaced with a patch layer, a normalization layer, and a first fully connected layer connected in sequence. The output layer of the base model is replaced with a second fully connected layer. This results in a pre-trained source-charge prediction model. During pre-training, the root mean square (RMS) normalization module of the decoder layer of the source-charge prediction model is fine-tuned, and the attention module and multilayer perceptron module of the decoder layer are frozen. The output dimension of the first fully connected layer is the same as the hidden layer dimension of the base model; the output dimension of the second fully connected layer is the same as the time scale of the source-charge prediction result.
2. The power system source-load prediction method according to claim 1, characterized in that, The normalization layer employs a reversible instance normalization layer.
3. The power system source-load prediction method according to claim 1, characterized in that, Also includes: The Shapley additive interpretation method is used to obtain the average marginal gain of each variable in the power system source load influence variables on the power system source load prediction results; And based on the average marginal gain of each variable on the power system's source load prediction results, the key variables for the power system's source load prediction results are obtained.
4. The power system source-load prediction method according to claim 3, characterized in that, The method of employing the Shapley additive interpretation to obtain the average marginal gain of each variable in the power system source load influence variables on the power system source load prediction results includes: The average marginal gain φ of each variable on the power system source load prediction results is obtained by the following formula. i : in, For the set of power system source-load influence variables; x i Let x be the i-th variable in the power system source-load influence variables; S is the set of power system source-load influence variables that does not include x. i The feature subset; m is the number of variables; f(x) is the pre-trained source load prediction large model based on The source load forecast results of the power system; f S (x) represents the source load prediction results of the pre-trained source load prediction large model based on the S power system.
5. A power system source-load prediction system, characterized in that, include: The data acquisition module is used to acquire the source-load influence variables of the power system; The source load prediction module is used to input the source load influencing variables of the power system into the pre-trained source load prediction model to obtain the source load prediction results of the power system. The pre-trained source-charge prediction model is obtained as follows: Using a Decoder-Only architecture-based large language model as the base model, the input layer of the base model is replaced with a patch layer, a normalization layer, and a first fully connected layer connected in sequence. The output layer of the base model is replaced with a second fully connected layer. This results in a pre-trained source-charge prediction model. During pre-training, the root mean square (RMS) normalization module of the decoder layer of the source-charge prediction model is fine-tuned, and the attention module and multilayer perceptron module of the decoder layer are frozen. The output dimension of the first fully connected layer is the same as the hidden layer dimension of the base model; the output dimension of the second fully connected layer is the same as the time scale of the source-charge prediction result.
6. The power system source-load prediction system according to claim 5, characterized in that, The normalization layer employs a reversible instance normalization layer.
7. The power system source-load prediction system according to claim 5, characterized in that, It also includes a model interpretation module; the model interpretation module is used for: The Shapley additive interpretation method is used to obtain the average marginal gain of each variable in the power system source load influence variables on the power system source load prediction results; and based on the average marginal gain of each variable on the power system source load prediction results, the key variables of the power system source load prediction results are obtained.
8. The power system source-load prediction system according to claim 7, characterized in that, The model interpretation module is specifically used for: The average marginal gain φ of each variable on the power system source load prediction results is obtained by the following formula. i : in, For the set of power system source-load influence variables; x i Let x be the i-th variable in the power system source-load influence variables; S is the set of power system source-load influence variables that does not include x. i The feature subset; m is the number of variables; f(x) is the pre-trained source load prediction large model based on The source load forecast results of the power system; f S (x) represents the source load prediction results of the pre-trained source load prediction large model based on the S power system.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the power system source-load prediction method as described in any one of claims 1 to 4.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the power system source-load prediction method as described in any one of claims 1 to 4.