A Cross-Regional Load Forecasting Method, System, Device and Storage Medium Based on a Shared Encoder
Through the shared encoder and multi-task learning framework, data sparsity, regional differences and data imbalance in cross-region load prediction are solved, and high-precision cross-region load prediction is achieved, which improves the generalization ability of the model.
Patent Information
- Application Number
- CN202411726152.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-11-28
AI Technical Summary
The existing load prediction methods face data sparsity, regional differences, long-term dependency capture and data imbalance when processing cross-region load data, resulting in reduced prediction accuracy and insufficient model generalization capabilities.
A cross-region load prediction method based on shared encoder is adopted, and a shared encoder with a shared covariate pre-training data set and a self-attention mechanism is constructed, combining a multi-task learning framework and a dynamic weighted loss function, the load data in different regions is predicted.
It effectively solves the problems of data sparsity, regional differences and data imbalance in cross-region load prediction, improves prediction accuracy and generalization capabilities of models, and significantly improves the accuracy of cross-region load prediction.
Smart Images

Figure CN119627886B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system load forecasting, and particularly to a cross-regional load forecasting method, system, device and storage medium based on a shared encoder. Background Art
[0002] Power system load forecasting is an important basis for power system planning and operation. With the continuous expansion of the scale of the power system and the increasing tightness of cross-regional interconnection, the importance of cross-regional load forecasting has become increasingly prominent. However, the existing load forecasting methods face the following challenges when dealing with cross-regional load data:
[0003] (1) Data sparsity: The historical data of some regions may be insufficient, resulting in a decrease in prediction accuracy.
[0004] (2) Regional differences: The load characteristics and influencing factors in different regions may vary significantly, making it difficult to accurately predict with a unified model.
[0005] (3) Long-term dependence relationship: Load data usually has complex long-term dependence relationships, which are difficult to effectively capture by traditional methods.
[0006] (4) Data imbalance: The distribution of different types of samples in the dataset may be uneven, affecting the generalization ability of the model.
[0007] Therefore, there is an urgent need for a cross-regional load forecasting method that can effectively handle the above problems. Summary of the Invention
[0008] Object of the Invention: A cross-regional load forecasting method, system, device and storage medium based on a shared encoder of the present invention can solve the problems of data sparsity, regional differences, capture of long-term dependence relationships and data imbalance existing in the existing load forecasting methods when dealing with cross-regional load data.
[0009] Technical Solution: A cross-regional load forecasting method based on a shared encoder of the present invention includes:
[0010] Collect the load historical data and covariate features of different regions, and perform data preprocessing. Use the preprocessed load historical data and covariate features of different regions to construct a specific regional dataset for each region; Considering the different data bases of different regions, select the intersection of the covariates of each region to construct a common covariate pre-training dataset;
[0011] Construct a time series prediction backbone network. The time series prediction backbone network includes an input encoding layer and a shared encoder based on the self-attention mechanism. Among them, the input encoding layer realizes the embedded encoding of time series data with different frequencies, and the shared encoder based on the self-attention mechanism is used to capture the long-term dependence relationships of time series data in different regions;
[0012] Design a multi-task learning framework, which includes a multi-task learning network architecture and a dynamic weighted loss function that automatically adjusts weights based on the rarity of samples;
[0013] Pre-train the shared encoder based on the self-attention mechanism using a common covariate pre-training dataset; after pre-training, use the specific regional datasets of each region, combined with the multi-task learning framework, to fine-tune the shared encoder based on the self-attention mechanism in the specific region to complete the training of the shared encoder based on the self-attention mechanism;
[0014] Use the trained shared encoder based on the self-attention mechanism to perform time series prediction on large-scale, cross-regional load data;
[0015] Evaluate the performance of the shared encoder based on the self-attention mechanism according to the time series prediction results. If the performance of the shared encoder based on the self-attention mechanism does not meet the standard, re-fine-tune the shared encoder based on the self-attention mechanism..
[0016] Furthermore, the expression of the self-attention mechanism is as follows:
[0017] Q = XW Q , K = XW K , V = XW V
[0018]
[0019] where X represents the input sequence; W Q , W K , W V represent learnable weight matrices; d k represents the dimension of the attention head; Q represents the query vector; K represents the key vector; V represents the value vector.
[0020] Furthermore, the improved position encoding method includes:
[0021]
[0022]
[0023] where pos represents the position index; i represents the dimension index, d model represents the dimension of the model; in addition, introduce additional periodic encoding:
[0024]
[0025] where T u represents different periods.
[0026] Furthermore, the multi-task learning network architecture includes:
[0027] Input encoding layer:
[0028] Divide the data into time series blocks of different lengths according to the frequency of the data, then perform a fully connected layer calculation on the divided time series blocks, and then output the time series encoding of the specified dimension.
[0029] Shared encoder:
[0030] Based on the time series prediction backbone network, extract the features shared by all tasks, and its mathematical representation is:
[0031] H = Encoder(X)
[0032] where X represents the input sequence; H represents the shared features;
[0033] Task-specific layer:
[0034] For each task, add a task-specific layer to capture the unique features of that task; for the i-th task, its specific layer can be expressed as:
[0035] H_i = f_i(H)
[0036] where f_i represents the task-specific transformation function, which adopts the cascade method of multiple convolutional layers and fully connected layers;
[0037] Output layer:
[0038] Each task has an output layer for generating the final prediction; for the i-th task, its output can be expressed as:
[0039] Y_i = g_i(H_i)
[0040] where g_i represents the output function, which adopts a fully connected layer and uses a linear activation.
[0041] Furthermore, the dynamic weighted loss function that automatically adjusts the weight based on the rarity of samples has the following expression:
[0042]
[0043] In the formula, l i represents the loss of the i-th sample; w i represents the sample weight, and N represents the total number of samples;
[0044] where the calculation formula of the sample weight w i is as follows:
[0045]
[0046] In the formula, f i represents the frequency of the i-th sample in the dataset; represents the sum of all sample frequencies;
[0047] To further balance the weights, the softmax function is used to normalize the sample weights w i as follows:
[0048]
[0049] Furthermore, the shared encoder based on the self-attention mechanism is pre-trained using the common covariate pre-training dataset, including:
[0050] The shared encoder based on the self-attention mechanism is pre-trained on the common covariate pre-training dataset, and the optimization objective is to minimize the prediction error:
[0051]
[0052] In the formula, y represents the true load value; represents the predicted value; L reg represents the regularization term; L pre represents the loss function; λ represents the regularization coefficient.
[0053] Furthermore, after pre-training, the shared encoder based on the self-attention mechanism is fine-tuned in specific regions using the specific region datasets of each region, in combination with the multi-task learning framework, to complete the training of the shared encoder based on the self-attention mechanism, including:
[0054] The shared encoder based on the self-attention mechanism is fine-tuned in specific regions using the specific region datasets of each region, in combination with the multi-task learning framework:
[0055]
[0056] In the formula, L r represents the task loss of the r-th region; L share represents the shared task loss; α r and β represent the balance coefficients.
[0057] Based on the same inventive concept, a cross-region load prediction system based on a shared encoder of the present invention includes:
[0058] A data acquisition and preprocessing module, which is used to collect the load historical data and covariate features of different regions, perform data preprocessing, and use the preprocessed load historical data and covariate features of different regions to construct a specific regional dataset for each region; considering the different data bases of different regions, select the intersection of covariates of each region to construct a common covariate pre-training dataset;
[0059] A model construction module, which is used to construct a time series prediction backbone network. The time series prediction backbone network includes an input encoding layer and a shared encoder based on the self-attention mechanism. Among them, the input encoding layer realizes the embedding encoding of time series data with different frequencies, and the shared encoder based on the self-attention mechanism is used to capture the long-term dependence relationship of time series data in different regions;
[0060] A multi-task learning framework establishment module, which is used to design a multi-task learning framework. The multi-task learning framework includes a multi-task learning network architecture and a dynamic weighted loss function that automatically adjusts weights based on the rarity of samples;
[0061] A model training module, which is used to pre-train the shared encoder based on the self-attention mechanism using the common covariate pre-training dataset; after pre-training, use the specific regional datasets of each region, combined with the multi-task learning framework, to fine-tune the shared encoder based on the self-attention mechanism in a specific region to complete the training of the shared encoder based on the self-attention mechanism;
[0062] A time series prediction module, which is used to perform time series prediction on large-scale, cross-regional load data using the trained shared encoder based on the self-attention mechanism;
[0063] A performance evaluation module, which is used to evaluate the performance of the shared encoder based on the self-attention mechanism according to the time series prediction results. If the performance of the shared encoder based on the self-attention mechanism does not meet the standard, re-fine-tune the shared encoder based on the self-attention mechanism.
[0064] Based on the same inventive concept, a cross-regional load prediction device based on a shared encoder according to the present invention is characterized in that it includes a processor and a memory. Computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the electronic device implements the steps of the above-mentioned cross-regional load prediction method based on the shared encoder.
[0065] Based on the same inventive concept, a computer-readable storage medium according to the present invention stores a computer program, and when the program is executed by a processor, it implements the steps of the above-mentioned cross-regional load prediction method based on the shared encoder.
[0066] Advantageous effects: Compared with the prior art, the remarkable technical effects of the present invention are:
[0067] By sharing the encoder and the pre-training strategy, the data sparsity problem existing in cross-regional load forecasting is effectively solved, and the prediction performance of the model in areas with scarce data is improved.
[0068] Adopting a multi-task learning framework, the adaptive learning of data features in different regions is realized, and the challenges brought by regional differences are overcome.
[0069] Introducing a shared encoder based on the self-attention mechanism can effectively capture the long-term dependence relationship of load data and improve the prediction accuracy.
[0070] Using a dynamic weighted loss function to solve the data imbalance problem and enhance the generalization ability of the model.
[0071] Using an improved position encoding method can better capture the periodic pattern of load data and further improve the prediction accuracy. Description of the Drawings
[0072] Figure 1 It is a schematic flowchart of a cross-regional load forecasting method based on a shared encoder disclosed in an embodiment of the present invention;
[0073] Figure 2 It is an effect verification diagram of a cross-regional load forecasting method based on a shared encoder disclosed in an embodiment of the present invention in different regions;
[0074] Figure 3 It is a schematic structural diagram of a cross-regional load forecasting system based on a shared encoder disclosed in an embodiment of the present invention;
[0075] Figure 4 It is a schematic structural diagram of a cross-regional load forecasting device based on a shared encoder disclosed in an embodiment of the present invention. Detailed Embodiments
[0076] The present invention will be described in detail below with reference to the drawings and specific embodiments. Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the specific beneficial effects described above, and the above and other purposes that the present invention can achieve will be more clearly understood according to the following detailed description.
[0077] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in combination with the embodiments disclosed in the present invention can be implemented in hardware, software, or a combination of both. Specifically, whether to execute in hardware or software depends on the specific application and design and tree conditions of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0078] As used in this invention, the term "embodiment" means that the specific features, structures, or characteristics described in connection with an embodiment may be included in at least one embodiment of the invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0079] Embodiment 1
[0080] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of a cross - regional load forecasting method based on a shared encoder disclosed in an embodiment of the present invention. Among them, Figure 1 The described cross - regional load forecasting method is applied in a power system, such as for power system load forecasting, etc., which is not limited in the embodiments of the present invention. As Figure 1 shown, the cross - regional load forecasting method based on a shared encoder may include the following operations:
[0081] S1. Collect the load historical data and covariate features of different regions, and perform data pre - processing. Use the pre - processed load historical data and covariate features of different regions to construct a specific regional data set for each region; Considering the different data bases of different regions, select the intersection of covariates of each region to construct a common covariate pre - training data set;
[0082] Among them, the covariate features include:
[0083] Weather features: temperature (T), humidity (H), wind speed (W)
[0084] Date type features: working day (D worl ), weekend (D weekend ), holiday (D holid )
[0085] Seasonal features: month (M), season index (S)
[0086] Considering the different data bases of different regions, select the intersection of covariates of multiple regions as the common covariate pre - training data set. For example, Region A has air temperature, humidity, and irradiance. While Region B only has air temperature and humidity. Then, when preparing the pre - training data set, the data only includes air temperature and humidity. The specific regional data set retains all its own data.
[0087] The feature vector for a certain moment t can be expressed as:
[0088]
[0089] where Lt Indicates the load value at that moment.
[0090] S2. Construct a time series prediction backbone network, which includes an input encoding layer and a shared encoder based on the self-attention mechanism. Among them, the input encoding layer realizes the embedding encoding of time series data with different frequencies, and the shared encoder based on the self-attention mechanism is used to capture the long-term dependencies of time series data in different regions.
[0091] The backbone network of the present invention adopts a pure encoder architecture. The core component is a shared encoder based on the self-attention mechanism, which is used to capture the long-term dependencies of time series data in different regions.
[0092] Step S2 specifically includes the following steps:
[0093] S2.1. Design the input encoding layer:
[0094] In this embodiment, according to the frequency of the data, different lengths are selected to divide the data into time series blocks, and then the divided time series blocks are calculated through a fully connected layer, and then time series encodings of specified dimensions are output. Among them, for hourly data, the data segmentation length is usually in [32, 64], and for minute-level data, the segmentation length is usually in [32, 64, 128].
[0095] S2.2. Design a shared encoder based on the self-attention mechanism.
[0096] In this embodiment, the expression of the self-attention mechanism is as follows:
[0097] Q = XW Q , K = XW K , V = XW V
[0098]
[0099] Among them, X represents the input sequence; W Q , W K , W V represents learnable weight matrices; d k represents the dimension of the attention head; Q represents the query vector, which is used to match with the key vector to find the most relevant value vector; K represents the key vector, and the key vector performs a dot product operation with the query vector to calculate the relevance between the query and each key; V represents the value vector, and the value vector contains the actual information. When the query matches a certain key, the corresponding value vector will be selected as part of the output.
[0100] S2.3. Design an improved position encoding method
[0101] In order to better capture the periodic patterns of load data, the present invention adopts an improved position encoding method.
[0102] An improved position encoding method, comprising:
[0103]
[0104]
[0105] where pos represents the position index; i represents the dimension index, and d model represents the dimension of the model; in addition, additional periodic encoding is introduced:
[0106]
[0107] where T i represents different periods (such as daily, weekly, monthly, yearly).
[0108] The shared encoder is used to extract the common features of load data in different regions, and includes a multi-layer self-attention mechanism and a feed-forward neural network.
[0109] The time series prediction backbone network can adaptively learn the general features and unique patterns of time series data in different regions; among them, the backbone network optimizes both the shared feature extraction and the region-specific prediction tasks through a multi-task learning framework to achieve adaptive learning of data features in different regions.
[0110] S3. Design a multi-task learning framework, which includes a multi-task learning network architecture and a dynamic weighted loss function that automatically adjusts weights based on the rarity of samples. Step S3 specifically includes the following:
[0111] S3.1. Establish a multi-task learning network architecture, including the following main parts:
[0112] (1) Input encoding layer:
[0113] According to the frequency of the data, time series blocks of different lengths are selected to divide the data, and then the divided time series blocks are calculated through a fully connected layer, and then time series encoding of a specified dimension is output. For hourly data, the data segmentation length is usually in [32, 64], and for minute-level data, the segmentation length is usually in [32, 64, 128].
[0114] (2) Shared encoder:
[0115] Based on the time series prediction backbone network constructed in step S2, extract the features shared by all tasks, and its mathematical representation is:
[0116] H = Encoder(X)
[0117] where X represents the input sequence; H represents the shared features.
[0118] (3) Task-specific layer:
[0119] For each task, a task-specific layer is added to capture the unique features of that task. For the i-th task, its specific layer can be expressed as:
[0120] H_i = f_i(H)
[0121] where f_i represents the task-specific transformation function, which adopts the cascade of multiple convolutional layers and fully connected layers.
[0122] (4) Output layer:
[0123] Each task has its own output layer for generating the final prediction result. For the i-th task, its output can be expressed as:
[0124] Y_i = g_i(H_i)
[0125] where g_i represents the output function, which uses a fully connected layer with linear activation.
[0126] S3.1. Establish a dynamic weighted loss function that automatically adjusts weights based on the rarity of samples.
[0127] To solve the problem of data imbalance, the present invention designs a dynamic weighted loss function that automatically adjusts weights based on the rarity of samples, and its expression is as follows:
[0128]
[0129] In the formula, l i represents the loss of the i-th sample; w i represents the sample weight, and N represents the total number of samples.
[0130] Among them, the calculation formula of the sample weight w i is as follows:
[0131]
[0132] In the formula, f i represents the frequency of the i-th sample in the dataset; represents the sum of the frequencies of all samples.
[0133] To further balance the weights, the softmax function is used to normalize the sample weight w i as follows:
[0134]
[0135] The multi-task learning network architecture allows different output layers to simultaneously learn shared features and region-specific features. Different output layers obtain the output of the shared encoder. The dynamic weighted loss function adaptively adjusts the weights of different samples according to the distribution of samples in the dataset, achieving unified training of cross-region load samples.
[0136] S4. Model training: Pre-train the shared encoder based on the self-attention mechanism using the common covariate pre-training dataset; after pre-training, use the specific region datasets of each region and combine the multi-task learning framework to fine-tune the shared encoder based on the self-attention mechanism in the specific region to complete the training of the shared encoder based on the self-attention mechanism.
[0137] The present invention adopts a phased training strategy, and step S4 specifically includes the following:
[0138] S4.1. Pre-training stage: Pre-train the shared encoder based on the self-attention mechanism using the common covariate pre-training dataset, including:
[0139] Pre-train the shared encoder based on the self-attention mechanism on the common covariate pre-training dataset, and the optimization objective is to minimize the prediction error:
[0140]
[0141] In the formula, y represents the true load value; represents the predicted value; L reg represents the regularization term; L pre represents the loss function; λ represents the regularization coefficient; λ represents the regularization coefficient.
[0142] S4.2. Fine-tuning stage: After pre-training, use the specific region datasets of each region and combine the multi-task learning framework to fine-tune the shared encoder based on the self-attention mechanism in the specific region to complete the training of the shared encoder based on the self-attention mechanism, including:
[0143] Use the specific region datasets of each region and combine the multi-task learning framework to fine-tune the shared encoder based on the self-attention mechanism in the specific region:
[0144]
[0145] In the formula, L r represents the task loss of the r-th region; L sha represents the shared task loss; α r and β represent the balance coefficients.
[0146] S5. Load prediction: As Figure 2As shown, a trained shared encoder based on the self-attention mechanism is used to perform time series prediction on large-scale, cross-regional load data.
[0147] Use a trained shared encoder based on the self-attention mechanism for load forecasting. For the input sequence X = [x1, x2,..., x T , predict the load values for the next H time steps:
[0148]
[0149] where f(·) represents the trained model and θ are the model parameters.
[0150] S6. Model evaluation: Evaluate the performance of the shared encoder based on the self-attention mechanism according to the time series prediction results. If the performance of the shared encoder based on the self-attention mechanism does not meet the standard, fine-tune the shared encoder based on the self-attention mechanism again.
[0151] Use multiple metrics to evaluate the model performance, including the mean absolute percentage error (MAPE) and the root mean square error (RMSE):
[0152]
[0153]
[0154] where y i represents the true value; represents the predicted value; n represents the number of samples.
[0155] Through the above steps, the present invention realizes efficient and accurate cross-regional load forecasting. This method effectively solves problems such as data sparsity, regional differences, capturing long-term dependencies, and data imbalance, and significantly improves the prediction accuracy and the model generalization ability.
[0156] The cross-regional load forecasting method based on a shared encoder of the present invention has the main task of predicting the load time series data of different regions. Regarding the load forecasting of each region as an individual task.
[0157] Embodiment 2
[0158] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a cross-regional load forecasting system based on a shared encoder disclosed in an embodiment of the present invention. This system can realize power system load forecasting and specifically includes:
[0159] The data acquisition and preprocessing module is used to collect the load historical data and covariate features of different regions, perform data preprocessing, and use the preprocessed load historical data and covariate features of different regions to construct a specific regional dataset for each region; considering the different data bases of different regions, select the intersection of covariates of each region to construct a common covariate pre-training dataset;
[0160] The model construction module is used to construct a time series prediction backbone network, which includes an input encoding layer and a shared encoder based on the self-attention mechanism. Among them, the input encoding layer realizes the embedded encoding of time series data with different frequencies, and the shared encoder based on the self-attention mechanism is used to capture the long-term dependence relationships of time series data in different regions;
[0161] The multi-task learning framework establishment module is used to design a multi-task learning framework, which includes a multi-task learning network architecture and a dynamic weighted loss function that automatically adjusts weights based on the rarity of samples;
[0162] The model training module is used to pre-train the shared encoder based on the self-attention mechanism using the common covariate pre-training dataset; after pre-training, use the specific regional datasets of each region, combined with the multi-task learning framework, to fine-tune the shared encoder based on the self-attention mechanism in a specific region to complete the training of the shared encoder based on the self-attention mechanism;
[0163] The time series prediction module is used to perform time series prediction on large-scale, cross-regional load data using the trained shared encoder based on the self-attention mechanism;
[0164] The performance evaluation module is used to evaluate the performance of the shared encoder based on the self-attention mechanism according to the time series prediction results. If the performance of the shared encoder based on the self-attention mechanism does not meet the standard, re-fine-tune the shared encoder based on the self-attention mechanism.
[0165] In an optional implementation manner, the cross-regional load prediction method based on the shared encoder includes: a) constructing a common covariate pre-training dataset and a specific regional dataset; b) constructing a time series prediction backbone network; c) designing a multi-task learning framework, which includes a multi-task learning network architecture and a dynamic weighted loss function that automatically adjusts weights based on the rarity of samples; d) pre-training the shared encoder based on the self-attention mechanism using the common covariate pre-training dataset; using the specific regional datasets of each region, combined with the multi-task learning framework, to fine-tune the pre-trained shared encoder based on the self-attention mechanism in a specific region; e) using the trained shared encoder based on the self-attention mechanism to perform time series prediction on large-scale, cross-regional load data; f) evaluating the performance of the shared encoder based on the self-attention mechanism according to the time series prediction results.
[0166] Example 3
[0167] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a cross - regional load prediction device based on a shared encoder disclosed in an embodiment of the present invention. Among them, Figure 4 the described device can be applied to a power system, such as for power system load prediction, etc., which is not limited in the embodiments of the present invention.
[0168] As Figure 4 shown, the device may include a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the electronic device implements the steps of the method as described in the above embodiments and can achieve the same technical effects as the above method.
[0169] The memory may include a computer system - readable medium in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The device may further include other removable / non - removable, volatile / non - volatile computer system storage media. By way of example only, the memory may be used to read and write non - removable, non - volatile magnetic media (commonly referred to as a "hard disk drive"). Programs / utilities with a set (at least one) of program modules may be stored in the memory, such as an operating system, one or more application programs, other program modules, and program data. Implementations of a network environment may be included in each or some combination of these examples. The program modules generally execute the functions and / or methods in the embodiments described in the present invention.
[0170] The processor executes various functional applications and data processing by running the programs stored in the memory, such as implementing the method provided in Embodiment 1 of the present invention.
[0171] Example 4
[0172] Embodiment 4 of the present invention further provides a computer - readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method as described in the above embodiments and can achieve the same technical effects as the above method.
[0173] The computer storage medium of an embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. The computer-readable storage media may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage media may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0174] The computer-readable signal media may include data signals propagated in a baseband or as part of a carrier wave, which carry computer-readable program codes. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal media may also be any computer-readable media other than the computer-readable storage media, and the computer-readable media may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0175] The program codes contained on the computer-readable media may be transmitted by any appropriate media, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0176] The computer program codes for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program codes may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0177] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the above method operations, and can also execute relevant operations in the methods provided by any embodiment of the present invention.
[0178] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cross-regional load forecasting method based on a shared encoder, characterized in that, Including: Collect the load historical data and covariate features of different regions, and perform data preprocessing. Using the preprocessed load historical data and covariate features of different regions, construct a specific regional dataset for each region; Considering the different data bases of different regions, select the intersection of covariates of each region to construct a common covariate pre-training dataset; Construct a time series prediction backbone network, which includes an input encoding layer and a shared encoder based on the self-attention mechanism. Among them, the input encoding layer realizes the embedded encoding of time series data with different frequencies, and the shared encoder based on the self-attention mechanism is used to capture the long-term dependence relationship of time series data in different regions; Use an improved position encoding method to capture the periodic pattern of load data; Design a multi-task learning framework, which includes a multi-task learning network architecture and a dynamic weighted loss function that automatically adjusts weights based on the rarity of samples; Use the common covariate pre-training dataset to pre-train the shared encoder based on the self-attention mechanism; After pre-training, use the specific regional datasets of each region, combined with the multi-task learning framework, to fine-tune the shared encoder based on the self-attention mechanism in a specific region to complete the training of the shared encoder based on the self-attention mechanism; Use the trained shared encoder based on the self-attention mechanism to perform time series prediction on large-scale, cross-regional load data; Evaluate the performance of the shared encoder based on the self-attention mechanism according to the time series prediction results. If the performance of the shared encoder based on the self-attention mechanism does not meet the standard, re-fine-tune the shared encoder based on the self-attention mechanism.
2. The cross-regional load forecasting method based on a shared encoder according to claim 1, wherein The expression of the self-attention mechanism is as follows: Q = XW Q ,K = XW K ,V = XW V Among them, X represents the input sequence; W Q , W K , W V represent learnable weight matrices; d k represents the dimension of the attention head; Q represents the query vector; K represents the key vector; V represents the value vector.
3. The cross-regional load forecasting method based on a shared encoder according to claim 1, wherein The improved position encoding method includes: where pos represents the position index; i represents the dimension index, and d model represents the dimension of the model; in addition, additional positional encodings are introduced: Among them, T i represents different periods.
4. The cross-regional load forecasting method based on a shared encoder according to claim 1, wherein The multi-task learning network architecture includes: Input encoding layer: According to the frequency of the data, select different lengths to divide the data into time series blocks, then perform full connection layer calculation on the divided time series blocks, and then output the time series encoding of the specified dimension; Shared encoder: Based on the time series prediction backbone network, extract the features shared by all tasks, and its mathematical expression is: H = Encoder(X) Where, X represents the input sequence; H represents the shared feature; Task-specific layer: For each task, add a task-specific layer to capture the unique features of the task; For the i-th task, its specific layer can be expressed as: H_i = f_i(H) Where, f_i represents the task-specific transformation function, and adopts the cascade method of multiple convolutional layers and full connection layers; Output layer: Each task has an output layer for generating the final prediction result; For the i-th task, its output can be expressed as: Y_i = g_i(H_i) Where, g_i represents the output function, adopts a full connection layer, and uses a linear activation.
5. The cross-regional load forecasting method based on a shared encoder according to claim 1, wherein The expression of the dynamic weighted loss function that automatically adjusts weights based on the rarity of samples is as follows: where \(l\) i represents the loss of the \(i\)-th sample; \(w\) i represents the sample weight, and \(N\) represents the total number of samples; Among them, the sample weight w i has the following calculation formula: where, f i represents the frequency of the i-th sample in the dataset; represents the sum of the frequencies of all samples; To further balance the weights, the softmax function is used to normalize the sample weights w i as follows:
6. The cross-regional load forecasting method based on a shared encoder according to claim 1, wherein Using the common covariate pre-training dataset to pre-train the shared encoder based on the self-attention mechanism includes: Pre-train the shared encoder based on the self-attention mechanism on the common covariate pre-training dataset, and the optimization objective is to minimize the prediction error: where y represents the true load value; represents the predicted value; L reg represents the regularization term; L pre represents the loss function; λ represents the regularization coefficient.
7. The cross-regional load forecasting method based on a shared encoder according to claim 1, wherein After pre-training, using the specific regional datasets of each region and combining with a multi-task learning framework to fine-tune the shared encoder based on the self-attention mechanism in the specific region, and completing the training of the shared encoder based on the self-attention mechanism, including: Using the specific regional datasets of each region and combining with a multi-task learning framework to fine-tune the shared encoder based on the self-attention mechanism in the specific region: where, L r represents the task loss of the r-th region; L shared represents the shared task loss; α r and β represent the balance coefficients.
8. A cross-region load forecasting system based on a shared encoder, characterized in that, Including: A data acquisition and preprocessing module, which is used to collect the load historical data and covariate features of different regions, and perform data preprocessing. Using the preprocessed load historical data and covariate features of different regions, a specific regional dataset is constructed for each region; considering the different data bases of different regions, the intersection of covariates of each region is selected to construct a common covariate pre-training dataset; A model construction module, which is used to construct a time series prediction backbone network. The time series prediction backbone network includes an input encoding layer and a shared encoder based on the self-attention mechanism. Among them, the input encoding layer realizes the embedding encoding of time series data with different frequencies, and the shared encoder based on the self-attention mechanism is used to capture the long-term dependence relationships of time series data in different regions; A multi-task learning framework establishment module, which is used to design a multi-task learning framework. The multi-task learning framework includes a multi-task learning network architecture and a dynamic weighted loss function that automatically adjusts weights based on the rarity of samples; A model training module, which is used to pre-train the shared encoder based on the self-attention mechanism using the common covariate pre-training dataset; after pre-training, using the specific regional datasets of each region and combining with a multi-task learning framework to fine-tune the shared encoder based on the self-attention mechanism in the specific region, and completing the training of the shared encoder based on the self-attention mechanism; A time series prediction module, which is used to use the trained shared encoder based on the self-attention mechanism to perform time series prediction on large-scale, cross-regional load data; A performance evaluation module, which is used to evaluate the performance of the shared encoder based on the self-attention mechanism according to the time series prediction results. If the performance of the shared encoder based on the self-attention mechanism does not meet the standard, the shared encoder based on the self-attention mechanism is re-fine-tuned.
9. A cross-region load forecasting device based on a shared encoder, characterized in that, Including a processor and a memory. Computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device realizes the steps of the cross-regional load prediction method based on the shared encoder as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the program is executed by the processor, the steps of the cross-regional load prediction method based on the shared encoder as described in any one of claims 1 to 7 are realized.
Citation Information
Patent Citations
Multi-energy load prediction method and system based on improved coding and decoding model
CN116777668A
Multi-element load prediction method based on ESAM-MTL model
CN117787078A