New energy power grid net load prediction method, model and system based on soft distillation mechanism and storage medium
Patent Information
- Application Number
- CN202611317978.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本申请的主要目的在于提供一种基于软蒸馏机制的新能源电网净负荷预测方法,旨在解决如何压缩预测模型规模的同时提升预测性能的问题
[0042]1.通过膨胀卷积注意力模块为电网负荷时序数据中的关键时间步分配权重,来增强模型对重要特征的提取能力;
Smart Images

Figure CN122840172A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular to a method, model, system and storage medium for predicting the net load of a new energy power grid based on a soft distillation mechanism. Background Technology
[0002] Against the backdrop of global energy transition, the rapid popularization of renewable energy and the widespread application of smart grid technology have brought greater challenges to power load forecasting. Therefore, achieving accurate load forecasting is crucial for ensuring the stable operation of the power system. Neural networks have shown significant advantages due to their powerful data fitting and feature learning capabilities. The commonly used Bidirectional Long Short-Term Memory (BiLSTM) network can capture sequence features from both directions, obtaining a more comprehensive temporal representation. However, BiLSTM cannot explicitly distinguish the differences in importance of different time steps to the prediction results. Existing research has introduced attention mechanisms to enhance the model's ability to extract important features by adaptively learning and assigning higher weights to key time steps. However, the large number of training parameters leads to long training cycles, overfitting, and increased computational burden, making it difficult for current prediction models to balance accuracy and lightweight design.
[0003] Among the relevant technical solutions, knowledge distillation is an effective model compression method. By transferring the knowledge of a complex teacher model to a lightweight student model, it can reduce the model complexity to a certain extent. However, existing knowledge distillation methods are prone to loss of boundary information and / or inaccurate matching during feature alignment, which limits the predictive performance of the student model when compressing the number of parameters.
[0004] In view of this, this application proposes a new method for predicting the net load of new energy power grids, which aims to improve prediction performance while reducing the scale. Summary of the Invention
[0005] The main purpose of this application is to provide a new energy grid net load forecasting method based on a soft distillation mechanism, which aims to solve the problem of how to improve forecasting performance while reducing the size of the forecasting model.
[0006] To achieve the above objectives, this application provides a new energy grid net load forecasting method based on a soft distillation mechanism, the method comprising:
[0007] S10, use the dilated convolutional attention module to extract the deep feature information of the power grid load time series data, and assign weights to the deep feature information;
[0008] S20, the output of the dilated convolutional attention module is used as the input of the first neural network and the second neural network, respectively, to construct the student model and the teacher model, wherein the structural complexity of the first neural network is lower than that of the second neural network;
[0009] S30, the teacher model and the student model are trained by knowledge transfer through a soft distillation mechanism to obtain a trained target student model, and net load prediction is performed based on the target student model. The soft distillation mechanism includes: updating the parameters of the student model based on the feature matching loss calculated based on the intermediate layer feature maps of the teacher model and the student model.
[0010] Optionally, the calculation steps of the feature matching loss include:
[0011] S31, the intermediate layer feature map is randomly divided into multiple small blocks;
[0012] S32, calculate the gradient representation of the boundary data of each of the small block regions with respect to itself and adjacent blocks, wherein the boundary data of the block is characterized by the data of the edge of each block within a predetermined width;
[0013] S33, normalize the gradient representation to generate boundary weights, and perform weighted fusion of the boundary data based on the boundary weights to obtain a soft boundary feature representation;
[0014] S34, Calculate the feature matching loss based on the soft boundary feature representation.
[0015] Optionally, S34 includes:
[0016] S341, Based on the soft boundary feature representation, calculate the feature matching loss of the teacher model and the student model at the corresponding positions, as the first loss term;
[0017] S342, perform cross-attention interaction between the intermediate layer features of the teacher model and the output layer of the teacher model, and perform cross-attention interaction between the intermediate layer features of the student model and the output layer of the student model, so as to correct the output layer representation of the student model;
[0018] S343, Calculate the output layer difference loss between the teacher model and the student model, as the second loss term;
[0019] S344, calculate the difference loss between the output results of the teacher model and the student model and the true label respectively, as the third loss term and the fourth loss term;
[0020] S345, the first loss term, the second loss term, the third loss term and the fourth loss term are added together to obtain the feature matching loss.
[0021] Optionally, the dilated convolutional attention module includes a first feature extraction path and a second feature extraction path configured in parallel, and S10 includes:
[0022] S11, Perform dilated causal convolution and pooling operations sequentially on the power grid load time series data through the first feature extraction path to extract first-scale features;
[0023] S12, perform multi-level pooling operation on the power grid load time series data through the second feature extraction path to extract second-scale features;
[0024] S13, the first scale feature, the second scale feature and the power grid load time series data are concatenated to obtain depth feature information and generate a weighted representation;
[0025] S14, the depth feature information is weighted based on the weight representation, and the weighted result and the depth feature information are both used as the output of the dilated convolutional attention module.
[0026] Optionally, the second neural network includes a deep bidirectional long short-term memory network, and the teacher model includes an encoder and a decoder. The construction process of the teacher model includes:
[0027] S21, the depth feature information is weighted by the dilated convolutional attention module to obtain the first weighted feature;
[0028] S22, the depth feature information is dimensionally compressed by the encoder and the decoder, and the data after dimensionality compression is weighted by the dilated convolutional attention module to obtain the second weighted feature;
[0029] S23, the first weighted feature and the second weighted feature are fused and input into the deep bidirectional long short-term memory network for prediction.
[0030] Optionally, the first neural network includes a lightweight long short-term memory network, and the construction process of the student model includes:
[0031] S24, the depth feature information is weighted and assigned through the dilated convolutional attention module to obtain the weighted assignment result;
[0032] S25, the weighted assignment result is input into the lightweight long short-term memory network for prediction.
[0033] Optionally, the data edges of the power grid load time series data are filled with zero-value cells, and an inflation rate is set to construct an inflationary causal convolution.
[0034] Furthermore, to achieve the above objectives, this application also provides a new energy power grid net load forecasting model, which includes:
[0035] The dilated convolutional attention module is used to extract deep feature information from the power grid load time series data and assign weights to the deep feature information.
[0036] A target student model is used for net load forecasting.
[0037] The target student model is obtained by knowledge transfer training of the teacher model and the student model through a soft distillation mechanism. The soft distillation mechanism includes: calculating the feature matching loss based on the intermediate layer feature maps of the teacher model and the student model, and updating the parameters of the student model.
[0038] The student model and the teacher model are constructed by inputting the output of the dilated convolutional attention module into the first neural network and the second neural network, respectively. The structural complexity of the first neural network is lower than that of the second neural network.
[0039] In addition, to achieve the above objectives, this application also provides a computer system, the computer system comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, it implements the steps of the new energy grid net load forecasting method based on the soft distillation mechanism as described in any of the preceding claims.
[0040] In addition, to achieve the above objectives, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the new energy grid net load forecasting method based on a soft distillation mechanism as described in any of the preceding claims.
[0041] This application has at least the following beneficial effects:
[0042] 1. By assigning weights to key time steps in the power grid load time series data using a dilated convolutional attention module, the model's ability to extract important features is enhanced;
[0043] 2. A soft distillation mechanism is designed, which uses the intermediate layer features of the teacher model and the student model as the carrier of knowledge transfer to calculate the feature matching loss. By minimizing the feature matching loss, the parameters of the student model are updated through backpropagation, so that it can approach the prediction performance of the teacher model as much as possible while maintaining lightweight design. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating the net load forecasting method for new energy power grids based on a soft distillation mechanism, as described in the embodiments of this application.
[0045] Figure 2 This is a schematic diagram of dilated convolution involved in an embodiment of this application;
[0046] Figure 3 This is a schematic diagram of the soft distillation mechanism involved in the embodiments of this application;
[0047] Figure 4 This is a schematic diagram of the dilated convolutional attention module architecture involved in an embodiment of this application;
[0048] Figure 5 This is a schematic diagram of the composition architecture of the teacher-student model involved in the embodiments of this application;
[0049] Figure 6 The above are simulation results of the model on the ENTSO-E platform involved in the embodiments of this application;
[0050] Figure 7 This is a schematic diagram of the hardware operating environment of the computer system involved in the embodiments of this application.
[0051] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0052] To better understand the above technical solutions, exemplary embodiments of this disclosure will be described in more detail below with reference to the accompanying drawings. While exemplary embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of this disclosure to those skilled in the art.
[0053] First Embodiment
[0054] Reference Figure 1 This embodiment provides a method for predicting the net load of a new energy power grid based on a soft distillation mechanism. The method includes the following steps:
[0055] S10, use the dilated convolutional attention module to extract the deep feature information of the power grid load time series data, and assign weights to the deep feature information;
[0056] In this embodiment, refer to Figure 2The diagram illustrates dilated convolution. The dilated convolution attention module, acting as a basic feature extraction unit, is configured to receive time-series load data from historical load data generated in the renewable energy grid as input. This module's weighting coefficient enhances key features in the deep feature information and suppresses redundant features. The core design concept of this module is to expand the receptive field using dilated convolution, thereby capturing longer-term dependencies without increasing computational load.
[0057] Specifically, weight assignment refers to the module's ability to dynamically generate weight values based on the importance of features at each time step in the deep feature information.
[0058] S20, the output of the dilated convolutional attention module is used as the input of the first neural network and the second neural network, respectively, to construct the student model and the teacher model, wherein the structural complexity of the first neural network is lower than that of the second neural network;
[0059] In this step, the teacher model, acting as the "teacher" in the knowledge distillation process, fully utilizes attention mechanisms to focus on key time steps, thereby learning more robust and generalizable feature knowledge during training. In contrast to the teacher model, the student model is used to meet the lightweight requirements of devices such as edge computing devices in practical applications, considering inference speed and deployment environment.
[0060] In this embodiment, both the student model and the teacher model integrate dilated convolutional attention modules, meaning they share the same attention module design, so that the student model can align with the feature space of the teacher model.
[0061] Compared to the second neural network of the teacher model, the first neural network of the student model has lower structural complexity, that is, fewer layers or a more compact structure.
[0062] S30, the teacher model and the student model are trained by knowledge transfer through a soft distillation mechanism to obtain a trained target student model, and net load prediction is performed based on the target student model. The soft distillation mechanism includes: updating the parameters of the student model based on the feature matching loss calculated based on the intermediate layer feature maps of the teacher model and the student model.
[0063] In this step, the soft distillation mechanism is one of the core innovations in this embodiment. Traditional knowledge distillation often only focuses on the alignment of the probability distribution of the output layer, and when processing feature maps, simple block operations can easily lead to the loss of boundary information between blocks, which can easily cause feature discontinuity.
[0064] Reference Figure 3The diagram shown illustrates the soft distillation mechanism. In this embodiment, the soft distillation mechanism extracts the intermediate layer feature maps of the teacher-student model as the carrier of knowledge transfer and calculates the feature matching loss. By minimizing the feature matching loss, the parameters of the student model are updated through backpropagation, so that while maintaining lightweight design, the predictive performance of the student model can approach or even surpass that of the teacher model.
[0065] Further, and optionally, the calculation steps for the feature matching loss include:
[0066] S31, the intermediate layer feature map is randomly divided into multiple small blocks;
[0067] S32, calculate the gradient representation of the boundary data of each of the small block regions with respect to itself and adjacent blocks, wherein the boundary data of the block is characterized by the data of the edge of each block within a predetermined width;
[0068] S33, normalize the gradient representation to generate boundary weights, and perform weighted fusion of the boundary data based on the boundary weights to obtain a soft boundary feature representation;
[0069] S34, Calculate the feature matching loss based on the soft boundary feature representation.
[0070] It is worth noting that the feature map segmentation in this embodiment is not a simple hard segmentation, but rather a soft fusion strategy based on gradient information is introduced. Gradient information can reflect the degree of change in feature data. In this embodiment, we use it to quantify the influence of the boundary on itself and its adjacent regions, thereby generating a soft boundary feature representation based on the gradient information, and thus compensating for the feature discontinuity caused by segmentation.
[0071] As an example, the intermediate layer feature maps of the teacher and student models are first exported. These intermediate layer feature maps are then randomly divided into blocks at corresponding spatial locations. Each complete intermediate layer feature map of both models is divided into multiple small regions, and each small region is assigned a location number. Next, based on the adjacent boundary regions of each small region, the gradient representation of each small region's boundary data with respect to itself and its adjacent regions is calculated. This yields all gradient representations of the small region's boundary data. Information fusion is then performed using matrix addition. Taking a case where there is only one adjacent small region, the absolute value of the gradient representations is processed to obtain a new information expression with unified dimensions. The magnitude of the absolute value reflects the degree of change in influence, as shown in the following formula:
[0072]
[0073]
[0074] In the formula, x represents the boundary data of the current small region, and y represents the boundary data of the adjacent small regions. , This represents the corrected gradient data. Gradient-self means the gradient of the current block boundary element is calculated with respect to all elements in the current block; Gradient-self means the gradient of the current block boundary element is calculated with respect to all elements in adjacent blocks. This represents absolute value operations.
[0075] The obtained absolute gradient data is normalized using the Softmax function to generate adaptive information weights for each small block of boundary data, as shown in the following two equations.
[0076]
[0077]
[0078] Subsequently, the weight is multiplied by the original boundary data of the corresponding small block using the Hadamard product to achieve an initial update of the boundary data; finally, the updated boundary data of adjacent regions are merged according to the following formula:
[0079]
[0080] In the formula, This represents the corrected boundary data, where x and y represent the boundary data. , Representing weights, Softmax represents the function transformation. Represents the Hadamard product.
[0081] Further and optionally, step S34 specifically includes the following steps:
[0082] S341, Based on the soft boundary feature representation, calculate the feature matching loss of the teacher model and the student model at the corresponding positions, as the first loss term;
[0083] It should be noted that the first loss term aims to force the intermediate features of the student model to approximate the intermediate features of the teacher model as closely as possible, thereby achieving knowledge alignment at the feature level.
[0084] S342, perform cross-attention interaction between the intermediate layer features of the teacher model and the output layer of the teacher model, and perform cross-attention interaction between the intermediate layer features of the student model and the output layer of the student model, so as to correct the output layer representation of the student model;
[0085] S343, Calculate the output layer difference loss between the teacher model and the student model, as the second loss term;
[0086] It should be noted that traditional distillation methods often neglect the relationship between the intermediate layer and the output layer. This embodiment utilizes a second loss term to construct a multi-party interaction mechanism: using the intermediate layer features of the teacher model as a key to interact with its output layer, it can "review" the intermediate states of the teacher's own learning process, thereby extracting the impact of the teacher model's learning states at each stage of training on the final learning result. On the other hand, the intermediate layer features of the student model interact with the output layer, allowing the student model to "review" its intermediate states when outputting prediction results, and similarly extracting the impact of the student model's learning states at each stage of training on the final learning result. The second loss term is used to constrain the output distribution of the student model to converge with that of the teacher model, ensuring the overall direction of knowledge distillation is correct.
[0087] S344, calculate the difference loss between the output results of the teacher model and the student model and the true label respectively, as the third loss term and the fourth loss term;
[0088] It is easy to understand that the third and fourth loss terms are used to ensure that the student model is always supervised by real data during the distillation process, so as to avoid over-reliance on the teacher model and deviating from the real task.
[0089] S345, the first loss term, the second loss term, the third loss term and the fourth loss term are added together to obtain the feature matching loss.
[0090] In some alternative implementations, the MSE loss function can be used as the corresponding loss function for the four loss terms mentioned above:
[0091]
[0092] Where n represents the number of data points, and yi represents the actual value. This represents the predicted value.
[0093] The first to fourth loss terms are added together to obtain the final feature matching loss. Based on this loss, the student model parameters are updated through the backpropagation algorithm. The training is iterated until the preset number of rounds is reached to obtain a target student model that can be used for ultra-short-term load prediction. Net load prediction is then performed based on the obtained target student model.
[0094] In the technical solution provided in this embodiment, weights are assigned to key time steps in the power grid load time series data by dilated convolutional attention modules to enhance the model's ability to extract important features. A soft distillation mechanism is designed, which uses the intermediate layer features of the teacher model and the student model as the carrier of knowledge transfer to calculate the feature matching loss. The student model parameters are updated by backpropagation by minimizing the feature matching loss, so that it can approach the prediction performance of the teacher model as closely as possible while maintaining lightweight design.
[0095] Second Embodiment
[0096] Based on the first embodiment, this embodiment provides a method for assigning weights to the depth feature information, which specifically includes the following steps:
[0097] S11, Perform dilated causal convolution and pooling operations sequentially on the power grid load time series data through the first feature extraction path to extract first-scale features;
[0098] S12, perform multi-level pooling operation on the power grid load time series data through the second feature extraction path to extract second-scale features;
[0099] S13, the first scale feature, the second scale feature and the power grid load time series data are concatenated to obtain depth feature information and generate a weighted representation;
[0100] S14, the depth feature information is weighted based on the weight representation, and the weighted result and the depth feature information are both used as the output of the dilated convolutional attention module.
[0101] Specifically, refer to Figure 4 The diagram shown illustrates the architecture of the dilated convolutional attention module. The following example, involving specific numerical values, further clarifies this:
[0102] The input grid load time-series data is padded with two zero-value cells at the edges, and an expansion rate of 2, a kernel size of 3, and a stride of 1 are set to construct a dilated causal convolution, thereby expanding its temporal receptive field. Subsequently, edge data is utilized by padding with 1 cell, and average pooling with a window size of 3 is performed to aggregate local information between adjacent elements. Next, the nonlinear learning capability of the model is enhanced by the Leaky ReLU activation function, the calculation formula of which is as follows:
[0103]
[0104] Where f(x) represents the Leaky ReLU activation function, x represents the input data, and a represents a fixed slope, which is set to 0.05 in this embodiment.
[0105] Next, the data after Leaky ReLU activation is padded with two zero-value units at the data edges, and the dilation rate is set to 2, the kernel size to 3, and the stride to 1, further expanding its receptive field. Subsequently, average pooling is performed with edge padding of 1 and a window size of 3 to capture the local dependencies between adjacent elements. The Leaky ReLU activation function enhances the model's non-linear learning capability.
[0106] After Leaky ReLU activation, the data edges are padded with two zero-value units for the third time, and the dilation rate is set to 2, the kernel size to 3, and the stride to 1, thus expanding the receptive field for the third time. Subsequently, by padding the edges with 1 and performing average pooling with a window size of 3, the local dependencies between adjacent elements are captured, resulting in the data information representation at the first scale.
[0107] In parallel, edge padding is performed on the input data, with a padding size of 1, a stride of 1, and a pooling window size of 3. Then, average pooling is performed to aggregate local feature information and increase the receptive field. Edge padding is then performed again on the pooled data, with a padding size of 3 and a stride of 1, followed by another average pooling operation to further increase the receptive field. A third edge padding is performed on the pooled data, with a padding size of 3 and a stride of 1, followed by another average pooling operation to further increase the receptive field, resulting in a data representation at the second scale.
[0108] The data from the first scale, the data from the second scale, and the original input data are concatenated, and the data size is transformed by a multilayer perceptron (MLP). The data is then input into a sigmoid function to obtain a weighted representation that integrates multi-view information.
[0109] Third Embodiment
[0110] Based on the first embodiment, in this embodiment, refer to Figure 5 The second neural network is configured as a deep bidirectional long short-term memory network, and this embodiment provides a process for constructing a teacher model, including the following steps:
[0111] S21, the depth feature information is weighted by the dilated convolutional attention module to obtain the first weighted feature;
[0112] S22, the depth feature information is dimensionally compressed by the encoder and the decoder, and the data after dimensionality compression is weighted by the dilated convolutional attention module to obtain the second weighted feature;
[0113] S23, the first weighted feature and the second weighted feature are fused and input into the deep bidirectional long short-term memory network for prediction.
[0114] Specifically, the deep bidirectional long short-term memory network BiLSTM can be set to have 5 hidden layers. After each hidden layer, an LN (layer normalization) normalization layer and a ReLU activation function are added to enhance the non-linear learning ability. Furthermore, a residual connection is added after the third hidden layer to avoid gradient vanishing or exploding. The ReLU activation function is calculated as follows:
[0115]
[0116] in, Indicates input data, This represents the ReLU activation function.
[0117] The encoder and decoder use the GELU function for nonlinear transformation. The GELU calculation formula is as follows:
[0118]
[0119] In this embodiment, the input data is first adaptively weighted by a dilated convolutional attention module to obtain the first weighted feature. Simultaneously, the input data undergoes feature compression via an encoder, followed by dimensionality enhancement via a decoder to reduce noise interference and enhance the multidimensional representation capability of the features. The data then passes through the dilated convolutional attention module again to extract multidimensional information representation, resulting in the second weighted feature. Next, the two weighted features are fused through a concatenation operation to form a richer comprehensive feature representation. Finally, the fused features are input into a deep bidirectional long short-term memory (BiLSTM) network, and the teacher model is constructed based on its output load prediction results.
[0120] Fourth embodiment
[0121] Based on any of the above embodiments, this embodiment also refers to... Figure 5 The first neural network is configured as a lightweight long short-term memory network (LSTM), and this embodiment provides a process for constructing a student model, including the following steps:
[0122] S24, the depth feature information is weighted and assigned through the dilated convolutional attention module to obtain the weighted assignment result;
[0123] S25, the weighted assignment result is input into the lightweight long short-term memory network for prediction.
[0124] In this embodiment, compared with the teacher model, the student model omits the encoder / decoder structure and only applies a dilated convolutional attention module once at the input end, in order to reduce the network complexity of the student model.
[0125] Specifically, the LSTM network can be simplified from a 5-layer structure of the teacher model to a 3-layer structure, and the number of internal stacking layers can be simplified from a 5-layer structure of the teacher model to a 1-layer structure, reducing the network structure by a total of 22 layers (i.e., 5*5-3*1=22).
[0126] The input data is weighted using a dilated convolutional attention module. The corrected data is then fed into a simple LSTM network to predict the student load, thus completing the student model construction.
[0127] Verification of Examples
[0128] To verify the effectiveness of this invention, simulations were conducted in the Python 3.8 environment, using mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), regression coefficient (R²), and prediction time as evaluation metrics. The GEFCom2017 dataset used originates from the global energy forecasting competition hosted by the IEEE Power Energy Forecasting Working Group and is a core data resource in this field, primarily used for hierarchical probabilistic power load forecasting research.
[0129] Figure 6 The simulation results using the Dutch National Electricity Dataset on the ENTSO-E platform are presented, and Table 1 quantifies the performance of the prediction model. (Combining Table 1 with...) Figure 6 It is evident that the addition of the dilated convolutional attention module further improves the predictive performance of the teacher model, reducing MAE by 5.17MW, RMSE by 4.73MW, MAPE by 0.07%, and R² by 0.02%. Furthermore, by implementing knowledge transfer between the teacher and student models through soft distillation, the student model, with its fewer parameters, achieves a comprehensive improvement in prediction accuracy while significantly reducing prediction time. Specifically, compared to the teacher model, the student model reduces MAE by 1.95MW, RMSE by 2.58MW, MAPE by 0.03%, R² by 0.01%, reduces prediction time by 343.26 seconds, and training time by 69.04%. Moreover, the number of parameters in the student model is drastically reduced from 21,567,273 in the teacher model to 256,372, a reduction of 98.81%, equivalent to only 1.19% of the teacher model's parameters. This smaller parameter count makes it easier to deploy on edge devices. The results demonstrate that the student model achieves higher prediction accuracy and faster inference speed with a significantly fewer parameter count.
[0130] Table 1. Performance of the Prediction Model
[0131]
[0132] Furthermore, as an implementation scheme, this embodiment also proposes a new energy power grid net load forecasting model, which includes:
[0133] The dilated convolutional attention module is used to extract deep feature information from the power grid load time series data and assign weights to the deep feature information.
[0134] A target student model is used for net load forecasting.
[0135] The target student model is obtained by knowledge transfer training of the teacher model and the student model through a soft distillation mechanism. The soft distillation mechanism includes: calculating the feature matching loss based on the intermediate layer feature maps of the teacher model and the student model, and updating the parameters of the student model.
[0136] The student model and the teacher model are constructed by inputting the output of the dilated convolutional attention module into the first neural network and the second neural network, respectively. The structural complexity of the first neural network is lower than that of the second neural network.
[0137] Furthermore, as an implementation scheme, Figure 7 This is a schematic diagram of the hardware operating environment of the computer system involved in the embodiments of this application.
[0138] like Figure 7 As shown, the computer system may include: a processor 1001, such as a CPU; a memory 1005; a user interface 1003; a network interface 1004; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0139] Those skilled in the art will understand that Figure 7 The computer system architecture shown does not constitute a limitation on the computer system and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0140] like Figure 7As shown, the memory 1005, as a storage medium, may include an operating system, a network communication module, a user interface module, and computer programs. The operating system is a program that manages and controls the hardware and software resources of the computer system, as well as the operation of the computer programs and other software or programs.
[0141] exist Figure 7 In the computer system shown, the user interface 1003 is mainly used to connect to the terminal and communicate with the terminal; the network interface 1004 is mainly used to communicate with the backend server; and the processor 1001 can be used to call the computer program stored in the memory 1005.
[0142] In this embodiment, the computer system includes: a memory 1005, a processor 1001, and a computer program stored in the memory and executable on the processor, wherein:
[0143] When processor 1001 calls the computer program stored in memory 1005, it executes each step of the new energy grid net load forecasting method based on soft distillation mechanism as described above.
[0144] Furthermore, those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in a computer system to implement the process steps of the embodiments of the above methods.
[0145] Therefore, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the various steps of the new energy grid net load forecasting method based on a soft distillation mechanism as described in the above embodiments.
[0146] The computer-readable storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0147] It should be noted that, since the storage medium provided in the embodiments of this application is the storage medium used to implement the methods of the embodiments of this application, those skilled in the art can understand the specific structure and variations of the storage medium based on the methods described in the embodiments of this application, and therefore will not be repeated here. All storage media used in the methods of the embodiments of this application fall within the scope of protection of this application.
[0148] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0149] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0150] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0151] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0152] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0153] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for predicting the net load of a new energy power grid based on a soft distillation mechanism, characterized in that, The method includes the following steps: S10, use the dilated convolutional attention module to extract the deep feature information of the power grid load time series data, and assign weights to the deep feature information; S20, the output of the dilated convolutional attention module is used as the input of the first neural network and the second neural network, respectively, to construct the student model and the teacher model, wherein the structural complexity of the first neural network is lower than that of the second neural network; S30, the teacher model and the student model are trained by knowledge transfer through a soft distillation mechanism to obtain a trained target student model, and net load prediction is performed based on the target student model. The soft distillation mechanism includes: updating the parameters of the student model based on the feature matching loss calculated based on the intermediate layer feature maps of the teacher model and the student model.
2. The method for predicting the net load of a new energy power grid based on a soft distillation mechanism as described in claim 1, characterized in that, The calculation steps for the feature matching loss include: S31, the intermediate layer feature map is randomly divided into multiple small blocks; S32, calculate the gradient representation of the boundary data of each of the small block regions with respect to itself and adjacent blocks, wherein the boundary data of the block is characterized by the data of the edge of each block within a predetermined width; S33, normalize the gradient representation to generate boundary weights, and perform weighted fusion of the boundary data based on the boundary weights to obtain a soft boundary feature representation; S34, Calculate the feature matching loss based on the soft boundary feature representation.
3. The new energy power grid net load forecasting method based on soft distillation mechanism as described in claim 2, characterized in that, S34 includes: S341, Based on the soft boundary feature representation, calculate the feature matching loss of the teacher model and the student model at the corresponding positions, as the first loss term; S342, perform cross-attention interaction between the intermediate layer features of the teacher model and the output layer of the teacher model, and perform cross-attention interaction between the intermediate layer features of the student model and the output layer of the student model, so as to correct the output layer representation of the student model; S343, Calculate the output layer difference loss between the teacher model and the student model, as the second loss term; S344, calculate the difference loss between the output results of the teacher model and the student model and the true label respectively, as the third loss term and the fourth loss term; S345, the first loss term, the second loss term, the third loss term and the fourth loss term are added together to obtain the feature matching loss.
4. The new energy power grid net load forecasting method based on soft distillation mechanism as described in claim 1, characterized in that, The dilated convolutional attention module includes a first feature extraction path and a second feature extraction path set in parallel. S10 includes: S11, Perform dilated causal convolution and pooling operations sequentially on the power grid load time series data through the first feature extraction path to extract first-scale features; S12, perform multi-level pooling operation on the power grid load time series data through the second feature extraction path to extract second-scale features; S13, the first scale feature, the second scale feature and the power grid load time series data are concatenated to obtain depth feature information and generate a weighted representation; S14, the depth feature information is weighted based on the weight representation, and the weighted result and the depth feature information are both used as the output of the dilated convolutional attention module.
5. The method for predicting the net load of a new energy power grid based on a soft distillation mechanism as described in claim 1, characterized in that, The second neural network includes a deep bidirectional long short-term memory network, and the teacher model includes an encoder and a decoder. The construction process of the teacher model includes: S21, the depth feature information is weighted by the dilated convolutional attention module to obtain the first weighted feature; S22, the depth feature information is dimensionally compressed by the encoder and the decoder, and the data after dimensionality compression is weighted by the dilated convolutional attention module to obtain the second weighted feature; S23, the first weighted feature and the second weighted feature are fused and input into the deep bidirectional long short-term memory network for prediction.
6. The method for predicting the net load of a new energy power grid based on a soft distillation mechanism as described in claim 1 or 5, characterized in that, The first neural network includes a lightweight long short-term memory network, and the construction process of the student model includes: S24, the depth feature information is weighted and assigned through the dilated convolutional attention module to obtain the weighted assignment result; S25, the weighted assignment result is input into the lightweight long short-term memory network for prediction.
7. The method for predicting the net load of a new energy power grid based on a soft distillation mechanism as described in claim 1, characterized in that, The data edges of the power grid load time series data are filled with zero-value cells, and an inflation rate is set to construct an inflationary causal convolution.
8. A new energy power grid net load prediction model, characterized in that, The net load forecasting model for the new energy power grid includes: The dilated convolutional attention module is used to extract deep feature information from the power grid load time series data and assign weights to the deep feature information. A target student model is used for net load forecasting. The target student model is obtained by knowledge transfer training of the teacher model and the student model through a soft distillation mechanism. The soft distillation mechanism includes: calculating the feature matching loss based on the intermediate layer feature maps of the teacher model and the student model, and updating the parameters of the student model. The student model and the teacher model are constructed by inputting the output of the dilated convolutional attention module into the first neural network and the second neural network, respectively. The structural complexity of the first neural network is lower than that of the second neural network.
9. A computer system, characterized in that, The computer system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the new energy grid net load forecasting method based on the soft distillation mechanism as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the new energy grid net load forecasting method based on a soft distillation mechanism as described in any one of claims 1 to 7.