Micro-grid energy management method and system based on double-sequence prediction model

By adopting an energy management method based on a dual-sequence prediction model in the microgrid, and using KL divergence to calculate the distribution difference between power generation and electricity consumption, the problems of low energy management efficiency and insufficient prediction accuracy in the prior art are solved, and high-precision energy prediction and dynamic adjustment are achieved.

CN119944615APending Publication Date: 2025-05-06ELECTRIC POWER RES INST OF GUANGXI POWER GRID CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411837683.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the intrinsic link between power generation and power consumption in microgrids, resulting in low energy management efficiency. Statistical methods or shallow artificial intelligence methods are limited in performance when processing nonlinear data, making it difficult to meet the needs of modern microgrids for high-precision prediction.

Method used

Using a microgrid energy management method based on a dual-sequence prediction model, a dual-sequence prediction model containing core modules and fully connected layers is constructed by acquiring and preprocessing historical electricity consumption and power generation data, and the distribution difference between power generation and power consumption is calculated using KL divergence to achieve accurate energy prediction and management.

Benefits of technology

It significantly improves the energy management efficiency in the microgrid, can more accurately capture the relationship between power generation and electricity consumption, improves prediction accuracy, meets the demand for high-precision prediction of modern microgrids, and optimizes energy distribution through dynamic adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119944615A_ABST
    Figure CN119944615A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent power grids, in particular to a micro-grid energy management method and system based on a double-sequence prediction model. The method comprises the following steps of: firstly, acquiring historical electricity consumption and generating capacity data in a micro-grid system, eliminating abnormal values by adopting a 3 sigma rule, filling missing values by adopting front and back mean value interpolation, and performing data preprocessing by adopting minimum-maximum normalization; dividing the processed data into a training set, a verification set and a test set according to the proportion of 7: 1.5: 1.5; the core is to construct a prediction model containing a double-branch structure, wherein a first branch is composed of a convolution layer, a batch normalization layer, a layer normalization layer, a ReLU activation function layer, a random inactivation layer and a multi-scale coordinate attention mechanism module; and the second branch forms residual connection by a one-dimensional convolution and a ReLU activation function. And outputting a power generation amount prediction value and a power consumption prediction value through a full connection layer of the three-layer neural network structure, calculating a distribution difference between the power generation amount prediction value and the power consumption prediction value by using KL divergence, and feeding the difference back to a micro-grid control center for energy adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart grid technology, and in particular to a microgrid energy management method and system based on a dual-sequence prediction model. Background Art

[0002] With the continuous growth of global energy demand and the widespread application of renewable energy, microgrids, as an important part of smart grids, play an increasingly critical role in energy management. Microgrids achieve reliable and stable transmission of electricity by integrating distributed renewable energy, energy storage units, and connections with the main grid platform, while reducing adverse impacts on the environment. However, the power supply in microgrids faces challenges such as the instability of renewable energy such as solar and wind power, and the volatility of user-side demand.

[0003] Traditional power forecasting methods mainly focus on the generation or consumption forecast of a single energy source, such as separate power generation forecast or power consumption forecast. These methods often ignore the inherent connection between power generation and power consumption in the power system, resulting in inefficient energy management. In addition, existing technologies mostly use statistical methods or shallow artificial intelligence methods for forecasting, which are limited in performance when processing nonlinear data and are difficult to meet the needs of modern microgrids for high-precision forecasting.

[0004] To overcome these challenges, in recent years, deep learning models, especially hybrid models that combine spatial and temporal feature extraction, have shown superior performance in power generation and consumption forecasting tasks. Although deep learning models have made significant progress in power forecasting, existing models still face challenges in efficiently processing spatial and temporal information simultaneously. These limitations stem from the inadequacy of the model structure, which cannot fully capture the inherent complexity of power data, and the lack of identification of key time dependencies, resulting in suboptimal prediction accuracy and efficiency. Summary of the invention

[0005] In view of the problems existing in the prior art, the present invention is proposed.

[0006] Therefore, the problem to be solved by the present invention is how to solve the problem that the intrinsic relationship between power generation and power consumption in the power system is often ignored, resulting in low energy management efficiency. In addition, the existing technologies mostly use statistical methods or shallow artificial intelligence methods for prediction, which are limited in performance when processing nonlinear data and are difficult to meet the needs of modern microgrids for high-precision prediction.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0008] In a first aspect, an embodiment of the present invention provides a microgrid energy management method based on a dual-sequence prediction model, which includes obtaining historical power consumption and power generation data in a microgrid system, and preprocessing the data, including outlier removal, missing value filling, and data normalization;

[0009] Split the preprocessed data into training set, validation set and test set;

[0010] A dual-sequence prediction model consisting of a core module and a fully connected layer is constructed, and the KL divergence is used to calculate the distribution difference between power generation and power consumption.

[0011] As a preferred solution of the microgrid energy management method based on the dual sequence prediction model described in the present invention, the core module includes: a first branch, consisting of a convolution layer, a batch normalization layer, a layer normalization layer, a ReLU activation function layer, a random inactivation layer and a multi-scale coordinate attention mechanism module; a second branch, a residual connection consisting of a one-dimensional convolution and a ReLU activation function.

[0012] As a preferred solution of the microgrid energy management method based on the dual sequence prediction model described in the present invention, the multi-scale coordinate attention mechanism module includes: a multi-scale feature extraction unit, used to obtain one-dimensional input features of different scales; a coordinate attention unit, used to encode feature information along the horizontal and vertical directions respectively; a residual connection unit, used to maintain the original feature information.

[0013] As a preferred solution of the microgrid energy management method based on the dual sequence prediction model described in the present invention, the multi-scale feature extraction unit uses one-dimensional convolution with different convolution kernel sizes for feature extraction, including convolution kernels of 1×1, 1×3 to 1×(2i+1), where i is a positive integer.

[0014] As a preferred solution of the microgrid energy management method based on the dual sequence prediction model described in the present invention, the coordinate attention unit includes: a horizontal pooling operation for obtaining the importance of information of different scales; a vertical pooling operation for obtaining the importance of information of different positions; and a feature fusion operation for combining attention information of different directions.

[0015] As a preferred solution of the microgrid energy management method based on the dual-sequence prediction model described in the present invention, the fully connected layer includes a three-layer neural network structure, which is used to: map the power consumption features extracted by the core module into power consumption prediction values; map the power generation features extracted by the core module into power generation prediction values.

[0016] As a preferred solution of the microgrid energy management method based on the dual-sequence prediction model described in the present invention, the KL divergence is used to calculate the distribution difference between the power generation prediction value and the power consumption prediction value; and feed back the calculated difference to the microgrid control center for energy regulation.

[0017] In a second aspect, an embodiment of the present invention provides a microgrid energy management system based on a dual-sequence prediction model, which includes a data acquisition and preprocessing module, which acquires historical power consumption and power generation data in the microgrid system, and preprocesses the data, including outlier removal, missing value filling and data normalization;

[0018] Model training module, which divides the preprocessed data into training set, validation set and test set;

[0019] The energy management module builds a dual-sequence prediction model consisting of a core module and a fully connected layer, and uses KL divergence to calculate the distribution difference between power generation and power consumption.

[0020] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, the steps of the microgrid energy management method based on the dual-sequence prediction model as described in the first aspect of the present invention are implemented.

[0021] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, the steps of the microgrid energy management method based on the dual-sequence prediction model as described in the first aspect of the present invention are implemented.

[0022] The beneficial effects of the present invention are: BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 It is a flow chart of the microgrid energy management method based on the dual sequence prediction model;

[0025] Figure 2 A computer device diagram of a microgrid energy management method based on a dual-sequence prediction model;

[0026] Figure 3 Schematic diagram of the overall structure of the microgrid energy management method based on the dual-sequence prediction model.

[0027] Figure 4 It is a schematic diagram of the overall design of the dual-sequence prediction model of the microgrid energy management method based on the dual-sequence prediction model;

[0028] Figure 5 It is a structural diagram of the core module of the dual-sequence prediction model of the microgrid energy management method based on the dual-sequence prediction model;

[0029] Figure 6 Module diagram of the multi-scale coordinate attention mechanism of the dual-sequence prediction model for the microgrid energy management method based on the dual-sequence prediction model. DETAILED DESCRIPTION

[0030] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0031] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0032] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments.

[0033] Example 1

[0034] Reference Figure 1-2 , which is the first embodiment of the present invention, and provides a microgrid energy management method based on a dual-sequence prediction model, comprising:

[0035] S100: Obtain historical power consumption and power generation data in the microgrid system, and preprocess the data, including outlier removal, missing value filling, and data normalization;

[0036] In the embodiment of the present application, the collection of historical data includes: electricity consumption data, power generation data and environmental data. Among them, the electricity consumption data includes information such as the real-time load power of the end user, historical power consumption curves and load characteristics; the power generation data includes data such as the output power, operating efficiency and equipment status of photovoltaic power generation, wind power generation and conventional power generation equipment; and the environmental data includes meteorological parameters such as temperature, humidity, light intensity and wind speed.

[0037] Specifically, the preprocessing of the collected raw data includes three steps:

[0038] Outlier Removal: Use the 3σ rule to identify and remove outliers in the data. First, calculate the mean μ and standard deviation σ of the data, and mark the data points outside the range of μ±3σ as outliers and remove them. This method assumes that the data follows a normal distribution and can remove 99.7% of the outliers.

[0039] Missing value filling: The missing data is processed by the mean interpolation method. For missing points in the time series data, the average value of the data before and after is taken to fill in the missing points. If multiple consecutive points are missing, the linear interpolation algorithm is used.

[0040] Data normalization: Use the minimum-maximum normalization method to map the data to the [0,1] interval. The normalization formula is:

[0041] x′=(x-xmin) / (xmax-xmin)

[0042] Where x is the original value, xmin and xmax are the minimum and maximum values ​​of the data respectively, and x' is the normalized value.

[0043] In an optional embodiment, the data collection method and frequency can be flexibly adjusted according to the data type. For example, for photovoltaic power generation and wind power generation with large fluctuations, a higher sampling frequency (such as 1 second / time) is used; for load data with relatively slow changes, a lower sampling frequency (such as 5 minutes / time) can be used. At the same time, according to the characteristics of different data sources, corresponding data quality control mechanisms are adopted.

[0044] In an optional embodiment, the preprocessing process may also introduce data compression and feature extraction steps. For example, downsampling or sliding average of high-frequency sampled data may be performed to reduce the amount of data while retaining key features; and fast Fourier transform (FFT) may be performed on the original data to extract frequency domain features for subsequent analysis.

[0045] It should be noted that this application adopts a distributed data acquisition architecture, and each measurement point is equipped with a local processing unit. These units are not only responsible for data acquisition, but also have preliminary data processing capabilities, which can perform data verification, compression and preprocessing operations. The system uses standardized communication protocols (such as Modbus TCP / IP) to ensure that data can be transmitted reliably.

[0046] S200: split the preprocessed data into a training set, a validation set and a test set;

[0047] In the embodiment of the present application, the data set is divided using the time series cross-validation method, with the specific ratio being: 70% training set, 15% validation set, and 15% test set. To ensure the temporal sequence and continuity of the data, continuous time periods are used for division rather than random sampling.

[0048] Specifically, assuming there are data at T time points, the division process is as follows:

[0049] Training set: data in the time period [1, 0.7T]; validation set: data in the time period (0.7T, 0.85T]; test set: data in the time period (0.85T, T]. In an optional embodiment, a sliding window method can be used to divide the data. Set a fixed-size time window (such as 24

[0050] Hours), each time sliding back a fixed step length (such as 1 hour), to generate multiple training samples. This method can make full use of the continuity characteristics of time series data.

[0051] In an optional embodiment, the seasonal characteristics of the data may also be considered for segmentation, for example, ensuring that the training set, validation set, and test set all contain sufficient peak and valley data to improve the generalization ability of the model.

[0052] It should be noted that the division of data sets needs to consider the balance and representativeness of samples. The consistency of data distribution is ensured by analyzing the statistical characteristics of each set (such as mean, variance, distribution, etc.). At the same time, the rationality of the division is verified by cross-validation method.

[0053] S300: Build a dual-sequence prediction model consisting of a core module and a fully connected layer, and use KL divergence to calculate the distribution difference between power generation and power consumption.

[0054] In the embodiment of the present application, the dual-sequence prediction model adopts a dual-branch structure to process the power generation sequence and the power consumption sequence respectively. Each branch includes two stages: feature extraction and feature mapping, and finally measures the distribution difference of the two sequences through KL divergence.

[0055] Specifically, the core module of each branch consists of the following components:

[0056] One-dimensional convolution layer: uses multiple convolution kernels of different scales to extract time series features; batch normalization layer: reduces internal covariate shift and accelerates training convergence; layer normalization layer: standardizes hidden outputs and improves model stability; ReLU activation layer: introduces nonlinear transformation capabilities; random inactivation layer: prevents overfitting; multi-scale attention module: enhances the model's ability to perceive important features.

[0057] In an optional embodiment, the input sequence length of the model can be flexibly adjusted according to the characteristics of the forecasting task. For example, for a day-ahead forecasting task, historical data from the past 7 days can be used as input; for a real-time forecast, data from the last 24 hours can be used.

[0058] In an optional embodiment, the model training process can adopt a transfer learning strategy. First, the model is pre-trained on a large-scale public dataset, and then fine-tuned using data from the target scenario, which can speed up the model convergence and improve prediction accuracy.

[0059] It should be noted that the design of the model fully considers the characteristics of the microgrid system. The dual-branch structure captures the correlation between power generation and power consumption sequences, multi-scale feature extraction is used to enhance the recognition of patterns at different time scales, and the attention mechanism is introduced to highlight the impact of important time points.

[0060] S301: The core modules include: the first branch, which consists of a convolutional layer, a batch normalization layer, a layer normalization layer, a ReLU activation function layer, a random dropout layer and a multi-scale coordinate attention mechanism module; the second branch, which consists of a residual connection composed of a one-dimensional convolution and a ReLU activation function.

[0061] In the embodiment of the present application, the core module adopts a dual-branch architecture:

[0062] The first branch includes: convolution layer: using {16, 32, 64} convolution kernels with kernel widths of {3, 5, 7}; batch normalization layer: parameters include learnable scaling factor γ and bias term β; layer normalization layer: normalization along the feature channel direction; ReLU activation function: max(0, x) nonlinear transformation; random inactivation layer: dropout rate is set to 0.3; multi-scale coordinate attention module.

[0063] The second branch includes: one-dimensional convolution: for residual connection; ReLU activation function: to ensure nonlinear mapping capability.

[0064] In an optional embodiment, the parameters of the convolution layer can be dynamically adjusted according to the characteristics of the input sequence. For example, for high-frequency fluctuating data, more small-scale convolution kernels can be used; for long-term trends, the proportion of large-scale convolution kernels can be increased.

[0065] In an optional embodiment, the parameters of the batch normalization layer and the layer normalization layer can adopt a dynamic update mechanism. In the early stage of training, a larger momentum parameter is used to accelerate convergence, and as the training progresses, the momentum parameter is gradually reduced to improve stability.

[0066] It should be noted that the design intention of the dual-branch structure is to enhance features while maintaining the original features. The first branch extracts complex features through a deep network, and the second branch retains the original information through residual connections. The two complement each other to improve the model's expressiveness.

[0067] S302: The multi-scale coordinate attention mechanism module includes: a multi-scale feature extraction unit, used to obtain one-dimensional input features of different scales; a coordinate attention unit, used to encode feature information along the horizontal and vertical directions respectively; and a residual connection unit, used to maintain the original feature information.

[0068] In the embodiment of the present application, the multi-scale coordinate attention mechanism module includes three key units:

[0069] Multi-scale feature extraction unit: Use 1×1 to 1×7 convolution kernels for feature extraction; use zero padding to keep the size of feature maps consistent; merge features of different scales through channel-wise concatenation.

[0070] Coordinate attention unit: horizontal pooling: captures the importance of different scales; vertical pooling: encodes the importance of different time steps; attention map calculation: generates two-dimensional attention weights based on horizontal and vertical features.

[0071] Residual connection unit: maintains the original feature path; prevents gradient disappearance / explosion; ensures model trainability.

[0072] In an optional embodiment, the calculation process of the attention mechanism can be parallelized. By parallelizing the batch data on the GPU, the processing speed is significantly improved.

[0073] In an optional embodiment, the attention weights may be calculated using a soft attention mechanism, and the weights may be normalized to the [0, 1] interval using a softmax function, so that the model can smoothly focus on features of different positions and scales.

[0074] It should be noted that the design of the multi-scale coordinate attention mechanism fully considers the characteristics of time series data. By combining feature representations and position encodings at different scales, the model's ability to understand time series patterns is enhanced. The use of residual connections ensures that the original feature information is not lost in the deep network.

[0075] S303: The multi-scale feature extraction unit uses one-dimensional convolution with different convolution kernel sizes to extract features, including convolution kernels of 1×1, 1×3 to 1×(2i+1), where i is a positive integer.

[0076] S304: The coordinate attention unit includes: a horizontal pooling operation for obtaining the importance of information at different scales; a vertical pooling operation for obtaining the importance of information at different positions; and a feature fusion operation for combining attention information at different directions.

[0077] S305: The fully connected layer includes a three-layer neural network structure, which is used to: map the power consumption features extracted by the core module into power consumption prediction values; and map the power generation features extracted by the core module into power generation prediction values.

[0078] S306: KL divergence is used to: calculate the distribution difference between the power generation forecast value and the power consumption forecast value; and feed back the calculated difference to the microgrid control center for energy regulation.

[0079] Furthermore, this embodiment also provides a microgrid energy management system based on a dual-sequence prediction model, comprising:

[0080] The data acquisition and preprocessing module acquires the historical power consumption and power generation data in the microgrid system and preprocesses the data, including outlier removal, missing value filling and data normalization;

[0081] Model training module, which divides the preprocessed data into training set, validation set and test set;

[0082] The energy management module builds a dual-sequence prediction model consisting of a core module and a fully connected layer, and uses KL divergence to calculate the distribution difference between power generation and power consumption.

[0083] In summary, high-quality cleaning of historical data is achieved by using a multi-layer preprocessing mechanism such as the 3σ rule to remove outliers, interpolation of the previous and next mean values ​​to fill missing values, and minimum-maximum normalization. This step breaks through the problem of unstable data quality in traditional methods, improves the reliability of subsequent predictions, and enables the model to be built on high-quality data.

[0084] By constructing a prediction model with a dual-branch structure, the feature extraction and prediction of power generation and power consumption sequences are performed simultaneously, overcoming the defect of traditional single-sequence prediction that ignores the correlation between power generation and power consumption. This innovative design enables the model to capture the intrinsic relationship between power generation and power consumption, significantly improving the prediction accuracy.

[0085] By introducing multi-scale feature extraction and residual connection in the attention mechanism, adaptive attention to features of different time scales is achieved. This mechanism can automatically identify and highlight the impact of key time points, which shows unexpected adaptability when dealing with the volatility of renewable energy.

[0086] By calculating the KL divergence of power generation and power consumption distribution, a novel energy supply and demand difference measurement method is established. This method can not only quantify the supply and demand difference, but also provide real-time adjustment signals for the microgrid control center, realizing dynamic optimization of energy distribution.

[0087] Through the organic combination of multi-dimensional data collection, dual-sequence deep learning and intelligent feedback mechanism, a closed-loop energy management system is formed. The system shows excellent robustness and adaptability in dealing with complex scenarios such as power generation fluctuations and load mutations, providing a new solution for the intelligent operation of microgrids.

[0088] Example 2

[0089] Reference Figure 2 - Figure 6 , which is the second embodiment of the present invention, provides a microgrid energy management method based on a dual-sequence prediction model. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0090] The microgrid is connected to the renewable energy power station, end users, power storage system and grid platform. If the renewable energy power generation is greater than the end user's power consumption, the microgrid will store the power in the power storage system and transmit the excess energy to the grid platform. If the power generation at the current time is lower than the power consumption, the microgrid will distribute power from the power storage system to the end user. If the storage system does not have enough power, the microgrid receives power from the connected grid platform. In order to achieve intelligent regulation of this behavior, a dual-sequence prediction model is established to simultaneously predict the power generation and power consumption, and the KL divergence is calculated to feed back the distribution difference between the two to the control center of the microgrid to achieve energy management of the microgrid.

[0091] according to Figure 4 ,The designed energy management method includes data collection, data preprocessing, data segmentation and dual ,sequence prediction model.

[0092] Specifically, historical data of user power consumption and power generation of power plants as well as weather data are collected, and data preprocessing is performed on the collected data, including removing outliers, filling missing values ​​and data normalization for effective training. Furthermore, the 3σ rule is used to remove outliers, the method of front and back mean interpolation is used to fill missing values, and the minimum-maximum scaling is used for data normalization.

[0093] Specifically, the preprocessed data is split into 70% training set, 15% validation set and 15% test set.

[0094] Specifically, the sequence prediction model includes a core module, a fully connected layer, and a KL divergence. The core module is used to extract the features of power consumption data and power generation data, the fully connected layer outputs the prediction results of power generation and power consumption, and the KL divergence is used to measure the difference between the two.

[0095] like Figure 5 As shown in the figure, the core module of the designed dual sequence prediction model includes two branches, one branch consists of a convolution layer, a batch normalization layer, a layer normalization, a ReLu activation function, a random dropout layer, and a multi-scale coordinate attention mechanism module. The other branch transforms the input by a one-dimensional convolution to achieve residual connection, followed by a ReLu activation function to achieve nonlinear output.

[0096] Specifically, for one-dimensional convolution, the input data sequence X, after the convolution layer, the output is expressed as,

[0097] Y c =WX+b (1)

[0098] Among them, W is the weight matrix of the convolutional layer, b is the bias term, and Y c is the output of the convolutional layer;

[0099] Batch Normalization layers are used to reduce internal covariance and distributional variation in each layer during training. Each input feature in a batch is normalized to obtain the mean and variance of the batch. Specifically, the BN operation converts the feature x i ∈Y c Translates to:

[0100]

[0101] where x i represents the i samples of the input data, E[x] and VAR[x] represent x i The mean and variance of y i represents the output of the batch normalization layer, γ i and β i denote the scale and translation parameters to be learned, respectively, and ε denotes a constant imposed to maintain training stability.

[0102] The ReLu activation function introduces nonlinearity into the neural network, enabling the network to learn and simulate complex function mappings. Its mathematical expression is:

[0103] f(y i )=max(0,y i ) (3)

[0104] The characteristic of the ReLU function is that when the input yi is less than 0, the output is 0; when the input yi is greater than or equal to 0, the output is yi itself. This function form makes ReLU linear in the positive interval and the output is 0 in the negative interval.

[0105] Layer normalization is a regularization mechanism that normalizes the distribution of hidden layers and is not limited by the batch size. Layer normalization can be defined as follows:

[0106]

[0107] Where x represents the feature map of the previous layer of one-dimensional convolution, γ and β represent trainable parameters for learning different distributions, E[x] and VAR[x] represent the mean and standard variation of x along the feature channel, and ε represents a small constant.

[0108] The random dropout layer is also a regularization technique, where the output of neurons with a certain probability is randomly set to zero.

[0109] The multi-scale coordinate attention mechanism module enhances the location information of the feature map by introducing a multi-scale method. A more effective feature representation can be obtained, which helps the model extract effective feature information.

[0110] The design diagram of the multi-scale coordinate attention mechanism module is as follows Figure 5 As shown in Figure 2. In time series prediction, different positions in a one-dimensional input signal contribute differently to the prediction. When the input dimension is large, the model is easily disturbed by irrelevant information, affecting the accuracy of the prediction. Therefore, the position information of the attention mechanism has an important impact on the prediction results.

[0111] Given that the existing channel attention mechanism uses global pooling to encode spatial information and describe the relationship between channels. Therefore, it is difficult to preserve position information. Coordinate attention decomposes global pooling into a pair of 1D feature encoding operations. For input X, coordinate attention uses two spatial range pooling kernels, namely (H, 1( and (1)W(, to encode each channel along) horizontal and vertical coordinates. The output of the cth channel with height h can be formulated as follows:

[0112]

[0113] The output of the cth channel with width w can be expressed as:

[0114]

[0115] These two transformations perform feature aggregation along two spatial directions and return a pair of direction-aware attention maps.

[0116] Considering the 1D characteristics of the current data, the coordinate attention is improved and a one-dimensional multi-scale coordinate residual attention model is proposed. A multi-scale method is introduced in the attention module to enhance the position information of the feature map. By embedding the operator for multi-scale feature extraction into the attention mechanism, a more effective feature representation can be obtained, which helps the model to locate the part of interest more accurately. Specifically, multi-scale convolutions with different convolution kernels are applied to the input to obtain multi-scale features, including 1×1, 1×3…1×(2i+1). It is worth noting that convolution operators with zero padding are used in this patent to keep the output and input sizes the same. After obtaining the multi-scale feature map of the current data, the feature maps are connected using a concatenation operation to form a 2D multi-scale feature map, which helps us use horizontal pooling and vertical pooling to capture different scale and position dependencies. The multi-scale feature map Xm can be further expressed as:

[0117] X m =[F1(x0,x1,…x w-1 ),F2(x0,x1,…x w-1 )…F n (x0,x1,…x w-1 )] (7)

[0118] Among them, F1(.), F2(.), F n (.) represents the dilated convolution operation with different kernels, [·,·] represents the cascade operation along the h dimension, and F(·)∈R C / r×(1+W) and F(·)∈R C / r×(H+W) Represent the single-scale feature map and the multi-scale feature superimposed by the single-scale feature map, r represents the reduction rate used to control the block size. The complexity of the attention mechanism can be controlled by reducing r. H represents the number of multi-scale kernels, and W represents the length of the input signal.

[0119] Specifically, feature maps in different directions are obtained using equations (5) and (6). Different feature maps are concatenated and sent to a shared 1×1 convolution transform F:

[0120]

[0121] Where δ represents the activation function. The position and scale information interact through shared convolution and nonlinear transformation to obtain f.

[0122] Furthermore, f is split into two independent tensors along the vertical and horizontal directions and two 1×1 convolutions are applied, namely, F h and F w , f h and f w Transform to the same as input X m The same characteristic channel size c.

[0123] g h =δ(F h (f h ))

[0124] g w =δ(F w (f w )) (9)

[0125] where g h ∈R C×(H+1) g w ∈R C×(1+W) Represents the importance of different scales and different positions. The output of the attention mechanism to redistribute weights can be expressed as

[0126]

[0127] Specifically, horizontal pooling is used to capture the importance of information at different scales, and vertical pooling is used to capture the importance of information at different positions.

[0128] Furthermore, simply stacking attention blocks will weaken the feature values ​​in deep networks and degrade model performance. At the same time, the good mapping that the model backbone has learned may be disturbed. Therefore, if the proposed attention mechanism can exploit the same mapping learned by the backbone model, the performance of the model with the attention mechanism will not be worse than that of the backbone model. Based on the above reasons, the structure of the residual attention network is introduced into the proposed attention mechanism.

[0129] Specifically, the output of the proposed attention mechanism can be expressed as:

[0130]

[0131] in Represents the pooling operation along the vertical direction to obtain a 1D output. A residual mechanism is introduced into the feature map of each scale, and then the output feature map is restored to the same size as the input feature map through vertical pooling.

[0132] It is worth mentioning that the multi-scale coordinate residual attention proposed in this paper is a general one-dimensional input feature extraction method, which can be easily integrated into other.

[0133] according to Figure 4 ,The fully connected layer of the designed dual sequence prediction model includes a 3-layer neural network structure, which is used to map the ,power consumption features and power generation features extracted by the core module into the ,power consumption prediction values ​​and power generation prediction values.

[0134] according to Figure 4, KL divergence is used to measure the difference between power generation EC and power consumption EG, and feed the difference back to the microgrid control center to achieve intelligent regulation between energy sources. The KL divergence algorithm is calculated according to the following formula:

[0135]

[0136] Where i represents the number of samples, EC(i) represents the electricity consumption of the i-th sample, and EG(i) represents the power generation of the i-th sample. KL divergence can effectively adjust energy supply to match the real-time power demand of consumers by quantifying the difference between power generation and consumption distribution. This dynamic adjustment ensures that the generated power is optimally utilized and distributed in the microgrid, reducing the risk of power depletion and improving the overall efficiency of the energy management system.

[0137] Model training. Based on the designed dual-sequence prediction model, first, the collected historical data is divided into training data, validation data, and test data after preprocessing. The training data is used for model training, the validation data is used for model validation, and the test data is used to test the generalization ability of the model. Secondly, here we use the mean square error as the loss function and use the Adam optimizer for back propagation optimization. The learning rate exponential decay method is adopted, that is, the learning rate is set high at the beginning of training so that the model can quickly converge to a better parameter space. As the training progresses, the learning rate is gradually reduced so that the model can make smaller adjustments when it is close to the optimal solution, thereby avoiding falling into the local minimum and improving the generalization ability of the model. The maximum number of iterations is set to 200, and the model training is completed when the maximum number of iterations is reached. Finally, the generalization performance of the trained model is tested using test data.

[0138] The second embodiment of the present invention is different from the first two embodiments in that it is used to verify the technical effects adopted in the present invention in order to verify the real effects of the method.

[0139] In order to evaluate the effectiveness of the proposed prediction model, the performance evaluation statistical indicators of the model, namely, root mean square error, mean square error, and mean absolute error, were selected. These indicators can measure the difference between the actual value and the predicted value of the model. The larger the value of root mean square error (RMSE), mean square error (MSE), and mean absolute error (MAE), the better the prediction effect of the model. The calculation formula is as follows:

[0140]

[0141] Among them, y i is the true value, is the predicted value, and m is the number of sample points.

[0142] The baseline models CNN-LSTM, CNN-GRU and CNN with bidirectional LST (CNN-BLSTM) are appropriately modified to be suitable for the binary sequence prediction task. The performance of the proposed model is compared with these baseline methods. Table 1 provides a detailed performance analysis of these models, showing the indicators of different models. In the baseline model, CNN-BLSTM achieves a root mean square error of 0.312, a mean square error of 0.097, and a mean absolute error of 0.291, where the proposed method further reduces these error indices to 0.236, 0.056, and 0.213.

[0143] Table 1 Comparison of prediction evaluation indicators of different models

[0144] Model Root mean square error Mean square error Mean absolute error CNN-LSTM 0.418 0.175 0.406 CNN-GRU 0.455 0.207 0.444 CNN-BLSTM 0.312 0.097 0.291 Method of the present invention 0.236 0.056 0.213

[0145] Example 3

[0146] This embodiment also provides a computer device, which is suitable for a microgrid energy management method based on a dual-sequence prediction model, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement a forced oscillation detection and positioning method for a distribution network as proposed in the above embodiment.

[0147] This embodiment further provides a storage medium on which a computer program is stored. When the program is executed by a processor, a forced oscillation detection and positioning method for a distribution network is implemented as proposed in the above embodiment.

[0148] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0149] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0150] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0151] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0152] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0153] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A microgrid energy management method based on a dual-sequence prediction model, characterized in that: It includes obtaining historical power consumption and power generation data in the microgrid system, and preprocessing the data, including outlier removal, missing value filling and data normalization; Split the preprocessed data into training set, validation set and test set; A dual-sequence prediction model consisting of a core module and a fully connected layer is constructed, and the KL divergence is used to calculate the distribution difference between power generation and power consumption.

2. The microgrid energy management method based on the dual sequence prediction model according to claim 1, characterized in that: The core module includes: a first branch, which is composed of a convolution layer, a batch normalization layer, a layer normalization layer, a ReLU activation function layer, a random inactivation layer and a multi-scale coordinate attention mechanism module; and a second branch, which is composed of a residual connection composed of a one-dimensional convolution and a ReLU activation function.

3. The microgrid energy management method based on the dual sequence prediction model according to claim 2, characterized in that: The multi-scale coordinate attention mechanism module includes: a multi-scale feature extraction unit, used to obtain one-dimensional input features of different scales; a coordinate attention unit, used to encode feature information along the horizontal and vertical directions respectively; and a residual connection unit, used to maintain the original feature information.

4. The microgrid energy management method based on the dual sequence prediction model according to claim 3, characterized in that: The multi-scale feature extraction unit uses one-dimensional convolution with different convolution kernel sizes to extract features, including convolution kernels of 1×1, 1×3 to 1×(2i+1), where i is a positive integer.

5. The microgrid energy management method based on the dual sequence prediction model according to claim 4, characterized in that: The coordinate attention unit includes: a horizontal pooling operation for obtaining the importance of information of different scales; a vertical pooling operation for obtaining the importance of information of different positions; and a feature fusion operation for combining attention information of different directions.

6. The microgrid energy management method based on the dual sequence prediction model according to claim 5, characterized in that: The fully connected layer includes a three-layer neural network structure, which is used to: map the power consumption features extracted by the core module into power consumption prediction values; and map the power generation features extracted by the core module into power generation prediction values.

7. The microgrid energy management method based on the dual sequence prediction model according to claim 6, characterized in that: The KL divergence is used to calculate the distribution difference between the power generation forecast value and the power consumption forecast value; and feed back the calculated difference to the microgrid control center for energy regulation.

8. A microgrid energy management system based on a dual-sequence prediction model, based on the microgrid energy management method based on a dual-sequence prediction model according to any one of claims 1 to 7, characterized in that: It also includes a data acquisition and preprocessing module, which acquires historical power consumption and power generation data in the microgrid system and preprocesses the data, including outlier removal, missing value filling and data normalization; Model training module, which divides the preprocessed data into training set, validation set and test set; The energy management module builds a dual-sequence prediction model consisting of a core module and a fully connected layer, and uses KL divergence to calculate the distribution difference between power generation and power consumption.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the microgrid energy management method based on the dual-sequence prediction model described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the microgrid energy management method based on the dual-sequence prediction model described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Micro-grid-oriented multi-scene electric power quantity prediction method and terminal equipment

    CN121813329A