Dual-channel wind power generation power prediction method and system

Through the dual-channel wind power generation power prediction method, combined with the improved Mamba module and LSTM module, the adaptive gradient perception EMA function and the multi-head cross-attention mechanism are used to solve the shortcomings of traditional models in data smoothing, adaptive and feature fusion, and improve the accuracy and stability of wind power generation power prediction.

CN120414534AActive Publication Date: 2025-08-01NANCHANG INST OF TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510905922.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Traditional wind power power prediction models perform poorly in data smooth noise reduction, adaptability, spatial and temporal feature fusion and frequency domain information utilization, resulting in incomplete feature extraction and affecting prediction accuracy.

Method used

The dual-channel wind power generation power prediction method is adopted, combined with the improved Mamba module and LSTM module, data smoothing is carried out through adaptive gradient-aware EMA function, multi-head cross-attention mechanism is introduced for feature fusion, and model capabilities are enhanced using frequency domain information.

Benefits of technology

It significantly improves the accuracy and stability of wind power generation power prediction, enhances the model's robustness and feature capture capabilities to noise, and achieves more comprehensive feature extraction and fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120414534A_ABST
    Figure CN120414534A_ABST
Patent Text Reader

Abstract

The invention discloses a two-channel wind power generation power prediction method and system, and the method comprises the steps: obtaining power load time sequence data, and processing the power load time sequence data according to a preset gradient perception index moving average EMA algorithm, and obtaining target power load time sequence data; inputting the target power load time sequence data into a pre-constructed Mama module and a pre-constructed LSTM module, the Mama module outputting to obtain a first result, and the LSTM module outputting to obtain a second result; and inputting the first result and the second result into a multi-head cross attention mechanism for adaptive adjustment, and outputting a final prediction result through an FFN feedforward network. And the precision and the stability of wind power generation power prediction are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wind power generation power prediction, and particularly relates to a dual-channel wind power generation power prediction method and system. Background Art

[0002] In wind power prediction technology, the traditional EMA function performs poorly in data smoothing and noise reduction, lacks adaptability, and is difficult to effectively handle noise and outliers in data. Also, the existing Mamba module design is relatively simple, unable to fully capture complex patterns and long-term dependencies in data, restricting the model's ability to capture time-series features. Moreover, the traditional LSTM module is inefficient in processing spatio-temporal feature fusion and lacks the utilization of frequency-domain information, resulting in the model being unable to comprehensively consider various influencing factors in wind power data. Additionally, the single-channel model architecture cannot fully utilize the advantages of different models, leading to incomplete feature extraction. Summary of the Invention

[0003] The present invention provides a dual-channel wind power generation power prediction method and system to solve the technical problem that the single-channel model architecture cannot fully utilize the advantages of different models, resulting in incomplete feature extraction.

[0004] In a first aspect, the present invention provides a dual-channel wind power generation power prediction method, including: Obtaining power load time-series data, and processing the power load time-series data according to a preset gradient-aware exponential moving average (EMA) algorithm to obtain target power load time-series data; Inputting the target power load time-series data into a pre-constructed Mamba module and an LSTM module respectively, where the Mamba module outputs a first result and the LSTM module outputs a second result; Inputting the first result and the second result into a multi-head cross-attention mechanism for adaptive adjustment, and outputting a final prediction result via a feed-forward neural network (FFN).

[0005] In a second aspect, the present invention provides a dual-channel wind power generation power prediction system, including: A processing module configured to obtain power load time-series data and process the power load time-series data according to a preset gradient-aware exponential moving average (EMA) algorithm to obtain target power load time-series data; An output module configured to input the target power load time-series data into a pre-constructed Mamba module and an LSTM module respectively, where the Mamba module outputs a first result and the LSTM module outputs a second result; An adjustment module configured to input the first result and the second result into a multi-head cross-attention mechanism for adaptive adjustment, and output a final prediction result via an FFN feed-forward network.

[0006] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the dual-channel wind power generation power prediction method according to any embodiment of the present invention.

[0007] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program instructions are executed by a processor, the processor is enabled to execute the steps of the dual-channel wind power generation power prediction method according to any embodiment of the present invention.

[0008] The dual-channel wind power generation power prediction method and system of the present application introduce an adaptive gradient-aware EMA function, and significantly improve the effect of data smoothing and noise reduction and enhance the robustness of the model to noise by dynamically adjusting the smoothing coefficient. By improving the Mamba module, short, medium, and long-term information in the data is extracted through convolution and feature fusion is performed. At the same time, forward and reverse Mamba modules and a chaotic oscillation gating mechanism are introduced to enhance the model's ability to capture data temporal features and dynamic feature fusion ability. The LSTM module is improved, and the spatio-temporal feature fusion efficiency is improved through spatio-temporal head division and frequency domain enhancement, and the model's ability to capture periodic features is enhanced using frequency domain information. Finally, a dual-channel model architecture is adopted, combining the advantages of the Mamba module and the LSTM module to achieve more comprehensive feature extraction and fusion. At the same time, a multi-head cross-attention mechanism is introduced to further optimize the feature weights and significantly improve the prediction accuracy. Description of the Drawings

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 It is a flowchart of a dual-channel wind power generation power prediction method provided by an embodiment of the present invention; Figure 2 It is a flowchart of smoothing data of the gradient-aware exponential moving average EMA function in a specific embodiment provided by an embodiment of the present invention; Figure 3The working flowchart of the Mamba module in a specific embodiment provided by an embodiment of the present invention; Figure 4 The structural flowchart of the LSTM module in a specific embodiment provided by an embodiment of the present invention; Figure 5 The structural block diagram of a dual-channel wind power prediction system provided by an embodiment of the present invention; Figure 6 It is the schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments

[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0012] Please refer to Figure 1 , which shows the flowchart of a dual-channel wind power prediction method of the present application.

[0013] As Figure 1 shown, the dual-channel wind power prediction method specifically includes the following steps: Step S101: Obtain power load time series data, and process the power load time series data according to a preset gradient-aware exponential moving average (EMA) algorithm to obtain target power load time series data.

[0014] In this step, please refer to Figure 2 , obtain power load time series data, perform normalization processing on the power load time series data, and scale the power load time series data to the interval [-1, 1]; Determine the optimal smoothing window size through a window selection function, set the window interval to [3, 5], calculate the information entropy of each window, and select the window size with the minimum information entropy as the final target window; Calculate the standard deviation of the data within the target window and the gradient variance of the model parameters, and calculate the dynamically adjusted smoothing coefficient according to the standard deviation of the data and the gradient variance of the model parameters. The expression is: , In the formula, is the dynamic smoothing coefficient, that is, the core parameter of EMA, is the time step, is the standard deviation of the data within the current window, used to measure the data volatility, is the gradient variance of the model parameters, which is used to measure the model update status. is the sensitivity adjustment factor. is the smoothing term, which is used to prevent the denominator from being zero. is the hyperbolic tangent function. The dynamic adjustment of the smoothing coefficient is carried out through the tanh function. is restricted to the interval [0, 0.96], and the EMA is calculated recursively to obtain the smoothed target power load time series data. Among them, the expression for recursively calculating the EMA is: , In the formula, is the exponentially weighted moving average calculated at time step t. is the input value at the current time step, that is, the original data of the data at time point t. is the exponentially weighted moving average calculated at time step t-1, that is, the exponentially weighted moving average at the previous moment.

[0015] Step S102: Input the target power load time series data into the pre-constructed Mamba module and LSTM module respectively. The Mamba module outputs the first result, and the LSTM module outputs the second result.

[0016] In this step, a dual-channel structure is adopted, which is composed of an improved Mamba module and a spatio-temporal decoupled multi-head LSTM module respectively, and an adaptive fusion is realized by combining the multi-head cross-attention mechanism.

[0017] The Mamba module extracts multi-scale time series features through three layers of depthwise separable convolutions with different dilation rates, and realizes the dynamic adjustment of multi-scale features by using a learnable linear weighted fusion mechanism. The module adopts forward and backward bidirectional state space mapping, and innovatively introduces a chaotic oscillation gating mechanism based on the Lorenz dynamical system to realize dynamic gating fusion at the time step level, effectively capturing non-stationary dynamics and long-term dependencies. The fused features are further abstracted through residual convolution and feed-forward networks, supporting efficient time series modeling and prediction.

[0018] The LSTM module first performs basic LSTM encoding on the input, and then divides the features into a time head and a space head according to the number of heads, and independently models the time step dependence and feature interaction. The time head combines the learnable decay factor through the autocorrelation matrix to dynamically strengthen the local time correlation; the space head screens important feature interactions through the gating mechanism to improve the spatial dependence modeling ability. This module also introduces frequency domain features, extracts periodic information through the fast Fourier transform, and fuses spatio-temporal-frequency multi-domain features to enhance the capture of complex periods and multi-scale changes.

[0019] The two-channel outputs are dynamically weighted and fused through a multi-head cross-attention mechanism module to achieve adaptive adjustment of the importance of different time steps and features.

[0020] Traditional Mamba modules usually rely on single-time-scale features, making it difficult to effectively capture multi-scale dynamics and lacking the adaptive ability for non-stationary sequences; traditional LSTM modules are prone to gradient vanishing in long sequences and cannot fully separate spatio-temporal features, resulting in insufficient short-term dependence and spatial relationship modeling. The traditional single-channel prediction method is difficult to balance long-term and short-term dependencies, with limited fusion ability, affecting the prediction accuracy. The present invention adopts a two-channel structure, combines an improved multi-scale Mamba module with chaotic gating to enhance the ability to capture long-term dependencies and dynamic adaptation; at the same time, The module enhances short-term and spatial dependence modeling through spatio-temporal decoupled multi-head design and frequency-domain features. The two channels cooperate and fuse to achieve complementary long-term and short-term features, adaptively adjust the weights, significantly improve the expressive ability and prediction accuracy of the model, and overcome the limitations of traditional single-channel methods.

[0021] Specifically, the improved Mamba module first performs depthwise separable convolutions with three different dilation rates on the input sequence to generate multi-scale features; subsequently, it uses learnable weights to linearly weight each scale feature in the channel dimension to obtain a unified time-frequency fusion representation. Subsequently, the module calculates two-way state space mappings in parallel for the forward and backward directions (implemented by forward Mamba and backward Mamba sub-units respectively) to capture forward and backward dependence information simultaneously. To achieve adaptive trade-off for non-stationary dynamics, this module introduces chaotic oscillation gating: iteratively generates a chaotic signal consistent with the time step length through the Lorenz dynamical system, projects it into two sequences of gating coefficients, and realizes the per-time-step dynamic weighted fusion of forward and backward features to output the final long-range representation vector.

[0022] Please refer to Figure 3 for the work process as follows: Input data: The data received by the Mamba module is in the form of a three-dimensional array with a shape of (B batch size × T time step length × D feature dimension). That is, each input sample includes data of several time steps and multiple features (such as wind speed, temperature, etc.).

[0023] Multi-scale convolution for feature extraction: For the input data with a shape of , three groups of parallel depthwise separable one-dimensional convolutional (DepthwiseConv1D) modules are adopted, using convolutional kernels with dilation rates of 1, 2, and 4 respectively, and the fixed size is 3. Through different dilation rates, these convolutional kernels can capture context dependencies at different scales in the time series, and the receptive field expands from 3 to 9, covering the sequence dynamics from the medium term to the long term, while keeping the computational complexity unchanged. Depthwise separable convolution enables each channel to operate independently, avoiding redundancy between channels and improving parameter efficiency. The shape of the feature tensor output by each group of convolutions, that is, the output data, is , representing the temporal features at the corresponding scale, and realizing multi-scale and efficient feature extraction.

[0024] Note: Each group corresponds to a different dilation rate (1, 2, 4) respectively, and the output of each convolutional group will generate a different feature tensor.

[0025] Learnable scale fusion: After completing multi-scale convolutional feature extraction, the model obtains three groups of feature tensors with different temporal receptive fields, that is, the shape of the data is [batch size, time step length, feature dimension (1, 2, 3)], representing the information of short-term (dilation = 1), medium-term (dilation = 2), and long-term (dilation = 4) time-dependent structures respectively. Specifically, the model equips two trainable scalar parameters for the output of each convolutional channel: (weight coefficient) and (bias term). The fusion process is as follows: , where, represents the feature map obtained by convolution at the -th scale, is the weight coefficient, controlling the contribution degree in the fusion process, is the bias term, and are automatically learned and optimized through the training process of the model; All and parameters will be automatically optimized during the training process to minimize the final prediction error. Here, the shape of the output data is , where the feature dimension = feature dimension + feature dimension + feature dimension .

[0026] Bidirectional Mamba state mapping: After completing multi-scale feature fusion, the model obtains a unified time series feature tensor with the shape of The forward path transforms the features at each time step through the Hierarchical Mamba sub-module (implemented as a Dense layer) to capture the causal dependencies of the sequence.

[0027] The forward module calculates: , where, represents the input features at the current time step, is the weight matrix, learned, is the bias term, learned, is the activation function, and the GELU activation function is used to introduce non-linear transformation.

[0028] So the output form of the forward is .

[0029] The reverse path first reverses the time dimension of the input sequence, then processes it through the Memory Efficient Mamba module (also a Dense layer), and finally reverses the output to align it with the forward output. The reverse path models the current state from a future perspective, supplementing the lag, anti-causal, and periodic information that is difficult to capture by the forward path, and enhancing the model's ability to express complex time series structures.

[0030] The reverse module calculates: , where, represents the input features at the reversed current time step.

[0031] So the output form of the reverse is .

[0032] The chaotic oscillator generates a gating signal: By introducing the Lorenz chaotic system, a dynamic gating signal is generated for each time step, which is used for adaptive fusion of forward and reverse features. First, a three-dimensional state vector is initialized for each sample. Here, the vector (1, 1, 1) is taken as the starting state of the Lorenz system; then its discrete iteration equation is used: , where, , , are the parameters of the Lorenz system, set as hyperparameters or trainable; is the time step, usually taken as 0.01; each sample is iterated independently to generate a chaotic trajectory of length L, which is the total length of the sample. Generating a chaotic trajectory of length L is the chaotic signal. So the output vector is , (1, 2, ..., L), a total of L time series are obtained.

[0033] Gated Projection and Weight Splitting: After generating the chaotic signal in the previous step, in order to convert this signal into gated weights that can be used for forward and backward feature fusion, a lightweight gated projection mechanism is introduced in this step. The model feeds it into a DenseLayer, projecting the chaotic signal at each time step onto a vector of length 2:

[0034] where and are learnable parameters. The sigmoid function restricts the output to the interval (0, 1), ensuring the interpretability and stability of its use as a weight coefficient. These two outputs respectively represent the gating strengths of the forward and backward features at the current time step. At this time, the corresponding .

[0035] Time-step Level Dynamic Fusion: After completing the generation of gated weights, this step is responsible for the time-step level adaptive fusion of bidirectional features. At each time step , the model already has the forward feature representation , the backward feature representation , as well as the corresponding gating coefficients , . The fusion operation is: , where is the element-wise multiplication, , are the gating coefficients of the forward and backward features at time step t respectively, is the feature representation after fusion at time step t, is the forward feature obtained by the forward path Mamba module at time step t, is the backward feature obtained by the backward path Mamba module at time step t.

[0036] This operation is performed over the entire sequence length, with each time step calculated independently, thereby generating a unified, fused feature sequence of shape .

[0037] The core advantage of this fusion strategy lies in "adaptive per time step". Different from traditional static weighting (such as 0.6 / 0.4) or fixed residual structures, this method can select the most reliable feature source for different time steps according to the changes in the chaotic state. For example, in the mutation region, it may rely more on the forward features of short-term changes; while in the stationary region, it may pay more attention to capturing the periodic backward features.

[0038] Output long-range representation: The fused temporal features are first fed into the MambaEncoder module, which contains a DepthwiseConv1D layer to extract local patterns of adjacent time steps, and a feed-forward network composed of two Dense layers (with GELU activation) for non-linear transformation. LayerNorm and residual connections ensure numerical stability. The processed features still maintain the original time step structure. Subsequently, the model flattens the three-dimensional tensor into a two-dimensional vector and generates the prediction result through the Dense output layer, completing the conversion from feature fusion to the final output. At this time, the output dimension is the same as the dimension of the input data, so the shape of the output data is . This step realizes the in-depth processing of the fused features and the docking of the task objectives.

[0039] Furthermore, in the LSTM module, the data first undergoes basic temporal encoding, and then through multi-head feature partitioning, it is decomposed into two parts: a time head and a space head, which are respectively used for time-dependent modeling and spatial feature interaction modeling. At the same time, the original input data will also undergo frequency domain enhanced feature extraction to enhance and capture periodic features in the frequency domain. Finally, in the spatio-temporal-frequency domain fusion processing step, the features obtained from the time head, space head, and frequency domain enhanced feature extraction will be fused together to form a unified feature representation, and then the final result is generated through output mapping.

[0040] Please refer to Figure 4 , and the workflow is as follows: Basic temporal encoding: Receive the input time series data and use the basic LSTM layer for preliminary feature extraction and encoding. The LSTM layer dynamically selects and retains the key information of the input sequence through gating units, captures the short-term dependencies and local change trends of the data, and the features of each time step in the output contain the encoding of the temporal information. The dimension of the encoded output is the product of the single-head feature dimension (d) and the total number of heads ( ), laying the foundation for subsequent processing. This step can be formally expressed as: , where (B batch size × T time step length × D feature dimension) is the input sequence, is the high-dimensional temporal feature data after encoding.

[0041] Multi-head feature partitioning: Structurally process the features after basic encoding. The total number of heads is , The single - head feature dimension is , then the total feature dimension is . Divide it into temporal modeling heads and spatial modeling heads, satisfying . Perform slicing on the last dimension of the temporal feature data ( ), and this division realizes the separation of temporal information and feature interaction information. Obtain the temporal feature tensor and the spatial feature tensor .

[0042] Temporal head dependency modeling: First, reshape the temporal head feature tensor into . Calculate the first - order temporal autocorrelation matrix for each temporal modeling head. This is to calculate the correlation between time steps. Introduce a learnable time decay factor vector , multiply it element - by - element with and sum them to obtain the temporal - domain energy feature , with a shape of . Among them, is the result obtained after the temporal feature tensor undergoes reshape / re - arrangement (exchanging the middle two dimensions), is represented as the transpose of , represents product, represents matrix multiplication, represents all - ones vector; The decay factor vector is a learnable parameter used to control the dependency between time steps. It helps the model to apply different weights to the correlations between different time steps when modeling time series. By learning the decay factor, the model can flexibly determine which dependencies between time steps are more important and which can be ignored.

[0043] Spatial head feature interaction modeling: Reshape the spatial head feature tensor into . Calculate the autocorrelation matrix for each spatial modeling head. Here, is represented as the transpose of . When calculating the spatial energy tensor, this is to obtain spatial dependency information by calculating the correlation strength of each time head.

[0044] The calculation formula is , where represents performing a maximum operation on the last dimension, with a shape of . Introduce a gating network for the original spatial head features Perform dynamic gating weight prediction: , where is the activation function, , are all learnable parameters. Finally, multiply the spatial energy tensor and the gating weights element-wise to obtain the weighted spatial dependence features , with the shape of , achieving accurate modeling of the interaction relationship between feature variables. Among them, and will first be initialized with initial values by the model.

[0045] Then, during the forward propagation of the model, the spatial head feature tensor will pass through the linear transformation defined by and , and then apply the Sigmoid activation function to generate the gating weight G.

[0046] Next, the output of the model will be compared with the true target value, and the prediction error or loss will be calculated through the loss function.

[0047] Then, the backpropagation algorithm will start from the loss function and calculate the gradients of the loss with respect to all the parameters of the model (including and ).

[0048] Finally, the optimizer will adjust the and values according to the direction and magnitude of the gradients, as well as hyperparameters such as the possible learning rate, to minimize the prediction error.

[0049] Frequency domain enhanced feature extraction: Perform on the high-dimensional time series feature data transformation to obtain the frequency domain , and take the real part of the frequency domain as the frequency domain feature tensor , where is the number of reserved frequency components; Transform the feature dimension of the frequency domain feature tensor to be consistent with the dimension of the spatio-temporal feature output through a linear projection layer to obtain , represents the projection operation, that is, project onto , and slice according to the multi-head division criterion to obtain the time modeling frequency domain feature and the spatial modeling frequency domain feature ; Spatio-temporal frequency domain fusion processing: Combine the time domain energy features Perform element-wise multiplication with the time-modeled frequency-domain features to obtain time-enhanced features , and perform element-wise multiplication of the spatial energy features with the spatial-modeled frequency-domain features to obtain spatial-enhanced features . Then concatenate the time-enhanced features and the spatial-enhanced features in the feature dimension to form a fused feature tensor , and output the second result through a fully connected neural network

[0050] Output mapping and result generation: First, apply a fully connected neural network (FFN) to the fused spatio-temporal-frequency feature tensor for non-linear transformation and feature reconstruction. The first layer of the FFN is , where , are learnable parameters, and the activation function enhances the non-linear expression ability of the features. The second layer is , further abstracting the features. Finally, map the features to the target prediction dimension through a linear output layer , and finally output the second result

[0051] Step S103: Input the first result and the second result into a multi-head cross-attention mechanism for adaptive adjustment, and output the final prediction result through the FFN feed-forward network

[0052] In this step, the introduced multi-head cross-attention mechanism can dynamically adjust the weights of the outputs of different modules according to the data features at each time step by establishing an adaptive fusion mechanism between the outputs of different modules (e.g., the Mamba module and the LSTM module in the present invention)

[0053] Workflow of the multi-head cross-attention mechanism Input data: The input data of the multi-head cross-attention mechanism comes from the outputs of two different modules: the output S of the Mamba module , representing features capturing dependencies. The output I of the LSTM module , representing features capturing dependencies

[0054] Feature Mapping and Spatial Transformation: The multi-head cross-attention mechanism first linearly maps the output features of the Mamba module and the LSTM module into multiple different feature subspaces (referred to as multiple "attention heads") to obtain Q, K, and V. Here, n1 attention heads are used to process the output of the Mamba module, and each head processes features with a dimension of m, so n1×m = D; similarly, n2 attention heads are used to process the output of the LSTM module, and each head processes features with a dimension of m, so n2×m = D.

[0055] Here, Q, K, and V are calculated as follows: Q represents the query, which is used to represent the query requirements of the current task. It comes from the output of the Mamba module, and the calculation formula is , is the linear transformation of the query, in the form of .

[0056] K represents the key, which is used to calculate the similarity with the query. It comes from the output of the LSTM module, and the calculation formula is , is the linear transformation of the key, in the form of .

[0057] V represents the value, which is the feature associated with the key. It also comes from the output of the LSTM module and is used for weighted summation to generate the final output feature. The calculation formula is , is the linear transformation of the value, in the form of .

[0058] Note: The purpose of these linear projections is to align the features from different modules so that they can interact effectively in the attention mechanism, calculate dynamic attention weights, and generate weighted output features.

[0059] Calculating the Similarity between the Query and the Key: The similarity between the query and the key is calculated through the dot product. This step is the core of the attention mechanism because it helps the model determine the strength of the relationship between the query at each time step and the keys at other time steps. The calculation formula is as follows:

[0060] The shape of the calculated similarity matrix is: . Here, there are two Ts, indicating the similarity between each query and the keys at all other time steps.

[0061] Note: Through the dot product operation, the higher the similarity score between the query and the key, the stronger the correlation between them. In this way, the model can judge which features at each time step are more important for the current task by calculating the similarity between each query and all keys.

[0062] Normalized attention weights: By function to normalize the calculated similarity scores and convert them into attention weights. In this way, higher similarities will get higher weights, while lower similarities will get lower weights. Normalization formula: , is the normalized attention weight, indicating the importance of each key for the query.

[0063] is the similarity score between the query and the key.

[0064] Normalize all keys so that the sum of the weights is 1.

[0065] Shape of the normalized attention weights: . This represents the weight relationship between each query and all keys.

[0066] Calculate the weighted output: Use the normalized attention weights to perform a weighted sum on the values to obtain the final output feature. The calculation formula is: . Shape of the weighted output: , and the results of each head will be concatenated together, finally obtaining an output with a shape of . The output form of the combination of multiple attentions is .

[0067] It should be noted that the FFN performs a deep non-linear transformation on the input data through multiple fully connected layers (Dense layers), helping the network learn complex feature representations. The FFN module maps in the feature space and can extract high-order features in the data through multi-level non-linear activation functions, thereby enhancing the model's expressive ability and improving the prediction accuracy.

[0068] Input the first result and the second result into the multi-head cross-attention mechanism, and the multi-head cross-attention mechanism fuses the first result and the second result to obtain the target feature; Input the target feature into the first fully connected layer, and map it to a new feature space through linear transformation to obtain the first output result after linear transformation, where the expression of the linear transformation mapping is: , In the formula, is the first output result, is the input data of the first fully connected layer, is the weight matrix of the first fully connected layer, is the bias vector of the first fully connected layer; After processing the first output result through the activation function, it is input into the second fully connected layer to perform a non-linear transformation on the first output result to obtain the second output result. The expression is: , In the formula, is the second output result, is the weight matrix of the second fully connected layer, is the bias vector of the second fully connected layer; The second output result is input into the third fully connected layer to perform feature mapping on the second output result as the final prediction result. The expression is: , In the formula, is the weight matrix of the third fully connected layer, is the bias vector of the third fully connected layer.

[0069] Note: The first two fully connected layers of the FFN module not only perform non-linear processing on the input data but also perform feature compression. By adjusting the output dimension layer by layer, FFN can effectively compress redundant information and retain the most important features, thereby improving the generalization ability of the model. The final output layer (dense3) of FFN compresses all the feature information in the network into a scalar value, and this value is the final prediction result. In the time series prediction task, this layer converts the internal feature representation of the model into the actual predicted value, such as the future wind power generation power.

[0070] In summary, the method of the present application captures short-term dependencies through an improved LSTM module and extracts long-term dependencies through an improved Mamba module, supplemented by a multi-head cross-attention mechanism for adaptive fusion, and combines a gradient-aware exponential moving average EMA function to perform more effective smoothing processing on the original data in the data preprocessing stage, thereby improving the prediction accuracy and significantly improving the accuracy and stability of wind power generation power prediction.

[0071] Please refer to Figure 5 , which shows the structural block diagram of a dual-channel wind power generation power prediction system of the present application.

[0072] As shown in Figure 5As shown in the figure, the dual-channel wind power prediction system 200 includes a processing module 210, an output module 220, and an adjustment module 230.

[0073] Among them, the processing module 210 is configured to obtain power load time series data and process the power load time series data according to a preset gradient-aware exponential moving average (EMA) algorithm to obtain target power load time series data; the output module 220 is configured to input the target power load time series data into a pre-constructed Mamba model and an LSTM model respectively. The Mamba model outputs a first result, and the LSTM model outputs a second result; the adjustment module 230 is configured to input the first result and the second result into a multi-head cross-attention mechanism for adaptive adjustment, and output a final prediction result via a feed-forward neural network (FFN).

[0074] It should be understood that Figure 5 the modules described in Figure 1 correspond to the respective steps in the method described in the reference Figure 5 Therefore, the operations, features, and corresponding technical effects described above for the method also apply to the modules in

[0075] In some other embodiments, the embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored. When the program instructions are executed by a processor, the processor executes the dual-channel wind power prediction method in any of the above method embodiments; As an implementation manner, the computer-readable storage medium of the present invention stores computer-executable instructions, which are set as follows: Obtain power load time series data, and process the power load time series data according to a preset gradient-aware exponential moving average (EMA) algorithm to obtain target power load time series data; Input the target power load time series data into a pre-constructed Mamba model and an LSTM model respectively. The Mamba model outputs a first result, and the LSTM model outputs a second result; Input the first result and the second result into a multi-head cross-attention mechanism for adaptive adjustment, and output a final prediction result via a feed-forward neural network (FFN).

[0076] A computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the dual-channel wind power prediction system, etc. In addition, the computer-readable storage medium may include high-speed random access memory, and may also include a memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories may be connected to the dual-channel wind power prediction system through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0077] Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, as Figure 6 shown, the device includes: a processor 310 and a memory 320. The electronic device may further include: an input device 330 and an output device 340. The processor 310, the memory 320, the input device 330, and the output device 340 may be connected through a bus or other means, Figure 6 taking the connection through the bus as an example. The memory 320 is the above-mentioned computer-readable storage medium. The processor 310 executes various functional applications and data processing of the server by running non-volatile software programs, instructions, and modules stored in the memory 320, that is, implementing the dual-channel wind power prediction method in the above method embodiment. The input device 330 may receive input digital or character information, and generate key signal inputs related to user settings and function controls of the dual-channel wind power prediction system. The output device 340 may include a display device such as a display screen.

[0078] The above electronic device may execute the method provided by the embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference may be made to the method provided by the embodiment of the present invention.

[0079] As an implementation manner, the above electronic device is applied to a dual-channel wind power prediction system and is used for a client, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Obtain power load time series data, and process the power load time series data according to a preset gradient-aware exponential moving average (EMA) algorithm to obtain target power load time series data; Input the target power load time series data into the pre-constructed Mamba model and LSTM model respectively. The Mamba model outputs the first result, and the LSTM model outputs the second result; Input the first result and the second result into the multi-head cross-attention mechanism for adaptive adjustment, and output the final prediction result via the FFN feed-forward network.

[0080] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A dual-channel wind power prediction method, characterized in that Including: Obtain the time series data of power load, and process the time series data of power load according to the preset gradient-aware exponential moving average (EMA) algorithm to obtain the target time series data of power load; Input the target time series data of power load into a pre-constructed Mamba module and an LSTM module respectively. The Mamba module outputs a first result, and the LSTM module outputs a second result; Input the first result and the second result into a multi-head cross-attention mechanism for adaptive adjustment, and output the final prediction result through a feed-forward neural network (FFN).

2. The dual-channel wind power generation power prediction method according to claim 1, wherein The process of processing the time series data of power load according to the preset gradient-aware exponential moving average (EMA) algorithm to obtain the target time series data of power load includes: Perform normalization processing on the time series data of power load, and scale the time series data of power load to the interval [-1, 1]; Determine the optimal smoothing window size through a window selection function, set the window interval as [3, 5], calculate the information entropy of each window, and select the window size with the minimum information entropy as the final target window; Calculate the standard deviation of the data within the target window and the gradient variance of the model parameters, and calculate the dynamically adjusted smoothing coefficient according to the standard deviation of the data and the gradient variance of the model parameters. The expression is: , In the formula, is the dynamic smoothing coefficient, which is the core parameter of EMA, is the time step, is the standard deviation of the data within the current window, which is used to measure the data volatility, is the gradient variance of the model parameters, which is used to measure the model update status, is the sensitivity adjustment factor, is the smoothing term, which is used to prevent the denominator from being zero, is the hyperbolic tangent function; Dynamically adjust the smoothing coefficient through the tanh function Restrict it to the interval [0, 0.96], and recursively calculate the EMA to obtain the smoothed target power load time series data. Among them, the expression for recursively calculating the EMA is: , Wherein, is the exponentially weighted moving average calculated at time step t, is the input value at the current time step, i.e., the original data of the data at time point t, is the exponentially weighted moving average calculated at time step t-1, i.e., the exponentially weighted moving average at the previous moment.

3. A dual-channel wind power generation power prediction method according to claim 1, characterized in that Where, Input the target time series data of power load into a pre-constructed Mamba module respectively. The Mamba module outputs a first result, specifically including: Perform depthwise separable convolution operations with three different dilation rates on the target time series data of power load respectively to obtain three groups of feature tensors with different temporal receptive fields; Perform linear weighting on each of the three groups of feature tensors with different temporal receptive fields in the channel dimension according to the learnable weights to obtain a fused feature tensor. The expression is: , In the formula, represents the feature map obtained by convolution at the -th scale, is the weight coefficient that controls the contribution degree in the fusion process, is the bias term, and are automatically learned and optimized through the training process of the model; Input the fused feature tensor into a forward path and a backward path respectively. The forward path outputs the forward features at each time step, and the backward path outputs the backward features at each time step; Iteratively generate a chaotic signal consistent with the time step length through a Lorenz chaotic system, and project the chaotic signal into a first gating coefficient sequence and a second gating coefficient sequence; Perform hourly dynamic weighted fusion on the forward features at each time step and the backward features at each time step according to the first gating coefficient sequence and the second gating coefficient sequence, and output the final long-range representation vector, that is, obtain the first result. Among them, the expression of hourly dynamic weighted fusion is: , wherein, is element-wise multiplication, , are the gating coefficients of the forward feature and the backward feature at time step t, respectively, is the feature representation after fusion at time step t, is the forward feature obtained by the forward path Mamba module at time step t, is the backward feature obtained by the backward path Mamba module at time step t.

4. A dual-channel wind power generation power prediction method according to claim 1, characterized in that, Where, Input the target time series data of power load into a pre-constructed LSTM module respectively. The LSTM module outputs a second result, specifically including: Perform preliminary feature extraction and encoding on the target time series data of power load according to the LSTM layer to obtain high-dimensional time series feature data. The expression is: , In the formula, is high-dimensional time-series feature data, , where d is the single-head feature dimension, h is the total number of heads, is the target power load time-series data, , where B is the batch size, T is the time step length, and D is the feature dimension, indicates that all elements of the output tensor Y are real numbers; Divide the high-dimensional time series feature data into time modeling heads and spatial modeling heads, where , and perform slicing on the last dimension of the high-dimensional time series feature data to obtain a time feature tensor and a spatial feature tensor ; Calculate the first-time autocorrelation matrix for each time modeling header , is reshaped, i.e., reshaped, and a learnable time decay factor vector is introduced, which is multiplied element-wise with the first-time autocorrelation matrix and summed to obtain the time-domain energy feature , where is denoted as the transpose of denotes the product,[[]] denotes matrix multiplication,[[]] denotes all vectors; Calculate the second autocorrelation matrix for each spatial modeling header and calculate the spatial energy tensor based on the second autocorrelation matrix , where is reshaped, i.e., dimensionally adjusted denoted as transposed denotes performing a maximum operation on the last dimension; Multiply the spatial energy tensor element-wise with the gating weights to obtain the weighted spatial energy features , , where is the activation function, , are both learnable parameters; For high-dimensional time-series feature data perform transformation to obtain the frequency domain , and take the real part of the frequency domain as the frequency domain feature tensor , where is the number of retained frequency components; Transform the feature dimension of the frequency-domain feature tensor to be consistent with the dimension of the spatio-temporal feature output, obtaining , indicating a projection operation, and slicing according to the multi-head division criterion, obtaining the time-modeling frequency-domain feature and the space-modeling frequency-domain feature ; Multiply the time-domain energy feature element-wise with the time-modeled frequency-domain feature to obtain the time-enhanced feature , and multiply the spatial energy feature element-wise with the spatial-modeled frequency-domain feature to obtain the spatial-enhanced feature ; Concatenate the time-enhanced feature and the space-enhanced feature in the feature dimension to form a fused feature tensor , and output a second result through a fully-connected neural network.

5. A dual-channel wind power generation power prediction method according to claim 1, characterized in that, The process of inputting the first result and the second result into a multi-head cross-attention mechanism for adaptive adjustment, and outputting the final prediction result through a feed-forward neural network (FFN) includes: Input the first result and the second result into a multi-head cross-attention mechanism, and the multi-head cross-attention mechanism fuses the first result and the second result to obtain a target feature; Input the target feature into a first fully connected layer, and map it to a new feature space through linear transformation to obtain a first output result after linear transformation, where the expression of the linear transformation mapping is: , Wherein, is the first output result, is the input data of the first fully connected layer, is the weight matrix of the first fully connected layer, is the bias vector of the first fully connected layer; After subjecting the first output result to processing by an activation function, it is input into a second fully connected layer to perform a non-linear transformation on the first output result, obtaining a second output result. The expression is: , wherein, is the second output result, is the weight matrix of the second fully-connected layer, is the bias vector of the second fully-connected layer; Input the second output result into a third fully connected layer, and map the second output result to a final prediction result, and the expression is: , In the formula, is the third output result, is the weight matrix of the third fully connected layer, is the bias vector of the third fully connected layer.

6. A dual-channel wind power generation power prediction system, characterized in that, Comprising: A processing module configured to acquire power load time series data and process the power load time series data according to a preset gradient-aware exponential moving average (EMA) algorithm to obtain target power load time series data; An output module configured to respectively input the target power load time series data into a pre-constructed Mamba module and an LSTM module, the Mamba module outputs a first result, and the LSTM module outputs a second result; An adjustment module configured to input the first result and the second result into a multi-head cross-attention mechanism for adaptive adjustment, and output a final prediction result via a feed-forward neural network (FFN).

7. An electronic device, characterized in that, Comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Ultra-short-term wind power segmentation prediction method based on turning time period identification

    CN114386324A

  • Two-channel wind power prediction method integrated with extrusion incentive attention mechanism

    CN115577748A

  • Wind power generation power prediction method and system

    CN119312168A

  • Photovoltaic generating capacity prediction method and device based on dynamic decomposition and digital-analog fusion

    CN119598387A

  • Power load prediction method based on dual-channel cross attention network

    CN119669732A