Wind power prediction method and device based on multivariable DMD and transfer learning, and medium
By combining multivariate DMD and transfer learning technology, the ETO-Medformer model was constructed, and the data inadequate and overfitting of wind power prediction in wind farms was solved, achieving higher prediction accuracy and model generalization ability.
Patent Information
- Application Number
- CN202510261162.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
The existing wind power prediction methods have problems such as insufficient data and overfitting, making it difficult to accurately predict the wind power fluctuation characteristics of different wind farms.
The wind power prediction method based on multivariate DMD and transfer learning is adopted, and the ETO-Medformer model is built through greedy feature selection, sparseness, promoting dynamic modal decomposition, exponential triangle optimization and parameter transfer learning strategies to improve the accuracy of wind power prediction and the generalization ability of the model.
The accuracy of wind power prediction and the generalization ability of the model are improved, the problems of insufficient data and overfitting are solved, the robustness and adaptability of the model are enhanced, and it is suitable for efficient operation of different wind farms and stable scheduling of power systems.
Smart Images

Figure CN120218312A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wind power prediction in wind farms, and particularly relates to a wind power prediction method, device and medium based on multivariate DMD and transfer learning. Background Art
[0002] Wind energy is a clean and renewable energy source and plays an important role in the energy transition. The reduction of the cost of wind power equipment has promoted its wide application. However, wind power is affected by multiple factors, with large fluctuations and high uncertainty. Accurate prediction of wind power is of great significance for power system dispatching, operation optimization of wind farms, and grid connection and consumption of wind power.
[0003] Traditional wind power prediction methods such as physical models and simple statistical methods have deficiencies. Although data-driven methods are widely used, they also face problems such as overfitting. Multivariate DMD can process multivariate data and extract key dynamic modes; transfer learning can share knowledge of different wind farms. The combination of the two can improve prediction accuracy, solve data insufficiency, and enhance the generalization ability of the model.
[0004] Newly built or wind farms with imperfect data often face the problem of data insufficiency. Transfer learning can make up for data insufficiency, and multivariate DMD can effectively utilize limited data. The wind power fluctuation characteristics of different wind farms are different. This combination enables the model to not only focus on specific characteristics but also absorb general laws, enhancing the generalization ability and improving the stability and reliability in practical applications. Summary of the Invention
[0005] Object of the Invention: The present invention proposes a wind power prediction method, device and medium based on multivariate DMD and transfer learning. By integrating multivariate DMD and transfer learning technologies, it improves the accuracy of wind power prediction and the generalization ability of the model, solves problems such as differences in data characteristics and insufficient data volume of different wind farms, and provides more accurate wind power prediction support for the efficient operation of wind farms and the stable dispatching of power systems.
[0006] Technical Solution: The wind power prediction method based on multivariate DMD and transfer learning according to the present invention includes the following steps:
[0007] Step 1: Collect relevant data information in the source domain and the target domain, and use greedy feature selection (GFS) to perform dimensionality reduction processing on the data to remove redundant features in the data;
[0008] Step 2: Process the data after feature selection based on the sparse-promoted dynamic mode decomposition algorithm (SPDMD) to reduce the complexity of the data; reconstruct the data to form an input data set;
[0009] Step 3: Build a wind power prediction model based on Medformer;
[0010] Step 4: Use the Exponential Triangle Optimization algorithm ETO to update the parameters of the Medformer model, obtain the optimized model ETO-Medformer, and obtain the optimal parameter set;
[0011] Step 5: Based on the parameter transfer learning strategy, train ETO-Medformer on the source domain, output the source domain wind power prediction result through the trained model, and use its parameters as the initial parameters of the same model on the target domain. Finally, further train ETO-Medformer on the target domain, and use the trained model to predict the power of wind power generation to obtain the final prediction result.
[0012] Furthermore, the relevant data information described in Step 1 includes wind speed, wind direction, temperature, humidity, air pressure, pitch angle, and nacelle yaw angle.
[0013] Furthermore, the implementation process of using Greedy Feature Selection GFS to perform dimensionality reduction processing on the data described in Step 1 is as follows:
[0014] Set the initial feature set F (0) = φ, set the iteration number k = 0, and set the total feature set F = {f1, f2, f3, f4, f5, f6, f7};
[0015] For each step k, evaluate the remaining feature f j ; realize the marginal contribution of f j by calculating the difference between the performance index after including this feature and the performance index without including this feature:
[0016] Δμ(f j |F (k) ) = μ(F (k) ∪{f j}) - μ(F (k) )
[0017] In the formula, j is the index of the feature; F (k) is the selected feature set at the k-th iteration; μ is the performance index function;
[0018] Among all the remaining features, select the feature that maximally improves the performance index
[0019]
[0020] In the formula, j * is the index of the maximum feature; argmax is the maximum value parameter; Δμ(·) is the change in the performance index;
[0021] Update the feature set and add the selected feature to the feature set:
[0022]
[0023] In the formula, F (k+1) represents the feature set after the (k + 1)-th iteration;
[0024] Update the iteration count, k = k + 1; check whether the stop condition is satisfied. Stop when the maximum number of iterations K is reached. The stop condition can be specified as:
[0025]
[0026] In the formula, τ is a preset threshold; the iteration ends, and the final feature set F (s) ={g1, g2,..., g s}.
[0027] Furthermore, the implementation process of step 2 is as follows:
[0028] Select a set of data from the data after feature selection to construct a set of snapshot sequences, forming a data matrix: (x1 x2... x m+1 ); then construct two snapshot matrices from the snapshot sequences:
[0029] X = [x1 x2... x m
[0030] X′ = [x2 x3... x m+1
[0031] And assume that the snapshots are generated by a discrete-time linear time-varying algorithm, which is expressed as follows:
[0032] x k+1 = Ax k , k = 1, 2,..., m
[0033] In the formula, A is a linear operator;
[0034] Connect the data matrices X and X' through the linear operator A, then the matrix X' has the following representation:
[0035] X′ = [x2 x3 … x m+1 = [Ax1 Ax2 … Ax m = AX
[0036] Perform singular value decomposition on the matrix X, select an appropriate truncation value r to reduce the dimension of the data matrix, and the redundant terms to be eliminated are represented by rem, and its dimension is m - r:
[0037]
[0038] In the formula, U is the left singular vector matrix; ∑ is the diagonal matrix; V* is the conjugate transpose of the right singular vector matrix; represents the submatrix composed of the first r left singular vectors selected from U; represents the diagonal submatrix composed of the first r singular values selected from ∑; represents selecting from V * the conjugate transpose of the submatrix composed of the first r right singular vectors; the linear operator A is expressed as: is expressed as:
[0039] Using the least squares method, minimize the Frobenius norm between X' and the difference of AX:
[0040] where there exist linearly independent eigenvectors (w1 w2 … w r ), and the corresponding eigenvalues are {λ1, λ2, …, λ r}, then transform it into the focal coordinate form:
[0041]
[0042] In the formula: Let and {z1, z2, …, z r} are eigenvectors, and the corresponding eigenvalues are Scale them appropriately to satisfy the following biorthogonality condition:
[0043]
[0044] Use a linear combination of DMD modes to approximate the numerical snapshots:
[0045]
[0046] In the formula: x k is the mode corresponding to the decomposition; α i is the amplitude corresponding to the DMD mode; expressed in matrix form as:
[0047]
[0048] The above formula shows that the dynamic mode evolution characteristics are controlled by the Vandermonde matrix, and the Vandermonde matrix is determined by the r complex eigenvalues of the matrix whose eigenvalues contain information on the relevant time frequencies and growth decay rates;
[0049] Introduce a sparsification method to adjust the amplitudes and eliminate non-critical modes; first construct the objective function regarding α i :
[0050] J(α) = α * pα - q * α - α * q + trace(Σ * Σ)
[0051] Where p is a complex vector; q is a right complex vector; trace(Σ * Σ) is the trace of the product of the conjugate transpose of Σ and Σ itself;
[0052] Add a regularization term card(z) to the objective function J(α) to solve the sparsity problem, and finally obtain the sparsified amplitude vector D α :
[0053] D α = min α J(α) + card(z)
[0054] This penalizes the number of non - zero elements in the amplitude vector D α and thus determines the finally optimized sparsified amplitude vector D α ;
[0055] Let Finally, there is an expression for the corresponding decomposition mode: X ≈ ΦD α V and ;
[0056] Set the stopping criterion for the final modal decomposition:
[0057] ||Ax k+1 + Bz k+1 - c||2 ≤ ε prim
[0058] ||z k+1 - z k+1 ||2 ≤ ε dual
[0059] Where A, B, C are constants; ∈ prim and ∈ dual represent the stopping thresholds; perform data iteration until the number of modes is decomposed to the set threshold, and finally it is expressed as: According to the above method, finally construct the final dataset: X = {g1', g2',..., g s '}.
[0060] Furthermore, the implementation process of step 3 is as follows:
[0061] Define the input data as multivariate time - series samples Let \(T\) denote the number of timestamps and \(C\) denote the number of channels. In the wind power prediction scenario, channel \(C\) contains: \(g1', g2',..., g\) s '
[0062] Introduce cross-channel multi-granularity patch embedding:
[0063] First, perform patch splitting. Given a list of different patch lengths \(\{L1, L2, …, L\}\) n For the \(i\)-th patch length \(L\) i , split the input sample into \(N\) i non-overlapping cross-channel patches Ensure that the timestamp \(T\) is divisible by \(L\) through zero-padding i such that \(N\) i = T / L i ;
[0064] Then, perform linear projection to map the patches to the latent embedding space, obtaining:
[0065]
[0066] where: is the non-overlapping patch; \(W\) (i) is the embedding space; is the result after mapping;
[0067] Finally, add the fixed positional embedding \(W\) pos and the learnable granularity embedding to obtain the final patch embedding:
[0068]
[0069] Next, introduce the multi-granularity self-attention mechanism and initialize the router. Initialize a router for each granularity:
[0070]
[0071] where; \(u\) (i) is one of the routers;
[0072] Inner-granularity self-attention adjustment. Vertically concatenate the patch embedding \(x\) (i) and the router embedding \(u\) (i) to form the intermediate sequence \(z\) (i) = [x (i) ‖ u (i) ; Perform self-attention operation on \(z\) (i) to update the patch embedding and the router embedding:
[0073] x (i) ← Attn Intra (x(i) , z (i) , z (i) )
[0074] u (i) ← Attn Intra (u (i) , z (i) , z (i) )
[0075] Where Attn Intra (·) is the self-attention mechanism calculation function;
[0076] Perform cross-granularity self-attention operation, connect all router embeddings into a sequence U = [u (1) ‖u (2) ‖...‖u (n) ; perform self-attention operation on the router embedding u (i) for each granularity i, and update the router embedding:
[0077] u (i) ← Attn Intra (u (i) , U, U)
[0078] Perform layer normalization on the obtained patch embedding x (i) :
[0079]
[0080] Where: μ B and σ B are the batch mean and standard deviation respectively, and γ and β are learnable parameters;
[0081] Apply a linear layer for prediction:
[0082]
[0083] Where is the predicted value of wind power; W P and b P are the output layer weights and biases; Var represents variance; E represents expectation.
[0084] Furthermore, the implementation process of step 4 is as follows:
[0085] Initialization stage: Initialize parameters, randomly generate a set of candidate solutions, each solution contains a scale parameter γ and an offset parameter β, and set the initial parameters of the algorithm, including population size and maximum number of iterations;
[0086] Evaluate the fitness of each candidate solution and find the current best solution;
[0087] Main loop phase: Starting from t = 1 until the maximum number of iterations T is reached; update CM, calculate the CM value of the current iteration, which is used to control the balance between exploration and exploitation, and CM is the conversion mechanism parameter;
[0088] If the current iteration number t is equal to CE i , then update the search space; where CE i is the i-th iteration count of the constrained exploration iteration;
[0089] Exploration phase, if CM > 1, then calculate α1 and α2, which are used to control the update step or search intensity of the solution in the control algorithm;
[0090] First stage of exploration:
[0091]
[0092] In the formula, rand() generates a random number in the interval [0, 1]; exp is the exponential function; d1, d2 are dynamically adjusted parameters; Max_Iter is the maximum number of iterations of the algorithm;
[0093] Update the individual position:
[0094]
[0095] In the formula, represents the value of the j-th parameter (dimension) of the i-th individual (solution) at the (t + 1)-th iteration; represents the value of the j-th parameter in the optimal solution found so far; q1 represents a number randomly generated in the interval [0, 1]; a1 is the exploration step size of the first stage;
[0096] Second stage of exploration:
[0097]
[0098] In the formula: tanh is the hyperbolic tangent function; Update the individual position:
[0099]
[0100] In the formula, a2 is the exploration step size of the second stage;
[0101] Exploitation phase, if CM ≤ 1, then calculate α3, which is used to control the local search of the solution:
[0102]
[0103] First stage of exploitation, update the individual position:
[0104]
[0105] In the formula, q4 and q3 represent a number randomly generated within the interval [0, 1]; α3 represents the step size in the development stage;
[0106] Second-stage development, update the individual position:
[0107]
[0108] Evaluate the fitness of the updated solution and update the global best solution (γ, β);
[0109] If the stopping condition is met, end the algorithm; otherwise, return to step 4.3 to continue the iteration;
[0110] Return the updated parameters (γ, β) to the Medformer model to form an optimized ETO-Medformer model for wind power prediction.
[0111] Furthermore, the implementation process of step 5 is as follows:
[0112] Divide the final dataset processed in the source domain into a training set, a test set, and a validation set; the divided training set and test set are used to train the wind power prediction model, and the divided validation set is used to validate the trained wind power prediction model and generate evaluation metrics;
[0113] Use the prediction data of the source region as the input of the trained wind power prediction model, and output the wind power prediction result of the source region through the trained wind power prediction model;
[0114] Migrate the trained wind power prediction model from the source domain to the target domain, obtain the real-time wind power data of the new wind farm area, and use the real-time wind power data as the target data of the transfer learning algorithm;
[0115] Combine most layers of ETO-Medformer, use the maximum reflection function to train some layers that need to be adjusted, gradually unfreeze some layers as needed, and fine-tune the entire model:
[0116]
[0117] In the formula: is the maximum mean difference between the source data and the target data; represents the model parameter matrix; μ, b are learnable parameters; T e is the constraint vector; and then generate an optimized wind power prediction model applicable to the new wind farm area;
[0118] Further train the ETO-Medformer model on the target domain, and use the trained model to predict the power of wind power generation to obtain the final prediction result.
[0119] An electronic device according to the present invention includes a memory and a processor, wherein:
[0120] The memory is used to store a computer program that can run on the processor;
[0121] The processor is used to execute the steps of the wind power prediction method based on multivariate DMD and transfer learning as described above when running the computer program.
[0122] A storage medium according to the present invention has a computer program stored thereon, and when the computer program is executed by at least one processor, it implements the steps of the wind power prediction method based on multivariate DMD and transfer learning as described above.
[0123] Advantageous effects: Compared with the prior art, the advantageous effects of the present invention are as follows:
[0124] The greedy feature selection method proposed by the present invention is simple, efficient and easy to implement; it optimizes the model performance step by step by selecting features step by step, while reducing the computational cost and the risk of overfitting; this method can evaluate the importance of features, reduce the model complexity and improve the interpretability, and is especially suitable for processing high-dimensional data sets;
[0125] The sparse-promoting dynamic mode decomposition (SPDMD) proposed by the present invention optimizes the mode extraction process of the traditional dynamic mode decomposition (DMD) by introducing sparsity constraints; SPDMD enhances the robustness of the algorithm to noise, helps reduce the risk of overfitting, and makes it more effective in analyzing complex dynamic systems;
[0126] The exponential triangle optimization (ETO) algorithm shows multiple advantages when updating the Medformer model; it achieves a faster convergence speed and higher model accuracy by optimizing the exponential form of the objective function; the ETO algorithm enhances the robustness of the model, improves the feature extraction process, is highly adaptable, and can effectively reduce the overfitting phenomenon; in addition, this algorithm improves the computational efficiency and is easy to implement and integrate;
[0127] Training the ETO-Medformer model using a parameter transfer learning strategy can make full use of the knowledge of the pre-trained model in the source domain, accelerate the convergence of the model in the target domain and improve the prediction accuracy; this method reduces the need for a large amount of labeled data in the target domain, enhances the generalization ability of the model, saves computational resources, and at the same time adapts to the data distribution differences between the source domain and the target domain, improving the robustness of the model; the transfer learning strategy provides flexibility for the model and helps the effective application of the model between different domains. Description of the Drawings
[0128] Figure 1Flow chart of a wind power prediction method based on multi-variable DMD and transfer learning. Detailed implementation manners
[0129] The present invention will be further described in detail below with reference to the accompanying drawings.
[0130] As Figure 1 shown, the present invention proposes a wind power prediction method based on multi-variable DMD and transfer learning, which specifically includes the following steps:
[0131] Step 1: Collect relevant data information in the source domain and the target domain, and use greedy feature selection (GFS) to reduce the dimension of the data and remove redundant features in the data.
[0132] Step 1.1: Set the initial feature set F (0) = φ, set the iteration number to k = 0, and set the total feature set F = {f1, f2, f3, f4, f5, f6, f7}, where the set corresponds to wind speed, wind direction, temperature, humidity, air pressure, pitch angle, and nacelle yaw angle respectively.
[0133] Step 1.2: For each step k, evaluate the remaining feature f j . The marginal contribution of f j is realized by calculating the difference between the performance index after including this feature and the performance index without including this feature:
[0134] Δμ(f j |F (k) ) = μ(F (k) ∪{f j ) - μ(F (k) )
[0135] In the formula: j is the index of the feature; F (k) is the feature set selected at the k-th iteration; μ is the performance index function.
[0136] Step 1.3: Among all the remaining features, select the feature that maximally improves the performance index
[0137]
[0138] In the formula: j * is the index of the maximum feature; argmax is the maximum value parameter; Δμ(·) is the change in the performance index.
[0139] Step 1.4: Update the feature set and add the selected feature to the feature set:
[0140]
[0141] In the formula: F(k+1) Denotes the feature set after the (k + 1)-th iteration.
[0142] Step 1.5: Update the iteration count, k = k + 1.
[0143] Step 1.6: Check if the stopping condition is met, reaching the maximum number of iterations K = 5 or the performance metric improvement is no longer significant. The stopping condition can be specified as:
[0144]
[0145] where τ is a preset threshold.
[0146] The iteration ends, and the final feature set F (s) = {g1, g2,..., g s}
[0147] The greedy feature selection method is simple, efficient, and easy to implement. It optimizes the model performance step by step by selecting features step by step, while reducing the computational cost and the risk of overfitting. This method can evaluate the importance of features, reduce the model complexity, and improve the interpretability, especially suitable for dealing with high-dimensional datasets.
[0148] Step 2: The process of processing using sparse-promoted dynamic mode decomposition (SPDMD) is as follows:
[0149] Step 2.1: From the data after feature selection, select one set of data to construct a set of snapshot sequences to form a data matrix: (x1 x2... x m+1 ). Then construct two snapshot matrices from the snapshot sequences:
[0150] X = [x1 x2... x m
[0151] X′ = [x2 x3... x m+1
[0152] And assume that the snapshots are generated by a discrete-time linear time-varying system, which is represented as follows:
[0153] x k+1 = Ax k , k = 1, 2,..., m
[0154] where A is a linear operator.
[0155] Connect the data matrices X and X' through the linear operator A, then the matrix X' has the following representation:
[0156] X′ = [x2 x3... x m+1=[Ax1 Ax2... Ax m =AX
[0157] Step 2.2: Perform singular value decomposition on the matrix X, select an appropriate truncation value r to reduce the dimension of the data matrix, and the eliminated redundant terms are represented by rem, and its dimension is m - r:
[0158]
[0159] where U is the left singular vector matrix; ∑ is the diagonal matrix; V * is the conjugate transpose of the right singular vector matrix; represents the submatrix composed of the first r left singular vectors selected from U; represents the diagonal submatrix composed of the first r singular values selected from ∑; represents the conjugate transpose of the submatrix composed of the first r right singular vectors selected from V * .
[0160] The linear operator A can be expressed as: It can be expressed as:
[0161] Step 2.3: Using the least squares method, by minimizing the Frobenius norm between the difference of X' and AX: where There exist linearly independent eigenvectors (w1 w2…w r ), and the corresponding eigenvalues are {λ1, λ2,…, λ r}), then it can be transformed into the focal coordinate form:
[0162]
[0163] In the formula: Let and {z1, z2,…, z r} is 's eigenvector, and the corresponding eigenvalue is Scale them appropriately to satisfy the following biorthogonality condition:
[0164]
[0165] Step 2.4: Use the linear combination of DMD modes to approximate the numerical snapshots:
[0166]
[0167] In the formula: x k is the mode corresponding to the decomposition; α i is the amplitude corresponding to the DMD mode. It can be expressed in matrix form as:
[0168]
[0169] The above equation shows that the dynamic modal evolution characteristics are controlled by the Vandermonde matrix, and the Vandermonde matrix is determined by the r complex eigenvalues of the matrix , and its eigenvalues contain information on the relevant time frequencies and growth / decay rates.
[0170] Step 2.5: Here, a sparsification method is introduced to adjust the amplitudes and eliminate these non-critical modes. First, construct the objective function with respect to α i :
[0171] J(α) = α * pα - q * α - α * q + trace(Σ * Σ)
[0172] where: p is the left complex vector; q is the right complex vector; trace(Σ * Σ) is the trace of the product of the conjugate transpose of Σ and Σ itself.
[0173] Step 2.6: Add a regularization term card(z) to the objective function J(α) to solve the sparsity problem, and finally obtain the sparsified amplitude vector D α :
[0174] D α = min α J(α) + card(z).
[0175] This penalizes the number of non-zero elements in the amplitude vector D α , thereby determining the finally optimized sparsified amplitude vector D α .
[0176] Step 2.7: Let Finally, there is an expression for the corresponding decomposed mode: X ≈ ΦD α V and .
[0177] Step 2.8: Set the stopping criterion for the final modal decomposition:
[0178] ||Ax k+1 + Bz k+1 - c||2 ≤ ε prim
[0179] ||z k+1 - z k+1 ||2 ≤ ε dual
[0180] where A, B, and C are constants; ∈ prim and ∈ dual represent the stopping threshold. Data iteration is performed until the number of modes is decomposed to the set threshold, and finally it is expressed as: According to the above method, finally construct the final dataset: X = {g1', g2',..., g s '}.
[0181] Sparse-promoted dynamic mode decomposition (SPDMD) optimizes the mode extraction process of traditional dynamic mode decomposition (DMD) by introducing sparsity constraints. SPDMD enhances the robustness of the algorithm to noise, improves feature selection, has strong adaptability, helps reduce the risk of overfitting, and makes it more effective in analyzing complex dynamic systems.
[0182] Step 3: Construct a wind power prediction model based on the Medformer model.
[0183] Step 3.1: Define the input data as multivariate time series samples where T represents the number of timestamps and C represents the number of channels. In the wind power prediction scenario, channel C includes: g1', g2',..., g s '.
[0184] Step 3.2: Introduce cross-channel multi-granularity patch embedding. First, perform patch segmentation. Given a list of different patch lengths {L1, L2,..., Ln}, for the i-th patch length Li, the input sample is segmented into Ni non-overlapping cross-channel patches By zero-padding, ensure that the timestamp T is divisible by Li, so that Ni = T / Li.
[0185] Then perform linear projection to map the patches to the latent embedding space, obtaining:
[0186]
[0187] In the formula: is the non-overlapping patch; W (i) is the embedding space; is the result after mapping.
[0188] Finally, add the fixed position embedding W pos and the learnable granularity embedding to obtain the final patch embedding:
[0189]
[0190] Step 3.3: Then introduce the multi-granularity self-attention mechanism and initialize the router. Initialize a router for each granularity:
[0191]
[0192] Where: u (i) is one of the routers.
[0193] Step 3.4: Inner-granularity self-attention adjustment. Vertically connect the patch embedding x (i) and the router embedding u (i) to form an intermediate sequence z (i) = [x (i) ‖ u (i) . Perform self-attention operation on z (i) to update the patch embedding and the router embedding:
[0194] x (i) ← Attn Intra (x (i) , z (i) , z (i) )
[0195] u (i) ← Attn Intra (u (i) , z (i) , z (i) )
[0196] Where: Attn Intra (·) is the self-attention mechanism calculation function.
[0197] Step 3.5: Finally, perform cross-granularity self-attention operation. Connect all the router embeddings into a sequence U = [u (1) ‖ u (2) ‖... ‖ u (n) . Perform self-attention operation on the router embedding u(i) of each granularity i to update the router embedding:
[0198] u (i) ← Attn Intra (u (i) , U, U)
[0199] Step 3.6: Perform layer normalization on the obtained patch embedding x (i) :
[0200]
[0201] Where: μ B and σ B are the batch mean and standard deviation respectively, and γ and β are learnable parameters.
[0202] Step 3.7: Apply a linear layer for prediction:
[0203]
[0204] In the formula, is the wind power prediction value; W P and b P are the weights and biases of the output layer; Var represents variance; E represents the expected value.
[0205] Medformer performs excellently in multi-task processing, robustness, flexibility, and model interpretability, can effectively process large-scale datasets, and supports cross-modal data analysis.
[0206] Step 4: Update the parameters of the Medformer model, the scale parameter γ and the offset parameter β, using the Exponential Triangular Optimization (ETO) algorithm. The specific process is as follows:
[0207] Step 4.1: Initialization phase. Initialize the parameters, randomly generate a set of candidate solutions, each solution containing the parameters γ and β, and set the initial parameters of the algorithm, including the population size and the maximum number of iterations.
[0208] Step 4.2: Evaluate the fitness of each candidate solution and find the current best solution.
[0209] Step 4.3: Main loop phase. Starting from t = 1 until the maximum number of iterations T is reached; update CM, calculate the CM value of the current iteration, which is used to control the balance between exploration and exploitation. CM is the transformation mechanism parameter.
[0210] Step 4.4: If the current iteration number t is equal to CE i , then update the search space. Where CE i is the i-th iteration count of the constrained exploration iteration.
[0211] Step 4.5: Exploration phase. If CM > 1, then calculate α1 and α2, which are used to control the update step or search intensity of the solutions in the control algorithm:
[0212] First-stage exploration:
[0213]
[0214] In the formula: rand() generates a random number in the interval [0, 1]; exp is the exponential function; d1, d2 are dynamically adjusted parameters; Max_Iter is the maximum number of iterations of the algorithm.
[0215] Update the individual position:
[0216]
[0217] In the formula: represents the value of the j-th parameter (dimension) of the i-th individual (solution) at the (t + 1)-th iteration; Represents the value of the j-th parameter in the optimal solution found so far; q1 represents a number randomly generated within the interval [0, 1]; a1 is the exploration step size in the first stage.
[0218] Second-stage exploration:
[0219]
[0220] In the formula: tanh is the hyperbolic tangent function.
[0221] Update the individual position:
[0222]
[0223] In the formula: a2 is the exploration step size in the second stage.
[0224] Step 4.6: In the exploitation stage, if CM ≤ 1, calculate α3 to control the local search of the solution.
[0225]
[0226] First-stage exploitation, update the individual position:
[0227]
[0228] In the formula: q4 and q3 represent numbers randomly generated within the interval [0, 1]; α3 represents the step size in the exploitation stage.
[0229] Second-stage exploitation, update the individual position:
[0230]
[0231] Step 4.7: Evaluate the fitness of the updated solution and update the global best solution (γ, β).
[0232] Step 4.8: If the stopping condition is met (such as reaching the maximum number of iterations or finding a satisfactory solution), end the algorithm; otherwise, return to Step 4.3 to continue the iteration.
[0233] Finally, return the updated parameters (γ, β) to the Medformer model to form the optimized ETO-Medformer model for subsequent wind power prediction.
[0234] The ETO algorithm exhibits multiple advantages when updating the Medformer model. It achieves a faster convergence speed and higher model accuracy by optimizing the exponential form of the objective function. The ETO algorithm enhances the robustness and adaptability of the model, effectively reducing the overfitting phenomenon. In addition, the algorithm improves the computational efficiency and is easy to implement and integrate.
[0235] Finally, the updated parameters (γ, β) are returned to the Medformer model to form the optimized ETO-Medformer model for subsequent wind power prediction.
[0236] Step 5: Based on the parameter transfer learning strategy, train the ETO-Medformer on the source domain. Output the wind power prediction results of the source domain through the trained model, and use its parameters as the initial parameters of the same model on the target domain. Finally, further train the ETO-Medformer on the target domain, and use the trained model to predict the power of wind power generation to obtain the final prediction result.
[0237] Step 5.1: Divide the final dataset processed in the source domain into a training set, a test set, and a validation set according to the ratio of 8:1:1. The divided training set and test set are used to train the wind power prediction model, and the divided validation set is used to validate the trained wind power prediction model and generate evaluation indicators.
[0238] Step 5.2: Use the prediction data of the source region as the input of the trained wind power prediction model, and output the wind power prediction results of the source region through the trained wind power prediction model.
[0239] Step 5.3: Transfer the trained wind power prediction model from the source domain to the target domain, obtain the real-time wind power data of the new wind farm area, and use the real-time wind power data as the target data of the transfer learning algorithm.
[0240] Step 5.4: Freeze most layers of the ETO-Medformer because it performs well in the source domain and keep their weights unchanged. Use the maximum reflection function to train some layers that need to be adjusted, and gradually unfreeze some layers as needed to fine-tune the entire model:
[0241]
[0242] In the formula: is the maximum mean difference between the source data and the target data; represents the model parameter matrix; μ, b are learnable parameters; T e is the constraint vector. Then generate an optimized wind power prediction model suitable for the new wind farm area.
[0243] Step 5.5: Finally, further train the ETO-Medformer model on the target domain, and use the trained model to predict the power of wind power generation to obtain the final prediction result.
[0244] Training the ETO-Medformer model using a parameter-based transfer learning strategy can fully utilize the knowledge of the pre-trained model in the source domain, accelerate the convergence of the model in the target domain, and improve the prediction accuracy. This method reduces the need for a large amount of labeled data in the target domain, enhances the generalization ability of the model, saves computing resources, adapts to the data distribution differences between the source domain and the target domain, and improves the robustness of the model. The transfer learning strategy provides flexibility for the model and helps the model to be effectively applied across different domains.
[0245] The present invention also provides an electronic device, including a memory and a processor, wherein: the memory is used for storing a computer program that can run on the processor; the processor is used for executing the steps of the wind power prediction method based on multivariate DMD and transfer learning as described above when running the computer program.
[0246] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by at least one processor, the steps of the wind power prediction method based on multivariate DMD and transfer learning as described above are implemented.
[0247] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A wind power prediction method based on multivariate DMD and transfer learning, characterized in that: The following steps are involved: Step 1: Collect relevant data information in the source domain and the target domain, and use greedy feature selection GFS to reduce the dimension of the data and remove redundant features in the data; Step 2: Process the data after feature selection based on the sparse promoted dynamic mode decomposition algorithm SPDMD to reduce the complexity of the data; reconstruct the data to form an input data set; Step 3: Construct a wind power prediction model based on Medformer; Step 4: Use the exponential triangular optimization algorithm ETO to update the parameters of the Medformer model, obtain the optimized model ETO-Medformer, and obtain the optimal parameter set; Step 5: Based on the parameter transfer learning strategy, ETO-Medformer is trained from the source domain. The source domain wind power prediction results are output through the trained model, and its parameters are used as the initial parameters of the same model in the target domain. Finally, ETO-Medformer is further trained in the target domain, and the trained model is used to predict the wind power generation to obtain the final prediction result.
2. A wind power prediction method based on multivariate DMD and transfer learning according to claim 1, characterized in that: The relevant data information in step 1 includes wind speed, wind direction, temperature, humidity, air pressure, pitch angle, and cabin yaw angle.
3. The wind power prediction method based on multivariate DMD and transfer learning according to claim 1, characterized in that: The implementation process of using greedy feature selection GFS to reduce the dimension of data in step 1 is as follows: Set the initial feature set F (0) =φ, set the number of iterations to k = 0, and set the total feature set F = {f1,f2,f3,f4,f5,f6,f7}; For each step k, evaluate the remaining features f j ; f is achieved by calculating the difference between the performance index after including the feature and the performance index without the feature j The marginal contribution is: Dm(f j |F (k) )=μ(F (k) ∪{f j })-μ(F (k) ) Where j is the index of the feature; F (k) is the feature set selected at the kth iteration; μ is the performance index function; Among all remaining features, select the feature that improves the performance index the most In the formula, j * is the index of the maximum feature; argmax is the maximum parameter; Δμ(·) is the change in performance index; Update a feature set to add the selected features to the feature set: In the formula, F (k+1) represents the feature set after the (k+1)th iteration; Update the iteration count, k=k+1; check whether the stopping condition is met, and stop when the maximum number of iterations K is reached. The stopping condition can be specified as: In the formula, τ is the preset threshold; after the iteration, the final feature set F is obtained (s) ={g1,g2,...,g s }.
4. The wind power prediction method based on multivariate DMD and transfer learning according to claim 1, characterized in that: The implementation process of step 2 is as follows: From the data after feature selection, select a set of data to build a set of snapshot sequences to form a data matrix: (x1x2...x m+1 ); then construct two snapshot matrices from the snapshot sequence: X=[x1 x2...x m ] X′=[x2 x3...x m+1 ] And assume that the snapshot is generated by a discrete-time linear time-varying algorithm, expressed as follows: x k+1 =Ax k ,k=1,2,...,m Where A is a linear operator; Connect the data matrices X and X' through the linear operator A, and the matrix X' is represented as follows: X′=[x2 x3...x m+1 ]=[Ax1 Ax2...Ax m ]=AX Perform singular value decomposition on the matrix X and select an appropriate cutoff value r to reduce the dimension of the data matrix. The eliminated redundant items are represented by rem, and their dimension is m-r: Where U is the left singular vector matrix; ∑ is the diagonal matrix; V * is the conjugate transpose of the right singular vector matrix; represents the submatrix consisting of the first r left singular vectors selected from U; represents the diagonal submatrix consisting of the first r singular values selected from ∑; Indicates that from V * The conjugate transpose of the submatrix consisting of the first r right singular vectors selected from; the linear operator A is expressed as: It is expressed as: Using the least squares method, minimize the Frobenius norm between the difference between X' and AX: in There are linearly independent eigenvectors (w1w2…w r ), the corresponding eigenvalues are {λ1,λ2,…,λ r }, then convert it into focus coordinate form: Where: Let W = [w1 w2 … w r ]; And {z1,z2,…,z r }yes The eigenvector of They are scaled appropriately to satisfy the following biorthogonal conditions: The numerical snapshot is approximated using a linear combination of the DMD modes: Where: x k is the corresponding decomposed mode; α i is the amplitude of the corresponding DMD mode; it is expressed in matrix form as: The above formula shows that the dynamic modal evolution characteristics are controlled by the Vandermonde matrix, and the Vandermonde matrix is composed of the matrix The r complex eigenvalues of , whose eigenvalues contain information about the relevant time frequency and growth decay rate; Introduce the sparse method to adjust the amplitude and eliminate non-critical modes; first construct i The objective function of J(a)=a * pα-q * a-a * q+trace(Σ * S) Where p is a complex vector; q is a right complex vector; trace(Σ * Σ) is the trace of the product of the conjugate transpose of Σ and Σ itself; Add a regularization term card(z) to the objective function J(α) to solve the sparsity problem, and finally get the sparse amplitude vector D α : D α =min α J(α)+card(z) This penalizes the amplitude vector D α The number of non-zero elements in the final optimized sparse amplitude vector D α ; make Finally, there is an expression for the corresponding decomposition mode: X≈ΦD α V and ; Set the stopping criterion for the final mode decomposition: In the formula, A, B, C are constants; ∈ prim and ∈ dual Represents the stopping threshold; iterate the data until the number of modes is decomposed to the set threshold, and the final expression is: According to the above method, the final data set is constructed: X = {g1', g2', ..., g s '}.
5. The wind power prediction method based on multivariate DMD and transfer learning according to claim 1, characterized in that: The implementation process of step 3 is as follows: Define the input data as a multivariate time series sample T represents the number of timestamps, and C represents the number of channels. In the wind power prediction scenario, channel C contains: g1', g2', ..., g s '; Introducing cross-channel multi-granularity patch embedding: First, patch segmentation is performed, given a list of different patch lengths {L1,L2,…,L n }, for the i-th patch length L i , split the input sample into N i non-overlapping patches of intersecting channels Zero padding ensures that the timestamp T can be i divisible by N i =T / L i ; Then a linear projection is performed to map the patch into the latent embedding space, resulting in: Where: are non-overlapping patches; W (i) For embedded space; is the result after mapping; Finally add the fixed position embed W pos and a learnable granular embedding W g (i) , and get the final patch embedding: Then introduce the multi-granularity self-attention mechanism, initialize the router, and initialize a router for each granularity: u (i) =W pos [N i +1]+W g (i) Where: u (i) For one of the routers; Inner-granularity self-attention adjustment, embedding the patch into x (i) and router embedded in u (i) Connect vertically to form the intermediate sequence z (i) =[x (i) ‖u (i) ]; for z (i) Perform self-attention operations and update patch embeddings and router embeddings: x (i) ←Attn Intra (x (i) ,z (i) ,z (i) ) at (i) ←Attn Intra (at (i) ,With (i) ,With (i) ) In the formula, Attn Intra (·) is the calculation function of the self-attention mechanism; Perform cross-granular self-attention operations and connect all router embeddings into a sequence U = [u (1) ‖u (2) ‖...‖u (n) ]; For each router of granularity i, embed u (i) Perform self-attention operation and update router embedding: in (i) ←Attn Intra (in (i) (U,U) Embed the obtained patch into x (i) Perform layer normalization: Where: μ B and σ B are the batch mean and standard deviation, respectively, γ and β are learnable parameters; Apply a linear layer to make predictions: In the formula, is the predicted wind power value; W P and b P is the output layer weight and bias; Var represents the variance; E stands for expected value.
6. A wind power prediction method based on multivariate DMD and transfer learning according to claim 1, characterized in that: The implementation process of step 4 is as follows: Initialization phase: Initialize parameters, randomly generate a set of candidate solutions, each solution contains a scale parameter γ and an offset parameter β, and set the initial parameters of the algorithm, including population size and maximum number of iterations; Evaluate the fitness of each candidate solution and find the current best solution; Main loop phase: starting from t=1 until the maximum number of iterations T is reached; Update CM and calculate the CM value of the current iteration to control the balance between exploration and development. CM is the conversion mechanism parameter. If the current iteration number t is equal to CE i , then update the search space; where CE i The i-th iteration count for the constraint exploration iteration; In the exploration phase, if CM>1, α1 and α2 are calculated to control the update step size or search intensity of the solution in the control algorithm; Phase I Exploration: In the formula, rand() generates a random number in the interval [0,1]; exp is the exponential function; d1 and d2 are dynamic adjustment parameters; Max_Iter is the maximum number of iterations of the algorithm; Update individual location: In the formula, Indicates the value of the jth parameter (dimension) of the i-th individual (solution) at the t+1 iteration; represents the value of the jth parameter in the optimal solution found so far; q1 represents a randomly generated number in the interval [0,1]; a1 is the exploration step length of the first stage; Second phase of exploration: Where: tanh is the hyperbolic tangent function; Update individual position: Where a2 is the exploration step length of the second stage; In the development phase, if CM≤1, α3 is calculated to control the local search for solutions: First phase of development, updating individual locations: In the formula, q4 and q3 represent a randomly generated number in the interval [0,1]; α3 represents the step length of the development phase; The second phase of development, updating individual locations: Evaluate the fitness of the updated solution and update the global optimal solution (γ, β); If the stopping condition is met, the algorithm ends; otherwise, return to step 4.3 to continue iterating; The updated parameters (γ, β) are returned to the Medformer model to form the optimized ETO-Medformer model for wind power prediction.
7. The wind power prediction method based on multivariate DMD and transfer learning according to claim 1, characterized in that: The implementation process of step 5 is as follows: The final processed data set in the source domain is divided into a training set, a test set and a validation set; the divided training set and the test set are used to train the wind power prediction model, and the divided validation set is used to verify the trained wind power prediction model and generate evaluation indicators; The prediction data of the source area is used as the input of the trained wind power prediction model, and the wind power prediction result of the source area is outputted through the trained wind power prediction model; Migrate the trained wind power prediction model from the source domain to the target domain, obtain the real-time wind power data of the new wind farm area, and use the real-time wind power data as the target data of the transfer learning algorithm; Combine most of the layers of ETO-Medformer, use the maximum reflection function to train some layers that need to be adjusted, gradually unfreeze some layers as needed, and fine-tune the entire model: Where: is the maximum mean difference between the source data and the target data; represents the model parameter matrix; μ and b are learnable parameters; T e is a constraint vector; and then generates an optimized wind power prediction model suitable for the new wind farm area; The ETO-Medformer model is further trained on the target domain, and the trained model is used to predict the power of wind power generation to obtain the final prediction result.
8. An electronic device, characterized in that: comprising a memory and a processor, wherein: A memory for storing computer programs that can be run on the processor; A processor, configured to execute the steps of the wind power prediction method based on multivariate DMD and transfer learning as described in any one of claims 1 to 7 when running the computer program.
9. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by at least one processor, the steps of the wind power prediction method based on multivariate DMD and transfer learning as described in any one of claims 1 to 7 are implemented.