A method for clustering wind-light-load joint scene using deep convolutional embedding with multi-head self-attention
Through the deep convolutional embedding clustering method of multi-head self-attention, combined with multi-strategy fusion to improve the slime mold algorithm and self-attention convolution autoencoder, the problem of difficult to capture the coupled feature information of the wind and light charge data is solved, and high-precision wind and light charge joint scene generation is achieved, supporting the optimized operation and planning of the power grid.
Patent Information
- Application Number
- CN202211176681.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-26
AI Technical Summary
When generating wind and light-collective coupled scenarios, the coupling feature information between wind and light-collective data cannot be accurately captured, and the traditional clustering method has decreased accuracy under high-dimensional data, resulting in an increase in uncertainty in power grid scheduling and planning.
The deep convolutional embedding clustering method of multi-head self-attention is adopted, and the VMD model parameters are optimized through multi-strategy fusion improvement algorithm, combined with the multi-head self-attention convolutional autoencoder and Kmeans clustering, a combined scenery and load scene is generated to accurately capture coupled feature information.
It improves the accuracy of extracting feature information of wind, light and charge data coupling, ensures the representativeness of embedded space features, and supports the optimized operation and planning of the power system.
Smart Images

Figure CN115496153B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grids, in particular to a method for jointly generating wind, light and load scenarios by deep convolutional embedding clustering with multi-head self-attention. Background Art
[0002] With the construction of a high-proportion renewable energy power system, the volatility and periodicity of wind power, photovoltaic power and load have brought challenges to power grid planning and dispatching operations.
[0003] If based on the scenario method, converting the uncertain scenarios of wind, light and load into multiple deterministic scenarios can lay a good foundation for optimizing power grid dispatching and planning.
[0004] Wind power, photovoltaic power output and electrical load show certain seasonal or daily periodicity with the change of time. At present, most scenario generation methods cannot fully explore the information value of power data, and there are also limitations in capturing the complementary relationship between wind and photovoltaic power generation and the energy coupling relationship between wind, light and load.
[0005] Currently, the methods for generating wind-light-load coupling scenarios mainly use clustering to extract the potential feature information between time series data and classify them. Traditional clustering models such as k-means clustering, spectral clustering, hierarchical clustering and Gaussian mixture clustering have been applied to optimize power grid dispatching and planning. However, traditional clustering cannot accurately extract the potential coupling features between time series data, and when facing large-scale high-dimensional data, the calculation accuracy of the clustering model decreases. To improve the clustering accuracy for high-dimensional data, generally, methods such as PCA and singular value decomposition are first used to reduce the data dimension and extract features, and then clustering is performed based on the feature information in the low-dimensional space. Although this method can effectively improve the clustering accuracy, the data dimension reduction process and the clustering process are independent of each other, usually resulting in the captured feature information not being able to truly reflect the inter-class structure of the data. In addition, some joint scenario generation methods based on deep embedding clustering do not consider the later clustering training, resulting in the distortion phenomenon of the low-dimensional embedding space, weakening the potential feature information between data and affecting the clustering accuracy. In addition, there are also defects in capturing feature information more one-sidedly. Summary of the Invention
[0006] The present invention proposes a method for jointly generating wind, light and load scenarios by deep convolutional embedding clustering with multi-head self-attention, which can accurately capture the coupling feature information between wind, light and load data, combines the feature extraction process and the clustering process, ensures the representativeness of the features in the embedding space, and establishes a deep convolutional embedding clustering with multi-head self-attention (DCEC-MS) model, which can generate wind-light-load joint scenarios and accurately capture the coupling feature information between wind, light and load data.
[0007] The present invention adopts the following technical solutions.
[0008] A multi-head self-attention based deep convolutional embedding clustering method for wind-solar-load combined scenarios, which is used to generate wind-solar-load coupling scenarios. By accurately capturing the coupling feature information between wind-solar-load data, the method combines the feature extraction process with the clustering process to ensure the representativeness of the features in the embedding space, and includes the following steps;
[0009] Step 1: Optimize the parameter combination of the VMD model with a multi-strategy fusion improved slime mould algorithm (SMA), and based on the optimal parameter combination, clean the wind-solar-load time series data to weaken the influence of noise signals on the data feature extraction process;
[0010] Step 2: Establish a convolutional autoencoder based on multi-head self-attention to extract the deep feature information of the processed wind-solar-load data. At the same time, use a convolutional decoder to reconstruct the original time series signal;
[0011] Step 3: Obtain the appropriate number of clustering clusters based on the elbow method, and use the Kmeans initial clustering center for the features; then, based on the joint loss function composed of the sum of the encoder reconstruction loss and the clustering loss, adjust the network structure parameters and update the clustering results, and obtain the clustering center of each type of scenario based on the mean method as the typical scenario of this type, providing a basis for the optimal operation and planning of the power system.
[0012] In step 1, abnormal data detection and cleaning are performed on the historical wind-solar-load data. The data is the wind-solar-load time series data f(t) within one year, with days as the unit, and each sample contains the wind power, photovoltaic output, and load data at 24 moments. The method is specifically as follows:
[0013] Step S1: Adopt a non-linear time-domain decomposition method - the VMD model to decompose the original wind-solar-load data f(t) into K intrinsic mode function components u k (t) with central frequencies, and at the same time calculate the sum of the finite bandwidths of the K u k (t) and make it minimum, so as to obtain the VMD model expression as:
[0014]
[0015] In the formula: ω k is the central frequency of the k-th u k (t); δ(t) is the unit impulse function; is the partial derivative operator; introduce the Lagrange operator λ and the quadratic term penalty factor α to streamline the solution process of formula (1), and the model expression after operation is:
[0016]
[0017] Based on the alternating direction multiplier method, solve Equation (5.2) and continuously optimize and iterate and λ, and the iteration expression is as follows:
[0018]
[0019]
[0020]
[0021] In the formula: n is the number of iterations; and are the Fourier transforms of u(t), λ(t) and f(t) respectively;
[0022] In step S2, when determining the preset values of parameters K and α in the VMD model, if the value of parameter K is too large, over-decomposition will occur, resulting in mode overlap; if the value of parameter α is too large, the loss of the central frequency will be caused. In this step, a multi-strategy fusion improved slime mold algorithm based on KL divergence is used to search for the optimal parameter combination of VMD. KL divergence is used to measure the similarity between the intrinsic mode component u k (t) and the original wind-solar-load data f(t), and the mathematical expression is:
[0023]
[0024] In the formula: the closer R is to 0, the higher the similarity between the intrinsic mode component u k (t) and the original wind-solar-load data f(t); as the fitness value of the objective function, when R is the smallest, it indicates that the parameters K and α are the optimal parameter combination at this time;
[0025] In step S3, the slime mold algorithm simulates the slime mold's foraging behavior by establishing a model, that is, when the slime mold initially approaches the food, it chooses whether to approach the food according to the food concentration, and the mathematical expression for position update is:
[0026]
[0027] In the formula: X(T) is the position of the slime mold in the T-th iteration; X b (T) is the best position found currently; W is the weight coefficient; X A (T) and X B (T) are the randomly selected slime mold individuals in the T-th iteration respectively; r is a random number in [0, 1]; v c is the feedback factor, and its value linearly decreases from 1 to 0; v b is the control parameter, and its value range is [-a, a]; p is the position update control parameter;
[0028] Among them, the mathematical models of p, a, and W are as follows:
[0029] p = tanh|S(i) - DF| Formula VIII;
[0030]
[0031]
[0032] Where: S(i) is the fitness value of the i-th slime mold individual; DF is the optimal fitness value in all iterations; T max is the maximum number of iterations; bF and wF are the optimal fitness value and the worst fitness value in the T-th iteration respectively; condition is the individuals with fitness values ranked in the first half;
[0033] In the slime mold algorithm, for the purpose of searching for high-quality food, the slime mold will split out some individuals to explore the remaining area, and the mathematical formula for position update is:
[0034]
[0035] Where: rand is a random number; ub and lb are the upper and lower bounds of the exploration area of the slime mold respectively; z is the proportion parameter of the slime mold individuals exploring the remaining area;
[0036] Step S4. To prevent the slime mold algorithm from being unable to accurately and timely feedback the concentration and quality of food due to the linear decrease of v c and resulting in a slow convergence speed in the early stage, an adaptive and adjustable v c is introduced to accelerate the decrease speed of v c in the early stage and improve the ability to search for the global optimal value; and v c is relatively stable in the later stage of iterative calculation to avoid falling into local optima; the mathematical expression of the adaptive and adjustable v c is:
[0037]
[0038] The adaptive opposition-based learning mechanism introduces a vector in the exploration area of the slime mold. This vector is opposite to the position of each slime mold individual and compares their fitness values to avoid falling into local optima; the expression of the position
[0039]
[0040]
[0041] Based on adaptive decision-making, when the i-th slime mold explores food, it will use the current fitness value Compare with the previous optimal fitness value S(X i (T)), determine whether to use the reverse learning mechanism for additional exploration, and update the position for the next iteration. The expression is as follows:
[0042]
[0043] Step S5: Select the mean absolute error (MAE) as the standard to measure the quality of data processing. The calculation method is as follows:
[0044]
[0045] In the formula: N is the total number of wind-solar-load data samples; is the wind-solar-load data after data processing;
[0046] MAE measures the ability of the data processing algorithm to retain the original features of the data on the premise of effectively removing the outliers in the wind-solar-load data; the smaller the MAE, the better the data processing effect.
[0047] In step two, each processed data sample is expanded into a tensor of size (9, 9) through normalization and zero-padding operations. Based on the multi-head self-attention improved convolutional encoder, the wind-solar-load data is dimensionally reduced to extract deep features. At the same time, the original time series signal is reconstructed using a convolutional decoder;
[0048] In the model of the convolutional autoencoder, the encoding process of its convolutional layer is as follows:
[0049] h conv = σ(X conv * ω conv + b conv ) Formula XVII;
[0050] In the formula: h conv is the output feature information of the convolutional layer; o is the Relu activation function; X conv is the wind-solar-load time series data after data processing; ω conv and b conv are the number of convolutional kernels and the bias of the convolutional layer respectively;
[0051] The feature information captured by the convolutional layer from the wind-solar-load time series data is sent to the multi-head self-attention layer for further feature extraction. Its encoding process is as described in Formula XVII;
[0052] The decoding process of the decoder for decoding the encoding result of Formula XVII is expressed by the formula as
[0053] h deconv = σ(h conv * ω deconv + b deconv) Formula XVIII;
[0054] Where: h deconv is the output feature information of the deconvolution layer; ω deconv and b deconv are the number of convolution kernels and the bias of the deconvolution layer, respectively;
[0055] The model of the convolutional autoencoder uses the mean square error (MSE) as the reconstruction loss function and minimizes it, continuously optimizing the network parameters of the encoder and decoder. Its expression is:
[0056]
[0057] Where: L r is the reconstruction loss function; N d is the number of days of wind, light, and load data;
[0058] In the model of the multi-head self-attention improved convolutional encoder, its multi-head self-attention is an improved type based on self-attention. It uses the method of multiple queries to parallelly capture multiple groups of feature information in different subspaces from the data and splices them by weight. The detailed calculation process is as follows:
[0059] Step A1: Based on linear transformation, the input data Y of the multi-head self-attention layer is transformed into the query matrix Q a , the key matrix K a and the value matrix V a , and the mathematical expression is:
[0060]
[0061] Where: W Q , W K and W V are transformation matrices;
[0062] Map Q a , K a and V a to θ feature subspaces to obtain the query matrix Q a6 , the key matrix K aθ and the value matrix V aθ in the θ-th subspace. The calculation method is:
[0063]
[0064] Where: W Qθ , W Kθ and W Vθ are the transformation matrices of the θ-th subspace;
[0065] Step A2: Calculate the self-attention values in θ feature subspaces based on scaled dot product and Softmax function. The calculation method is as follows:
[0066]
[0067] In the formula: d is the scaling factor; head θ is the self-attention value in the θ-th feature subspace;
[0068] Finally, fuse the self-attention in θ feature subspaces:
[0069] M multi-head = C Concat (head1, head2,..., head θ )W o Formula 23;
[0070] In the formula: M multi-head is the fused self-attention value; C Concat is the matrix concatenation operation; W o is the parameter matrix.
[0071] Step 3 includes the following steps
[0072] Step B1: Based on the low-dimensional feature information obtained from the autoencoder, use the elbow method, i.e., the elbow method, to observe the elbow values for different numbers of cluster classes, so as to determine the optimal number of cluster classes. Initialize the Kmeans initial clustering centers, and based on the joint loss function, use the Adam optimizer to fine-tune the entire network parameters and ensure the representativeness of the embedded space features, thereby obtaining the optimal clustering result;
[0073] The joint loss function is expressed by the formula as
[0074] L = L r + γL c Formula 24;
[0075] In the formula: L is the joint loss function; γ is the coefficient for controlling the distortion degree of the embedding space, taking 0.1; L c is the clustering loss function;
[0076] Use the clustering center μ j as the connection weight between it and the low-dimensional space feature Z i , and map each low-dimensional space feature Z i to the soft label; at the same time, to improve the adaptation accuracy between Z i and μ i , use the Gaussian distribution as the ideal target distribution. In this step, the clustering loss describes the KL divergence between the distribution on the soft label and the Gaussian distribution, which is used to measure the similarity between the two; the specific process is expressed by the formula as follows;
[0077]
[0078]
[0079]
[0080] Where: q ij is the probability of the low-dimensional space feature Z i belonging to the clustering center μ j ; b ij is the target distribution auxiliary function;
[0081] Step B2: Use the KL divergence as the auxiliary clustering objective function to optimize the stack encoder, and use the Adam optimizer to adjust the parameters to obtain an encoder structure suitable for scene clustering; Adam integrates the advantages of the first-order momentum of stochastic gradient descent (SGD) and the second-order momentum of root mean square prop (RMSprop); Based on the moment means of the two, while maintaining the learning rate of each parameter, it gives full play to the performance of sparse gradients, making the algorithm have high robustness in non-steady-state problems. The specific calculation process is as follows:
[0082]
[0083] M tt = β1·M tt-1 +(1 - β1)·g tt Formula 29;
[0084]
[0085]
[0086]
[0087]
[0088] Where: t is the time interval; g tt is the gradient; M tt is the first-order moment estimate of g tt ; is the model parameter; η tt is the second-order moment estimate of g tt ; and are their corresponding network outputs respectively; ψ is the network step value; β1 is the exponential decay rate of M tt , taking 0.9; β2 is the exponential decay rate of η ttThe exponential decay rate is taken as 0.999; ε is a constant used to ensure the robustness of the algorithm;
[0089] Step B3: Through the joint training of the encoder and the clustering layer, obtain the optimal annual scenario clustering result, evaluate the clustering result using three evaluation indicators, CHI, SC, and DBI, and use the mean method to obtain the clustering center of each type of scenario to obtain the typical scenario of each type.
[0090] In the above content, DCEC is the deep convolutional embedded clustering algorithm DCEC; VMD is the abbreviation of Variational mode decomposition for parameter optimization.
[0091] The present invention can accurately capture the coupled feature information among wind, light, and load data, combine the feature extraction process with the clustering process, and ensure the representativeness of the features in the embedding space. Brief Description of the Drawings
[0092] The following further details the present invention in conjunction with the drawings and specific embodiments:
[0093] Att Figure 1 is a schematic flow chart of the method of the present invention;
[0094] Att Figure 2 is a schematic algorithm flow chart of the present invention;
[0095] Att Figure 3 is a schematic diagram of the network structure of the deep convolutional embedded clustering model DCEC-MS improved based on multi-head self-attention. Specific Embodiments
[0096] As shown in the figure, the multi-head self-attention based deep convolutional embedded clustering method for wind, light, and load joint scenarios is used to generate wind, light, and load coupled scenarios. The method accurately captures the coupled feature information among wind, light, and load data, combines the feature extraction process with the clustering process to ensure the representativeness of the features in the embedding space, and includes the following steps;
[0097] Step 1: Optimize the parameter combination of the VMD model with a multi-strategy fusion improved slime mould algorithm (SMA), and based on the optimal parameter combination, clean the wind, light, and load time series data to weaken the influence of noise signals on the data feature extraction process;
[0098] Step 2: Establish a convolutional autoencoder based on multi-head self-attention to extract the deep feature information of the processed wind, light, and load data. At the same time, use the convolutional decoder to reconstruct the original time series signal;
[0099] Step 3: Determine the appropriate number of clustering clusters based on the elbow method, and use the Kmeans initial clustering centers for the features; then, based on the joint loss function composed of the sum of the encoder reconstruction loss and the clustering loss, adjust the network structure parameters and update the clustering results, and obtain the clustering centers of various scenarios based on the mean method as the typical scenarios of each class, providing a basis for the optimal operation and planning of the power system.
[0100] In Step 1, perform abnormal data detection and cleaning on the historical wind, light, and load data. The data is the time-series data f(t) of wind, light, and load within one year, with a daily unit, and each sample contains the wind power, photovoltaic output, and load data at 24 moments; the specific method is as follows:
[0101] Step S1: Adopt the non-linear time-domain decomposition method - VMD model to decompose the original wind, light, and load data f(t) into K intrinsic mode function components u k (t) with central frequencies, and at the same time obtain the sum of the finite bandwidths of the K u k (t) and make it minimum, so as to obtain the VMD model expression as:
[0102]
[0103] In the formula: ω k is the central frequency of the k-th u k (t); δ(t) is the unit impulse function; is the partial derivative operator;
[0104] Introduce the Lagrange operator λ and the quadratic term penalty factor α to streamline the solution process of formula (1), and the model expression after operation is:
[0105]
[0106] Based on the alternating direction multiplier method, solve formula (5.2), and continuously optimize and iterate and λ, and the iteration expression is:
[0107]
[0108]
[0109]
[0110] In the formula: n is the number of iterations; and are the Fourier transforms of u(t), λ(t) and f(t) respectively;
[0111] Step S2. When the VMD model determines the preset values of parameters K and α, if the value of parameter K is too large, over-decomposition occurs, resulting in modal overlap; if the value of parameter α is too large, the loss of the center frequency occurs. In this step, a multi-strategy fusion improved slime mold algorithm based on KL divergence is used to search for the optimal parameter combination of VMD. KL divergence is used to measure the similarity between the intrinsic mode component u k (t) and the original wind-solar-load data f(t). The mathematical expression is:
[0112]
[0113] In the formula: The closer R is to 0, the higher the similarity between the intrinsic mode component u k (t) and the original wind-solar-load data f(t); as the fitness value of the objective function, when R is the smallest, it indicates that the parameters K and α at this time are the optimal parameter combination;
[0114] Step S3. The slime mold algorithm simulates the foraging behavior of slime mold by establishing a model, that is, when the slime mold initially approaches food, it decides whether to approach the food based on the food concentration. The mathematical expression for position update is:
[0115]
[0116] In the formula: X(T) is the position of the slime mold in the T-th iteration; X b (T) is the best position found currently; W is the weight coefficient; X A (T) and X B (T) are randomly selected slime mold individuals in the T-th iteration respectively; r is a random number in [0, 1]; v c is the feedback factor, and its value linearly decreases from 1 to 0; v b is the control parameter, and its value range is [-a, a]; p is the position update control parameter;
[0117] Among them: The mathematical models of p, a, and W are:
[0118] p = tanh|S(i) - DF| Formula VIII;
[0119]
[0120]
[0121] In the formula: S(i) is the fitness value of the i-th slime mold individual; DF is the optimal fitness value in all iterations; T max is the maximum number of iterations; bF and wF are the optimal fitness value and the worst fitness value in the T-th iteration respectively; condition is the individuals with the fitness value ranked in the first half;
[0122] In the slime mold algorithm, for the purpose of searching for high-quality food, the slime mold will divide some individuals to explore the remaining area, and the mathematical formula for position update is:
[0123]
[0124] In the formula: rand is a random number; ub and lb are the upper and lower bounds of the exploration area of the slime mold; z is the proportion parameter of the slime mold individuals exploring the remaining area;
[0125] Step S4. To prevent the slime mold algorithm from being unable to accurately and timely feedback the concentration and quality of food due to the linear decreasing manner of v c and resulting in a slow convergence speed in the early stage, an adaptive and adjustable v c is introduced to accelerate the descending speed of v c in the early stage and improve the ability to search for the global optimal value; and v c in the later stage of iterative calculation is relatively stable to avoid falling into local optima; the mathematical expression of the adaptive and adjustable v c is:
[0126]
[0127] The adaptive opposition-based learning mechanism introduces a vector in the exploration area of the slime mold. This vector is opposite to the position of each slime mold individual, and their fitness values are compared to avoid falling into local optima; the expression of the position of the i-th slime mold individual at the T-th iteration is:
[0128]
[0129]
[0130] Based on the adaptive decision-making, when the i-th slime mold explores food, it compares the current fitness value with the previous optimal fitness value S(X i (T)) to determine whether to use the opposition-based learning mechanism for additional exploration and update the position for the next iteration. The expression is:
[0131]
[0132] Step S5. Select the mean absolute error (MAE) as the standard to measure the quality of data processing. The calculation method is:
[0133]
[0134] In the formula: N is the total number of samples of wind-solar-load data; is the wind-solar-load data after data processing;
[0135] The MAE measures the ability of the data processing algorithm to retain the original features of the data on the premise of effectively removing outliers in the wind, light, and load data; the smaller the MAE, the better the data processing effect.
[0136] In step two, each processed data sample is expanded into a tensor of size (9, 9) through normalization and zero-padding operations. Based on the multi-head self-attention improved convolutional encoder, the wind, light, and load data is dimensionally reduced to extract deep features. At the same time, the original time series signal is reconstructed using a convolutional decoder;
[0137] In the model of the convolutional autoencoder, the encoding process of its convolutional layer is as follows:
[0138] h conv = σ(X conv * ω conv + b conv ) Formula XVII;
[0139] Where: h conv is the output feature information of the convolutional layer; o is the Relu activation function; X conv is the wind, light, and load time series data after data processing; ω conv and b conv are the number of convolutional kernels and the bias of the convolutional layer respectively;
[0140] The feature information captured by the convolutional layer from the wind, light, and load time series data is sent to the multi-head self-attention layer for further feature extraction, and its encoding process is as described in Formula XVII;
[0141] The decoding process of the decoder for decoding the encoding result of Formula XVII is expressed by the formula as
[0142] h deconv = σ(h conv * ω deconv + b deconv ) Formula XVIII;
[0143] Where: h deconv is the output feature information of the transposed convolutional layer; ω deconv and b deconv are the number of convolutional kernels and the bias of the transposed convolutional layer respectively;
[0144] The model of the convolutional autoencoder uses the mean square error (MSE) as the reconstruction loss function and minimizes it to continuously optimize the network parameters of the encoder and decoder. Its expression is:
[0145]
[0146] Where: L ris the reconstruction loss function; N d is the number of days of wind, light and load data;
[0147] In the model of the multi-head self-attention improved convolutional encoder, its multi-head self-attention is an improved type based on self-attention. It captures multiple groups of feature information in different subspaces from the data in a parallel query manner and splices them according to weights. The detailed calculation process is as follows:
[0148] Step A1: Based on linear transformation, the input data Y of the multi-head self-attention layer is transformed into a query matrix Q a , key matrix K a and value matrix V a , and the mathematical expression is:
[0149]
[0150] where: W Q , W K and W V are transformation matrices;
[0151] Map Q a , K a and V a to θ feature subspaces to obtain the query matrix Q a6 , key matrix K aθ and value matrix V aθ in the θ-th subspace, and the calculation method is:
[0152]
[0153] where: W Qθ , W Kθ and W Vθ are the transformation matrices of the θ-th subspace;
[0154] Step A2: Based on the scaled dot product and the Softmax function, calculate the self-attention values in θ feature subspaces, and the calculation method is:
[0155]
[0156] where: d is the scaling factor; head θ is the self-attention value in the θ-th feature subspace;
[0157] Finally, fuse the self-attention in θ feature subspaces:
[0158] M multi-head = C Concat (head1, head2,..., head θ )W o Formula XXIII;
[0159] Where: M multi-head is the self-attention value after fusion; C Concat is the matrix concatenation operation; W o is the parameter matrix.
[0160] Step three includes the following steps
[0161] Step B1: Based on the low-dimensional feature information obtained in the autoencoder, use the elbow method, that is, the elbow method, to observe the elbow values of different cluster numbers, so as to determine the optimal cluster number, initialize the Kmeans initial clustering center, and based on the joint loss function, use the Adam optimizer to finely adjust the entire network parameters and ensure the representativeness of the embedded space features, so as to obtain the optimal clustering result;
[0162] The joint loss function is expressed by the formula
[0163] L = L r + γL c Formula twenty-four;
[0164] Where: L is the joint loss function; γ is the coefficient for controlling the distortion degree of the embedding space, taking 0.1; L c is the clustering loss function;
[0165] Using the clustering center μ j as the connection weight between it and the low-dimensional space feature Z i , and mapping each low-dimensional space feature Z i to the soft label; at the same time, in order to improve the adaptation accuracy between Z i and μ i , a Gaussian distribution is used as the ideal target distribution. In this step, the clustering loss describes the KL divergence between the distribution on the soft label and the Gaussian distribution, which is used to measure the similarity between the two; the specific process is expressed by the formula as follows;
[0166]
[0167]
[0168]
[0169] Where: q ij is the probability that the low-dimensional space feature Z i belongs to the clustering center μ j ; b ij is the target distribution auxiliary function;
[0170] Step B2: Use the KL divergence as an auxiliary clustering objective function to optimize the stack encoder, and use the Adam optimizer to adjust the parameters to obtain an encoder structure suitable for scenario clustering; Adam integrates the advantages of the first-order momentum of stochastic gradient descent (SGD) and the second-order momentum of root mean square prop (RMSprop); based on the moment means of the two, while maintaining the learning rate of each parameter, give full play to the performance of sparse gradients, making the algorithm have high robustness in non-steady-state problems. The specific calculation process is as follows:
[0171]
[0172] M tt = β1·M tt-1 +(1 - β1)·g tt Formula 29;
[0173]
[0174]
[0175]
[0176]
[0177] Where: tt is the time interval; g tt is the gradient; M tt is the first-order moment estimate of g tt ; is the model parameter; η tt is the second-order moment estimate of g tt ; and are the corresponding network outputs respectively; ψ is the network step value; β1 is the exponential decay rate of M tt and is taken as 0.9; β2 is the exponential decay rate of η tt and is taken as 0.999; ε is a constant used to ensure the robustness of the algorithm;
[0178] Step B3: Obtain the optimal scenario clustering results within the year through the joint training of the encoder and the clustering layer, use the three evaluation indicators of CHI, SC, and DBI to evaluate the clustering results, and use the mean method to obtain the clustering centers of various scenarios to obtain the typical scenarios of each category.
[0179] In the above content, DCEC is the deep convolutional embedding clustering algorithm DCEC; VMD is the abbreviation of Variational mode decomposition for parameter optimization.
[0180] In this example, the actual wind-solar-load time series data with a sampling time of 1 h is used as a sample for clustering to construct a combined wind-solar-load scenario, so as to fully consider the coupling characteristics of the three in the process of power grid dispatching and planning.
[0181] Combined with specific embodiments, the comparison of the clustering evaluation index results of different methods is shown in the following table.
[0182] Table.1 Comparison of clustering evaluation index results of different methods
[0183]
Claims
1. A deep convolutional embedding clustering method for wind-solar-load combined scenarios with multi-head self-attention, which is used to generate wind-solar-load coupling scenarios, and is characterized in that: The method combines the feature extraction process with the clustering process by accurately capturing the coupling feature information among the wind, light, and load data to ensure the representativeness of the features in the embedding space. It includes the following steps: Step 1: Optimize the parameter combination of the VMD model with a multi-strategy fusion improved slime mold algorithm. Based on the optimal parameter combination, clean the wind, light, and load time-series data to weaken the influence of noise signals on the data feature extraction process. Step 2: Establish a convolutional autoencoder based on multi-head self-attention to extract the deep feature information of the processed wind, light, and load data. At the same time, use a convolutional decoder to reconstruct the original time-series signal. Step 3: Obtain the appropriate number of clustering clusters based on the elbow method, and use the Kmeans initial clustering center for the features. Then, based on the joint loss function composed of the sum of the encoder reconstruction loss and the clustering loss, adjust the network structure parameters and update the clustering results. Calculate the clustering center of each scenario based on the mean method as the typical scenario of this class, providing a basis for the optimal operation and planning of the power system. In Step 1, perform abnormal data detection and cleaning on the historical wind, light, and load data. The data is the wind, light, and load time-series data f(t) within one year, with days as the unit. Each sample contains the wind power, photovoltaic output, and load data at 24 moments. The specific method is as follows: Step S1: Using a non-linear time-domain decomposition method - the VMD model, decompose the original wind-solar-load data f(t) into K intrinsic mode function components u k (t) with central frequencies. At the same time, calculate the sum of the finite bandwidths of the K u k (t) and minimize it, so as to obtain the VMD model expression as follows: where: ω k is the center frequency of the k-th u k (t); δ(t) is the unit impulse function; is the partial derivative operator; introducing the Lagrange operator λ and the quadratic term penalty factor α to streamline the solution process of Equation (1), the model expression after operation is: Based on the alternating direction multiplier method, solve Equation (5.2) and continuously optimize and iterate and λ, and the iteration expression is as follows: where: n is the number of iterations; and are the Fourier transforms of u(t), λ(t) and f(t), respectively; In Step S2, when determining the preset values of the parameters K and α of the VMD model, if the value of the parameter K is too large, over-decomposition will occur, resulting in mode overlap. If the parameter α takes too large a value, it will lead to the loss of the center frequency. This step uses the improved slime mold algorithm with multi-strategy fusion based on KL divergence to search for the optimal parameter combination of VMD. KL divergence is used to measure the similarity between the intrinsic mode function u k (t) and the original wind-solar-load data f(t). The mathematical expression is as follows: Where: the closer R is to 0, the higher the similarity between the intrinsic mode function u k (t) and the original wind-solar-load data f(t); as the fitness value of the objective function, when R is the smallest, it indicates that the parameters K and a are the optimal parameter combination at this time; In Step S3, the slime mold algorithm simulates the foraging behavior of slime mold by establishing a model. That is, when the slime mold initially approaches food, it decides whether to approach the food based on the food concentration. The mathematical expression for position update is: Where: X(T) is the position of the slime mold in the T-th iteration; X b (T) is the currently discovered best position; W is the weight coefficient; X A (T) and X B (T) are the slime mold individuals randomly selected in the T-th iteration respectively; r is a random number in [0, 1]; v c is the feedback factor, whose value linearly decreases from 1 to 0; v b is the control parameter, and its value range is [-a, a]; p is the position update control parameter; Among them: The mathematical models of p, a, and W are: p = tanh|S(i) - DF| Formula VIII; Where: S(i) is the fitness value of the i-th slime mold individual; DF is the optimal fitness value in all iterations; T max is the maximum number of iterations; bF and wF are the optimal fitness value and the worst fitness value in the T-th iteration, respectively; condition is the individuals whose fitness values are ranked in the first half; In the slime mold algorithm, for the purpose of searching for high-quality food, the slime mold will divide some individuals to explore the remaining area. The mathematical formula for position update is: In the formula: rand is a random number; ub and lb are the upper and lower bounds of the exploration area of the slime mold respectively; z is the proportion parameter of the slime mold individuals exploring the remaining area. Step S4. To prevent the slime mold algorithm from being unable to accurately and timely feedback the concentration and quality of food due to the linearly decreasing v c and resulting in a slow convergence speed in the early stage, an adaptive and adjustable v c is introduced to accelerate the decline speed of v c in the early stage and improve the ability to search for the global optimal value; and the v c in the later stage of iterative calculation is relatively stable to avoid falling into a local optimum; the mathematical expression of the adaptive and adjustable v c is as follows: The adaptive reverse learning mechanism introduces a vector into the exploration area of the slime mold This vector is opposite to the positions of each slime mold individual and compares the fitness values of the two to avoid falling into local optima; the expression for the position of the i-th slime mold individual at the T-th iteration is as follows: Based on the adaptive decision, when the i-th slime mold explores food, it compares the current fitness value with the previous optimal fitness value S(X i (T)) to determine whether to use the reverse learning mechanism for additional exploration and update the position for the next iteration. The expression is as follows: In Step S5, select the mean absolute error as the standard to measure the quality of data processing. The calculation method is: Where: N is the total number of samples of wind, light and load data; is the wind, light and load data after data processing; MAE measures the ability of the data processing algorithm to retain the original features of the data on the premise of effectively removing the outliers in the wind, light, and load data. The smaller the MAE, the better the data processing effect. Step 3 includes the following steps Step B1: Based on the low-dimensional feature information obtained from the autoencoder, use the elbow method, that is, the elbow method, to observe the elbow values of different numbers of cluster classes, so as to determine the optimal number of cluster classes. Initialize the Kmeans initial clustering center, and based on the joint loss function, use the Adam optimizer to slightly adjust the parameters of the entire network and ensure the representativeness of the features in the embedded space, thereby obtaining the optimal clustering result. The joint loss function is expressed by the formula L = L r + γL c Formula 24 Where: L is the combined loss function; γ is the coefficient for controlling the distortion degree of the embedding space, taking 0.1; L c is the clustering loss function; With the clustering center μ j as its connection weight with the low-dimensional space feature Z i and map each low-dimensional space feature Z i to the soft label; meanwhile, in order to improve the adaptation accuracy between Z i and μ j , a Gaussian distribution is adopted as the ideal target distribution. In this step, the clustering loss describes the KL divergence between the distribution on the soft label and the Gaussian distribution, which is used to measure the similarity between the two; the specific process is expressed by the formula as follows; where: q ij is the probability that the low-dimensional space feature Z i belongs to the clustering center μ j ; b ij is the target distribution auxiliary function; Step B2: Use the KL divergence as an auxiliary clustering objective function to optimize the stack encoder, and use the Adam optimizer to adjust the parameters to obtain an encoder structure suitable for scenario clustering. The specific calculation process is: M tt = β1·M tt-1 + (1 - β1)·g tt Formula 29; where: \(t_t\) is the time interval; \(g\) tt is the gradient; \(M\) tt is the first moment estimate of \(g\) tt ; is the model parameter; \(\eta\) tt is the second moment estimate of \(g\) tt ; and are the corresponding network outputs respectively; \(\psi\) is the network step value; \(\beta_1\) is the exponential decay rate of \(M\) tt and is taken as 0.9; \(\beta_2\) is the exponential decay rate of \(\eta\) tt and is taken as 0.999; \(\varepsilon\) is a constant used to ensure the robustness of the algorithm; Step B3: Obtain the optimal annual scenario clustering result through the joint training of the encoder and the clustering layer. Evaluate the clustering result using three evaluation metrics, namely CHI, SC, and DBI, and use the mean method to obtain the clustering centers of various scenarios to obtain the typical scenarios of each category.
2. The method for jointly clustering the scenery, wind, and load in the deep convolutional embedding of multi-head self-attention according to claim 1, wherein: In step two, each processed data sample is expanded into a tensor of size (9, 9) through normalization and zero-padding operations. Based on the multi-head self-attention improved convolutional encoder, the wind, light, and load data is dimensionally reduced to extract deep features. At the same time, the original time series signal is reconstructed using a convolutional decoder. In the model of the convolutional autoencoder, the encoding process of its convolutional layer is as follows: h conv = σ(X conv * ω conv + b conv ) Formula XVII; Where: h conv is the output feature information of the convolutional layer; σ is the Relu activation function; X conv is the time series data of wind, light and load after data processing; ω conv and b conv are the number of convolutional kernels and the bias of the convolutional layer respectively; The feature information captured by the convolutional layer from the wind, light, and load time series data is sent to the multi-head self-attention layer for further feature extraction. Its encoding process is expressed as Equation (17). The decoding process of the decoder for decoding the encoding result of Equation (17) is expressed as h deconv = σ(h conv * ω deconv + b deconv ) Formula XVIII; where: h deconv is the output feature information of the deconvolution layer; ω deconv and b deconv are the number of convolution kernels and the bias of the deconvolution layer, respectively; The model of the convolutional autoencoder uses the mean square error (MSE) as the reconstruction loss function and minimizes it to continuously optimize the network parameters of the encoder and the decoder. Its expression is: where: L r is the reconstruction loss function; N d is the number of days of wind, light and load data; In the model of the multi-head self-attention improved convolutional encoder, its multi-head self-attention is an improved type based on self-attention. It uses the method of multiple queries to parallelly capture multiple groups of feature information in different subspaces from the data and splices them according to weights. The detailed calculation process is as follows: Step A1: Based on linear transformation, the input data Y of the multi-head self-attention layer is transformed into a query matrix Q a , a key matrix K a and a value matrix V a , and the mathematical expression is as follows: Where: W Q , W K and W V are transformation matrices; Map Q a , K a and V a to θ feature subspaces to obtain the query matrix Q a6 , key matrix K aθ and value matrix V aθ , and the calculation method is as follows: Where: W Qθ , W KB and W Vθ are the transformation matrices of the θ-th subspace; Step A2: Calculate the self-attention values in θ feature subspaces based on the scaled dot product and the Softmax function. The calculation method is: where: d is the scaling factor; head θ is the self-attention value in the θ-th eigen-subspace; Finally, fuse the self-attention in θ feature subspaces: M multi-head = C Concat (head1, head2,..., head θ )W o Formula XXIII; Where: M multi-head is the self-attention value after fusion; C Concat is the matrix concatenation operation; W o is the parameter matrix.