Bridge wind field change mode prediction method based on LSTM depth embedded clustering model
By introducing the LSTM depth embedded clustering model in the wind field prediction, extracting the wind speed sequence characteristics and using the softmax function to calculate the soft allocation, the problem of poor prediction of the daily change mode of wind speed in the existing technology is solved, and more efficient and accurate wind field prediction is achieved.
Patent Information
- Application Number
- CN202510452130.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing wind field prediction methods perform poorly in complex terrain areas, making it difficult to effectively predict the daily change patterns of wind speed, especially in mountainous canyon environments.
The bridge wind field change mode prediction method based on the LSTM depth embedded clustering model is adopted, and the characteristics of the wind speed sequence are extracted through the LSTM autoencoder layer, and the softmax function is used to calculate the soft allocation in the deep embedded clustering layer to realize the prediction of the wind speed daily change mode.
This method can effectively extract the characteristics of the time series, improve the prediction accuracy of the daily change mode of the wind farm, and is suitable for complex wind farm environments, reducing calculation costs and training time.
Smart Images

Figure CN119989932A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bridge wind field prediction, and specifically relates to a bridge wind field change pattern prediction method based on an LSTM deep embedded clustering model. Background Art
[0002] Since the stiffness and damping of long-span bridges are relatively low, wind-induced vibration is prone to occur under wind loads. Therefore, the wind-resistant design of long-span bridge structures usually plays a key control role. In mountainous canyon environments, the wind field is also complex and changeable due to the influence of the terrain. Long-term monitoring of the wind field at the bridge site and prediction of the change pattern of the wind field through monitoring data will help the normal construction, operation management and daily maintenance decision-making of the bridge.
[0003] Existing wind field prediction methods include numerical weather forecasting: by collecting meteorological data for numerical calculations, solving fluid mechanics and thermodynamic differential equations, and predicting future wind speeds; statistical forecasting methods: establishing a predictive regression model, using historical data as model input, and training to obtain specific model parameters for wind speed prediction, specifically including time series method and machine learning method.
[0004] However, there are many problems with existing wind field prediction methods. Numerical weather forecasting methods have significant computational complexity and are highly sensitive to initial conditions. Numerical models cannot adapt well to the climate characteristics of complex terrain areas such as mountainous canyons. A single statistical forecasting method is difficult to adapt to the prediction of complex and changeable wind fields: the time series method cannot perfectly fit the nonlinear relationship between variables, and the performance of time series methods used for medium-term or long-term predictions is poor; supervised machine learning algorithms require a large amount of data for training and prediction, which is difficult and time-consuming to train. The discreteness of bridge wind field data is large, and the existing methods have poor prediction effects on the high-frequency part of wind field data, and there is a lack of a method for predicting the daily variation pattern of wind speed. Summary of the invention
[0005] For the prediction of the daily variation pattern of wind speed, clustering method can be used to form a library of daily variation patterns of wind speed, and the prediction of daily variation patterns of wind speed can be realized through feature matching. Clustering can realize the automatic grouping of samples in the data set according to a certain similarity measure, and is widely used in data mining, pattern recognition, image processing, market analysis and other fields. Especially in the absence of pre-labeling, clustering can help discover the potential structure or pattern in the data. However, traditional clustering algorithms have some significant limitations, including sensitivity to initial values and noise, simple assumption of cluster shape, and poor high-dimensional data processing capabilities. In order to overcome these shortcomings, many new clustering algorithms have been produced in recent years, such as clustering methods based on deep learning, which can better handle complex data and high-dimensional data, and have stronger flexibility and robustness. However, the existing technology does not fully extract the features of the data, and ignores the relationship between feature extraction and clustering, resulting in the fact that the improvement of clustering effect by deep learning is not obvious. In order to solve the above problems, this paper proposes a prediction method for the daily variation pattern of bridge wind farm based on LSTM deep embedded clustering model, improves the deep embedded clustering algorithm, and realizes the prediction of the daily variation pattern of wind farm based on this algorithm.
[0006] The present invention provides a bridge wind field change pattern prediction method based on LSTM deep embedded clustering model, comprising the following steps: S1. Data preparation: Obtain wind speed and direction information at the mountain valley bridge site, perform preprocessing and generate a wind speed sequence data set; S2. Model construction: Construct a fused LSTM deep embedded clustering model, which includes an LSTM autoencoder layer and a deep embedded clustering layer; S3, model training: inputting the wind speed sequence data set into the LSTM deep embedded clustering model for training until the loss function is stable; S4. Construction of daily variation pattern library: Using the trained LSTM deep embedded clustering model, the cluster center is reconstructed through the decoder to form a wind speed daily variation pattern library and divide the wind speed sequence through soft allocation; S5. Daily variation pattern prediction: match the wind speed data to be predicted with the wind speed daily variation pattern library, and take the most similar wind speed daily variation pattern as the prediction result.
[0007] Preferably, the S1 specifically includes: S11, dividing the wind speed and wind direction data at predetermined time intervals, and eliminating the measured data that does not meet the preset length and standard deviation; S12, decomposing the wind speed and wind direction into wind speed components in specific directions, and calculating the average wind speed and wind direction angle; S13, dividing the average wind speed according to the same time interval, using a filter to obtain the change trend, and performing standardization processing to form a wind speed sequence data set.
[0008] Preferably, the S2 specifically includes: S21. Construct an LSTM autoencoder layer, including an encoder composed of an LSTM layer and a decoder composed of an LSTM layer and a fully connected layer; S22. Build a deep embedded clustering layer, classify the latent variables, use k-means to initialize the cluster centers, calculate the soft allocation through the softmax function and build the auxiliary allocation, and use the KL divergence of the two as the clustering loss; S23, merging the reconstruction loss and clustering loss into the LSTM deep embedded clustering model loss function; S24. Set the learning rate, initialize the clustering loss weight and the clustering loss weight growth rate.
[0009] Preferably, the soft allocation is specifically expressed as: ; In the formula, express Belong to j The probability of the class, express Features , is the temperature parameter, k is the number of clusters; m is the serial number of the cluster center, from 1 to k change; c m It is m Probability of class; c j It is j The cluster center of the class; The auxiliary allocation is specifically expressed as: ; In the formula, for Belong to j Probability of class; for Belong to m Probability of class; for Belong to m Probability of class; x is the sequence number of the wind speed sequence; n is the number of wind speed sequences.
[0010] This formula is given calculate , that is, Features , belongs to jThe probability of a class, once given, cannot be changed. In the summation formula, we need to traverse All elements belong to j The probability of the class, where x is a variable. x = i situation, that is to say, z i It is certain. z x is changing, the two sets are the same, but the ranges they represent are different. z x Indicates x The characteristics of a variable.
[0011] Preferably, the LSTM deep embedded clustering model loss function is specifically expressed as: ; In the formula, Represents the weight of clustering loss; ; To initialize the weight of clustering loss, t is the iteration round, is the weight growth rate of clustering loss; is the reconstruction loss function; is the clustering loss.
[0012] Preferably, S3 specifically includes: The wind speed sequence data set is input into the LSTM deep embedded clustering model, the number of clusters and the total number of training rounds are set, and training is performed; wherein, in each training, the wind speed sequence data set is passed through the LSTM deep embedded clustering model to obtain the LSTM deep embedded clustering model loss function, and the LSTM network parameters and the clustering center of the deep embedded clustering layer are optimized by the stochastic gradient descent method with momentum, until the training is stopped after the set total number of rounds are executed, ensuring that the number of rounds when the total loss is reduced to a stable state is less than the specified total number of training rounds, so that the total loss is stable after all rounds of training are executed. The total loss here represents the LSTM deep embedded clustering model loss function.
[0013] The total loss stability is determined by the following formula: ; In the formula, is the total loss in round t, is the total loss in the t+1th round. When the change in loss is less than 0.0002, the loss can be considered stable.
[0014] Preferably, the S4 specifically includes: S41, using the optimized LSTM decoder to reconstruct the optimized cluster center to form a wind speed daily variation pattern library; S42, calculating the probability that the wind speed sequence in the latent variable belongs to each cluster center through the soft allocation, and dividing the wind speed sequence into the cluster center with the calculated maximum probability.
[0015] Preferably, the S5 specifically includes: S51, for any section of wind speed data divided at the same time interval as in S1, obtain the average wind speed sequence according to step S1 and perform maximum value standardization to obtain a standardized average wind speed; S52. Standardize the wind speed daily variation pattern in the wind speed daily variation pattern library, extract the wind speed in the same time interval as the standardized average wind speed sequence to be predicted from the standardized wind speed daily variation pattern to obtain a standardized wind field pattern fragment, and find the wind speed variation pattern with the smallest Euclidean distance to the standardized average wind speed as the closest wind speed variation pattern in the same time interval from the standardized wind field pattern fragment and perform denormalization processing to obtain the final wind speed daily variation pattern prediction data.
[0016] Preferably, the normalization of the diurnal variation pattern of wind speed is specifically expressed as follows: ; In the formula, To obtain the variable Maximum value function within a period; For the j A standardized daily variation pattern of wind speed; For the j Diurnal variation pattern of wind speed.
[0017] Preferably, the final wind speed daily variation pattern prediction data is specifically expressed as: ; In the formula, is the predicted diurnal variation pattern of normalized wind speed; is the average wind speed series; The average wind speed series The maximum value of .
[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. This solution realizes the prediction of the daily variation pattern of the wind farm. It innovatively integrates LSTM into the deep embedded clustering model, which can well extract the characteristics of the time series, and innovatively proposes to use the softmax function to calculate the soft allocation in the deep embedded clustering layer to achieve end-to-end efficient training of the neural network. The clustering loss weight is introduced into the model loss function to dynamically adjust the proportion of the reconstruction task and the clustering task. It can be seen from the embodiments that the bridge wind farm daily variation pattern prediction method based on the LSTM deep embedded clustering model has a good effect in the prediction of the daily variation pattern of the wind farm.
[0019] 2. This scheme predicts the daily variation pattern of wind speed at the bridge site of a large-span bridge, obtains the daily variation trend of wind speed, and provides support for bridge construction and operation and maintenance decisions. The proposed LSTM deep embedded clustering algorithm, which adds LSTM autoencoders to deep embedded clustering, can better extract the characteristics of time series and realize end-to-end time series clustering, simplifying the process, reducing human intervention, and is very convenient; in addition, this method has strong versatility and can be widely used in various wind farm monitoring tasks, and can adapt to complex wind farm environments; it is easy to program, more efficient to train, and has low time cost.
[0020] 3. This application realizes the dynamic fusion of autoencoder and clustering algorithm: The LSTM deep embedded clustering model in this paper integrates feature extraction and clustering tasks to achieve end-to-end training, and one model completes both tasks at the same time. However, the two tasks have a sequence. Clustering is performed on the features extracted by the autoencoder, so feature extraction should be implemented before clustering. Therefore, a parameter is set , achieving different emphases on the two tasks at the beginning and end of the task. Compared with the static fusion loss, we focus on different tasks at different stages, propose a dynamically changing fusion loss, optimize the allocation of computing resources, speed up the convergence of the model, and improve the efficiency of model training. At the same time, focusing on different tasks at different stages improves the performance of the model on different tasks.
[0021] 4. This application uses the softmax function to calculate auxiliary allocation: The general clustering algorithm is to deterministically give a set of data belonging to a certain class, which does not have the property of smooth probability distribution, and it is difficult or impossible to realize the construction of loss function. The soft assignment strategy can realize the training of clustering tasks, but the general soft assignment method such as Student's t distribution is not stable enough, the calculation is complex, and the convergence speed is slow. Therefore, it is proposed to introduce the method of calculating soft assignment with softmax function into the model. By using soft assignment, the loss function can be constructed to realize the training of clustering tasks, thereby realizing end-to-end integrated model training. Compared with the traditional Student's t distribution, the calculation of soft assignment using softmax function is simpler and more efficient. It is highly compatible with common loss functions (such as cross entropy loss), and the gradient calculation is relatively easy. For large data sets and high-dimensional data, Softmax calculation is very efficient. At the same time, Softmax can avoid the problem of gradient vanishing or gradient exploding, and is more suitable for the training of deep neural networks. In the Softmax method, the membership of each data point to each cluster center is expressed by probability, which makes the model results more interpretable and intuitive.
[0022] 5. Improved deep embedded clustering - using LSTM autoencoders instead of general autoencoders: For the clustering task of time series, it is necessary to pay attention to its temporal features, while general autoencoders cannot extract temporal feature information well. By constructing an LSTM autoencoder layer to replace the autoencoder layer, the deep embedded clustering method is improved. In deep embedded clustering, using LSTM autoencoders instead of ordinary autoencoders can significantly improve the model's ability to process time series data. LSTM can capture long-term dependencies and temporal relationships in data, generate potential representations with more temporal information, is suitable for variable-length inputs, and can better filter out noise. In contrast, ordinary autoencoders mainly focus on static features and are difficult to effectively model the dynamic changes of time series data. Therefore, when processing data with time series dependencies, LSTM autoencoders perform better in feature extraction, noise robustness, and clustering accuracy.
[0023] 6. Prediction method - holistic matching prediction The discreteness of bridge wind farm data is large, and the existing methods have poor prediction effects on the high-frequency part of wind farm data, and there is a lack of a method for predicting the daily variation pattern of wind speed. The results obtained by integrating the LSTM deep embedded clustering method can well realize the prediction of daily variation patterns. Commonly used sliding prediction methods (such as ARIMA, LSTM, etc.) require the assumption of data stability, which may not be applicable for some complex time series data. At the same time, in ordinary prediction methods, the latter step prediction depends on the previous step prediction, so the cumulative error is large and the calculation is more complicated. Time series matching prediction can effectively capture long-term dependencies and complex dynamic changes by finding historical similar patterns to make predictions, avoiding the error accumulation problem in sliding prediction methods. It can adaptively process sequence data of different time scales, enhance sensitivity to sudden events and trend changes, and reduce the interference of noise on prediction results. In addition, the matching prediction method does not depend on the previous step prediction, can avoid gradually accumulated errors, has strong adaptability, and is particularly suitable for time series data with complex nonlinear and time-varying characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A schematic diagram of a flow chart of an embodiment of the present invention; Figure 2 This is a schematic diagram of the arrangement of wind farm data acquisition equipment in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the LSTM deep embedded clustering model network integrated in an embodiment of the present invention; Figure 4 This is a schematic diagram of the results of clustering and reconstruction in an embodiment of the present invention; Figure 5 A comparison diagram of the wind speed daily variation pattern prediction result and the measured result in the embodiment of the present invention; Figure 6 It is a schematic diagram of the prediction result of a certain day of the daily variation pattern of wind speed in an embodiment of the present invention. DETAILED DESCRIPTION
[0025] Embodiment 1: This application solves the problem that the existing methods have poor prediction effect on the high-frequency part of wind field data. In addition, the present invention also has the following purposes: First of all, it is the needs of the project: the change in wind speed trend can provide a preliminary reference for technicians in the construction, operation and maintenance of the bridge, and make corresponding decisions based on the daily wind speed. Moreover, different change trends, such as sudden rise and slow fall, slow rise and slow fall, can cause great differences in the response of the bridge.
[0026] Secondly, in the study, it was found that in the complex mountainous canyon areas, when analyzing the observed wind field data, it was found that the change trend of wind speed data has a certain periodicity. One of the obvious regular fragments was found: Taking a day as a cycle, the wind speed is relatively low in the early morning, then increases, reaches a maximum in the afternoon, and then decreases to a lower value. There are many similar segments. Therefore, it is possible to summarize the daily change of wind speed on a daily basis, and summarize the particularities and related laws.
[0027] The present invention obtains wind speed and wind direction information at a bridge site in a mountain canyon, performs preprocessing and generates a wind speed sequence data set; constructs a fused LSTM deep embedded clustering model, inputs the wind speed sequence data set into the LSTM deep embedded clustering model for training until the loss function is stable; then uses the trained LSTM deep embedded clustering model to reconstruct cluster centers through a decoder to form a wind speed daily variation pattern library and divide the wind speed sequence through soft allocation; finally, matches the wind speed data to be predicted with the wind speed daily variation pattern library, and takes the most similar wind speed daily variation pattern as the prediction result.
[0028] Based on the aforementioned inventive concept, the implementation idea of this application is: Average wind speed extraction is a mature data preprocessing method for wind field research. In engineering, 10 minutes is often used as the basic time interval, which means that the data originally collected may be one per second, but it is too discrete. In engineering, it is often divided into 10 minutes to obtain an average wind speed within each 10 minutes, and then many 10-minute data sets are combined to form an average wind speed sequence for research.
[0029] During the research, it was found that even the 10-minute average wind speed still has strong volatility, but although the data is fluctuating, it also has a very obvious trend. The research direction of this application is to focus on this trend and extract it. Here, the wind speed trend is extracted by Gaussian filtering, high-frequency signals are removed, and only the trend of wind speed change is focused on.
[0030] To summarize the daily variation of this time series, the first thing that comes to mind is the clustering algorithm, which is also a commonly used and highly efficient method. The more common data for clustering is two-dimensional, which is displayed on the drawing as a group of scattered points. Each point can be represented by a coordinate ( , ) means that each point is described by two values. The time series is ( , , , …, ), each point is represented by There are usually hundreds of n, so this point can be called a "bar". A piece of data contains many values. dimensional data. There are obvious problems in clustering high-dimensional data such as time series: the amount of data is huge, and computer storage takes up a large space. At the same time, when clustering high-dimensional data, it is very complicated to calculate the distance between two pieces of data, which will lead to a very low convergence speed of clustering. Therefore, it is necessary to reduce the dimension of the data. There are many methods of dimensionality reduction, and the autoencoder is very suitable, with the characteristics of high efficiency, low error, and easy operation. Therefore, the autoencoder is used to reduce the data dimension, and then the reduced-dimensional data is clustered. After clustering, a library of types with different change rules can be obtained. In this way, the wind speed change trend of each library is consistent and very representative, because the clustering uses a lot of historical data, which basically includes all types of changes. Therefore, the future daily change trend of wind speed is basically included in these libraries, and the library can be used to predict the wind speed.
[0031] After you have the initial ideas, you need to continue to refine them.
[0032] First of all, the combination of autoencoder and clustering realizes the summary of wind speed library. The autoencoder can reduce the length of a wind speed daily sequence of length 144, which is convenient for clustering. However, it was found in the research process that when performing this step, it is always necessary to train the autoencoder first, extract the features, and then take this feature to cluster. Therefore, when there is new data, it is still necessary to proceed step by step. Such an operation process (processing process) is very inconvenient. There are too many human interventions in the processing process, which increases the potential risk of errors, increases the workload and working time of operators, and reduces efficiency. For this reason, this application puts the autoencoder and clustering together, and finally can be integrated into a model. It only needs input to directly output the clustering results to me, which greatly improves the efficiency, and is also convenient for developing programs, facilitating the repetition of experiments and use with other technicians. Therefore, it is thought of using a deep embedded clustering method, putting clustering in the training, and clustering is trained with the autoencoder and outputs the clustering results at the end of the training.
[0033] Among them, how to achieve fusion is a very important issue, because the result of traditional clustering is a deterministic value, so the error can only be expressed as "yes" or "no". Specifically, if a piece of data is judged to be the first category, if it is really the first category, the error is 0, otherwise the error is 100%, there are only these two possibilities. In simulation training, taking the autoencoder as an example, its error is expressed as the gap between the reconstructed data and the original data. This gap can be calculated by a formula, such as MSE, calculating the square root of the difference between the original data and the reconstructed data. The calculated value can be any value, so this loss function can be continuous and smooth, and the derivative can be calculated. Calculating the derivative is a necessary step to update the aforementioned established model during the training process. Therefore, the autoencoder can be trained using MSE, while the traditional clustering cannot be differentiated to determine the discrete results of "yes" or "no", so the model cannot be updated. This application introduces the softmax function for classification tasks in machine learning, and defines a certain data as belonging to a certain class by a probability method, and quantifies the clustering results, that is: output the probability that a piece of data belongs to each class, rather than assuming that the probability of belonging to one of the classes is 1 and the rest is 0. Since the clustering results are quantified, the calculation errors are also quantified. The change of the calculation error result is no longer a transformation between two values, but a process that can be described by a certain function. Through the gradient descent method, the parameters of the clustering model (the parameters in the method of this application are the cluster centers) can be changed by derivation to continuously optimize the results. This application completely changes the original clustering judgment mode and judgment results.
[0034] However, this application clusters the features extracted by the autoencoder after all. In the prior art, such feature extraction and clustering tasks cannot be performed at the same time, and they must be focused on at different stages. This application proposes a dynamic loss function for the embedded clustering algorithm. In the first stage, the clustering loss is basically 0, so the loss will be more of the reconstruction error of the autoencoder, mainly to optimize the autoencoder. In the second stage, the reconstruction error of the autoencoder is already very small. At this time, the autoencoder already has very good performance, and the clustering loss of the overall model becomes the dominant loss. At this time, the clustering center will be continuously optimized to eventually obtain the best clustering result.
[0035] Since this application is to extract features from time series, this application uses the long short-term memory neural network LSTM, which can well extract time-related information from the data and improve the accuracy of the autoencoder's reconstruction of the original data. If the data can be well reconstructed, it means that the autoencoder can extract key information. Therefore, the LSTM deep embedded clustering method is formed, the wind speed daily variation pattern library is obtained, and the law of wind speed daily variation is summarized.
[0036] After obtaining the wind speed daily variation pattern library, it can be seen that the wind speed changes in each category are similar, and the categories are not similar. At the same time, the time series has a strong autocorrelation. The data of the previous period will affect the data of the next period, and the changes of the next period have a certain dependence on the previous period. Therefore, after the previous period is determined, the change trend of the whole day can be basically determined. Therefore, for a wind speed sequence of less than one day (such as the first 14 hours), it is effective to match the wind speed data of the previous period (the first 14 hours) of the wind speed daily variation pattern library and use the complete wind speed daily variation pattern to represent the wind speed change of the day. Due to the integration of more historical data, the wind speed daily variation pattern library can more comprehensively summarize the law of wind speed changes, and it is effective to use the matching method to predict the wind speed change pattern.
[0037] In the prediction method of the diurnal variation pattern of wind speed mentioned in this method, the diurnal variation pattern of wind speed refers to the trend of wind speed, such as Figure 6 The red curve in (blue is the original data, green is the fluctuation, and red is the trend item). Therefore, this method can be used as a basis for preliminary engineering decision-making and provided to technical personnel for reference to determine the trend of wind speed changes. This method has reached a high level. Of course, the prediction of the fluctuation item is not considered, and it is impossible to predict when the wind speed has a short-term sudden change.
[0038] like Figure 1-6 As shown in the figure, the bridge wind field change pattern prediction method based on the LSTM deep embedded clustering model includes: S1. Data preparation: Obtain wind speed and direction information at the mountain valley bridge site, perform preprocessing and generate a wind speed sequence data set; S2. Model construction: Construct a fused LSTM deep embedded clustering model, which includes an LSTM autoencoder layer and a deep embedded clustering layer; S3, model training: inputting the wind speed sequence data set into the LSTM deep embedded clustering model for training until the loss function is stable; S4. Construction of daily variation pattern library: Using the trained LSTM deep embedded clustering model, the cluster center is reconstructed through the decoder to form a wind speed daily variation pattern library and divide the wind speed sequence through soft allocation; S5. Daily variation pattern prediction: match the wind speed data to be predicted with the wind speed daily variation pattern library, and take the most similar wind speed daily variation pattern as the prediction result.
[0039] Furthermore, the specific method of data preparation in step S1 is to convert wind speed Corresponding wind direction Split the data into 10-minute intervals. For any data, its length should be:
[0040] In the formula, For any data length, The minimum number of samples required for sampling equipment in ten minutes.
[0041] The measured data should meet the following requirements:
[0042]
[0043] In the formula represents the average wind speed. is the standard deviation of the measured wind speed, Indicates the average value of the measured wind direction. is the standard deviation of the measured wind direction.
[0044] The measured data that do not meet the above inequality conditions are eliminated. With wind direction Decompose along Direction and Direction, that is:
[0045]
[0046] In the formula and Respectively indicate along Direction and The above two directions are the specific directions mentioned in this application, and the specific directions here are commonly used in the vector decomposition method of wind speed; in the specific research process, in order to facilitate the analysis of the data obtained by the Wind3D 6000 wind laser radar, the instrument is set to the north as the wind direction angle 0° direction, when rotating clockwise, the wind direction angle Increase. The general method of wind field research, vector decomposition method (Xiang Haifan, Ge Yaojun, Zhu Ledong. Modern Bridge Wind Resistance Theory and Practice [M]. Beijing: People's Communications Press, 2005) is used to decompose the wind speed into two mutually orthogonal directions; after this treatment, Along the =0° direction, Along the =90° direction. After decomposing the instantaneous wind speed into these two directions, take 10 minutes as the basic time interval, and calculate two average values of the wind speed in the two directions in each interval. , then square and sum to synthesize the average wind speed, and get the wind direction angle. The average wind speed obtained in this way is the main wind speed, and the average wind direction is the main wind direction. This is the vector decomposition method of wind speed.
[0047] Then calculate the ten-minute average wind speed and wind direction ,Right now:
[0048]
[0049] In the formula and Respectively indicate along Direction and The ten-minute average wind speed is then The data is divided into 24-hour units. The starting time of each data is 00:00:00. For any data, the following conditions must be met:
[0050] In the formula It represents the length of each data in 24 hours. Data that does not meet the above formula is directly eliminated. Then a Gaussian filter is used to obtain the change trend of each data. The maximum normalization method is used to standardize the data, that is:
[0051] In the formula represents the normalized wind speed, The ten-minute average wind speed. is the maximum value function.
[0052] Integrate all Forming a wind speed series data set , .
[0053] Contains n sequences, each of which is a time series , represents the normalized wind speed data for one day, each sequence Contains 144 normalized values of the maximum ten-minute average wind speed, arranged in chronological order from the beginning to the end.
[0054] Furthermore, the model building process in step S2 constructs a fused LSTM deep embedded clustering model, which includes an LSTM autoencoder layer and a deep embedded clustering layer, wherein the specific structure of the LSTM autoencoder layer includes: Encoder: It consists of LSTM layers to extract input features. The latent variables of its output represent the key features of the input data and are expressed as:
[0055] In the formula represents the encoder, Represents the input of the model, which is a set of wind speed series data. represents the encoder model parameters of the LSTM autoencoder, is the output of the model, representing the wind speed series data set The feature set extracted from There are n sequences, each of which are all one-dimensional vectors with length less than 144, express Features; Decoder: It consists of LSTM layers and fully connected layers, which are used to reconstruct the original data from the latent variables, expressed as:
[0056] In the formula Represents the set of wind speed sequences reconstructed by the decoder by decoding the features extracted by the encoder. , Represents a decoder, is the input of the decoder model, representing the wind speed sequence data set The feature set extracted from Represents the decoder model parameters of the LSTM autoencoder; Latent space: The latent variables output by the encoder serve as a low-dimensional feature representation of the input data.
[0057] Furthermore, the model building process in step S2 constructs a fused LSTM deep embedded clustering model, which includes an LSTM autoencoder layer and a deep embedded clustering layer, wherein the specific behaviors of the deep embedded clustering layer include: Will gather n sequences in Divided into k classes, the class center is Indicates that , use k-means to initialize the cluster center ,calculate Soft assignment to cluster centers and auxiliary distribution , and the KL divergence composed of soft assignment and auxiliary assignment is used as the clustering loss.
[0058] Furthermore, in the model building process in step S2, the fused LSTM deep embedded clustering model is constructed, including an LSTM autoencoder layer and a deep embedded clustering layer, and the reconstruction loss and clustering loss fusion are used as the loss function of the model. The reconstruction loss is expressed as:
[0059] In the formula, represents the reconstruction loss function, , . Use the softmax function to calculate Soft assignment to cluster centers ,Right now:
[0060] In the formula express The probability of belonging to the jth class, , is the temperature parameter that controls the hardness or softness of clustering. When the temperature is high, the samples are evenly distributed among the clusters, while when the temperature is low, the samples are more concentrated in a certain cluster. , which can also be called target allocation, that is:
[0061] In the formula, express The probability of belonging to the jth class is given by Calculated, with higher credibility, requires Getting closer ,when equal The clustering loss is zero when , so the clustering loss can be defined as the KL divergence between the soft assignment and the auxiliary assignment, that is:
[0062] In the formula Indicates the calculation of the KL divergence between P and Q. The LSTM autoencoder reconstructs the loss With Deep Embedded Clustering Loss Fusion constitutes the loss function of the fused LSTM deep embedded clustering model, namely:
[0063] In the formula represents the weight of clustering loss, It has dynamic changing characteristics, namely:
[0064] In the formula is the weight for initializing clustering loss, t is the iteration round, is the weight growth rate of clustering loss.
[0065] Furthermore, in the model building process in step S2, the fused LSTM deep embedded clustering model is constructed, and the LSTM network parameters are optimized by the stochastic gradient descent method with momentum. , and the cluster centers of the deep embedded clustering layer LSTM network parameters , The optimization calculation is as follows:
[0066]
[0067] In the formula is the encoder model parameter in the t-th round of LSTM autoencoder, is the encoder model parameter in the LSTM autoencoder of the t+1th round, is the decoder model parameter in the t-th round of LSTM autoencoder, is the decoder model parameter in the LSTM autoencoder of the t+1th round, represents the learning rate, Indicates about The gradient operator of Indicates about The gradient operator of . Cluster centers of deep embedded clustering layer The optimization calculation is as follows:
[0068] In the formula is the deep embedded cluster center of the tth round, is the deep embedded clustering center of the t+1th round, represents the learning rate, Indicates about The gradient operator.
[0069] Furthermore, in the model building process in step S2, the constructed model includes an LSTM autoencoder layer and a deep embedded clustering layer, each of which is built by code, including all the calculation processes described in step S2. Each part is connected together by code to form a trainer for the entire model. The overall model needs to input a wind speed sequence data set, a selected number of clusters, and a total number of training rounds, and the output is an encoder, a decoder, and a cluster center.
[0070] Furthermore, in the model training process in step S3, the data prepared in step S1 is input into the model constructed in step S2, the selected number of clusters and the total number of training rounds are input, and the code is run to start training. In each round, the total loss formula (16) described in S2 is calculated after the data set passes through the model, and the model and cluster centers are updated using formulas (18)-(20) to make the model and cluster centers more accurate. The round is ended and the next round is carried out. After the specified total number of rounds is executed, the training is stopped. The total loss represents the error. As the number of rounds increases, the total loss decreases. It should be ensured that the number of rounds when the total loss is reduced to a stable state and no longer changes significantly is less than the specified total number of training rounds.
[0071] Furthermore, the daily variation pattern library construction process in step S4 uses an optimized LSTM decoder The optimized cluster centers Reconstruction is performed to form a wind speed daily variation pattern library, namely:
[0072] In the formula , represents the wind speed daily variation pattern library, which contains k wind speed sequences, each sequence is a one-dimensional vector of length 144, represents the network parameters of the optimal LSTM decoder. All sequences in , each sequence are all one-dimensional vectors with length less than 144, express The characteristics of each cluster center are calculated by formula (13) Probability , assign the class label with the maximum probability to the sequence: for example The largest, then Belongs to the jth category.
[0073] Furthermore, the daily variation pattern prediction process in step S5 is performed on the basis of steps S1-S4. Each sequence in the data set obtained in step S1 All wind speed data start at 00:00:00, so the wind speed data for model prediction should also start at 00:00:00. For any wind speed data starting at 00:00:00 and lasting no longer than 24 hours, the start and end times are , ,use Time period display to First, calculate the ten-minute average wind speed sequence for period T according to equations (1)-(6) in step S1 , is a one-dimensional vector with a length of N and N is less than 144, and is calculated using formula (9) Perform maximum value standardization to obtain standardized average wind speed . Then , and standardize, namely:
[0074] In the formula, To obtain the variable A function of the maximum value within a period; is the jth normalized daily variation pattern of wind speed; is the jth daily wind speed variation pattern; from the standardized daily wind speed variation pattern Intercept Wind speed during the period , The length of is also N, and One-to-one correspondence, and then according to the Euclidean distance from Find and The closest Wind speed change mode during the period, if , which corresponds to , and Perform denormalization to obtain the final wind speed daily variation pattern prediction data ,Right now:
[0075] In the formula, is the predicted diurnal variation pattern of normalized wind speed; is the average wind speed series; The average wind speed series The maximum value of .
[0076] Embodiment 2: A method for predicting bridge wind field change patterns based on an LSTM deep embedded clustering model, comprising: S1 Data preparation: Obtain the original information of wind speed and direction at the mountain canyon bridge site, perform data preprocessing, and generate a data set.
[0077] This embodiment uses a laser radar anemometer to measure wind field data, such as Figure 2 The laser radar anemometer is 80m away from the bridge tower on one side in the direction of the bridge axis, 40m away from the bridge tower in the direction perpendicular to the bridge axis, and has a vertical elevation difference of 4m with the bridge deck. The laser radar anemometer can measure wind speeds at all heights at the same time. In this embodiment, the wind speed at a height of 56m is selected as the measured wind speed. ,wind direction The bridge axis is 0° and clockwise is positive.
[0078] First, the wind speed Corresponding wind direction Split the data into 10-minute intervals. For any data, its length should be:
[0079] In the formula, For any data length, is the minimum sampling number requirement for the sampling device in ten minutes. Take 100, and remove the data with a length less than 100. The measured data should meet the following requirements:
[0080]
[0081] In the formula represents the average wind speed. is the standard deviation of the measured wind speed, Indicates the average value of the measured wind direction. is the standard deviation of the measured wind direction. The measured data that do not meet the above inequality conditions are eliminated. With wind direction Decompose along The bridge axis direction and The direction perpendicular to the bridge axis, that is
[0082]
[0083] In the formula and Respectively indicate along Direction and The wind speed component in the direction of the wind. Then calculate the ten-minute average wind speed and wind direction ,Right now:
[0084]
[0085] In the formula and Respectively indicate along Direction and The average value of the wind speed components in the direction.
[0086] The ten-minute average wind speed The data is divided into 24-hour units. The starting time of each data is 00:00:00. For any data, the following conditions must be met:
[0087] In the formula Indicates the length of each data in 24 hours. Data that does not meet the above formula is directly eliminated, and finally 130 data are obtained. One data represents a wind speed data sequence. A sequence contains 144 values, and one value represents the average wind speed for ten minutes. Here, a data of July 1, 2023 is used as an example, as shown in Table 1: Table 1 Average wind speed data on July 1, 2023
[0088] Then use Gaussian filter to obtain the change trend of each data and set parameters in Gaussian filter =8.
[0089] The above data changes to the data shown in Table 2: Table 2 Average wind speed data after using Gaussian filter
[0090] The maximum normalization method is used to standardize the data, that is:
[0091] In the formula represents the normalized wind speed, The ten-minute average wind speed. is the maximum value function.
[0092] The data in Table 2 above changes to the data in Table 3 below : Table 3 Standardized average wind speed data
[0093] Integrate all Forming a wind speed series data set , . Contains 130 sequences, each of which is a time series , represents the normalized wind speed data for one day, each sequence Contains 144 normalized values of the maximum ten-minute average wind speed, arranged in chronological order from front to back. Wind speed series data set , where each row represents a ,therefore There are 130 rows and 144 columns.
[0094] S2 model construction: This model is a fused LSTM deep embedded clustering model, which includes two parts: LSTM autoencoder and deep embedded clustering layer.
[0095] After getting the processed data set, we then use Python code to build a fused LSTM deep embedded clustering model, which includes an LSTM autoencoder layer and a deep embedded clustering layer. The specific structure of the LSTM autoencoder layer includes: Encoder: It is composed of LSTM layers and is used to extract input features. The potential variables of its output represent the key features of the input data, which can be expressed as:
[0096] In the formula represents the encoder, Represents the input of the model, which is a set of wind speed series data. represents the encoder model parameters of the LSTM autoencoder, is the output of the model, representing the wind speed series data set The feature set extracted from There are 130 sequences in are all one-dimensional vectors with a length of 32. express Features, and One-to-one correspondence; decoder: consists of LSTM layer and fully connected layer, used to reconstruct the original data from the latent variables, expressed as:
[0097] In the formula Represents the set of wind speed sequences reconstructed by the decoder by decoding the features extracted by the encoder. , Represents a decoder, is the input of the decoder model, representing the wind speed sequence data set The feature set extracted from Represents the decoder model parameters of the LSTM autoencoder; Latent space: the latent variables output by the encoder, which serve as the low-dimensional feature representation of the input data. The structure of the entire LSTM autoencoder is as follows Figure 3 shown.
[0098] The reconstruction loss and clustering loss are fused as the loss function of the model. The reconstruction loss is expressed as:
[0099] In the formula, represents the reconstruction loss function, , The code is as follows: reconstruction_loss =tf.reduce_mean(tf.square(X_train -reconstructed)) After constructing the LSTM autoencoder structure, a deep embedded clustering layer is constructed. The specific data processing process of the clustering layer is as follows: The 130 sequences in are divided into 8 clusters. The silhouette coefficient method is used to select the number of clusters. The silhouette coefficient method can measure the compactness and separation of each data point after clustering, and then evaluate the quality of the entire clustering. The calculation formula of the silhouette coefficient s(i) is as follows:
[0100] In the formula, represents the average distance from the data point 𝑖 to the nearest other clusters (i.e., the degree of separation between clusters), Represents the average distance from the data point 𝑖 to its cluster (i.e., the closeness within the cluster). The silhouette coefficients of all data points are added together to obtain the total silhouette coefficient S. The larger S is, the better the clustering effect is. Different cluster numbers are used for clustering and S is calculated. The cluster number when S is the largest is selected as the cluster number of this embodiment. The cluster center is Indicates that , use k-means to initialize the cluster center , the code is as follows: # Initialize cluster centers using K-means from sklearn.cluster import KMeans kmeans =KMeans(n_clusters=num_clusters) encoded_features = encoder.predict(X_train) # Use encoder to extract features initial_centers =kmeans.fit(encoded_features).cluster_centers_ calculate Soft assignment to cluster centers and auxiliary distribution , and the KL divergence composed of soft allocation and auxiliary allocation is used as the clustering loss. Use the softmax function to calculate Soft assignment to cluster centers ,Right now:
[0101] In the formula express The probability of belonging to the jth class, , is the temperature parameter, which is used to control the hardness or softness of clustering. When the temperature is high, the samples are evenly distributed to each cluster, while when the temperature is low, the samples are more concentrated in a certain cluster. Here, the temperature parameter is 0.993. Construct a high-confidence allocation as an auxiliary allocation , which can also be called target allocation, that is:
[0102] In the formula express The probability of belonging to the jth class is given by Calculated, with higher credibility, requires Getting closer ,when equal The clustering loss is zero when , so the clustering loss can be defined as the KL divergence between the soft assignment and the auxiliary assignment, that is:
[0103] In the formula Indicates calculating the KL divergence between P and Q.
[0104] Reconstructing the LSTM Autoencoder Loss With Deep Embedded Clustering Loss Fusion constitutes the loss function of the fused LSTM deep embedded clustering model, namely:
[0105] In the formula represents the weight of clustering loss, It has dynamic changing characteristics, namely:
[0106] In the formula is the weight for initializing clustering loss, taking 0.01, t is the iteration round, is the weight growth rate of clustering loss, which is 1.005.
[0107] S3 model training: Input data into the model instance and train until the loss function is stable.
[0108] A total loss is calculated in each round, and the LSTM network parameters need to be optimized by stochastic gradient descent with momentum. , and the cluster centers of the deep embedded clustering layer , so that the total loss of the next round is reduced. LSTM network parameters , The optimization calculation is as follows:
[0109]
[0110] In the formula is the encoder model parameter in the t-th round of LSTM autoencoder, is the encoder model parameter in the LSTM autoencoder of the t+1th round, is the decoder model parameter in the t-th round of LSTM autoencoder, is the decoder model parameter in the LSTM autoencoder of the t+1th round, represents the learning rate, which is 0.001. Indicates about The gradient operator of Indicates about The gradient operator.
[0111] Cluster centers of deep embedded clustering layer The optimization calculation is as follows:
[0112] In the formula is the deep embedded cluster center of the tth round, is the deep embedded clustering center of the t+1th round, represents the learning rate, which is 0.001. Indicates about The gradient operator.
[0113] The constructed model includes LSTM autoencoder layer and deep embedded clustering layer. Each part is built by code, including all the calculation processes described in step S2. Each part is connected together by code to form the trainer of the whole model. The overall model needs to input the wind speed sequence data set , as well as the selected number of clusters, and the total number of training rounds, the output is the encoder, decoder and cluster centers.
[0114] After building the overall model trainer, prepare the data , after adding a dimension, the shape becomes (130, 1, 144) to adapt it to the input of the model.
[0115] The data set is input into the constructed overall model trainer, the number of clusters is defined as 8, the total number of training rounds is 600, and the output is LSTM autoencoder (autoencoder), encoder (encoder), decoder (decoder) and cluster centers (cluster_centers).
[0116] When running the trainer of the above code, the total loss formula (16) described in S2 is calculated after each round of the data set passes through the model, and the model and cluster center are updated using formulas (18)-(20) to make the model and cluster center more accurate. The round ends and the next round is carried out. The training stops after the specified total number of rounds is completed. The total loss represents the error. As the number of rounds increases, the total loss decreases. The total loss of each round is printed. When the total loss reaches 600 rounds, the change in the loss is less than 0.0002, so it can be judged that the loss function value has stabilized at this time. The condition for stopping training here is that the change in the loss is less than 0.0002 (customized according to the specific project) or the set training rounds are completed. In either case, it is necessary to ensure that the number of rounds when the total loss is reduced to a stable state and no longer changes significantly is less than the specified total number of training rounds.
[0117] After training, we can get the LSTM autoencoder, encoder, decoder and cluster center. The encoder is used to train a dataset with 130 rows and 144 columns. Encode to get , There are 130 rows and 32 columns. Each line is from The features extracted from each row of .
[0118] The length of each cluster center is 32, 8 cluster centers .
[0119] S4 Daily Variation Pattern Library Construction: The cluster centers are reconstructed using the LSMT decoder to form a wind speed daily variation pattern library.
[0120] For the above 8 cluster centers Using optimized LSTM decoder Reconstruction is performed to form a wind speed daily variation pattern library, namely:
[0121] In the formula , represents the wind speed daily variation pattern library, which contains 8 wind speed sequences, each sequence is a one-dimensional vector of length 144, Represents the network parameters of the optimal LSTM decoder. There are 8 rows and 144 columns.
[0122] Will All sequences in , each sequence are all one-dimensional vectors with length less than 144, express The characteristics of each cluster center are calculated by formula (13) Probability , assign the class label with the maximum probability to the sequence: for example The largest, then Belongs to the jth category, thus dividing All sequences in The class to which it belongs.
[0123] The data belonging to the same class Corresponding Draw it in the same picture as the center of this class, such as Figure 4 As shown, there are eight subgraphs, representing eight classes respectively. is a set of wind speed series data The length of the sequence is 144. Is a collection A sequence of length 32, express Features, and One to one correspondence. Figure 4 In each sub-graph, the horizontal axis represents time, and the vertical axis represents the maximum normalized wind speed data. The thin dotted line in each sub-graph is the wind speed series data in this category. , the thick solid line is the diurnal variation pattern of wind speed of this category. So there are eight diurnal variation patterns of wind speed in total. The wind speed series data in each sub-graph are similar, and each category of diurnal variation pattern of wind speed can represent the diurnal variation law of this category of wind speed.
[0124] S5 Daily Variation Pattern Prediction: Match the known wind speed data with the corresponding area of the wind speed daily variation pattern, and find the most similar wind speed daily variation pattern as the prediction result.
[0125] Finally, the daily variation pattern prediction is carried out. The daily variation pattern prediction is carried out on the basis of steps S1-S4. Each sequence in the data set obtained in S1 All of them start from 00:00:00 (start at the same time), so the wind speed data that needs to be predicted by the model should also start from 00:00:00. For any section of wind speed data starting from 00:00:00 and lasting no longer than 24 hours (24 hours is for daily change prediction), the start and end times are , ,use Time period display to Time, here , First, calculate the ten-minute average wind speed sequence for period T according to equations (1)-(6) in step S1 , is a one-dimensional vector with a length of 84, and is calculated using formula (9) Perform maximum value standardization to obtain standardized average wind speed . Then , and standardize, namely:
[0126] In the formula To obtain the variable The diurnal variation pattern of wind speed after normalization is Intercept Wind speed during the period , The length of is also 84, and One-to-one correspondence, and then according to the Euclidean distance from Find and The closest Wind speed change mode during the period, if , which corresponds to , and Perform denormalization to obtain the final wind speed daily variation pattern prediction data ,Right now:
[0127] You only need to input the data (new_data_cut) of the prediction pattern to complete the prediction of the daily variation pattern of wind speed and draw the prediction results.
[0128] In order to observe whether the prediction result is accurate, this embodiment uses the complete wind speed daily variation pattern data of a certain day, intercepts the data of the 00:00:00-14:00:00 segment, the length of which is 84, and uses the above method to make a prediction. The prediction result is plotted on Figure 5 The intercepted data to be predicted and the predicted data are as follows: Example 1: To be predicted ; Prediction results .
[0129] In order to verify the prediction effect, The data from 00:00:00 to 14:00:00 are intercepted from the complete wind speed sequence of a known day. By comparing the predicted value with the actual value, the effect of the prediction can be known. Figure 5The solid line in the figure is the curve of the complete wind speed series data. Figure 5 The middle dotted line corresponds to the curve drawn based on the prediction results of the wind speed series. Figure 5 The horizontal axis represents time, and the vertical axis represents wind speed. Figure 5 It can be seen that the wind speed sequence from 14:00:00 to 24:00:00 that needs to be predicted is very consistent with the real sequence, the calculated MAPE is equal to 0.1138, and the prediction error is very small, indicating that this method is effective in predicting the daily variation pattern of wind speed.
[0130] The MAPE calculation formula is as follows:
[0131] In the formula, is the number of data points, is the forecast data, It is real data. The value of MAPE is between 0 and 1. The smaller the MAPE value, the smaller the prediction error.
[0132] This solution realizes the prediction of the daily variation pattern of the wind farm. It innovatively integrates LSTM into the deep embedded clustering model, which can well extract the characteristics of the time series, and innovatively proposes to use the softmax function to calculate the soft allocation in the deep embedded clustering layer to achieve end-to-end efficient training of the neural network. The clustering loss weight is introduced into the model loss function to dynamically adjust the proportion of the reconstruction task and the clustering task. It can be seen from the embodiments that the bridge wind farm daily variation pattern prediction method based on the LSTM deep embedded clustering model has a good effect in the prediction of the daily variation pattern of the wind farm.
Claims
1. A bridge wind field change pattern prediction method based on LSTM deep embedded clustering model, characterized in that: The following steps are involved: S1. Data preparation: Obtain wind speed and direction information at the mountain valley bridge site, perform preprocessing and generate a wind speed sequence data set; S2. Model construction: Construct a fused LSTM deep embedded clustering model, which includes an LSTM autoencoder layer and a deep embedded clustering layer; S3, model training: inputting the wind speed sequence data set into the LSTM deep embedded clustering model for training until the loss function is stable; S4. Construction of daily variation pattern library: Using the trained LSTM deep embedded clustering model, the cluster center is reconstructed through the decoder to form a wind speed daily variation pattern library and divide the wind speed sequence through soft allocation; S5. Daily variation pattern prediction: match the wind speed data to be predicted with the wind speed daily variation pattern library, and take the most similar wind speed daily variation pattern as the prediction result.
2. The method according to claim 1, characterized in that The S1 specifically includes: S11, dividing the wind speed and wind direction data at predetermined time intervals, and eliminating the measured data that does not meet the preset length and standard deviation; S12, decomposing the wind speed and wind direction into wind speed components in specific directions, and calculating the average wind speed and wind direction angle; S13, dividing the average wind speed according to the same time interval, using a filter to obtain the change trend, and performing standardization processing to form a wind speed sequence data set.
3. The method according to claim 1, characterized in that The S2 specifically includes: S21. Construct an LSTM autoencoder layer, including an encoder composed of an LSTM layer and a decoder composed of an LSTM layer and a fully connected layer; S22. Build a deep embedded clustering layer, classify the latent variables, use k-means to initialize the cluster centers, calculate the soft allocation through the softmax function and build the auxiliary allocation, and use the KL divergence of the two as the clustering loss; S23, merging the reconstruction loss and clustering loss into the LSTM deep embedded clustering model loss function; S24. Set the learning rate, initialize the clustering loss weight and the clustering loss weight growth rate.
4. The method according to claim 3, characterized in that The soft allocation is specifically expressed as: ; In the formula, express Belong to j The probability of the class, express Features , is the temperature parameter, k is the number of clusters; m is the serial number of the cluster center, ranging from 1 to k; c m is the cluster center of the mth class; c j is the cluster center of the jth class; The auxiliary allocation is specifically expressed as: ; In the formula, for Belong to j Probability of class; for Belong to m Probability of class; for Belong to m Probability of class; x is the sequence number of the wind speed sequence; n is the number of wind speed sequences.
5. The method according to claim 4, characterized in that The LSTM deep embedded clustering model loss function is specifically expressed as: ; In the formula, Represents the weight of clustering loss; ; To initialize the weight of clustering loss, t is the iteration round, is the weight growth rate of clustering loss; is the reconstruction loss function; is the clustering loss.
6. The method according to claim 1, characterized in that The S3 specifically includes: The wind speed sequence data set is input into the LSTM deep embedded clustering model, the number of clusters and the total number of training rounds are set, and training is performed; wherein, in each training, the wind speed sequence data set obtains the LSTM deep embedded clustering model loss function through the LSTM deep embedded clustering model, and optimizes the LSTM network parameters and the clustering center of the deep embedded clustering layer by the stochastic gradient descent method with momentum, and stops training after the set total number of rounds are executed, ensuring that the number of rounds when the total loss is reduced to a stable state is less than the specified total number of training rounds, so that the total loss has stabilized after all rounds of training are executed.
7. The method according to claim 1, characterized in that The S4 specifically includes: S41, using the optimized LSTM decoder to reconstruct the optimized cluster center to form a wind speed daily variation pattern library; S42. Calculate the probability that the wind speed sequence in the latent variable belongs to each cluster center through the soft allocation, and divide the wind speed sequence into the cluster center with the calculated maximum probability.
8. The method according to claim 1, characterized in that The S5 specifically includes: S51, for any section of wind speed data divided at the same time interval as in S1, obtain an average wind speed sequence according to step S1 and perform maximum value standardization to obtain a standardized average wind speed; S52. Standardize the wind speed daily variation pattern in the wind speed daily variation pattern library, extract the wind speed in the same time interval as the standardized average wind speed sequence to be predicted from the standardized wind speed daily variation pattern to obtain a standardized wind field pattern fragment, and find the wind speed variation pattern with the smallest Euclidean distance to the standardized average wind speed as the closest wind speed variation pattern in the same time interval from the standardized wind field pattern fragment and perform denormalization processing to obtain the final wind speed daily variation pattern prediction data.
9. The method according to claim 8, characterized in that The normalization of the diurnal variation pattern of wind speed is specifically expressed as: ; In the formula, To obtain the variable Maximum value function within a period; For the j A standardized daily variation pattern of wind speed; For the j Diurnal variation pattern of wind speed.
10. The method according to claim 8, characterized in that The final wind speed daily variation pattern prediction data is specifically expressed as: ; In the formula, is the predicted diurnal variation pattern of normalized wind speed; is the average wind speed series; The average wind speed series The maximum value of .
Citation Information
Patent Citations
Method for embedding and clustering depth self-coding based on Sliced-Waserstein distance
CN111178427A
LSTM fiber-optic gyroscope temperature compensation modeling method based on deep embedded clustering
CN111238462A
Wind speed interpolation method based on deep learning network
CN117874424A
Deep clustering method for bridge time sequence anomaly classification
CN118656752A
Hydroelectric equipment on-line monitoring and diagnosis system
CN119179919A