A method for predicting missing wave data based on machine learning
Through the machine learning-based wave detection data forecasting method, k-means clustering analysis and improved LSTM algorithm are used to solve the problem of inaccurate prediction model caused by the loss of wave data, and efficient forecasting of wave detection data is achieved.
Patent Information
- Application Number
- CN202310675335.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-06-08
AI Technical Summary
The lack of wave data caused by natural environment or man-made factors leads to inaccurate wave prediction model.
The machine learning-based wave detection data forecasting method is used to build a missing data forecast model by obtaining float information, preprocessing data, using k-means clustering analysis and improved LSTM algorithm to predict the missing data.
Effectively explore and analyze the changing patterns of ocean data, improve the accuracy of forecasting of wave missing data, and avoid the problem of unsatisfactory research results caused by missing data.
Smart Images

Figure CN116595442B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ocean wave missing data forecasting, and in particular to an ocean wave missing data forecasting method based on machine learning. Background Art
[0002] Ocean waves are one of the most common marine phenomena and are closely related to human life. Nearshore waves often destroy coastal dams and other buildings, hindering the development of coastal economies. Far offshore waves can also pose threats and disasters to maritime transportation, military activities, and fishing. Therefore, it is very important to understand the distribution characteristics and changing patterns of ocean waves. People can obtain authentic and reliable sea surface data through observation methods such as buoys and radars, and then restore the temporal and spatial distribution and changing process of ocean waves.
[0003] However, in the actual observation process, some wave data are often missing due to bad weather, instrument damage and other reasons, resulting in incomplete data. This has caused certain difficulties in studying the laws of wave changes, and also made the forecast results of many wave prediction models less than ideal.
[0004] The existing method for processing missing wave data is usually to eliminate the missing wave data or to take the average of the wave data measured by adjacent buoys to replace the missing wave data. Although the above processing method avoids the impact of missing wave data to a certain extent, it will cause the basic data of wave research to be incomplete and even destroy the wave characteristics contained in the original wave data, resulting in unsatisfactory results of subsequent research. Summary of the invention
[0005] The purpose of the present invention is to provide a method for predicting missing wave data based on machine learning, which effectively mines and autonomously analyzes the changing patterns of data through artificial intelligence forecasting methods, obtains a series of complex and nonlinear ocean characteristics through training and learning, and realizes the prediction of missing wave data, so as to solve the problem mentioned in the above background technology that the wave data measured by the buoy is missing due to natural environment or human factors, thereby causing the established prediction model to be inaccurate.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for predicting missing wave data based on machine learning, comprising:
[0007] S1. Acquire buoy information, including buoy position information and buoy data; the buoy position information is the latitude and longitude information of the buoy; the buoy data is the wave data measured by the buoy;
[0008] S2, preprocessing the buoy data;
[0009] S3. Classify the buoys using the k-means clustering analysis algorithm based on the buoy position information. Based on the preprocessed buoy data that has been classified, determine the training set and the validation set; use the first 70%-80% of the buoy data within the same cluster as the training set for the deep learning model, and the remaining as the test set for the model;
[0010] S4. Based on the training set and the validation set, use the improved LSTM method to construct a missing data prediction model;
[0011] S5. Determine the optimal buoy through the distance formula between two points, and substitute the data of the optimal buoy into the missing data prediction model to obtain the missing data prediction result;
[0012] S6. Use RMSE and MAPE to verify the missing data prediction result.
[0013] According to the above technical solution, the above preprocessing includes buoy abnormal data detection and buoy abnormal data correction. The accuracy of the buoy data is enhanced through preprocessing.
[0014] Due to natural and human factors, errors will occur in the monitoring data. If these abnormal data are not corrected, the subsequent work results will not be ideal. The buoy abnormal data detection includes the Grubbs criterion, the 3-sigma rule, etc., and the buoy abnormal data correction includes wavelet denoising method, maximum likelihood estimation method, etc.
[0015] According to the above technical solution, the buoy classification steps are as follows:
[0016] Determine the position information of all buoys within the predicted destination range to determine the buoy set;
[0017] Initialize k clustering centers, and calculate the Euclidean distance from each buoy to each clustering center;
[0018] Compare the distance of each object to each clustering center in turn, and assign the object to the cluster of the nearest clustering center to obtain k clusters.
[0019] Among them, the Euclidean distance calculation formula from each buoy to each clustering center is:
[0020]
[0021] Among them, X i represents the i-th buoy (1 ≤ i ≤ n), C j represents the j-th clustering center (1 ≤ j ≤ n), X it represents the t-th attribute of the i-th buoy, C jt represents the t-th attribute of the j-th clustering center, and n represents the number of buoys.
[0022] Among them, the k-means algorithm defines the cluster center as the mean value of all objects in the cluster in each dimension, and its calculation formula is as follows:
[0023]
[0024] C l represents the l-th clustering center, where 1 ≤ l ≤ k, k represents the number of clusters, and |S l | represents the number of buoys in the l-th cluster, and M i represents the i-th buoy in the l-th cluster, where 1 ≤ i ≤ |S l |.
[0025] According to the above technical solution, the improved LSTM method includes:
[0026] An input layer that inputs a training set into the improved LSTM method;
[0027] An expansion layer that expands the training set using an expansion model; since the data between waves in the same cluster is limited and uncertain, but a large amount of data is required during training to ensure the accuracy of the training result, it is necessary to expand the training set;
[0028] A screening layer that screens the data of the training set in the expansion layer using a correlation model and eliminates buoy data with a correlation less than α; there is some wave data with a low correlation with the wave data of the to-be-measured buoy in the expanded training set, and this data with a low correlation will reduce the accuracy of the training model, so it is necessary to screen the data in the expanded training set;
[0029] An LSTM layer that is used to learn the data feature relationship of the training set in the screening layer;
[0030] An output layer that is used for the predicted value of missing wave measurements.
[0031] According to the above technical solution, the steps for establishing the expansion model are as follows:
[0032] Determine the Q table according to the training set;
[0033] Determine the expansion times, reward function, attenuation value, and learning rate according to the actual situation of the waves;
[0034] Update the Q table using the Q-learning method until the expansion times are reached and stop updating the Q table;
[0035] Among them, the final result of the Q table update is the expanded data set.
[0036] Use the Q-learning method to expand the training set so that the expanded data can contain the distinct features of the waves.
[0037] According to the above technical solution, the relevance model is as follows:
[0038]
[0039] where n is the number of buoys, x i is the data of the buoy closest to the clustering center, is the average data of the buoy closest to the clustering center within time T, u i is the buoy data to be evaluated, is the average data of the buoy to be evaluated within time T.
[0040] The LSTM layer has two transmission states, one C t , and one h t ; C t changes very slowly during network propagation and represents a long-term and relatively stable piece of information; while h t changes very quickly during network propagation and represents short-term local information; each layer of the LSTM network needs to update the cell state C t representing long-term memory based on the current input x t and the short-term memory h t from the previous moment. The update is achieved through three gate structures, where:
[0041] Forget gate:
[0042] f t = σ(W f ·[h t-1 , x t +b f );
[0043] The σ layer sequentially uses h t-1 and x t as inputs and outputs f t , which is the probability of forgetting the cell state of the previous layer. The sigmoid activation function outputs a value ranging from 0 to 1. If the value is 0, the information of the previous state is completely deleted; if it is 1, the information is completely retained. W f represents the weight matrix of the forget gate, and b f represents the bias matrix of the forget gate;
[0044] Input gate:
[0045]
[0046] h t-1 and x t are activated together in the σ layer to obtain the new memory weight i t, determine which information needs to be updated, which corresponds to the part selected to be forgotten in the forget gate; then combine h t-1 with x t to obtain a new candidate vector through the tanh activation function f t and multiply to achieve the forgetting of long-term memory, i t and are combined using the Hadamard product operator to achieve the update of c t-1 to ; W i represents the weight matrix of the input gate, b i represents the bias matrix of the input gate, W c represents the weight matrix of the output gate, b c represents the bias matrix of the output gate;
[0047] Output gate:
[0048]
[0049] The σ layer obtains the weight O t , determining which part of the output unit state. After processing the cell state c t through the tanh activation function, the state value is mapped between -1 and 1, and then it is combined with the O t output by the σ layer using the Hadamard product operator to obtain the updated short-term memory h t , where W λ represents the weight assigned to each layer, b θ represents the bias matrix assigned to each layer, λ ∈ {f, i, c, o}, θ ∈ {f, i, c, o}.
[0050] According to the above technical solution, the coordinates of all the buoys within the same cluster as the buoy to be measured are respectively substituted into the distance formula between two points together with the coordinates of the buoy to be measured, and screening is performed to obtain the shortest distance from the buoy to be measured; among them, the buoy with the shortest distance from the buoy to be measured is the optimal buoy.
[0051] According to the above technical solution, the RMSE is the root mean square error used to characterize the deviation between the simulation result and the measured value, and is relatively sensitive to extreme values. Its calculation formula:
[0052]
[0053] where N represents the number of elements in the validation set, y i represents the predicted value, Y i represents the true value;
[0054] The MAPE is the mean absolute percentage error, which represents the degree of deviation of the predicted value from the measured value in percentage. Its calculation formula:
[0055]
[0056] where N represents the number of elements in the validation set, y i represents the predicted value, and Y i represents the true value.
[0057] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0058] 1. The present invention adopts an improved LSTM algorithm, including an input layer, an expansion layer, a screening layer, an LSTM layer, and an output layer. The expansion layer is used to expand the data of the training set, which to a certain extent avoids the problem of low accuracy of the training result caused by less data. The screening layer uses a correlation model to screen the data of the training set in the expansion layer, and eliminates the buoy data with a correlation less than α, so that the data in the expanded training set has distinct ocean wave characteristics, and the accuracy of the training model is improved by using the improved LSTM algorithm;
[0059] 2. The present invention effectively mines and autonomously analyzes the variation law of data through the improved LSTM algorithm, obtains a series of complex and non-linear ocean characteristics through training and learning, realizes the prediction of missing ocean wave data, and to a certain extent avoids the problem that the subsequent research results are not ideal due to the missing ocean wave data. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0061] Figure 1 is a schematic flow chart of a method for predicting missing ocean wave data based on machine learning according to the present invention;
[0062] Figure 2 is a schematic flow chart of the k-means clustering analysis algorithm in the present invention;
[0063] Figure 3 is the structure diagram of the LSTM model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0065] Please refer to Figures 1 - 3 , taking the wave data obtained from 8 buoys in the North Pacific as an example, classifying the buoys and using the buoy data within the same cluster as the training set and the validation set, establishing a missing data prediction model using an improved LSTM algorithm, and performing wave missing data prediction, which shows that a method for predicting wave missing data based on machine learning of the present invention includes:
[0066] S1. Obtain buoy information, including buoy position information and buoy data. Specifically, obtain the wave height data of 8 buoys in the North Pacific from 2019 to 2020, with a time interval of 1 hour, and the time is selected from 0:00 on January 1, 2019 to 23:00 on December 31, 2019. The position information of its buoy information is shown in Table 1:
[0067] Table 1
[0068]
[0069] S2. Preprocess the buoy data. Specifically, query the critical value of the corresponding Grubbs discrimination method according to the actual situation of the buoy. When |G1|≥G0(n, μ) or |G n |≥G0(n, μ), determine that the buoy data is an outlier, and use the wavelet denoising analysis method to correct the outlier;
[0070] Among them, G1 represents the Grubbs statistic for judging whether the minimum value in the observed data is an outlier, and G n represents the Grubbs statistic for judging whether the maximum value in the observed data is an outlier, and n represents the n buoy data in the buoy data L;
[0071] The steps of correcting the outlier by the wavelet denoising analysis method include: selecting a suitable wavelet basis function, the number of decomposition layers and related parameters according to the actual situation of the buoy, and further determining the threshold standard for wavelet reconstruction, so as to remove the error noise and correct the buoy data.
[0072] S3. Classify the buoys using the k-means clustering analysis algorithm based on the buoy position information, and determine the training set and the validation set based on the preprocessed buoy data that has been classified. Specifically, input the distance data of the 8 stations after the above preprocessing into the clustering analysis model with K = 3 to obtain 3 clusters, and the data of the 8 buoy stations are divided into 3 categories. Classify according to the process of the k-means clustering analysis algorithm of Figure 2 ;
[0073] Among them, the steps for buoy classification are as follows:
[0074] Determine the position information of all buoys within the predicted destination range to determine the buoy set;
[0075] Initialize k clustering centers and calculate the Euclidean distance from each buoy to each clustering center;
[0076] Compare the distance of each object to each clustering center in turn, and assign the object to the cluster of the nearest clustering center to obtain k clusters;
[0077] Among them, the calculation formula for the Euclidean distance from each buoy to each clustering center is:
[0078]
[0079] Among them, X i represents the i-th buoy (1 ≤ i ≤ n), C j represents the j-th clustering center (1 ≤ j ≤ n), X it represents the t-th attribute of the i-th buoy, and C jt represents the t-th attribute of the j-th clustering center, and n represents the number of buoys.
[0080] S4. Based on the training set and the validation set, use the improved LSTM method to construct a missing data prediction model;
[0081] Among them, the improved LSTM method includes:
[0082] Input layer, input the training set into the improved LSTM method;
[0083] Expansion layer, expand the training set using the expansion model. The steps for establishing the expansion model are as follows: Determine the Q table according to the training set; Determine the expansion times, reward function, decay value, and learning rate according to the actual situation of the sea waves; Update the Q table using the Q learning method until the expansion times are reached and stop updating the Q table; Among them, the final result of Q table update is the expanded data set;
[0084] Screening layer, screen the data of the training set in the expansion layer using the correlation model, and eliminate the buoy data with a correlation less than α; The correlation model is:
[0085]
[0086] Among them, n is the number of buoys, and x i is the data of the buoy closest to the clustering center, is the average data of the buoy closest to the clustering center within time T, u i is the data of the buoy to be evaluated, is the average data of the buoy to be evaluated within time T;
[0087] The LSTM layer is used to learn the data feature relationship of the training set of the screening layer;
[0088] The output layer is used to output the predicted value of the missing wave measurement.
[0089] S5. Substitute the coordinates of all the buoys within the same cluster as the buoy 51101 and the coordinates of the buoy 51101 into the distance formula between two points, and perform screening to obtain the shortest distance from the buoy 51101; among them, the buoy with the shortest distance from the buoy 51101 is the optimal buoy, and substitute the data of the optimal buoy into the missing data prediction model to obtain the missing data prediction result.
[0090] S6. Use RMSE and MAPE to test the missing data prediction result. The RMSE is the root mean square error, which is used to describe the deviation between the simulation result and the measured value, and is more sensitive to extreme values. Its calculation formula:
[0091]
[0092] Among them, N represents the number of elements in the validation set, y i represents the predicted value, and Y i represents the true value;
[0093] The MAPE is the mean absolute percentage error, which is used to represent the degree of deviation of the predicted value from the measured value in percentage. Its calculation formula:
[0094]
[0095] Among them, N represents the number of elements in the validation set, y i represents the predicted value, and Y i represents the true value.
[0096] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0097] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for predicting missing wave data based on machine learning, characterized in that, It includes the following steps: Obtain buoy information, where the buoy information includes buoy position information and buoy data; Preprocess the buoy data; Classify the buoys using the k-means clustering analysis algorithm based on the buoy position information to generate classified preprocessed buoy data, and determine a training set and a validation set based on the classified preprocessed buoy data; Based on the training set and the validation set, construct a missing data prediction model using an improved LSTM neural network; The improved LSTM neural network includes: An input layer that inputs the training set into the improved LSTM neural network; An expansion layer that expands the training set using an expansion model; A screening layer that screens the data of the training set in the expansion layer using a correlation model and eliminates buoy data with a correlation less than α; An LSTM layer for learning the data feature relationships of the training set in the screening layer; An output layer for outputting the predicted value of missing ocean wave data; Determine the optimal buoy through the distance formula between two points, and substitute the data of the optimal buoy into the missing data prediction model to obtain the missing data prediction result; Use RMSE and MAPE to test the missing data prediction result.
2. The method for predicting missing wave data based on machine learning according to claim 1, characterized in that: The preprocessing includes buoy abnormal data detection and buoy abnormal data correction.
3. A method for predicting missing wave data based on machine learning according to claim 1, characterized in that, The steps for classifying the buoys are as follows: Determine the position information of all buoys within the predicted destination range to determine a buoy set; Initialize k clustering centers and calculate the Euclidean distance from each buoy to each clustering center; Compare the distance of each object to each clustering center in turn, and assign the object to the cluster of the clustering center with the closest distance to obtain k clusters; Among them, the formula for calculating the Euclidean distance from each buoy to each clustering center is: Among them, X i represents the i-th buoy, where 1 ≤ i ≤ n, and C j represents the j-th cluster center, where 1 ≤ j ≤ n, and X it represents the t-th attribute of the i-th buoy, and C jt represents the t-th attribute of the j-th cluster center, and n represents the number of buoys.
4. A method for predicting missing wave data based on machine learning according to claim 1, characterized in that The steps for establishing the expansion model are as follows: Determine the Q table based on the training set; Determine the expansion times, reward function, decay value, and learning rate according to the actual situation of ocean waves; Update the Q table using the Q-learning method until the expansion times are reached and stop updating the Q table; Among them, the final result of Q table update is the expanded data set.
5. A method for predicting missing wave data based on machine learning according to claim 1, characterized in that, The correlation model is: Among them, n is the number of buoys, and x i is the data of the buoy closest to the clustering center, is the average data of the buoy closest to the clustering center within time T, u i is the data of the buoy to be evaluated, is the average data of the buoy to be evaluated within time T.
6. The method for predicting missing wave data based on machine learning according to claim 1, characterized in that The specific process of determining the optimal buoy through the distance formula between two points includes: substituting the coordinates of all buoys in the same cluster as the buoy to be measured and the coordinates of the buoy to be measured into the distance formula between two points, and performing screening to obtain the shortest distance from the buoy to be measured; among them, the buoy with the shortest distance from the buoy to be measured is the optimal buoy.
7. A method for predicting missing wave data based on machine learning according to claim 1, characterized in that: The RMSE is the root mean square error, which is used to describe the deviation between the simulation result and the measured value, and its calculation formula: Among them, N represents the number of elements in the validation set, y i represents the predicted value, and Y i represents the true value; The MAPE is the mean absolute percentage error, which represents the degree of deviation of the predicted value from the measured value in percentage, and its calculation formula: Among them, N represents the number of elements in the validation set, y i represents the predicted value, and Y i represents the true value.
Citation Information
Patent Citations
Wave missing measurement data forecasting method based on bivariate long-short term memory algorithm
CN116245018A