Ocean buoy data filling method based on space-time neural network

By combining the graph convolution network and the spatiotemporal neural network of the gated recurrent unit, the problem of insufficient spatiotemporal correlation modeling of ocean buoy data is solved, and high-precision data filling and real-time data processing are realized, which is suitable for marine monitoring networks.

CN120508757APending Publication Date: 2025-08-19SOUTHEAST UNIV

Patent Information

Application Number
CN202510622141.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The prior art fails to effectively capture the spatial and temporal correlation between buoys in the data processing of marine buoys, resulting in insufficient data filling accuracy, affecting subsequent analysis and prediction effects.

Method used

A spatiotemporal neural network combining graph convolutional network (GCN) and gated recurrent unit (GRU) is used to achieve high-precision filling of marine buoy data through data preprocessing, missing pattern analysis and sample construction, spatiotemporal feature modeling and model training.

Benefits of technology

It significantly improves the accuracy of ocean buoy data filling and the generalization ability of the model. It is suitable for real-time data processing of large-scale marine monitoring networks and has high engineering application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508757A_ABST
    Figure CN120508757A_ABST
Patent Text Reader

Abstract

The invention discloses an ocean buoy data filling method based on a space-time neural network. In order to solve the problems of insufficient spatial-temporal dynamic correlation modeling and low filling precision in the prior art, high-precision reconstruction of ocean buoy missing data is realized by coupling a graph convolutional network GCN and a gating cycle unit GRU in combination with spatial-temporal feature modeling. The method comprises the steps of data preprocessing; analyzing a data missing mode, and constructing a data sample by utilizing double mask matrix construction and a time window division strategy; based on multi-scale spatial feature extraction of GCN and long-term time-dependent modeling of GRU, model performance is improved through hyper-parameter optimization and Bayesian search; and carrying out model training, and evaluating the filling effect by adopting indexes such as mean square error. According to the method, the reconstruction precision of the space-time correlation missing value is remarkably improved, the generalization ability for a real missing mode is enhanced, and the method is suitable for real-time data processing of a large-scale ocean monitoring network and has high engineering application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of ocean monitoring data processing, and in particular relates to an ocean buoy data filling method based on a spatiotemporal neural network, which is particularly suitable for high-precision reconstruction of spatiotemporal correlation missing values in ocean buoy monitoring data. Background Art

[0002] Ocean buoys are important tools for obtaining marine environmental parameters, providing critical information for navigation safety, marine engineering design, and coastal disaster warnings. However, due to the complex marine environment, buoy data often suffers from data loss due to equipment failure or communication interruptions, and the loss patterns exhibit spatiotemporal coupling.

[0003] To conduct buoy data statistics and prediction tasks, it is necessary to fill in missing buoy data. Traditional filling methods (such as linear interpolation and K-nearest neighbor) rely only on single-dimensional features of time or space, and cannot effectively capture the dynamic spatiotemporal correlations between multiple buoys. This leads to insufficient filling accuracy and poor prediction results.

[0004] With the development of artificial intelligence technology, deep learning has shown great potential in data prediction. Existing deep learning models (such as CNN and LSTM) can process spatial or temporal features separately, but they do not fully integrate the synergy between the two and have limited ability to model non-Euclidean spatial relationships between buoys. Graph convolutional networks (GCNs), as an emerging deep learning architecture, can effectively process graph-structured data and mine complex relationships between nodes. The gated recurrent unit (GRU), a recurrent neural network (RNN) architecture designed for processing sequential data, has shown excellent performance in time series prediction tasks. GCNs are combined with GRUs for wave data prediction.

[0005] The present application differs from the prior art in the following ways:

[0006] Technical comparison with patent CN115935139A "A spatial field interpolation method for ocean observation data"

[0007] Patent CN115935139A proposes a spatial field interpolation method for ocean observation data to interpolate missing buoy data, mainly targeting temperature and salinity fields and ocean current structure data. This patent mentions the spatial field, but in its specific implementation, it does not focus on the integration of spatial information and spatiotemporal characteristics. It only fills in some missing data to construct complete spatial distribution data.

[0008] This patent proposes a method for filling ocean buoy data based on spatiotemporal deep learning, focusing on using deep learning technology to extract the spatiotemporal features between each buoy to obtain more reliable filling results and build a high-quality data set for use in various subsequent research work.

[0009] There are essential differences between the two in terms of technical paths and research objectives.

[0010] Patent CN115935139A uses TCN and Attention in its specific implementation. The output of TCN is input into Attention, and the features that have an important impact on the predicted interpolation value are extracted, and then the final output is obtained through multiple fully connected layers.

[0011] This patent integrates the GCN and GRU models, fully considering the spatial correlation and temporal evolution characteristics between buoy nodes, and effectively models the complex spatiotemporal dependencies in ocean observation data. This design not only improves the accuracy of data interpolation, but also provides new technical ideas and solutions for dealing with missing data in distributed sensor networks.

[0012] There are essential differences between the two in system design and specific algorithms.

[0013] Patent CN115622684B does not clearly explain the specific processing mechanism for missing values, especially how to handle missing values in the training set construction process to adapt to the model input requirements, and how to select evaluation points in the test set to objectively evaluate model performance.

[0014] This patent introduces the missing value processing part in detail: first, the original time series is interpolated, and then two mask matrices are used to mask the data set, so that the model in the training set can learn the real data, and the artificial missing situation is consistent with the real missing situation.

[0015] There are essential differences in the technical solutions between the two. Summary of the Invention

[0016] In response to the problems in the existing ocean buoy data filling methods that are insufficient in modeling spatiotemporal dynamic correlations and have low filling accuracy, the present invention aims to provide an ocean buoy data filling method based on spatiotemporal neural networks, which achieves robust reconstruction of missing data by coupling a graph convolutional network (GCN) with a gated recurrent unit (GRU).

[0017] To achieve the above objectives, the present invention provides a method for filling ocean buoy data based on a spatiotemporal neural network, which is characterized by comprising the following steps:

[0018] S1: data preprocessing;

[0019] Collect raw monitoring data from multiple buoys, identify and remove noise and outliers by setting physical thresholds and statistical standards, and smooth the data based on the sampling frequency;

[0020] S2: Missing pattern analysis and sample construction;

[0021] The missing patterns are classified into spatial block missing, temporal block missing, and general missing. The time series is divided into fixed windows using a time window partitioning strategy. A mask matrix is constructed to simulate real missing scenarios and generate the occlusion data required for training and evaluation.

[0022] S3: Spatiotemporal feature modeling and optimization;

[0023] The buoy network is mapped into a graph structure, an adjacency matrix is constructed based on geographic distance or Pearson correlation coefficient, and multi-scale spatial features are extracted through a multi-layer GCN. GRU is used to capture long-term dependencies in time series. Model performance is improved through hyperparameter optimization and Bayesian search.

[0024] S4: Model training and evaluation;

[0025] The model is trained on the preprocessed dataset and the filling error is evaluated on an independent test set.

[0026] As a further improvement of the present invention, in step S2, the fixed size of the time window is 168 time steps, and 56 time steps of complete data are retained before and after the padding segment.

[0027] As a further improvement of the present invention, in step S2, the mask matrix includes a training mask mask and an evaluation mask eval_mask, wherein the training mask marks the occluded data as 0, and the evaluation mask marks the occluded data as 1.

[0028] As a further improvement of the present invention, the adjacency matrix calculates the dynamic spatial correlation between buoys through a Gaussian kernel function or an exponential decay function.

[0029] As a further improvement of the present invention, in step S3, the update gate and reset gate of the GRU capture the long-term dependency of the time series and combine the input gate and forget gate functions to update the hidden state by the following formula:

[0030] Update Gate:

[0031] z t =σ(W z [x t ,h t-1 ]+b t )

[0032] Reset Gate:

[0033] r t =σ(W r [x t ,ht-1 ]+b r )

[0034] Candidate Hidden State:

[0035] c t =tanh(W h [x t ,r t ⊙h t-1 ]+b h )

[0036] Hide status update:

[0037]

[0038] As a further improvement of the present invention, in step S4, the loss function is defined based on the mean square error MSE as:

[0039]

[0040] As a further improvement of the present invention, in step S3, the hyperparameter optimization includes adjusting the learning rate, batch size, hidden layer dimension and regularization coefficient.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] 1. By coupling GCN and GRU, the dynamic spatial correlation and temporal evolution of the buoy network are modeled simultaneously, significantly improving the filling accuracy.

[0043] 2. The proposed double masking mechanism and time window division strategy enhance the model's generalization ability for real missing patterns.

[0044] 3. It is suitable for large-scale ocean monitoring networks, supports real-time data processing, and has high engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is an overall flow chart of the method of the present invention;

[0046] Figure 2 This is the Pearson correlation coefficient matrix diagram of some buoy data;

[0047] Figure 3 Schematic diagram of data organization and processing using mask matrix;

[0048] Figure 4 Schematic diagram of the spatiotemporal neural network architecture;

[0049] Figure 5Schematic diagram of the model prediction effect when there are 12 (left) and 5 (right) buoys around;

[0050] Figure 6 Schematic diagram comparing the prediction effects of the present invention and the BRITS model. DETAILED DESCRIPTION

[0051] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings.

[0052] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0053] A method for filling ocean buoy data based on spatiotemporal neural network, such as Figure 1 Shown, including:

[0054] Step S1: data preprocessing;

[0055] First, raw data, including but not limited to temperature, wind speed, and significant wave height, is collected from multiple buoys deployed across a vast ocean area. Due to the complex and ever-changing ocean environment, the collected data may contain noise or outliers. Therefore, rigorous preprocessing of the raw data is required.

[0056] Specifically, it is necessary to determine an appropriate sampling frequency, identify and remove abnormal data points by setting reasonable physical thresholds and statistical standards, and perform denoising on the raw data if necessary. This process not only improves data quality but also provides a reliable foundation for subsequent analysis.

[0057] A set of characteristic wave height data from adjacent buoys is selected, and the time average value every three hours is calculated to construct a relatively stable time series to meet the interpolation needs. Subsequently, the Pearson Correlation Coefficient (Pearson Correlation Coefficient) between the characteristic wave height time series of each buoy is calculated to quantify the strength and direction of the linear correlation between these time series. Figure 2 This coefficient is defined as the ratio of the covariance of two variables to the product of their standard deviations, which is used as the adjacency matrix of the graph convolutional network (GCN) input.

[0058] In order to simplify the data processing process, only the degree of linear correlation between these time series is evaluated instead of considering more complex nonlinear relationships.

[0059] Step S2: Missing pattern analysis and sample construction:

[0060] In data analysis, common missing patterns include missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). Obviously, the missingness of ocean buoy data belongs to missing at random.

[0061] For the specific buoy data collected, missing patterns can be divided into three types: spatial block missing, temporal block missing, and general block missing. Spatial block missing refers to missing data from multiple buoys at the same time; temporal block missing refers to missing data from the same buoy at multiple consecutive time steps; and general block missing refers to missing data other than these two types.

[0062] By statistically analyzing the length and frequency of missing data segments, a corresponding probability distribution model is constructed. This step is intended to ensure that the generated mask matrix can more accurately simulate the actual situation of missing data.

[0063] Specifically, in the process of making the mask matrix, a method of randomly blocking consecutive time steps is adopted, and the specific number of blocked time steps is determined according to the probability distribution model established above.

[0064] When constructing the data samples required for this spatiotemporal filling model, an innovative data organization and processing method was proposed to enhance the model's learning ability and evaluation accuracy. This method generates two mask matrices through two random occlusions and incorporates a time window mechanism to construct training samples. The specific steps are as follows:

[0065] 1. Random occlusion and mask matrix generation

[0066] In order to simulate the data missing situation in the real world and enable the model to learn more robust feature representation, the present invention only randomly masks the non-missing data in the input data twice, generating two mask matrices respectively:

[0067] Mask: This matrix indicates which data points are blocked and the model needs to learn these blocked data. In this matrix, the position marked as 1 represents that the data here is not blocked, and the position marked as 0 represents that it is blocked, which is recorded as M1. Figure 3 As shown in (a).

[0068] eval_mask: This matrix indicates which data points are blocked and used to evaluate the performance of the model. In this matrix, the position marked with 0 represents that the data here is not blocked; the position marked with 1 represents that it is blocked and needs to be evaluated, which is recorded as M2. Figure 3 (b) shown.

[0069] That is, the model learns based on the occluded data points in the mask during training, and uses the specific occluded points in eval_mask to verify the model's filling ability during the evaluation phase.

[0070] Assume that the original data is X and the obscured data is X1, then the obscuration process is as follows:

[0071] X1=[X⊙M1⊙(1-M2)]

[0072] Where ⊙ represents the bitwise multiplication of matrices.

[0073] 2. Time window division

[0074] In order to effectively capture the dynamic changes in the time series, the entire time series is divided into multiple fixed-size time windows. Each time window is used as an independent training sample, which is defined as follows:

[0075] Time window size: Let the time window size be T. Then each sample contains data from t to t+T. That is, each sample dimension is [T,N,n], that is, each sample contains T time steps, N nodes, and each node has n variables. The purpose of this setting is to allow the model to learn how to recover occluded data points while maintaining an understanding of the complete data structure.

[0076] When addressing the issues of defining time window size and data set partitioning, this paper specifically addresses short-term missing data. By performing a precise analysis of historical data, we define missing data within a timeframe of one week (i.e., 56 time steps) as short-term missing data. Based on this definition, the time window set by this paper includes the 56 unobstructed time steps before and after the infill segment, ensuring the data integrity of these time periods to provide data support for the infill segment.

[0077] The size of the intermediate imputation segments is determined using the aforementioned probability distribution sampling method. To standardize and simplify the processing, the time window size is fixed at 168 time steps. This approach not only effectively handles short-term missing data but also ensures the flexibility and accuracy of the imputation model in practical applications. This strategy optimizes the performance of the imputation model and improves the prediction accuracy for missing data while maintaining data continuity and integrity. Furthermore, the standardized time window size facilitates application and comparison across different datasets, enhancing the model's versatility and practicality.

[0078] Step S3: Spatiotemporal feature modeling and optimization:

[0079] Spatiotemporal neural network architectures such as Figure 4 shown.

[0080] 1. Spatial Modeling:

[0081] A graph convolutional network (GCN) is a deep learning model specifically designed to process graph-structured data. Unlike traditional convolutional neural networks (CNNs), GCNs can directly operate on data in non-Euclidean spaces, updating node feature representations by aggregating information about nodes and their neighbors, thereby capturing complex structural patterns in the graph.

[0082] The buoy network is mapped into a graph structure, where nodes represent the spatial positions of the buoys and edge weights reflect the dynamic spatial correlation between the buoys. In order to establish a spatial adjacency matrix that describes the relationship between the buoys, it is first necessary to determine the relative distance and interaction strength between the buoys. Commonly used calculation methods include Gaussian kernel functions or exponential decay functions based on geographic distance.

[0083] After constructing the adjacency matrix, a graph convolutional network (GCN) is used to extract multi-scale spatial features. GCN aggregates information about adjacent nodes at each layer through a message passing mechanism, thereby extracting local and global spatial patterns. Specifically, the node representation of each layer is updated using the following formula:

[0084]

[0085] Where, Right now Add self-loops (identity matrix I) to the adjacency matrix A; for The degree matrix (diagonal matrix); H (l) is the node feature matrix of the lth layer, H (0) = X(input features); W (l) is the trainable parameter matrix of the current layer convolution transformation; σ is the nonlinear activation function (such as ReLU); H (l+1) is the eigenvector matrix after one convolution operation.

[0086] 2. Time Modeling:

[0087] A time series is constructed for the historical monitoring data of the buoy, and the gated recurrent unit (GRU) is used to extract the long-term dependencies in the time dimension to capture the temporal change patterns driven by environmental factors such as tides and ocean currents.

[0088] Traditional recurrent neural networks (RNNs) are prone to vanishing or exploding gradients when processing long time series, making them ineffective in capturing long-term dependencies. Long Short-Term Memory (LSTM) networks are an effective architecture for addressing this vanishing gradient problem. They control the flow of information through input, forget, and output gates. While LSTMs excel at capturing complex temporal dependencies, they are relatively complex, requiring more parameters and requiring more computation.

[0089] Compared to long short-term memory (LSTM) networks, the GRU reduces the number of gates and merges the functions of the input gate and forget gate in the LSTM into the update gate, simplifying the model structure. This change not only significantly reduces the number of model parameters and computational overhead, but also maintains performance comparable to or even better than that of the LSTM in various tasks. In addition, due to its low complexity and high training efficiency, the GRU is particularly suitable for large-scale datasets and application scenarios with high real-time requirements, demonstrating wide applicability and practicality. The specific state update formula of the GRU is as follows:

[0090] Update Gate:

[0091] z t =σ(W z [x t ,h t-1 ]+b t )

[0092] z t : The output of the update gate, at time step t, determines how much of the previous state to keep.

[0093] W z : Update the gate weight matrix.

[0094] h t-1 : The hidden state at the previous moment.

[0095] x t : Input for the current time step.

[0096] b t : Update gate bias.

[0097] σ: Sigmoid activation function.

[0098] Reset Gate:

[0099] r t =σ(W r [x t ,h t-1 ]+br )

[0100] The parameters have the same meaning as above.

[0101] Candidate Hidden State:

[0102] c t =tanh(W h [x t ,r t ⊙h r-1 ]+b h )

[0103] c t : Candidate hidden state, ready to be mixed with the current hidden state.

[0104] tanh: Hyperbolic tangent activation function.

[0105] Hide status update:

[0106] h t =(1-z t )⊙h t-1 +z t ⊙c t

[0107] h t : The final hidden state at time step t.

[0108] 3. Model optimization:

[0109] After the architecture is designed, the next step is to select and optimize hyperparameters. These include, but are not limited to, the learning rate, batch size, number of epochs, hidden layer dimensions, and regularization coefficients. The selection of these parameters is crucial to model performance. If necessary, Bayesian optimization methods can be used to systematically explore the optimal hyperparameter combination.

[0110] Step S4: Model training and evaluation:

[0111] After completing the above steps, the model training is performed using the pre-processed dataset. During this process, the dynamics of the loss function should be closely monitored to avoid overfitting or underfitting. The loss function is defined based on the mean square error (MSE) as

[0112]

[0113] During the evaluation phase, we use the mask matrix M2 to select specific data points for evaluation and calculate the filling error of the model on these occluded points to evaluate the model performance.

[0114] In addition, the model parameters need to be saved periodically, and the training plan needs to be optimized and adjusted based on the performance feedback on the validation set, including timely adjustment of the learning rate or implementation of early stopping strategies to ensure the effectiveness and generalization ability of the model.

[0115] Furthermore, after the model training reaches the expected goal, a comprehensive evaluation should be conducted on an independent test set to verify the actual performance and generalization ability of the model. For this filling task, the evaluation indicators can be selected as correlation coefficient (r), determination coefficient (R 2 ), mean absolute percentage error (MAPE), and root mean square error (RMSE). At the same time, relevant effect graphs are drawn, such as a time series comparison graph of the imputed value and the true value, and a residual distribution graph, to intuitively demonstrate the performance of the model and its deviation at different data points.

[0116] Example 1:

[0117] The data of a buoy in 2024 is blocked and filled, and the spatial information of 12 and 5 surrounding buoys is added respectively. The effect is as follows Figure 4 As shown in the figure, the correlation coefficient, determination coefficient, MAPE, and RMSE of the predicted values all reach good levels compared to the real values. This shows that the model proposed in this invention works well, proving the effectiveness of spatial information in filling in buoy data.

[0118] Example 2:

[0119] The data of the same buoy in 2024 are blocked and filled differently from those in Example 1, providing spatial information of 12 surrounding buoys. The model effect of the present invention is compared with the Bidirectional Recurrent Imputation for Time Series (BRITS) model that does not fully consider spatial information. The effect is as follows: Figure 5 and Figure 6 As shown in the figure, the model proposed in this paper outperforms BRITS, with the correlation coefficient, determination coefficient, MAPE, and RMSE increased by 2.22%, 5.59%, 24.33%, and 29.89%, respectively. This further proves the effectiveness of spatial information in filling in buoy data.

[0120] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for filling ocean buoy data based on spatiotemporal neural network, characterized in that: The following steps are involved: S1: data preprocessing; Collect raw monitoring data from multiple buoys, identify and remove noise and outliers by setting physical thresholds and statistical standards, and smooth the data based on the sampling frequency; S2: Missing pattern analysis and sample construction; The missing patterns are classified into spatial block missing, temporal block missing, and general missing. The time series is divided into fixed windows using a time window partitioning strategy. A mask matrix is constructed to simulate real missing scenarios and generate the occlusion data required for training and evaluation. S3: Spatiotemporal feature modeling and optimization; The buoy network is mapped into a graph structure, an adjacency matrix is constructed based on geographic distance or Pearson correlation coefficient, and multi-scale spatial features are extracted through a multi-layer GCN. GRU is used to capture long-term dependencies in time series. Model performance is improved through hyperparameter optimization and Bayesian search. S4: Model training and evaluation; The model is trained on the preprocessed dataset and the filling error is evaluated on an independent test set.

2. The method for filling ocean buoy data based on spatiotemporal neural network according to claim 1, characterized in that: In step S2, the fixed size of the time window is 168 time steps, and 56 time steps of complete data are retained before and after the padded segment.

3. The method for filling ocean buoy data based on spatiotemporal neural network according to claim 1, characterized in that: In step S2, the mask matrix includes a training mask mask and an evaluation mask eval_mask, wherein the training mask marks the occluded data as 0, and the evaluation mask marks the occluded data as 1.

4. The method for filling ocean buoy data based on spatiotemporal neural network according to claim 1, characterized in that: The adjacency matrix calculates the dynamic spatial correlation between buoys through a Gaussian kernel function or an exponential decay function.

5. The method for filling ocean buoy data based on spatiotemporal neural network according to claim 1, characterized in that: In step S3, the GRU captures the long-term dependencies of the time series by combining the update gate and the reset gate with the input gate and the forget gate functions, and updates the hidden state using the following formula: Update Gate: z t =σ(W z [x t ,h t-1 ]+b t ) Reset Gate: r t =σ(W r [x t ,h t-1 ]+b r ) Candidate Hidden State: c t =tanh(W h [x t ,r t ⊙h t-1 ]+b h ) Hide status update:

6. The method for filling ocean buoy data based on spatiotemporal neural network according to claim 1, characterized in that: In step S4, the loss function is defined based on the mean square error MSE:

7. The method for filling ocean buoy data based on spatiotemporal neural network according to claim 1, characterized in that: In step S3, the hyperparameter optimization includes adjusting the learning rate, batch size, hidden layer dimension and regularization coefficient.

Citation Information

Patent Citations

  • Traffic data restoration method based on spatial self-attention graph convolutional recurrent neural network

    CN112988723A

  • Deep learning-based population distribution completion method and system

    CN115526362A

  • Spatial field interpolation method for ocean observation data

    CN115935139A

  • Ocean observation data anomaly detection method based on space-time correlation of adjacent sites

    CN117972604A

  • Meteorological observation data complementing method for polar orbit satellite

    CN118366049A

Cited By

  • Robust lightweight time sequence prediction method based on bidirectional filling and geometric attention

    CN121435109A

  • Marine phytoplankton abundance monitoring method based on data driving and underwater acoustic network

    CN121542644A