An interpolation method, system, and device for spatio-temporal data based on inductive learning
By building complex neural network models, using deep learning and graph neural network technologies, the shortcomings of traditional interpolation methods in capturing nonlinear relationships in spatiotemporal data are solved, and high-precision and adaptive spatiotemporal data interpolation are achieved.
Patent Information
- Application Number
- CN202411629126.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-15
AI Technical Summary
The traditional spatiotemporal data interpolation method has shortcomings in capturing complex dependency features and nonlinear relationships, especially in areas where data is scarce and changes significantly, the interpolation effect is not ideal.
Using an inductive learning method, a complex neural network model is constructed, and nonlinear relationships in spatiotemporal data are captured through deep learning, graph neural networks and time encoders are constructed, combining aggregate functions and scalers to enhance the learning ability of the model.
High-precision interpolation of spatiotemporal data is realized, which improves the adaptability and versatility of interpolation, and can accurately predict data at unsampled locations when data is scarce.
Smart Images

Figure CN119150919B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing, and specifically relates to an interpolation method, system and device for spatio-temporal data based on inductive learning. Background Art
[0002] With the development of global climate change and urbanization, the monitoring, analysis and prediction of spatio-temporal data have gradually become important research fields. Spatio-temporal data plays a crucial role in multiple fields. For example, air and water pollution in environmental monitoring, congestion prediction in traffic management, spread of infectious diseases in public health, resource distribution in geographic information systems, and population flow in smart city planning, etc. The richness and diversity of spatio-temporal data make it have irreplaceable application value in the decision-making and management processes of various fields.
[0003] However, due to the complexity of data collection, spatio-temporal data often shows the characteristics of sparsity and uneven distribution. Especially in complex or vast environments, such as large-scale ocean, atmosphere and other fields, the phenomenon of data gaps is more significant. Although traditional interpolation methods (such as Kriging method, inverse distance weighting method, etc.) can fill some missing data, their ability to capture complex dependence features in spatio-temporal data is limited, so there are obvious deficiencies in accuracy. These methods are usually based on preset assumptions and cannot flexibly adapt to the non-linear and non-stationary features existing in the data. Especially in areas with significant spatio-temporal changes, the interpolation effect is often not ideal. Summary of the Invention
[0004] The present invention provides an interpolation method for spatio-temporal data based on inductive learning to solve the above technical problems. The method constructs a more complex neural network model and uses deep learning to capture the non-linear relationships in spatio-temporal data, thereby achieving accurate prediction of unsampled positions.
[0005] The present invention is implemented through the following technical solutions:
[0006] An interpolation method for spatio-temporal data based on inductive learning, comprising the following steps:
[0007] Step 1: Use corresponding instruments to collect spatio-temporal data, and then perform a preprocessing stage on the spatio-temporal data set;
[0008] Furthermore, the data set must cover the necessary timestamps and spatial information, and then perform data cleaning and normalization processing. The processed data is organized into a historical data matrix. At the same time, an initial all-zero mask matrix is constructed to mark the missing values in the subsequent data and simulate unsampled nodes. According to the data missing in the historical data matrix, the data missing matrix is updated, and the historical data matrix is divided into a training set, a validation set and a test set according to a ratio.
[0009] Step 2: Use the dataset in Step 1 to construct training samples for the spatio-temporal data interpolation model, so as to generate training samples that can generalize the model to unknown nodes and graph structures;
[0010] Furthermore, the method of Step 2 is as follows: First, randomly select time points within the time range of the historical data matrix, and extract sub-matrices from the complete historical data matrix according to the selected time nodes; the sub-matrices contain the observed data of all nodes at specific time points; randomly select 10%-20% of the nodes in the sub-matrix to simulate un-sampled point nodes; set the positions corresponding to the set un-sampled point nodes to 1 in the mask matrix, and reconstruct the data of these un-sampled nodes through model training.
[0011] Step 3: Use the training samples of the spatio-temporal data interpolation model constructed in Step 2 to construct a spatial aggregation network belonging to the graph neural network;
[0012] Furthermore, regard the spatial positions in the spatio-temporal data as nodes in the graph, and regard reachability and distance relationships as connection relationships in the graph. Extract the spatial dependencies of the spatio-temporal data through the spatial aggregation network, use the aggregation function to integrate the features of neighbor nodes, and learn and update the features of the nodes;
[0013] Furthermore, based on the graph neural network, the spatial aggregation network uses the aggregation function and distance information in the same layer to capture complex spatial dependencies.
[0014] The described aggregation functions include average aggregation, weighted average aggregation, Softmax aggregation, Softmin aggregation, standard deviation aggregation, max pooling, min pooling, mean distance aggregation, standard distance deviation aggregation, or self-attention aggregation;
[0015] Furthermore, the spatial aggregation network uses a scaler to consider the influence of different features, adjust the feature value range or distribution, and combines the aggregation function and the scaler using the tensor product to enhance the learning ability of the model.
[0016] Step 4: Construct a time encoder;
[0017] Furthermore, the time encoder in Step 4 is constructed as follows: Use one or more one-dimensional convolutional kernels of different sizes to perform convolutional operations on the time series data of a single node in parallel, extract local and long-term time features, and perform zero-padding before the convolutional operation; concatenate all the feature channels generated by the convolution together, and use a gating mechanism to transmit information crucial to the task. At the same time, the time encoder uses different activation functions, residual connections, and skip connections to enhance the learning ability of the network and avoid the problem of gradient disappearance.
[0018] Step 5: Train the spatio-temporal data interpolation model;
[0019] Furthermore, determine the hyperparameter settings of the model. At the beginning of training, alternately stack the spatial aggregation network layer and the temporal encoder layer to learn the spatial and temporal features of spatio-temporal data. In each layer, the spatial aggregation network layer synthesizes the information of neighbor nodes using more than one aggregation function, while the temporal encoder captures the dynamic changes of the time series through multi-scale convolution operations. Introduce residual connections into the model. In each iteration, randomly extract training samples from the historical data according to the number of samples determined by the training batch parameter, and apply the masking strategy to simulate the data missing situation. Adjust the model parameters by minimizing the loss function, i.e., root mean square error or mean absolute error. Use the Adam optimizer gradient descent method to update the learnable parameters of the model according to the gradients calculated by the backpropagation algorithm.
[0020] During the training process, adopt an early stopping mechanism to monitor the loss on the validation set. When the loss does not decrease significantly for consecutive multiple iterations, terminate the training in advance to avoid overfitting. When the performance of the model on the validation set reaches stability or meets the preset training conditions, obtain the final target model.
[0021] Furthermore, the hyperparameters include the number of network layers, the training batch size, the maximum number of training epochs, the proportion of missing nodes, the input sequence length of the training data, the length of the temporal convolution kernel in the temporal encoder, and the number of neighbor nodes.
[0022] Step 6: The target model obtained from Step 5 is used to simulate and generate new sensor data and generate data for unsampled nodes. Use the dataset to be interpolated as the basis for training the model, determine the area to be interpolated, including the positions and time periods of unsampled nodes, process the data into the input that the interpolation model described in Step 5 can accept. The adjacency matrix needs to include the information of unsampled nodes, and at the same time, set zeros in the masking matrix to identify the unsampled nodes, and use the trained interpolation model described in Step 5 for interpolation.
[0023] Furthermore, in the meteorological or oceanographic field, fill in the gaps in the monitoring network by generating virtual data to improve the spatial coverage of the data. In traffic monitoring, generate traffic flow data for unsampled road sections to help with more comprehensive traffic analysis and decision-making. In natural disaster monitoring, generate data for unsampled areas to help better assess risks and formulate response measures.
[0024] A spatio-temporal data interpolation system based on inductive learning, the system includes a data input and processing module, a training sample module for constructing a spatio-temporal data interpolation model, a spatial aggregation network construction module, a temporal encoder construction module, a spatio-temporal data interpolation model training module, and a data interpolation module.
[0025] The described data input and processing module is used to obtain data and preprocess the data, and the module runs the step (1);
[0026] The described training sample module for constructing the spatio-temporal data interpolation model runs the step (2) of the method;
[0027] The described spatial aggregation network construction module runs the step (3) of the method;
[0028] The described time encoder construction module runs the step (4) of the method;
[0029] The described spatio-temporal data interpolation model training module runs the step (5) of the method;
[0030] The described data interpolation module runs the step (6) of the method.
[0031] The present invention also provides a spatio-temporal data interpolation device based on inductive learning, and the device is equipped with the system.
[0032] The present invention has beneficial effects compared with the prior art:
[0033] The method of the present invention not only improves the accuracy of spatio-temporal interpolation, but also has better adaptability and generality. The model capable of making high-precision interpolation by fully utilizing spatio-temporal dependence features can achieve accurate interpolation of unobserved points in the case of scarce data, and has important application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic flow chart of a spatio-temporal data interpolation method based on inductive learning provided by an embodiment of the present invention;
[0035] Figure 2 is a model structure diagram of a spatio-temporal data interpolation method based on inductive learning provided by an embodiment of the present invention;
[0036] Figure 3 is a spatial aggregation network structure diagram of a spatio-temporal data interpolation method based on inductive learning provided by an embodiment of the present invention;
[0037] Figure 4 is a time encoder structure diagram of a spatio-temporal data interpolation method based on inductive learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The present invention will be further described below with reference to specific embodiments.
[0039] Embodiment 1
[0040] A spatio-temporal data interpolation method based on inductive learning includes the following steps:
[0041] Step 1: Use corresponding instruments to collect spatio-temporal data, and then enter the preprocessing stage of the spatio-temporal data set;
[0042] As a preferred implementation, it is first necessary to ensure that the collected data set covers the necessary timestamps and spatial information. Subsequently, data cleaning is performed to identify and process outliers caused by measurement errors, equipment failures, or data recording errors, so as to improve the data quality and ensure the accuracy of subsequent analysis. To ensure that all data points are on the same scale and eliminate the influence of different dimensions and magnitudes, the data is normalized. Normalization can be achieved through various methods, such as min-max scaling or Z-score standardization. Organize the cleaned and normalized data into a historical data matrix, which is a multivariate time series matrix containing spatio-temporal data of multiple location nodes within a specific time range. At the same time, construct an initial all-zero mask matrix for subsequent marking of missing values in the data and simulating unsampled nodes. Update the data missing matrix according to the data missing in the historical data matrix. The setting of the data missing matrix makes the model robust. The construction of the adjacency matrix can be based on geographical distance, network connection, accessibility, or other correlation metrics, and appropriate formulas are used to quantify the geographical location relationship, such as Gaussian kernel function or other spatial weight functions. Divide the historical data matrix into training set, validation set, and test set according to a certain proportion, ensuring the continuity of the data in the time series and avoiding temporal overlap to ensure the effectiveness of model training.
[0043] Step 2: Construct training samples for the spatio-temporal data interpolation model; In this step, the goal is to generate training samples that can generalize the model to unknown nodes and graph structures.
[0044] As a preferred implementation, randomly select time points within the time range of the historical data matrix, and extract sub-matrices from the complete historical data matrix according to the selected time nodes. The sub-matrix contains the observed data of all nodes at a specific time point. Randomly select about 10%-20% of the nodes in the sub-matrix to simulate unsampled point nodes. Set the corresponding positions of the unsampled point nodes in the mask matrix to 1. The goal of model training is to reconstruct the data of these unsampled nodes. Since the time points and masked nodes selected each time are random, the model not only learns the spatio-temporal dependencies in the data, but also has good generalization ability and inductive ability, enabling the model to accurately perform data interpolation when encountering new and unseen spatio-temporal configurations.
[0045] Step 3: Construct a spatial aggregation network; As Figure 3 shown, the spatial aggregation network is a deep learning model belonging to graph neural networks, and graph neural networks are deep learning models specifically used to process graph-structured data.
[0046] As one of the specific implementation manners, the spatial location in the spatio-temporal data is regarded as a node in the graph, and the reachability and distance relationship are regarded as the connection relationships in the graph. The spatial dependence of the spatio-temporal data is extracted through a graph neural network, and an aggregation function (such as average, sum, etc.) is used to integrate the features of neighbor nodes, and the feature representation of the nodes is learned and updated:
[0047] ;
[0048] Among them, represents the node feature matrix of the l-th layer, represents the weight matrix of the l-th layer, is the result of adding a self-loop to the adjacency matrix, is the degree matrix of.
[0049] Based on the graph neural network, the spatial aggregation network uses multiple aggregation functions and distance information in the same layer to capture complex spatial dependencies, such as average aggregation, weighted average aggregation, Softmax aggregation, Softmin aggregation, standard deviation aggregation, max pooling, min pooling, mean distance aggregation, standard distance deviation aggregation, self-attention aggregation, etc., so that each node can collect information of neighbor nodes from different perspectives; at the same time, the spatial aggregation network uses a scaler to consider the influence of different features, adjust the feature value range or distribution, and combines the aggregation function and the scaler using the tensor product to enhance the learning ability of the model.
[0050] Step 4: Construct a time encoder; as Figure 4 shown, another major feature of the spatio-temporal data is the temporal continuity. The data at the current time point of a certain node is related to the data in the past and future. Temporal feature extraction can help the model capture this dependence and perform interpolation and prediction more accurately. At the same time, there may be important information at different time scales in the spatio-temporal data, and it is necessary to capture multi-scale information from short-term fluctuations to long-term trends. The time encoder is designed to perform convolution operations on the time series data of a single node in parallel using multiple one-dimensional convolution kernels of different sizes to extract local and long-term temporal features.
[0051] As one of the preferred implementation manners, to maintain the consistency of the input and output data dimensions, zero padding is performed before the convolution operation. All the feature channels generated by the convolution are concatenated together to form a richer feature representation, and a gating mechanism is used to selectively transmit the information crucial for the task. At the same time, the time encoder uses different activation functions, residual connections, and skip connections to enhance the learning ability of the network and avoid the problem of gradient disappearance.
[0052] Step 5: Train the spatio-temporal data interpolation model; To train the spatio-temporal data interpolation model, it is first necessary to determine the hyperparameter settings of the model, including the number of network layers, the training batch size, the maximum number of training epochs, the proportion of missing nodes, the input sequence length of the training data, the length of the time convolution kernel in the time encoder, and the number of neighbor nodes.
[0053] As one of the specific implementation manners, at the beginning of training, a multi-layer spatial aggregation network layer and a time encoder layer are alternately stacked to learn the spatial and temporal features of the spatio-temporal data. In each layer, the spatial aggregation network layer synthesizes the information of neighbor nodes using multiple aggregation functions, while the time encoder captures the dynamic changes of the time series through multi-scale convolution operations. Residual connections are introduced into the model to facilitate the flow of information in the deep network and prevent the problem of gradient vanishing or explosion during the training process. In each iteration, according to the number of samples determined by the training batch parameter, training samples are randomly drawn from the historical data, and a masking strategy is applied to simulate the data missing situation. The model parameters are adjusted by minimizing the loss function, the root mean square error or the mean absolute error. Using the Adam optimizer gradient descent method, the learnable parameters of the model are updated according to the gradients calculated by the backpropagation algorithm.
[0054] During the training process, an early stopping mechanism is adopted to monitor the loss on the validation set. When the loss does not decrease significantly for consecutive multiple iterations, the training is terminated early to avoid overfitting. When the performance of the model on the validation set reaches stability or meets the preset training conditions, the final target model is obtained.
[0055] Step 6: The obtained target model can be used to simulate and generate new sensor data and generate data for unsampled nodes.
[0056] As one of the specific implementation manners, the dataset to be interpolated is used as the basis for training the model. The area to be interpolated is determined, including the positions and time periods of the unsampled nodes. The data is processed into an input that the model can accept. The adjacency matrix needs to contain information about the unsampled nodes, and at the same time, zeros are set in the masking matrix to identify the unsampled nodes, and the trained target model is used for interpolation. In fields such as meteorology and oceanography, virtual data can be generated to fill the gaps in the monitoring network and improve the spatial coverage of the data; in traffic monitoring, traffic flow data for unsampled sections can be generated to help with more comprehensive traffic analysis and decision-making; in natural disaster monitoring, generating data for unsampled areas can help better assess risks and formulate response measures.
[0057] Embodiment 2
[0058] An embodiment of the present invention provides a method for interpolating spatio-temporal data, and the method can be applied to the interpolation process of ocean surface temperature datasets.
[0059] See Figure 1, Figure 2 which is a schematic flowchart of an interpolation method for spatio-temporal data provided by an embodiment of the present invention. The interpolation process of the method for the ocean surface temperature dataset includes the following steps:
[0060] Step 1: Normalize the collected spatio-temporal data as a historical data matrix, mark the missing values in the historical matrix, generate a data missing matrix, identify the missing data in the historical data matrix according to the data missing matrix, and divide the training set and the test set. Construct an adjacency matrix A, considering the geographical location relationship or reachability between position nodes. The specific formula for the geographical location relationship is as follows:
[0061] ;
[0062] ;
[0063] where the latitude and longitude information of position node i is (lat1, lon1), and the latitude and longitude information of position node j is (lat2, lon2), is a hyperparameter that controls the scale;
[0064] Specifically, when implemented, the ocean surface temperature dataset uses OISST observation data from 2014 to 2023 from different platforms including satellites, ships, buoys, and Argo floats. The data has been processed, and the valid data among them is used as historical data X. The numerical range is scaled to the [-1, 1] interval using the maximum absolute value normalization method. X contains n position nodes and has a time range of [1, p], and exists in the form of a univariate time series matrix. The training set and the test set are divided in a ratio of 7:3. Define a data missing matrix of all 1s, set the corresponding points to 0 according to the null value points in the historical data matrix, and construct an adjacency matrix according to the geographical location relationship of the nodes in X.
[0065] Step 2: Randomly select different time points in the historical data matrix to generate multiple training samples. In each sample subset, randomly select some nodes as the missing position set, define a full-zero mask matrix, set a masking ratio of 10% - 20% of the nodes, and mask the corresponding positions in the mask matrix;
[0066] Based on the above embodiment, step 2 specifically includes:
[0067] Step 2.1: Randomly select a time point v within the time range {1 + s, p - h}, and obtain the training sample submatrix X of the ocean surface temperature according to the selected time node sample =X[:,V sample ;
[0068] where X sampleThe size is n×(h + s), where s is the reduction in time length due to temporal convolution, and h is the time length of the interpolation target, i.e., the length of the input sequence, which is set to 7.
[0069] Step 2.2: Select about 10% - 20% of the nodes in the training sample sub - matrix to generate a random set S as the masked node set to simulate the situation of un - sampled nodes. Define a all - zero mask matrix and set the corresponding positions in the mask matrix to 1.
[0070] Specifically, when implementing, randomly select different time points in the historical data matrix X to generate multiple training samples X of sea surface temperature sample . In each sample subset, randomly select 20% of the nodes to simulate un - sampled nodes and mask the corresponding positions in the mask matrix.
[0071] Step 3: Construct a spatial aggregation network. The spatial aggregation network is a deep - learning model specifically designed to process graph - structured data. By effectively aggregating the features of nodes and their neighbors, it captures and utilizes the complex dependencies in spatial data. The model can capture spatial dependencies at different scales from local to global. The spatial aggregation network uses multiple aggregation functions, combines the adjacency matrix and the masking strategy, and collects diverse information from the k nearest neighbor nodes of each node.
[0072] Based on the above - mentioned embodiment, the specific steps of step 3 include:
[0073] Step 3.1: Use multiple aggregation functions based on the adjacency matrix in each layer to more comprehensively capture the influence of neighbor nodes on the current node. The aggregation operations include mean, weighted mean, Softmax, Softmin, standard deviation, and standard distance deviation. The specific explanations are as follows:
[0074] The mean aggregation treats all neighbors of a node equally, regardless of the influence of distance.
[0075] The feature value of node i aggregating all neighbor nodes j at the l - th layer using mean aggregation is represented by :
[0076] ;
[0077] The weighted - mean aggregation considers the distance weights of neighbor nodes. The feature value of node i aggregating all neighbor nodes j at the l - th layer using weighted - mean aggregation is represented by : ;
[0078] Softmax and Softmin aggregations indirectly measure the maximum and minimum values of received messages, providing better generalization ability for graph neural networks. The feature value of node i aggregating all neighbors j at the l - th layer using Softmax aggregation is represented by Denote:
[0079] ;
[0080] Node i at layer l uses Softmax to aggregate the feature values of all neighbor nodes j, denoted by Denote:
[0081] ;
[0082] The average distance and standard deviation distance deviation aggregate characterize the spatial distance distribution from the current node to adjacent nodes. Node i at layer l uses the average distance to aggregate the feature values of all neighbor nodes j, denoted by Denote:
[0083] ;
[0084] Node i at layer l uses the standard deviation distance deviation to aggregate the feature values of all neighbor nodes j, denoted by Denote
[0085] ;
[0086] Where, represents the feature vector of node j at layer l, k is the number of neighbors of node i, is the adjacency matrix, represents the distance or weight or reachability metric from node j to node i in the adjacency matrix, is the metric of all nodes connected to i in the adjacency matrix, d is the distance or weight or reachability metric between nodes, (x) is defined as , used to ensure the squared difference, non - negative ϵ is a very small positive number, used to prevent division by zero;
[0087] Step 3.2: The aggregated information is adjusted by a sampling scaler and a saturation scaler. The sampling scaler uses a logarithmic function to amplify smaller aggregated values while compressing larger ones, increasing the model's sensitivity to subtle changes. The saturation scaler uses the inverse logarithmic function to attenuate larger aggregated values while amplifying smaller ones, paying more attention to changes in small values. The scaler formula is expressed as:
[0088] ;
[0089] ;
[0090] Where, is the aggregated feature value of node i's neighbors, is the metric of all nodes connected to i in the adjacency matrix;
[0091] Step 3.3: The spatial aggregation network combines the aggregation function and the scaler, and integrates different aggregation results and scale adjustments through the tensor product to form a comprehensive representation of neighbor information;
[0092] ;
[0093] where I is an identity matrix, is obtained by multiplying all scalers and aggregators and stacking them together, which are the result matrices of mean aggregation, weighted mean aggregation, Softmax and Softmin aggregation, mean distance and standard deviation distance deviation aggregation respectively;
[0094] Step 3.4: The spatial aggregation network introduces a masking strategy to prevent unobserved nodes from transmitting information during the aggregation process through a masking matrix. The spatial aggregation network contains multiple spatial aggregation layers, and each layer independently applies the above-mentioned aggregation function and scaler, enabling the model to capture spatial dependencies at different abstraction levels and deepening the understanding of the spatial structure layer by layer:
[0095] ;
[0096] where, is the node feature matrix of the l-th layer at the t-th time point, is the activation function, is the weight matrix of the l-th layer for spatial aggregation, represents the bias vector of the l-th layer;
[0097] In specific implementation, a spatial aggregation network is constructed. Combining the adjacency matrix and the masking matrix obtained in Step 1 and Step 2, the aggregation functions of mean, weighted mean, Softmax, Softmin, standard deviation, and standard distance deviation are applied to aggregate the diverse information of the 5 nearest neighbor nodes of each node in the historical feature matrix X sample . Further, a sampling scaler and a saturation scaler are applied to adjust these feature values to improve the sensitivity and robustness of the model. Finally, the spatial feature information of the nodes is updated through the spatial feature extraction and aggregation of multiple aggregation functions and scalers of the spatial aggregation network.
[0098] Step 4: Use a time encoder based on one-dimensional convolution to perform multi-scale convolution on the time series data of each spatial node to extract local and long-term dependence features in the time dimension;
[0099] In specific implementation, the time encoder performs zero-padding on the input data before the convolution operation. Convolution kernels of 1×2, 1×3, 1×5, and 1×7 are applied to the time axis of the input data for each block. After convolution operations of different scales, the outputs of all convolutional layers are concatenated along the channel dimension, and finally, through a gating mechanism, information crucial to the task is selectively transmitted:
[0100] ;
[0101] where is the feature obtained by concatenating multi-scale convolutions, (·) represents the Sigmoid activation function, and tanh(·) represents the Tanh activation function.
[0102] Step 5: Alternately use the spatial aggregation network and the time encoder layer to capture the spatio-temporal dependence features of the nodes layer by layer, and train the model to minimize the error between the estimated value of the output and the true label. The model is trained until the learnable parameters in the network reach the preset conditions to obtain the target model;
[0103] Based on the above embodiments, step 5 specifically includes:
[0104] Step 5.1, set the hyperparameters of the model, including the number of network layers, the training batch size, the maximum number of training epochs, the proportion of missing nodes, the input sequence length of the training data, and the number of neighbor nodes. Define the loss function as the error between the estimated value of the model output and the true label, and use the root mean square error RMSE and the mean absolute error MAE as the error metrics;
[0105] Step 5.2, alternately use the spatial aggregation network and the time encoder layer, and introduce residual connections in the model to improve the stability during the training process and the effect of the deep network;
[0106] Step 5.3, during the model training process, adopt an early stopping mechanism. When the loss on the validation set no longer decreases significantly within the set number of iterations, terminate the training in advance;
[0107] Step 5.4, transmit the calculated loss to the learnable parameters of the model through the backpropagation algorithm, and use the Adam optimizer to update the value of each learnable parameter. The training process continues until the learnable parameters in the network meet the preset training conditions to obtain the final target model;
[0108] During specific implementation, set the hyperparameters of the model and the number of spatial neighbor nodes, alternately use the spatial aggregation network and the temporal encoder layer constructed in Steps 3 and 4 for three layers, capture the spatio-temporal dependence features of the nodes layer by layer, introduce residual connections to prevent the problem of gradient disappearance or explosion in the deep network during training, train the model to minimize the error between the estimated value of the output and the true label, and train the model until the learnable parameters in the network reach the preset conditions to obtain the target model.
[0109] Step 6: Use the target model to simulate and generate new sensor data, i.e., the data of unsampled nodes, use the dataset to be interpolated as the basis for training the model, determine the area to be interpolated, and perform spatio-temporal interpolation using the trained target model;
[0110] During specific implementation, to obtain a complete gridded ocean surface temperature dataset, an iterative interpolation strategy is adopted. Use the target model to interpolate and predict the ocean surface temperature of unsampled points in the ocean surface temperature dataset. The iterative interpolation strategy starts from the vicinity of known nodes and gradually expands outwards. The initially generated interpolation results can be used to augment the training set, and the interpolation accuracy of the model is further optimized through multiple iterations. Finally, a gridded ocean surface temperature dataset with a daily time resolution and a spatial resolution of 0.25°×0.25° without missing data is obtained. The target interpolation model achieves high-precision interpolation for unobserved points in the case of scarce data and has achieved good technical effects. Taking the sea temperature data in the Bohai Sea area on January 15, 2022 as an example, the data missing rate is 30%, about 15% of the shielding nodes are set, and the target model is used for prediction interpolation. The root mean square error (RMSE) of the model prediction is 0.2021, and the mean absolute error (MAE) is 0.1794, indicating that the present invention can still provide accurate temperature prediction under limited data conditions and has high applicability and practical value.
Claims
1. A spatiotemporal data interpolation method based on inductive learning, the method is applied to the interpolation of sea surface temperature datasets, characterized in that: The method comprises the following steps: Step 1: Use platforms including satellites, ships, buoys, and Argo buoys to collect sea surface temperature data sets, normalize the collected sea surface temperature data as a historical data matrix, mark the missing values in the historical matrix, generate a data missing matrix, identify the missing data of the historical data matrix according to the data missing matrix, and divide it into training set, validation set, and test set; Step 2: In the historical data matrix of step 1, randomly select different time points to generate multiple training samples. In each sample subset, randomly select some nodes as the missing position set, define the all-zero mask matrix, set the shielding nodes with a ratio of 10%-20%, and shield the corresponding positions in the mask matrix; specifically include: Step 3: Use the training samples of the spatiotemporal data interpolation model constructed in step 2 to build a spatial aggregation network belonging to the graph neural network; regard the spatial position in the spatiotemporal data as a node in the graph, and the reachability and distance relationship as the connection relationship in the graph. Use the spatial aggregation network to extract the spatial dependency of the spatiotemporal data, use the aggregation function to integrate the features of neighbor nodes, and learn and update the features of the nodes; Based on graph neural networks, spatial aggregation networks use aggregation functions and distance information in the same layer to capture complex spatial dependencies. Step 4: Construct a temporal encoder; use more than one one-dimensional convolution kernel of different sizes to perform convolution operations on the time series data of a single node in parallel to extract local and long-term temporal features, and perform zero padding before the convolution operation; splice all feature channels generated by the convolution together, and use a gating mechanism to pass information that is critical to the task. At the same time, the temporal encoder uses different activation functions, residual connections, and skip connections to enhance the network's learning ability and avoid the gradient vanishing problem; Step 5: Train the spatiotemporal data interpolation model; determine the model's hyperparameter settings. At the beginning of training, use spatial aggregation network layers and temporal encoder layers to alternately stack to learn the spatial and temporal features of spatiotemporal data; in each layer, the spatial aggregation network layer uses more than one aggregation function to integrate the information of neighboring nodes, while the temporal encoder captures the dynamic changes of the time series through multi-scale convolution operations; introduce residual connections into the model; in each iteration, randomly extract training samples from historical data according to the number of samples determined by the training batch parameter, and apply a masking strategy to simulate data missing situations; adjust the model parameters by minimizing the root mean square error or mean absolute error of the loss function; use the Adam optimizer gradient descent method to update the learnable parameters of the model according to the gradient calculated by the back-propagation algorithm; During the training process, an early stopping mechanism is used to monitor the loss on the validation set. When the loss does not decrease significantly over multiple consecutive iterations, the training is terminated early to avoid overfitting. When the performance of the model on the validation set reaches stability or meets the preset training conditions, the final target model is obtained. The hyperparameters include the number of network layers, the training batch size and the maximum training cycle, the ratio of missing nodes, the input sequence length of the training data, the length of the temporal convolution kernel in the temporal encoder, and the number of neighbor nodes; Step 6: Use the target model trained in step 5 to simulate the generation of new sensor data and the data of unsampled nodes; use the data set to be interpolated as the basis for the training model, determine the area to be interpolated, including the location and time period of the unsampled nodes, and process the data into input acceptable to the interpolation model described in step 5. The adjacency matrix needs to contain the information of the unsampled nodes, and set zeros in the mask matrix to identify the unsampled nodes. Use the trained interpolation model described in step 5 for interpolation; the specific method is as follows: The target model is used to interpolate and predict the sea surface temperature of unsampled points in the sea surface temperature dataset. The iterative interpolation strategy starts from the vicinity of known nodes and gradually expands outward. The preliminary interpolation results are used to expand the training set. The interpolation accuracy of the model is further optimized through multiple iterations, and finally a gridded sea surface temperature dataset with a temporal resolution of daily and a spatial resolution of 0.25°×0.25° and no missing data is obtained.
2. The spatiotemporal data interpolation method based on inductive learning according to claim 1, characterized in that: The sea surface temperature dataset described in step 1 must include timestamp and spatial information.
3. The spatiotemporal data interpolation method based on inductive learning according to claim 1, characterized in that: The aggregation functions include average aggregation, weighted average aggregation, Softmax aggregation, Softmin aggregation, standard deviation aggregation, maximum pooling, minimum pooling, mean distance aggregation, standard distance deviation aggregation or self-attention aggregation. The spatial aggregation network uses a scaler to consider the influence of different features, adjust the feature value range or distribution, and use tensor product to combine the aggregation function and the scaler together to enhance the learning ability of the model.
4. A spatiotemporal data interpolation system based on inductive learning, characterized in that: The system includes a data input and processing module, a training sample module for constructing a spatiotemporal data interpolation model, a module for constructing a spatial aggregation network, a module for constructing a time encoder, a module for training a spatiotemporal data interpolation model, and a module for interpolating data; The data input and processing module is used to acquire data and pre-process the data, and the module runs step 1 of the spatiotemporal data interpolation method based on inductive learning described in claim 1; The training sample module for constructing a spatiotemporal data interpolation model Execute step 2 of the spatiotemporal data interpolation method based on inductive learning described in claim 1; The spatial aggregation network building module runs step 3 of the spatiotemporal data interpolation method based on inductive learning described in claim 1; The construction time encoder module runs step 4 of the spatiotemporal data interpolation method based on inductive learning described in claim 1; The training spatiotemporal data interpolation model module executes step 5 of the spatiotemporal data interpolation method based on inductive learning described in claim 1; The data interpolation module runs step 6 of the spatiotemporal data interpolation method based on inductive learning described in claim 1.
5. A spatiotemporal data interpolation device based on inductive learning, characterized in that: The device is equipped with the system according to claim 4.
Citation Information
Patent Citations
Spatio-temporal data prediction method based on graph convolution network
CN111639787A
Graph neural network classification method and device based on small sample learning
CN112633403A