Spatial-temporal data completion method with uncertainty perception

Through the graph Transformer neural process model, the problem of insufficient accuracy and reliability of spatiotemporal data completion in the existing technology is solved, and more efficient spatiotemporal data completion and uncertainty management are achieved.

CN120104982APending Publication Date: 2025-06-06JINLING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108397.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate time and spatial information in spatiotemporal data completion, and lacks uncertain perception ability when facing sparse data and fuzzy signals, resulting in insufficient completion accuracy and reliability.

Method used

The graph Transformer neural process model is used to represent the spatial and temporal dependencies between the target node and the context node in the learning stage through space-time joint representation, and the random distribution derivative is performed by graph summing aggregation and Gaussian prior distribution in the random generation stage to complete the missing attribute values ​​of the spatiotemporal data.

Benefits of technology

It realizes a more comprehensive feature representation and completion of spatiotemporal data, improves completion accuracy and reliability, can effectively manage data uncertainty, and helps make more reliable decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104982A_ABST
    Figure CN120104982A_ABST
Patent Text Reader

Abstract

The invention discloses a spatio-temporal data completion method with uncertainty perception, which comprises the following steps of: S1, preprocessing acquired spatio-temporal data, and constructing a training data set; step S2, constructing a graph Transform neural process model, and training the graph Transform neural process model by using the training data set; and S3, complementing the missing attribute value of the target node in the original spatio-temporal data by using the trained Transform neural process model of the graph. According to the method, time and space features are captured by using a graph neural network and Transform, complete spatial and temporal feature joint representation can be obtained, approximate data of unknown nodes can be complemented, the deployment and maintenance cost of a sensor is saved, in addition, uncertainty estimation of data complementation is obtained by using an improved neural process, and the accuracy of data complementation is improved. And management personnel can make more reliable decisions such as sensor deployment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a spatiotemporal data completion method with uncertainty perception. Background Art

[0002] As an important part of data preprocessing, data completion technology has attracted more and more attention from researchers. Although there are some outstanding sequence completion methods in the existing technology, such as traditional statistical methods such as historical average, these statistical methods are simple and convenient to implement, but the interpolation accuracy of spatiotemporal data missing for a long period of time is not ideal. Machine learning methods such as K nearest neighbor, matrix method, Gaussian process, etc. can improve the prediction accuracy through theoretical technology, but the prediction results of such methods are greatly affected by the setting of hyperparameters. The third type of method is deep learning methods, such as recurrent neural networks, Transformer, etc. Although such methods can capture nonlinearity and dynamics well in the time dimension, they have the following shortcomings when completing spatiotemporal data: First, there are multiple correlations in spatiotemporal data, such as time correlation, spatial correlation and spatiotemporal correlation. However, the existing methods are either unable to fuse time information or spatial information, and the representation of spatiotemporal features is not complete. Second, the deployment and maintenance of sensors for collecting spatiotemporal data requires high costs, so the data in many application scenarios is sparse, which affects the quality of data mining and decision-making solutions. Therefore, completing the approximate data of unknown nodes is an urgent problem to be solved. Third, sensors that collect spatiotemporal data usually produce fuzzy signals to varying degrees. For example, the uncertainty of the collected data caused by errors in hardware equipment or network reasons. Data uncertainty estimation helps managers make more reliable decisions, but there is currently little research on the uncertainty of complete data. Summary of the invention

[0003] Purpose of the invention: The technical problem to be solved by the present invention is to provide a spatiotemporal data completion method with uncertainty perception in view of the deficiencies in the prior art.

[0004] In order to solve the above technical problems, the present invention discloses a spatiotemporal data completion method with uncertainty perception, comprising the following steps:

[0005] Step S1, preprocessing the collected spatiotemporal data to construct a training data set, wherein the training data set includes a context node data set and a target node data set;

[0006] Step S2, constructing a graph Transformer neural process model, and training the graph Transformer neural process model using the training data set;

[0007] Step S3, using the trained graph Transformer neural process model to complete the missing attribute values ​​of the target nodes in the original spatiotemporal data.

[0008] Furthermore, the graph Transformer neural process model in step S2 includes a spatiotemporal joint representation learning stage and a random generation stage.

[0009] The spatiotemporal joint representation learning phase is used to learn the spatial relationship between the target node and the context node, extract the temporal dependency of the spatiotemporal sequence, and obtain the spatiotemporal learning sequence of the context node and the spatiotemporal learning sequence of the target node in combination with the features of the context node;

[0010] The random generation stage is used to aggregate the spatiotemporal coding information of the context nodes through graph summation aggregation, and derive a random distribution describing the target to be completed by combining Gaussian prior distribution and likelihood estimation, thereby completing the missing attribute value of the target node.

[0011] Furthermore, the spatiotemporal joint representation learning stage includes a spatial encoding layer and a temporal encoding layer. The spatial encoding layer inputs the attribute values ​​of the context node and the target node, learns the spatial relationship between the target node and the context node by local graph convolution, and obtains the spatiotemporal sequence that has captured the spatial features; the temporal encoding layer inputs the spatiotemporal sequence that has captured the spatial features, extracts the temporal dependency of the spatiotemporal sequence by simplifying the Transformer, and obtains the context node spatiotemporal learning sequence and the target node spatiotemporal learning sequence.

[0012] Furthermore, the spatial encoding layer learns the spatial relationship between the target node and the context node by local graph convolution to obtain a spatiotemporal sequence that has captured spatial features, including: updating the representation of the target node with K-hop neighbors, as follows:

[0013]

[0014] in, It is the spatiotemporal sequence of captured spatial features obtained through the spatial encoding layer. L represents the total number of graph convolution layers, and l represents the index of the graph convolution layer. represents the K-hop neighbors of the node m to be completed, represents the K-hop adjacency matrix weight value of the target node m to be completed and the context node n, Refers to the encoding result of the context node information, W l is the learnable weight matrix of the lth layer of graph convolution; the computational complexity is Reduce to N represents the number of context nodes, M represents the number of target nodes to be completed, 1≤n≤N, 1≤m≤M.

[0015] Furthermore, the time coding layer includes a position coding module, a multi-head self-attention layer, a layer normalization module, a fully connected layer and an output layer. The position coding module is used to perform position coding on the spatiotemporal sequence of the captured spatial features, and the sequence obtained by splicing the coding result with the spatiotemporal sequence of the captured spatial features is input into the multi-head self-attention layer; the multi-head self-attention layer is used to compare each element in the sequence with other elements to calculate different weight values, and use different weight values ​​to generate a comprehensive representation vector to capture complex time dependencies; the residual connection is to connect the input information to the output of the multi-head attention layer, so that the model learning process is more stable; the layer normalization module normalizes the output of all neurons in the layer so that the output distribution has a stable mean and variance; in the fully connected layer, each node is connected to all nodes in the previous layer to increase the fitting ability of the model; the output layer is used to generate context node spatiotemporal learning sequences and target node spatiotemporal learning sequences through linear transformation and activation functions.

[0016] Using local graph convolutional networks and simplified Transformers to perform complete spatiotemporal joint representation learning can reduce the computational complexity. Complete spatiotemporal joint representation learning can improve the accuracy of spatiotemporal data completion.

[0017] Furthermore, the random generation stage aggregates the spatiotemporal coding information of the context nodes through graph summation aggregation, and the calculation formula is as follows:

[0018] in, Refers to the one-hop neighbor of the target node to be completed, represents the spatiotemporal learning sequence of context nodes obtained after the spatiotemporal joint representation learning phase of context node n, Express Perform a separate encoding operation, A m,n Represents the value of the adjacency matrix between the target node m to be completed and the context node n; r represents the representation of the target node after aggregation.

[0019] By aggregating the information of context nodes, the target latent variable describing the random process of the target node is derived T represents the time step, d l Represents the dimension of the target latent variable at each time step.

[0020] Furthermore, the random generation stage learns the mean and variance of the conditional prior distribution through a neural network, and the conditional prior of the target node is recorded as is a decomposable Gaussian distribution:

[0021]

[0022] in, represents the context node set, μ Z and It is the parameter learned by the neural network Encoder;

[0023] The target latent variable Z is composed of the representation of the target node m and the spatiotemporal learning sequence of the context nodes Determine, given the target node spatiotemporal learning sequence and a sample of p(Z), then the neural network It is possible to learn a conditional prior for the target node The representation of the time series obtained by the context node n after the spatiotemporal joint representation learning phase is used to learn the potential observations and then updated by the above formula Parameters.

[0024] Furthermore, the generation process of the graph Transformer neural process model is expressed as:

[0025]

[0026] Where A represents the adjacency matrix, represents the known features of the target node T time steps, represents the attribute value of the target node T time steps. The first term on the right side of the equation is the likelihood function, assuming that the likelihood is a decomposable Gaussian distribution; the second term is the conditional prior learned by graph summation aggregation;

[0027] The distribution of the attribute value of the target node to be completed is obtained according to the above formula, and the mean of the distribution is taken as the final completion prediction value.

[0028] Furthermore, step S3 also includes variable estimation by maximizing ELBO:

[0029]

[0030] in, represents the known features of the context node T time steps, represents the attribute value of the context node T time steps, represents the marginal likelihood, represents the alternative distribution, p(y m |Z,x m ) represents the posterior distribution, where y m represents the attribute value of the target node m, x m represents the known features of the target node m, Z represents the target latent variable, represents the joint distribution, represents the divergence, Express expectations.

[0031] Combining the neural process of graph summation aggregation to perform random generation process, the missing data can be completed, and the uncertainty estimation of the completion can be obtained, which helps to improve reliability. Further, step S1 includes: filtering and deleting the data samples with missing values ​​from the collected spatiotemporal data, and retaining the spatiotemporal samples without missing items; dividing the retained samples into context node data sets and target node data sets according to different spatial points, each node data includes known features and features to be completed, and the context node data set is recorded as The target node dataset is represents the known features in the context node dataset, Represents the features to be completed in the context node dataset; represents the known features in the target node dataset, Represents the features to be completed in the target node dataset; randomly replaces the features to be completed of the context node according to the set missing rate Set to missing as training data; features to be completed for the target node Problems with unknown initial values, defining a learnable target node encoding e represents the encoding dimension.

[0032] Beneficial effects: The present application provides a spatiotemporal data completion method with uncertainty perception. The method uses graph neural networks and Transformer to capture temporal and spatial features, and can obtain a complete joint representation of spatiotemporal features. It can complete the approximate data of unknown nodes and save the deployment and maintenance costs of sensors. In addition, the uncertainty estimation of data completion is obtained by improving the neural process, which helps managers make more reliable decisions such as sensor deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0034] Figure 1 This is a structural diagram of the GraphformerNP model in a spatiotemporal data completion method with uncertainty awareness proposed in an embodiment of the present application.

[0035] Figure 2 This is a structural diagram of the time coding layer of the GraphformerNP model in the uncertainty-aware spatiotemporal data completion method proposed in an embodiment of the present application.

[0036] Figure 3A schematic diagram of the PM2.5 data completion results in a spatiotemporal data completion method with uncertainty perception proposed in an embodiment of the present application. DETAILED DESCRIPTION

[0037] The embodiments of the present invention will be described below in conjunction with the accompanying drawings.

[0038] The present application provides an uncertainty-aware spatiotemporal data completion method, which can be applied to spatiotemporal data completion in air quality datasets.

[0039] The present application embodiment discloses a method for completing spatiotemporal data with uncertainty perception, comprising the following steps:

[0040] Step S1, preprocessing the collected spatiotemporal data to construct a training data set, wherein the training data set includes a context node data set and a target node data set;

[0041] Step S1 includes: filtering the collected spatiotemporal data to delete data samples with missing values ​​and retaining spatiotemporal samples without missing items; dividing the retained samples into context node data sets and target node data sets according to different spatial points, each node data includes known features and features to be completed, and the context node data set is recorded as The target node dataset is represents the known features in the context node dataset, Represents the features to be completed in the context node dataset; represents the known features in the target node dataset, Represents the features to be completed in the target node dataset; randomly replaces the features to be completed of the context node according to the set missing rate Set to missing as training data; features to be completed for the target node Problems with unknown initial values, defining a learnable target node encoding e represents the encoding dimension. In the specific implementation process, the target encoding that can be learned is A tensor parameter of a specified dimension e can be randomly initialized or initialized to 0, the gradient is calculated through back-propagation, and the parameters are updated through the optimizer to finally determine the learnable target node encoding.

[0042] Step S2, constructing a graph Transformer neural process model, and training the graph Transformer neural process model using the training data set;

[0043] The graph transformer neural process model (hereinafter referred to as GraphformerNP) integrates graph neural network, transformer and neural process with uncertainty perception to complete the missing values ​​of spatiotemporal series. The overall network model architecture is as follows Figure 1 shown.

[0044] The graph Transformer neural process model includes a spatiotemporal joint representation learning stage and a random generation stage.

[0045] The spatiotemporal joint representation learning stage is used to learn the spatial relationship between the target node and the context node, extract the temporal dependency of the spatiotemporal sequence, and obtain the spatiotemporal learning sequence of the context node by combining the features of the context node. and target node spatiotemporal learning sequence

[0046] The random generation stage is used to aggregate the spatiotemporal coding information of the context nodes through graph summation aggregation, and derive a random distribution describing the target to be completed by combining the Gaussian prior distribution and likelihood estimation, thereby completing the missing attribute value of the target node.

[0047] In the spatiotemporal joint representation learning phase, the incomplete spatiotemporal sequence is used as input, all the features, attribute values ​​of the context nodes and the known features of the target node are encoded into a latent space by one-dimensional convolution, and then the attribute value of the target node (learnable target node encoding) is used to generate the target node. ) and the attribute values ​​of the context nodes to learn the spatial relationship between the target node and the context nodes, and then learn the temporal features of each target node, and combine the features of the context nodes to obtain a deterministic spatiotemporal representation encoding.

[0048] The spatiotemporal joint representation learning stage mainly consists of two parts, namely the spatial encoding layer and the temporal encoding layer. First, the spatial encoding layer is used to learn the spatial relationship between the target node to be completed and the context node, and then the temporal encoding layer is used to capture the temporal dependency of the node sequence. In order to solve the problem that the initial value of the target node to be completed is unknown, a learnable target node encoding layer is defined. Set the initial value of all context node information according to the parameter Encode, d y Indicates the feature number of the context node. In the embodiment of the present application, the value of e is 16.

[0049] The spatial encoding layer inputs the attribute values ​​of the context node and the target node, and mainly uses the local graph convolution method to learn the spatial relationship between the target node and the context node, and obtains the spatiotemporal sequence of the captured spatial features. In the task of completing the spatiotemporal features of unknown nodes, since the method proposed in this application ultimately focuses on the data completion of the target node to be completed, it is only necessary to consider the spatial relationship between the target node and the context node. It is believed that the spatial relationship between the context nodes alone has no effect on the target node. Therefore, a local graph convolution method is proposed, which only focuses on the spatial dependency between the nodes in the two sets of context node set and target node set. This can reduce the computational complexity without affecting performance. Specifically, as shown in Formula 1, K-hop neighbors are used to update the representation of the target node. The computational complexity is reduced from Reduce to N represents the number of context nodes, and M represents the number of target nodes to be completed.

[0050]

[0051] Among them, L represents the total number of graph convolution layers, l represents the index of graph convolution layer number, represents the K-hop neighbors of the node m to be completed, represents the K-hop adjacency matrix weight value of the target node m to be completed and the context node n. In this embodiment, it refers to the longitude and latitude distance calculated by the Haversine formula. Refers to the encoding result of the context node information, W l is the learnable weight matrix of the lth layer of graph convolution. In the specific implementation process, the parameter W is updated by the gradient descent algorithm during the training process. l It is a learnable parameter matrix that can be initialized by random initialization and then updated by gradient descent. It is a spatiotemporal sequence of captured spatial features obtained through the spatial encoding layer.

[0052] The temporal encoding layer takes the temporal sequence of the captured spatial features as input, extracts the temporal dependencies of the temporal sequence by simplifying the Transformer, and obtains the temporal learning sequence of the context node and the temporal learning sequence of the target node. This application adopts the simplified Transformer layer for temporal encoding, and only uses the Encoder module in modeling, such as Figure 2 shown.

[0053] The time encoding layer includes a position encoding module, a multi-head self-attention layer, a layer normalization module, a fully connected layer, and an output layer. First is the position encoding module, which is used to perform position encoding on the spatio-temporal sequence of the captured spatial features. The sequence obtained after concatenating the encoding result with the spatio-temporal sequence of the captured spatial features is input into the multi-head self-attention layer. Specifically, the time series after local graph convolution is used as the input vector. Considering the order of the time series, the embedding vector itself does not contain position information. Therefore, the position encoding module in NLP (Natural Language Processing) is adopted. Specifically, each element in a sequence is mapped into a vector in one-dimensional space, as shown in Formulas 2 and 3.

[0054]

[0055]

[0056] Among them, i represents the position of the object in the input time series, j has no actual meaning and is used to map to the column index k (that is, 2j = k or 2j + 1 = k, 0 <= i < d / 2), d represents the dimension of the output embedding space, and PE represents the position function, which is used to map the element at position i in the input sequence to the (i, k) position of the position matrix.

[0057] Position encoding is performed using sine and cosine functions to achieve a better understanding of the sequence information.

[0058] To allow the model to focus on different elements in the time series in a dynamic manner to better capture the dependencies between elements, a multi-head self-attention layer is adopted. Different from traditional CNN (Convolutional Neural Network) or RNN (Recurrent Neural Network), the self-attention mechanism in Transformer is not restricted by a fixed window or order, enabling it to capture global temporal patterns. Figure 2 The multi-head attention mechanism in it compares each element in the sequence with other elements to calculate different weight values, representing the degree of importance. Finally, a comprehensive representation vector is generated using different weight values, which can capture complex time dependencies. Residual connection connects the input information to the output of the multi-head attention layer, making the model learning process more stable; the layer normalization module normalizes the outputs of all neurons within the layer, making the distribution of the outputs have stable mean and variance; the fully connected layer, where each node is connected to all nodes in the previous layer, is used to increase the fitting ability of the model; the output layer is used to generate the spatio-temporal learning sequence of context nodes through linear transformation and activation functions. and target node spatiotemporal learning sequence The Relu activation function is selected in this application.

[0059] In the random generation stage, the N context nodes obtained through the spatiotemporal joint spatiotemporal representation learning are first encoded separately, that is, Figure 1 In The encoded context node information is then aggregated. In order to aggregate effective information more accurately and reduce computational complexity, this application proposes a new graph summation aggregator, that is, only the spatiotemporal coding information of the context nodes that are connected to the target node to be completed is aggregated, as shown in Formula 4.

[0060]

[0061] in, Refers to the one-hop neighbor of the target node to be completed, represents the spatiotemporal learning sequence of context node n obtained after the spatiotemporal joint representation learning phase, A m,n Represents the value of the adjacency matrix between the target node m to be completed and the context node n, specifically the longitude and latitude distance between the nodes calculated by the Haversine formula; r represents the representation of the target node after aggregation.

[0062] By aggregating the information of context nodes, the target latent variable describing the random process of the target node is derived In the specific implementation process, it refers to the target hidden variable corresponding to a target node, T represents the time step, and d l Represents the dimension of the target latent variable at each time step.

[0063] Then, the representation of the target node m after combining the target latent variable Z and the spatiotemporal joint representation learning is The mean and variance of the conditional prior distribution are learned through a neural network. Assume that the conditional prior of the target node is is a decomposable Gaussian distribution:

[0064]

[0065] in, represents the context node set, μ Z and It is the parameter learned by the neural network Encoder.

[0066] The target latent variable Z is composed of the representation of the target node m and the spatiotemporal learning sequence of the context nodes Determine, given the target node spatiotemporal learning sequence and a resampling of p(Z), then the neural network It is possible to learn a conditional prior for the target node The representation of the time series of context node n obtained after the spatiotemporal joint representation learning phase. The spatiotemporal learning sequence of context node n is used to learn the latent observations and then updated by Formula 5 Parameters.

[0067] Finally, the likelihood model uses the target latent variable Z and the covariates of the target node, such as weather characteristics, to predict the The generation process of GraphformerNP can be expressed as:

[0068]

[0069] Where A represents the adjacency matrix, represents the known features of the target node T time steps, represents the attribute value to be completed for the target node T time steps. The first term on the right side of the equation is the likelihood function, assuming that the likelihood is also a decomposable Gaussian distribution. The second term is the conditional prior learned by graph summation and aggregation.

[0070] The distribution of the attribute value of the target node to be completed is obtained according to Formula 6, and the mean of the distribution is taken as the final completion prediction value.

[0071] Since closed-form solutions of nonlinear transformations and likelihoods are difficult to obtain, variable estimation can be performed by maximizing ELBO (Evidence Lower Bound).

[0072]

[0073] in, represents the known features of the context node T time steps, represents the attribute value of the context node T time steps, represents the marginal likelihood, represents the alternative distribution, p(y m |Z,x m ) represents the posterior distribution, where y m represents the attribute value of the target node m, x m represents the known features of the target node m, Z represents the target latent variable, represents the joint distribution, represents the divergence, Express expectations.

[0074] Example:

[0075] In order to verify the effectiveness of the graph transformer neural process model (hereinafter referred to as: GraphformerNP) proposed in the embodiment of the present application, the embodiment of the present application uses the Beijing air quality data set, which contains hourly air quality data of 36 stations and meteorological data such as temperature, wind speed, wind direction, humidity, air pressure and weather in the same area. The goal of the embodiment of the present application is to complete the PM2.5 and CO air quality index of unobserved nodes. Among these features, wind direction and weather are categorical variables, and the others are continuous features. Wind direction contains 10 categories (including 4 basic directions, 4 secondary directions, and unstable and no direction). Weather has 17 categories, including but not limited to rainy days, foggy days, sunny days and dusty days. The dataset is publicly available through the following website (http: / / research.microsoft.com / apps / pubs / ?id=246398). The time span of the dataset is one year (May 1, 2014 to April 30, 2015). In order to solve the missing items in the dataset, the embodiment of the present application removes sites with a large number of missing values. 50% of the stations in the dataset had at least 60% missing values ​​for the pressure feature. Therefore, pressure was removed from the meteorological variables. In addition, five stations (station IDs: 1009, 1013, 1015, 1020, 1021) had only 35% of the data available for meteorological data, so these five stations were removed from the experiment. In the remaining data, at least 85% of the data were available for all variables. To fill the missing data for the actual value variables (PM2.5, temperature, humidity, and wind speed), average interpolation was performed over time.

[0076] The experimental environment configuration is: Ubuntu 18.04.6; CPU: LTS Intel(R) Core(TM) i9-9980HK CPU2.40GHz*2CPU; Memory: 32G; Graphics card: GeForce RTX 2060Mobile. The programming language is Python, using the Pytorch computing framework, ADAM (Adaptive Moment Estimation) optimization is selected when training the model, and the hyperparameters are set to batch_size=128, epoch=80, learning_rate=0.001, dropout=0.1, and the number of graph nodes is 12. In order to verify the classification accuracy of the graph Transformer neural process model proposed in the embodiment of the present application, traditional classification methods such as HA (Historical Average), KNN (K-nearst neighbor), RF (Random Forest), MICE (Multiple Imputation by Chained Equations) and deep learning methods such as RNN (Recurrent Neural Network) are used as comparative experiments. In each iteration, 4 nodes are randomly selected as target nodes to be completed, and the remaining nodes are used as context training nodes.

[0077] In order to compare the advantages and disadvantages of the uncertainty-aware spatiotemporal data completion method proposed in the embodiment of the present application and the baseline method in many aspects, the embodiment of the present application adopts three evaluation indicators, namely, mean absolute error (MAE), root mean square error (RMSE) and mean absolute percentage error (MAPE), and the calculation formula is as follows:

[0078]

[0079]

[0080]

[0081] Where M is the number of all target missing values, 1≤m≤M, is the estimated value, and y is the true value. The smaller the three indicators are, the closer the estimated value is to the true value.

[0082] 1. Comparative Experiments of Different Missing Rates

[0083] In order to verify the generalization ability of the model, comparative experiments with different missing rates of 0.2, 0.4 and 0.6 were carried out respectively. In order to ensure fairness, this experiment first randomly sets a fixed missing rate, and then saves the data. Different baseline methods use the same missing data. In order to reduce the randomness of the experiment, the experimental results take the average of 5 experiments. It can be seen that for all experimental methods, the error of data completion increases with the increase of missing rate. Although MICE has greatly improved the completion accuracy compared with traditional HA and KNN methods, as the missing rate increases, the data completion accuracy of MICE is greatly reduced compared with GraphformerNP. For the two data sets of PM2.5 and CO, the GraphformerNP proposed in the embodiment of the present application achieved good results when the missing rates were 0.2, 0.4 and 0.6, proving the effectiveness of GraphformerNP.

[0084] Table 1 Comparison of GraphformerNP and baseline methods

[0085]

[0086]

[0087] 2. Uncertainty Analysis

[0088] Figure 3 The PM2.5 data completion results of one of the stations from May 3, 2014 to May 4, 2014 are given in the figure. Assuming that a variance of σ 2 The uncertainty is estimated by the Gaussian likelihood, where the shaded part represents the 1σ range centered on the target node completion data. It can be seen from the shape of the solid line and the curve in the figure that the GraphformerNP model has a more accurate completion accuracy. In addition, it can be found from the shaded part that during this time period, all the true values ​​of the site are within the uncertainty estimation range of the method proposed in the embodiment of the present application, indicating that the method proposed in the embodiment of the present application can provide high-quality uncertainty estimation. In practical applications, whether the data status of unknown areas can be accurately inferred is crucial to making better decisions. For example, if the completion results of all methods at a certain unknown location are not very accurate, then the method of deploying sensors at this location in advance to collect data can be adopted to achieve accurate data analysis, and the purpose of cost saving can be achieved.

[0089] 3. Ablation Experiment

[0090] Table 2 shows the ablation experiment results when the missing rate is 0.2. During the experiment, only one of the element modules was changed while the other parts remained unchanged. The embodiment of the present application conducted ablation experiments on two data sets to evaluate the effectiveness of the added elements. GraphformerNP is the method proposed in the embodiment of the present application. w / o local graph convolution, w / o Transformer, and w / o graph sum aggregation respectively represent the completion results obtained by removing local graph convolution, removing Transformer, and removing graph sum aggregation. r / p graph convolution represents the experimental results of replacing local graph convolution with graph convolution. It can be seen from the experimental results that the model without graph sum aggregation has a significant decrease in completion accuracy compared with GraphformerNP. This shows that the use of graph sum aggregation for context node aggregation is crucial to the results of the model of the embodiment of the present application. The results of removing the local graph convolution and Transformer modules are also lower than the GraphformerNP method. At the same time, it can be found that even if the full graph convolution network is used, the accuracy obtained has not improved, indicating that local graph convolution can not only save computational complexity, but also more accurately capture the features of important neighbor nodes.

[0091] Table 2 Ablation experiment

[0092]

[0093]

[0094] 4. Hyperparameter Comparison Experiment

[0095] In order to verify the impact of different variables on the experimental results, the impact of different parameters on the experimental results was analyzed when the missing rate was 0.4 on the PM2.5 dataset. There are mainly 5 hyperparameters in the GraphformerNP model, of which there are three hyperparameters in the Transformer encoding module, namely, the embedding dimension, the number of attention heads, and the number of encoding layers. There are two parameters for the entire model, namely, the learning rate and the batch size. The above hyperparameters were experimentally compared and analyzed using the control variable method.

[0096] Table 3 Hyperparameter comparison experiment

[0097]

[0098] As can be seen from the table, in the Transformer encoding module, the fewer the embedding dimensions, the better the effect. However, when the embedding dimension is adjusted to 32, it does not give the best results. Therefore, it is most appropriate to choose an embedding dimension of 64. The number of attention heads is actually sometimes redundant, so it is better to choose an appropriate number of attention heads. As the number of encoding layers increases, the completion accuracy decreases, so the number of encoding layers is selected as 1. Experiments show that too large a learning rate may cause failure to converge and affect the final result. Therefore, in this experiment, the learning rate is set to 0.001, which is the most appropriate, and the batch size is set to 128, which has the highest completion accuracy.

[0099] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit, wherein the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, the invention content of a spatiotemporal data completion method with uncertainty perception provided by the present invention and some or all steps in each embodiment can be executed. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0100] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of computer programs and their corresponding general hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention are essentially or partly contributed to the prior art can be embodied in the form of a computer program, i.e., a software product, which can be stored in a storage medium and includes several instructions for enabling a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, a MUU or a network device, etc.) to execute the methods described in various embodiments of the present invention or certain parts of the embodiments.

[0101] The present invention provides a spatiotemporal data completion method with uncertainty perception. There are many methods and ways to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention. All components not specified in this embodiment can be implemented by existing technologies.

Claims

1. A method for completing spatiotemporal data with uncertainty perception, characterized in that: The following steps are involved: Step S1, preprocessing the collected spatiotemporal data to construct a training data set, wherein the training data set includes a context node data set and a target node data set; Step S2, constructing a graph Transformer neural process model, and training the graph Transformer neural process model using the training data set; Step S3, using the trained graph Transformer neural process model to complete the missing attribute values ​​of the target nodes in the original spatiotemporal data.

2. The uncertainty-aware spatiotemporal data completion method according to claim 1, characterized in that: The graph Transformer neural process model in step S2 includes a spatiotemporal joint representation learning stage and a random generation stage. The spatiotemporal joint representation learning stage is used to learn the spatial relationship between the target node and the context node, extract the temporal dependency of the spatiotemporal sequence, and obtain the context node spatiotemporal learning sequence and the target node spatiotemporal learning sequence in combination with the characteristics of the context node; the random generation stage is used to aggregate the spatiotemporal coding information of the context node through graph summation aggregation, and derive the random distribution describing the target to be completed in combination with the Gaussian prior distribution and likelihood estimation, thereby completing the missing attribute value of the target node.

3. The uncertainty-aware spatiotemporal data completion method according to claim 2, characterized in that: The spatiotemporal joint representation learning stage includes a spatial encoding layer and a temporal encoding layer. The spatial encoding layer inputs the attribute values ​​of the context node and the target node, learns the spatial relationship between the target node and the context node by local graph convolution, and obtains a spatiotemporal sequence that has captured spatial features; the temporal encoding layer inputs the spatiotemporal sequence that has captured spatial features, extracts the temporal dependency of the spatiotemporal sequence by simplifying the Transformer, and obtains the context node spatiotemporal learning sequence and the target node spatiotemporal learning sequence.

4. The uncertainty-aware spatiotemporal data completion method according to claim 3, characterized in that: The spatial encoding layer learns the spatial relationship between the target node and the context node by local graph convolution to obtain a spatiotemporal sequence that has captured spatial features, including: using K-hop neighbors to update the representation of the target node, the formula is: in, It is the spatiotemporal sequence of captured spatial features obtained through the spatial encoding layer. L represents the total number of graph convolution layers, and l represents the index of the graph convolution layer. represents the K-hop neighbors of the node m to be completed, represents the K-hop adjacency matrix weight value of the target node m to be completed and the context node n, Refers to the encoding result of the context node information, W l It is the learnable weight matrix of the lth layer of graph convolution; 1≤n≤N, 1≤m≤M, N represents the number of context nodes, and M represents the number of target nodes to be completed.

5. The uncertainty-aware spatiotemporal data completion method according to claim 4, characterized in that: The temporal coding layer includes a position coding module, a multi-head self-attention layer, a layer normalization module, a fully connected layer and an output layer. The position coding module is used to perform position coding on the spatiotemporal sequence of the captured spatial features, and the sequence obtained by splicing the coding result with the spatiotemporal sequence of the captured spatial features is input into the multi-head self-attention layer; the multi-head self-attention layer is used to compare each element in the sequence with other elements to calculate different weight values, and use different weight values ​​to generate a comprehensive representation vector to capture complex time dependencies; The residual connection connects the input information to the output of the multi-head attention layer; the layer normalization module normalizes the output of all neurons in the layer; each node in the fully connected layer is connected to all nodes in the previous layer to increase the fitting ability of the model; the output layer is used to generate context node spatiotemporal learning sequences and target node spatiotemporal learning sequences through linear transformation and activation functions.

6. The uncertainty-aware spatiotemporal data completion method according to claim 5, characterized in that: The random generation stage aggregates the spatiotemporal coding information of context nodes through graph summation aggregation, and the calculation formula is as follows: in, Refers to the one-hop neighbor of the target node to be completed, represents the spatiotemporal learning sequence of context nodes, Express Perform a separate encoding operation, A m,n represents the value of the adjacency matrix between the target node m to be completed and the context node n; r represents the representation of the target node after aggregation; By aggregating the information of context nodes, the target latent variable describing the random process of the target node is derived T represents the time step, d l Represents the dimension of the target latent variable at each time step.

7. The uncertainty-aware spatiotemporal data completion method according to claim 6, characterized in that: The random generation stage uses a neural network to learn the mean and variance of the conditional prior distribution. The conditional prior of the target node is recorded as is a decomposable Gaussian distribution: in, represents the context node set, μ Z and It is the parameter learned by the neural network Encoder; The target latent variable Z is composed of the representation of the target node m and the spatiotemporal learning sequence of the context nodes Determine, given the target node spatiotemporal learning sequence and a sample of p(Z), then the neural network It is possible to learn a conditional prior for the target node Contextual node spatiotemporal learning sequence is used to learn the potential observations and then updated by the above formula Parameters.

8. The uncertainty-aware spatiotemporal data completion method according to claim 7, characterized in that: The generation process of the Transformer neural process model is expressed as: Where A represents the adjacency matrix, represents the known features of the target node T time steps, represents the attribute value of the target node T time steps. The first term on the right side of the equation is the likelihood function, assuming that the likelihood is a decomposable Gaussian distribution; the second term is the conditional prior learned by graph summation and aggregation; according to the above formula, the distribution of the attribute value of the target node to be completed is obtained, and the mean of the distribution is taken as the final completion prediction value.

9. The uncertainty-aware spatiotemporal data completion method according to claim 8, characterized in that: Step S3 also includes variable estimation by maximizing the ELBO: in, represents the known features of the context node T time steps, represents the attribute value of the context node T time steps, represents the marginal likelihood, represents the alternative distribution, p(y m |Z,x m ) represents the posterior distribution, where y m represents the attribute value of the target node m, x m represents the known features of the target node m, represents the joint distribution, represents the divergence, Express expectations.

10. The uncertainty-aware spatiotemporal data completion method according to claim 9, characterized in that: Step S1 includes: filtering the collected spatiotemporal data to delete data samples with missing values ​​and retaining spatiotemporal samples without missing items; dividing the retained samples into context node data sets and target node data sets according to different spatial points, each node data includes known features and features to be completed, and the context node data set is recorded as The target node dataset is represents the known features in the context node dataset, Represents the features to be completed in the context node dataset; represents the known features in the target node dataset, Represents the features to be completed in the target node dataset; randomly replaces the features to be completed of the context node according to the set missing rate Set to missing as training data; features to be completed for the target node Problems with unknown initial values, defining a learnable target node encoding e represents the encoding dimension.