IoT Multi-Sensor Time Series Missing Value Completion Method Based on Graph Learning

By introducing copula function and graph learning in multivariate time series completion, combined with Gaussian copula function and graph structure, the problem of ignoring variable relationships and spatial structures in the existing technology is solved, and the accurate completion of missing values ​​of multi-sensor time series is achieved.

CN118861537BActive Publication Date: 2025-06-20DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410363783.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-06-20
Estimated Expiration
2044-03-28

AI Technical Summary

Technical Problem

The prior art ignores the relationship between variables in the completion of missing values ​​of multiple time series, and deep learning methods fail to make full use of the spatial structure of time series.

Method used

A method of combining copula function and graph learning is introduced, and the dependencies between different variables in multiple time series are extracted through the Gaussian copula function, and the graph structure is used to integrate time information and spatial information for missing value completion.

Benefits of technology

Accurate completion of missing values ​​of multi-sensor time series is achieved, making full use of the dependence between spatial information of the time series and variables, and improving the accuracy and reliability of completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118861537B_ABST
    Figure CN118861537B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for filling missing values in multi-sensor time series of the Internet of Things based on graph learning, belonging to the technical field of time series analysis. Aiming at the missing values in multi-sensor time series, the present invention studies how to use a bidirectional GRU module based on graph learning for filling, including the following steps: 1) Using copula analysis technology to automatically extract the correlation between different sensors, and better explore the spatial correlation contained in different sensors in the Internet of Things; 2) Through the message passing mechanism MPNN module and the gated recurrent unit GRU, the spatio-temporal information of the entire multi-sensor time series of the Internet of Things is fully extracted, and the missing values are more comprehensively filled; 3) Through the forward module and the reverse module, the entire multi-sensor time series of the Internet of Things is analyzed from different time dimensions, fully considering the past and future details of the missing value points, and a comprehensive analysis of the missing values is made through MLP combined with a large number of collected features, and finally the filling is carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer graph learning and time series analysis, and more specifically, to a method for completing missing values in multi-sensor time series of the Internet of Things based on graph learning. Background Art

[0002] The Internet of Things usually relies on multiple sensors to collect and monitor real-time data, so as to realize the perception, connection and control of the physical world. Sensors play a key role in the Internet of Things, providing various data that can be used for various applications such as environmental monitoring, object tracking, parameter measurement, etc. However, due to problems such as hardware failures and transmission errors, some values in the data collected by sensors are often lost. Therefore, from the perspective of improving data availability and ensuring decision-making accuracy, the completion algorithm for such multivariate time series has important application value.

[0003] Traditional multivariate time series completion methods mainly rely on statistics and time series analysis techniques. For example, the mean, median or mode of other variables at the same time point are used to fill in the missing values, or the observed values at the previous time point or the next time point are used to fill in the missing values. This is a simple and intuitive method, but it ignores the relationships between variables. In addition, for time series with autoregressive and moving average structures, the autoregressive integrated moving average model (ARIMA) can be used for interpolation and prediction. On the other hand, many deep learning-based methods have also been widely applied to solve this problem, such as means like RNN, LSTM, and graph learning.

[0004] For the problem of time series completion, there are already many means to solve it using traditional time series analysis techniques, but such methods have various problems. For example, the moving average method only considers the linear relationship of the time series and only uses the observed values in the past of the missing points for calculation; the exponential smoothing method is sensitive to the selection of the initial value; the effect of the arima model is easily affected by outliers.

[0005] In recent years, deep learning techniques have also begun to be widely applied in this field. First of all, the deep autoregressive method based on RNN has achieved extensive success. For example, the GRU-D model, which can learn how to handle sequences with missing data by controlling the decay of the hidden state of the gated RNN. There are also some methods that use generative adversarial networks (GAN) to complete the missing values. However, the model structures adopted by these methods cannot extract the dependencies between different variables in the time series and ignore the utilization of the spatial structure of the multivariate time series.

[0006] To utilize the spatial information of multivariate time series, some methods employ graph learning to model the dependencies between different variables in multivariate time series using a graph structure. For example, the GRIN method represents the correlations between various data based on the distances between different sensors. The drawback of such methods is that they only consider partial information in multivariate time series and do not fully extract and utilize the spatial structure. Summary of the Invention

[0007] In view of the problems existing in the above background technology, the present invention introduces the copula function into graph learning to construct a method that can fully utilize the temporal and spatial information of multi-sensor time series to complete the missing values in multivariate time series. On the one hand, the copula function is a mathematical tool based on Sklar's theorem that can connect the multivariate distribution with the marginal distributions corresponding to its individual variables. Specifically in the application scenario of the present invention, it can fully extract the dependency relationships between different variables in multivariate time series. On the other hand, the graph learning structure incorporates both temporal and spatial information into the model, analyzes from the overall multi-sensor time series, and finally accurately completes the missing values.

[0008] To achieve the above object, the present invention provides a method for completing missing values in multi-sensor time series of the Internet of Things based on graph learning, including the following steps:

[0009] S1. Mark missing values: Select a multi-sensor time series S of the Internet of Things with partial data missing, and import it into this algorithm through the PyTorch toolkit to generate a marking matrix M;

[0010] S2. Data processing: Normalize the multi-sensor time series S of the Internet of Things to obtain a normalized data set X;

[0011] S3. Data sampling: Use the Python tool to screen out the complete data in the normalized data set X to obtain the Internet of Things data set X';

[0012] S4. Calculate the correlation coefficient: Use the Gaussian copula function to model the Internet of Things data set X', calculate the correlation coefficients between each sensor, form a correlation matrix R, and use the correlation coefficient matrix R to complete the Internet of Things data set;

[0013] S5. Spatiotemporal encoding: Input the Internet of Things data set X' and the marking matrix M into the spatiotemporal encoder for spatiotemporal encoding to obtain a hidden state matrix;

[0014] S6. Spatial decoding: Use the spatial decoder to decode the spatiotemporal encoding sequence to obtain the algorithm prediction value at time t, use the algorithm prediction value to preliminarily fill the missing values to obtain a data set, and then aggregate the data set to obtain a feature matrix;

[0015] S7, Two-way time prediction: According to the forward and reverse order of time, perform steps S5 to S6 respectively to obtain the forward hidden state matrix, forward feature matrix, reverse hidden state matrix, and reverse feature matrix. Use the MLP to combine the forward hidden state matrix, forward feature matrix, reverse hidden state matrix, and reverse feature matrix to complete the missing values.

[0016] Under the preferred scheme, the elements in the annotation matrix M are either 0 or 1. 0 indicates that the data at the corresponding position in the IoT multi-sensor time series S is missing, and 1 indicates that the data at the corresponding position in the IoT multi-sensor time series S is complete.

[0017] Under the preferred scheme, calculating the correlation coefficient in step S4 includes the following steps:

[0018] S41, Use the Gaussian copula function to model the IoT dataset X':

[0019] C(x ·1 ,x ·2 ,……,x ·d ; ∑) = Φ d (Φ -1 (x ·1 ), Φ -1 (x ·2 ),……, Φ -1 (x ·d ); 0, R) (4-1)

[0020] Among them, C(x ·1 ,x ·2 ,……,x ·d ; ∑) represents the copula function, Σ represents the covariance matrix of the copula function, x ·1 ,x ·2 ,……,x ·d is the set of observed values of each sensor in the IoT dataset X', Φ d is a multivariate Gaussian distribution with a mean of 0, R represents the correlation coefficient matrix of Φ d , the dimension of R is d×d, and Φ -1 is the quantile function of the standard Gaussian distribution;

[0021] S42, Take the derivative of the established model to obtain the copula density function of C:

[0022]

[0023] Among them, detR represents the determinant of R, exp(·) represents the exponential function, and I d represents the d-dimensional identity matrix;

[0024] S43. Calculate the probability density function:

[0025]

[0026] where f(·) is the density function, c is the copula density function, and f i (x ·i ) represents the distribution function of each variable x ·i .

[0027] S44. Find the contingency function:

[0028]

[0029] S45. Find the correlation coefficient matrix: Substitute X' into formula (4-4), and combine equations according to formula (4-2) and formula (4-4) to find the correlation coefficient matrix R.

[0030] Under the preferred solution, the spatio-temporal coding described in step S5 includes the following steps:

[0031] S5. Spatial coding: Use the pytorch toolkit in python for spatial coding, and the formula is as follows:

[0032]

[0033] where MPNN() represents the message passing function, x ki is the observation value obtained by sensor i at time k; x kj is the observation value obtained by sensor j at time k; R is the correlation coefficient matrix, γ, ρ are two learnable multi-layer perceptrons MLP; r ij represents the correlation coefficient between sensor i and sensor j;

[0034] S52. Use the gated recurrent unit to perform temporal coding on formula (5-1), and the formula is as follows:

[0035]

[0036]

[0037]

[0038]

[0039] where are the reset gate and the update gate respectively, is the hidden state of the Internet of Things data at time point t, is the hidden state of the Internet of Things data at time point t-1, is the value for this time step, is the element in the annotation matrix M corresponding to to obtain the spatio-temporal coding sequence

[0040] In the preferred solution, the spatial decoding described in step S6 includes the following steps:

[0041] S61. Use the spatial decoder to decode the spatio-temporal coding sequence to obtain the algorithm prediction value at time t, as follows:

[0042]

[0043] where, is the prediction value generated by the algorithm for time t, h t-1 is the hidden state at time t-1 output by the spatio-temporal encoder, V is a learnable matrix, and b is a learnable bias vector;

[0044] S62. Initially fill the original missing values with the prediction values:

[0045]

[0046] where, represents the dataset at time t after initial filling, m tj , x tj is the value in the annotation matrix M and the Internet of Things multi-sensor time series X corresponding to time t, represents the bitwise inversion of m tj ;

[0047] S63. Aggregate the data at time t using the MPNN function, as shown in formula (6-3):

[0048]

[0049] where, represents the feature matrix after initial completion of each node, MPNN is the same as formula (5-1), m ti , respectively represent whether the data of the i-th sensor is missing at time t (m ti being 0 indicates missing) and its hidden state, m ti represents the element in the annotation matrix.

[0050] In the preferred solution, the two-way time prediction described in step S7 has the following formula:

[0051]

[0052] where, denotes the filled value for the vacancy of the i-th sensor at the t-th time point. MLP stands for Multi-Layer Perceptron. respectively represent the forward feature matrix and the backward feature matrix of the node. respectively represent the forward hidden state matrix and the backward hidden state matrix of the node.

[0053] Under the preferred solution, the data processing in step S2 is to substitute the Internet of Things time series S into formula (2-1) as follows:

[0054]

[0055] where s ij represents the data observed by sensor j at the i-th moment in S, and s ·j represents all the data observed by sensor j. respectively represent the maximum and minimum values among all the data observed by sensor j. x ij represents a normalized data, and X represents the set of all normalized data.

[0056] Under the preferred solution, a method for filling missing values in the Internet of Things multi-sensor time series based on graph learning includes the following steps:

[0057] S1. Mark the missing values: Select a part of the Internet of Things multi-sensor time series S with missing data and generate a marking matrix M as follows:

[0058] S=(s ij ) t×d (1-1)

[0059] M = numpy.where(S == None, 0, 1) (1-2)

[0060] where s ij represents the data observed by sensor j at the i-th moment, t represents the observation time, and d represents the number of sensors; None represents the missing value in the Internet of Things multi-sensor time series S, 0 represents the data missing at the corresponding position in the Internet of Things multi-sensor time series S, and 1 represents the data complete at the corresponding position in the Internet of Things multi-sensor time series S. The data in the Internet of Things multi-sensor time series S is recorded in a csv file, and the where function in the numpy toolkit in python is used to generate the marking matrix M.

[0061] S2. Data standardization processing: Perform normalization processing on the Internet of Things multi-sensor time series, that is, substitute the Internet of Things multi-sensor time series S into formula (2-1) to obtain the normalized data set X. The formula is as follows:

[0062]

[0063] Among them, s ij represents the data observed by sensor j in S at the i-th moment, and s ·j represents all the data observed by sensor j, respectively represent the maximum and minimum values among all the data observed by sensor j, and x ij represents a normalized data, and X represents the set of all normalized data.

[0064] In some embodiments of the present invention, the normalized data set X is divided, and the training set and the test set are divided by using the random_split function in the pytorch toolkit as follows:

[0065] Dividing the data set: Divide the normalized data set X into a training set and a test set:

[0066] dataset=(X,M) (8-1)

[0067] train,test=random_split(datast,[a,b]) (8-2)

[0068] Among them, train represents the training set, test represents the test set, a represents the division ratio of the training set, b represents the division ratio of the test set, and a + b = 1; the training set is used for algorithm training; the test set is used to verify whether the algorithm is effective, and the effectiveness of this algorithm will be verified by the experimental results of different algorithms in the test set later.

[0069] In some embodiments of the present invention, the training set and the test set are divided according to the ratio of 0.8:0.2, and the formula is as follows:

[0070] train,test=random_split(dataset,[0.8,0.2]).

[0071] In some embodiments of the present invention, the training set and the test set are divided according to the ratio of 0.7:0.3.

[0072] S3. Data sampling: Use python to screen out the complete data of X in the normalized Internet of Things multi-sensor time series to obtain the Internet of Things data set X'.

[0073] S4. Calculate the correlation coefficient: Use the Gaussian copula function to model the Internet of Things data set X', and calculate the correlation coefficients between sensors to form a correlation matrix R.

[0074] Specifically, it includes the following steps:

[0075] S41. Model the Internet of Things dataset X' using the Gaussian copula function:

[0076] C(x ·1 ,x ·2 , ……,x ·d ; ∑) = Φ d (Φ -1 (x ·1 ), Φ -1 (x ·2 ), ……, Φ -1 (x ·d ); 0, R) (4 - 1)

[0077] Among them, C(x ·1 ,x ·2 , ……,x ·d ; Σ) represents the copula function, Σ represents the covariance matrix of the copula function, x ·1 ,x ·2 , ……,x ·d is the set of observed values of each sensor in the Internet of Things dataset X', Φ d is a multivariate Gaussian distribution with a mean of 0, R represents the correlation coefficient matrix of Φ d , the dimension of R is d×d, Φ -1 is the quantile function of the standard Gaussian distribution.

[0078] S42. Use the numpy toolkit to take the derivative of formula (4 - 1) to obtain the copula density function of C:

[0079]

[0080] Among them, detR represents the determinant of R, exp(·) represents the exponential function, I d represents the d-dimensional identity matrix;

[0081] S43. Further, use the numpy toolkit to find the probability density function according to formula (5 - 2), and the formula is as follows:

[0082]

[0083] Among them, f(·) is the density function, c is the copula density function, f i (x ·i ) represents the distribution function of each variable x ·i ;

[0084] S44. Find the contingency function:

[0085]

[0086] S45. Calculate the correlation coefficient matrix: Substitute X' into formula (4-4), combine equations according to formulas (4-2) and (4-4), and obtain the correlation coefficient matrix R as the adjacency matrix E.

[0087] S5. Space-time coding, including the following steps:

[0088] S51. Space coding: Use the pytorch toolkit in python for space coding, and input the labeled matrix M of the IoT dataset X' into the encoder according to time points for space coding. The formula is as follows:

[0089]

[0090] where x ki is the observation value obtained by sensor i at time k; x kj is the observation value obtained by sensor j at time k; R is the correlation coefficient matrix, γ, ρ are two learnable multi-layer perceptrons MLP, and γ, ρ are built through the pytorch toolkit; r ij is the corresponding element in matrix R, representing the correlation coefficient between sensor i and sensor j;

[0091] S52. Use the gated recurrent unit to perform time coding on formula (5-1), and the formula is as follows:

[0092]

[0093]

[0094]

[0095]

[0096] where, are the reset gate and update gate respectively, is the hidden state of the IoT data at time point t, is the hidden state of the IoT data at time point t-1. In the initial stage of this algorithm, a 0 vector will be input as the initial After that, will be obtained according to the of the previous time step through formula (5-5); is the value of this time step, obtained from X', is the element in the labeled matrix obtained in S1 and corresponding to; is the output of the previous time step. The symbols ⊙ and || represent the Hadamard product and concatenation operator respectively to obtain the space-time coding sequence

[0097] S6, Spatial Decoding:

[0098] First, according to the spatio-temporal coding sequence obtained in step S52, decode the algorithm prediction value of the point with missing value at time t:

[0099]

[0100] Among them, is the prediction value for time t generated by the algorithm, h t-1 is the hidden state at time t-1 output by the spatio-temporal encoder, obtained from the spatio-temporal coding sequence H[t,t+T] above, V is a learnable matrix, b is a learnable bias vector, and V and b are tensors initialized by pytorch;

[0101] Next, preliminarily fill the original missing values with the prediction values:

[0102]

[0103] Among them, represents the dataset at time t after preliminary filling, m tj , x tj is the value in the annotation matrix M and the IoT dataset X' corresponding to time t, represents the bitwise negation of the original m tj .

[0104] Finally, use the MPNN function to aggregate the data at time t, as shown in formula (6-3):

[0105]

[0106] Among them, represents the feature matrix after preliminary completion of each node through the spatio-temporal encoder, MPNN is the same as formula (5-1), m ti , respectively represent whether the data of the i-th sensor is missing and its hidden state at time t.

[0107] To make full use of the information of the missing value points in the past and future, this model adopts a bidirectional structure. The two modules have the same structure, both including the spatio-temporal encoder and the spatial decoder mentioned above. The difference is that the input order of the forward module is x t,j , x t+1,j , …, x t+T,j , while the input order of the reverse module is the opposite. In the final stage of this algorithm, the information obtained from the forward module is combined with the information obtained from the reverse module through the MLP to complete the original missing values.

[0108] S7. Bidirectional Time Prediction: Combine the information obtained from the forward module and the reverse module through an MLP to complete the original missing values:

[0109]

[0110] Among them, represents the completed value of the missing value of the i-th sensor at the t-th time point by the algorithm. MLP represents a multi-layer perceptron, which is implemented using the corresponding functions in the pytorch package. It should be noted that this algorithm includes a forward module and a reverse module, and the results obtained from the two modules are distinguished by the superscripts fwd and bwd. respectively represent the feature matrices obtained by this node through the forward module and the reverse module; respectively represent the hidden state matrices obtained by this node i through the forward module and the reverse module.

[0111] Detailed Overview of the Present Invention

[0112] A method for filling missing values in multi-sensor time series of the Internet of Things based on graph learning, comprising the following steps:

[0113] S1: Mark Missing Values: Generate a marking matrix M through the where function in the numpy toolkit in python. The formula is as follows:

[0114] M = numpy.where(X == None, 0, 1) (1 - 2)

[0115] Among them, M is the marking matrix, and S is the Internet of Things dataset. Usually, the observed data of each sensor in the Internet of Things can be recorded in a csv file. This file (contains a total of t * d data observed by d sensors after t time points, and each data is a real number. From a computer science perspective, it can be regarded as a multi-sensor time series. Read this csv file into the program of the present invention using python tools to form the Internet of Things dataset S. The data of S can be expressed as S = (s ij ) t×d , which is a matrix. In addition, some data in S are missing and are marked with None. None represents the missing value in the Internet of Things dataset; 1 represents that the data at the corresponding position in the Internet of Things dataset S is complete; 0 represents that there is data missing at the corresponding position in the Internet of Things multi-sensor time series.

[0116] S2: Data Normalization: Substitute the Internet of Things dataset S into the normalization formula (2 - 1) for standardization processing. The calculation process is as follows to form the normalized dataset X:

[0117]

[0118] Among them, \(x\) ij represents the data observed by sensor \(j\) at the \(i\)-th moment in the normalized Internet of Things dataset \(X\); \(x\) ·j represents all the data observed by sensor \(j\); respectively represent the maximum and minimum values among all the data observed by sensor \(j\);

[0119] Dividing the training set and the test set: Using the PyTorch toolkit, the data in the Internet of Things multi-sensor time series are divided into a training set and a test set according to a certain ratio. Conventionally, it is generally divided according to a ratio of 8:2 or 7:3;

[0120] train, test = random_split(dataset, [0.8, 0.2])

[0121] Among them, train and test represent the training set and the test set respectively. dataset contains the normalized dataset \(X\) and the annotation matrix \(M\), that is, dataset = (\(X\), \(M\)), and random_split is a function in the PyTorch toolkit for dividing the dataset.

[0122] S3: Data sampling: Use Python to screen out the complete data that does not need to be filled in the Internet of Things dataset \(X\) to form the dataset \(X'\);

[0123] S4: Calculate the correlation coefficient: Use the copulae tool in Python to model the complete data set \(X'\) screened in step S3 through the Gaussian copula function as:

[0124]

[0125] Among them, \(C\) represents the copula function, and the right side of formula (4-1) can be abbreviated as \(C\), \(\sum\) is its covariance matrix; \(x\) ·1 , \(x\) ·2 , … \(x\) ·d is the set of the observed values of each sensor in the Internet of Things data \(X'\) obtained through step S3; \(\varPhi\) d is a multivariate Gaussian distribution with a mean of 0, and its correlation coefficient matrix \(R\) is a square matrix of dimension \(d\times d\), which is an unknown quantity; \(\varPhi\) -1 is the quantile function of the standard Gaussian distribution;

[0126] Use the numpy toolkit to take the derivative of formula (4-1) to obtain the copula density function of \(C\) as:

[0127]

[0128] where detR is the determinant of R; X` is the Internet of Things dataset obtained in step S4, as above; I d is the d-dimensional identity matrix;

[0129] Further, the probability density function is obtained through the numpy toolkit according to (4-2) as follows:

[0130]

[0131] where f(·) is the density function, c is the copula density function, and the meanings of other symbols are as above;

[0132] Further, its likelihood function is obtained, as shown in formula (5-4):

[0133]

[0134] where F is the likelihood function. Through the numpy toolkit, the dataset X` in step S4 is substituted into formula (4-4), and the correlation coefficient matrix R is obtained as the adjacency matrix E according to the system of equations formed by formulas (4-2) and (4-4);

[0135] S5: Spatiotemporal Encoding: The dataset X and the annotation matrix M are input into the encoder step by step according to time points through the pytorch toolkit; The message passing function is built using the pytorch toolkit in Python, and spatial encoding is realized through the message passing function MPNN, as shown in formula (5-1):

[0136]

[0137] where x ki is the observation value obtained by sensor i at time k; R is the correlation coefficient matrix obtained in step S5, γ, ρ are two learnable multi-layer perceptrons MLP, built through the pytorch toolkit; r ij represents the correlation coefficient between sensor i and sensor j, which is the corresponding element in matrix R;

[0138] Then, the gated recurrent unit is used to perform time encoding on formula (6-1):

[0139]

[0140]

[0141]

[0142]

[0143] where, where, They are the reset gate and the update gate respectively. is the hidden state of the Internet of Things data at time point t. is the output of the previous time step. The symbols ⊙ and || represent the Hadamard product and the concatenation operator respectively. The spatio-temporal coding sequence is obtained.

[0144] S6: Spatial decoding: First, according to the spatio-temporal coding sequence obtained in step S6, decode the algorithm prediction value of the points with missing values at time t:

[0145]

[0146] Among them, is the prediction value generated by the algorithm for time t, h t-1 is the hidden state at time t-1 output by the spatio-temporal encoder. V is a learnable matrix, b is a learnable bias vector. V and b are tensors initialized by pytorch.

[0147] Next, preliminarily fill the original missing values with the prediction values:

[0148]

[0149] Among them, θ(y t ) represents the data set at time t after preliminary filling, m tj , x tj is the annotation matrix corresponding to time t and the value in the Internet of Things multi-sensor time series. represents the bitwise inversion of the original m tj . After that, aggregate the data at this time using the MPNN function, as shown in formula (6-2):

[0150]

[0151] Among them, represents the feature matrix after preliminary completion of each node through the spatio-temporal encoder. MPNN is the same as formula (5-1), m ti , respectively represent whether the data of the i-th sensor is missing at time t and its hidden state.

[0152] S7: Bidirectional time prediction: Combine the information obtained by the forward module with the information obtained by the reverse module through the MLP to complete the original missing values:

[0153]

[0154] Among them, denotes the filled value of the algorithm for the i-th sensor at the t-th time point. MLP stands for Multi-Layer Perceptron, which is implemented using the corresponding functions in the pytorch package. respectively represent the feature matrices obtained by the node through the forward module and the backward module; respectively represent the hidden state matrices obtained by the node through the forward module and the backward module.

[0155] The beneficial effects of the present invention are:

[0156] For the missing values in multi-sensor time series, the present invention studies how to use the filling based on the bidirectional GRU module of graph learning. It has the following advantages: 1) Using copula analysis technology to automatically extract the correlation between different sensors, and better explore the spatial correlation contained in different sensors in the Internet of Things; 2) Through the message passing mechanism MPNN module and the gated recurrent unit GRU, the spatio-temporal information of the multi-sensor time series of the entire Internet of Things is fully extracted, and the missing values are more comprehensively filled; 3) Through the forward module and the backward module, the multi-sensor time series of the entire Internet of Things is analyzed from different time dimensions, fully considering the past and future details of the missing value points, and a comprehensive analysis of the missing values is made through MLP combined with a large number of collected features, and finally filled. Brief Description of the Drawings

[0157] Figure 1 is a schematic diagram of the algorithm of the present invention. Detailed Embodiment

[0158] To make the technical solution of the present invention clearer, the following further describes the specific embodiments of the present invention in detail with reference to the drawings and specific embodiments. The multi-sensor time series filling method based on graph learning involved in the present invention uses the spatio-temporal information of the data before and after the missing value to fill it through methods such as copula function, graph learning, and GRU. Specifically, the method is as described below.

[0159] As Figure 1 shown, the method for filling missing values of multi-sensors in the Internet of Things based on graph learning includes the following steps:

[0160] I. Data Preprocessing

[0161] The task of this stage is to convert the Internet of Things data into a format that can be processed by the present algorithm invention, including operations such as annotating missing values and dividing the data set.

[0162] S1: Marking missing values: The IoT multi-sensor time series processed by the present invention can be represented as X, which is a matrix of dimension t*d. Generally, the data observed by each sensor in the IoT can be recorded in a csv file, which contains a total of t*d data observed by d sensors after t time points. Each data is a real number. From the perspective of computer science, it can be regarded as a multi-sensor time series. Among them, t represents the number of time points recorded in the IoT dataset X; d represents the number of sensors in the IoT dataset X. In order to mark the positions of the missing values in the dataset, the present invention generates a marking matrix M in the data preprocessing stage through the where function in the numpy toolkit in python. The dimension of M is the same as that of X. M is a 0-1 matrix, that is, the value of each element in the matrix M is either 0 or 1. As shown in the following formula:

[0163] M = numpy.where(X == None, 0, 1) (1-2)

[0164] Among them, M is the marking matrix, X is the IoT dataset (including a total of t*d data observed by d sensors after t time points. Each data is a real number. Among them, some data are missing and are marked with None), None represents the missing value, indicating that the data at the corresponding position in the IoT dataset X is complete; 0 represents that there is data missing at the corresponding position in the IoT multi-sensor time series and needs to be complemented by this algorithm.

[0165] S2: Data normalization: Substitute the above-mentioned IoT dataset X into the normalization formula (2-1) for standardization processing to ensure the consistency of the dimensions of the training samples, so as to better train the model and obtain the normalized training samples.

[0166]

[0167] Among them, x ij represents the data observed by sensor j at the i-th moment in the IoT dataset X; x ·j represents all the data observed by sensor j; respectively represent the maximum and minimum values among all the data observed by sensor j.

[0168] Then, divide the training set and the test set: Use the pytorch toolkit to divide the data in the IoT multi-sensor time series into the training set and the test set, and the ratio between the two can be 8:2. It should be noted that there is no strict requirement for the ratio of the training set to the test set. By convention, it is divided according to the ratio of 8:2 or 7:3.

[0169] train, test = random_split(dataset, [0.8, 0.2])

[0170] Among them, train and test represent the training set and the test set respectively. Dataset contains the normalized dataset X and the annotation matrix M, that is, dataset = (X, M). random_split is a function for dividing the dataset built into the pytorch toolkit.

[0171] II. Constructing the adjacency matrix using the copula function

[0172] The purpose of this stage is to use the copula function to extract the correlation between the data observed by different sensors in the IoT multi-sensor time series. The present invention uses a graph structure to model the correlation information between multiple sensors, as shown in formula (4):

[0173] G = (V, E) (4)

[0174] Among them, G represents a graph; V represents the set of points of the graph, which are the data of each sensor in the dataset X in the present invention and are known quantities; E represents the relationship between different points in the graph. In the present invention, E is a symmetric matrix of size d * d (d is the number of sensors mentioned above), and the elements in E are represented as e mn , representing the weight relationship between sensor m and sensor n in the graph structure. In this stage, in order to solve the correlation information between different sensors, first sample the original time series, and then use the copula toolkit in python to calculate the correlation coefficients between each sensor, and form a correlation coefficient matrix R as the adjacency matrix E.

[0175] S4: Data sampling: For the IoT multi-sensor time series, the sampling interval of the sensor is short, which results in the values of adjacent time points being relatively similar. On the other hand, the IoT multi-sensor time series processed by the present invention contains missing values. To ensure that the finally obtained correlation coefficients are accurate and effective, in this step, the time points with complete observed data will be screened out through python tools for calculation, and the complete data that does not need to be filled in the IoT dataset X will be screened out through python to form the dataset X`.

[0176] S4: Calculating the adjacency matrix E: The present invention uses the numpy toolkit in python to calculate the correlation coefficient matrix R equivalent to the adjacency matrix E. Specifically, the complete data set X` screened in S4 can be modeled as a Gaussian copula function through the numpy toolkit:

[0177]

[0178] Among them, C represents the copula function, and the right side of formula (4-1) can be abbreviated as C, Σ is its covariance matrix; x ·1 , x ·2 , … x ·d is the set of observed values of each sensor in the Internet of Things data X` obtained through S4; Φ d is a multivariate Gaussian distribution with a mean of 0, and its correlation coefficient matrix R is a square matrix with a dimension of d*d, which is an unknown quantity; Φ -1 is the quantile function of the standard Gaussian distribution.

[0179] Next, use numpy to obtain the copula density function of C according to formula (4-2) as follows:

[0180]

[0181] Among them, detR is the determinant of R; X` is the Internet of Things data set obtained in step S3, the same as above; I d is the d-dimensional identity matrix. Then, the density function of formula (4-2) obtained by numpy is shown in formula (4-3):

[0182]

[0183] Among them, f(·) is the density function, c is the copula density function, and the meanings of other symbols are the same as above.

[0184] Further, find its likelihood function, as shown in formula (4-4):

[0185]

[0186] Among them, F is the likelihood function. Substitute the processed Internet of Things multi-sensor time series X` in S3 into it, and use the numpy toolkit to solve the covariance matrix Σ and the correlation coefficient matrix R according to formulas (4-2) and (4-4). In the present invention, the correlation coefficient matrix R is used as the adjacency matrix E of the graph G.

[0187] III. Bidirectional spatio-temporal encoder

[0188] The purpose of this step is to use the normalized data set X in step one, the annotation matrix M and the adjacency matrix E in step two to complete the missing values in the original data set X through a spatio-temporal encoder, a spatial decoder and an MLP (multi-layer perceptron).

[0189] S5: Spatiotemporal Encoder: This step is to combine the temporal information and spatial information of the dataset X for encoding. Specifically, the dataset X and the annotation matrix M are input into the encoder step by step according to time points through the pytorch toolkit. First, spatial encoding is achieved through the message passing function MPNN. The MPNN is built using the pytorch toolkit in python, as shown in Equation (5-1):

[0190]

[0191] where, x ki is the observation value obtained by sensor i at time k; R is the correlation coefficient matrix obtained in Step 2, and γ, ρ are two learnable multi-layer perceptrons MLP, which are implemented through the pytorch toolkit; r ij represents the correlation coefficient between sensor i and sensor j, which is the corresponding element in matrix R.

[0192] Next, the gated recurrent unit is used for temporal encoding. The GRU is implemented through the message passing layer of Equation (5-1). Specifically for each sensor, it can be described by Equations (5-2)(5-3)(5-4)(5-5):

[0193]

[0194]

[0195]

[0196]

[0197] where, are the reset gate and update gate respectively, is the hidden state of the IoT data at time point t, is the output of the previous time step. The symbols ⊙ and || represent the Hadamard product and concatenation operator respectively. The initial value of can be initialized either to a constant or to a learnable embedding. Note that for steps with missing input data, the encoder will get predictions from the decoder block, which will be explained in the next step. Through the above calculations, we obtain the spatiotemporal encoding sequence

[0198] S6: Spatial Decoder: First, according to the spatiotemporal encoding sequence obtained above, the algorithm prediction value of the point with missing value at time t is decoded, as shown in Equation (6-1):

[0199]

[0200] where, is the predicted value for time t generated by the algorithm, h t-1 is the hidden state at time t-1 output by the spatio-temporal encoder, V is a learnable matrix, b is a learnable bias vector, and V and b are tensors initialized by pytorch.

[0201] Next, the original missing values are preliminarily filled with the predicted values to obtain the input for the next time step, as shown in Equation (6-2):

[0202]

[0203] where, θ(y t ) represents the dataset at time t after preliminary filling, m tj , x tj is the annotation matrix corresponding to time t and the value in the IoT multi-sensor time series, represents the bitwise inversion of the original m tj . Then, the data at this time step is aggregated using the MPNN function mentioned above, so that the data in each point contains the information in other nodes, as shown in Equation (6-3):

[0204]

[0205] where, represents the feature matrix after preliminary completion of each node through the spatio-temporal encoder, MPNN is the same as Equation (5-1), m ti , respectively represent whether the data of the i-th sensor is missing at time t and its hidden state.

[0206] In summary, through the spatio-temporal encoder and the spatial decoder, the hidden state matrix H[t,t+T] and the feature matrix P[t,t+T] of each sensor at each time point are obtained.

[0207] S7: Bidirectional time prediction: In order to make full use of the information of the missing value points in the past and the future, this model adopts a bidirectional structure, as Figure 1 shown. The two modules have the same structure, both including the spatio-temporal encoder and the spatial decoder mentioned above. The difference is that the input order of the forward module is x t,j , x t+1 , j, …, x t+T,j , while the input order of the reverse module is the opposite. In the final stage of this algorithm, the information obtained from the forward module and the information obtained from the reverse module are combined through MLP to complete the original missing values, as shown in Equation (6-3):

[0208]

[0209] Among them, represents the complement value of the algorithm for the vacancy of the i-th sensor at the t-th time point. MLP represents a multi-layer perceptron, which is implemented using the corresponding functions in the pytorch package. respectively represent the feature matrices obtained by this node through the forward module and the reverse module; similarly, respectively represent the hidden state matrices obtained by this node through the forward module and the reverse module.

[0210] Example 1

[0211] The dataset is the ETTh dataset, which contains 8 sensors. Each sensor records information every 1 hour, with a total of two years of detection data.

[0212] According to the specific steps mentioned above, the present invention conducted corresponding experiments on the IoT dataset with the structure described above and compared it with common completion algorithms. Specifically, in this experiment, part of the data in the complete dataset was deliberately covered to simulate the situation of data loss. The algorithm was used to complete specific positions and compared with the true values. Taking mse (mean square error) as the measurement index, the smaller the MSE, the better the effect. The results are as follows:

[0213]

[0214] It can be clearly seen from the above table that the error of the present invention is smaller than that of several common algorithms, and the accuracy is better, which proves that the model has achieved good results in the aspect of IoT dataset completion.

[0215] The present invention uses a graph structure to complete the IoT multi-sensor time series, and its significance is relatively large. The present invention fully extracts and analyzes the spatio-temporal feature information of the IoT, and uses related technologies such as graph learning and gated recurrent units to use spatio-temporal information to complete missing values. The algorithm proposed in the present invention fully extracts the spatio-temporal information in the IoT multi-sensor time series, making the completion result more accurate.

[0216] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A method for completing missing values ​​of multi-sensor time series in the Internet of Things based on graph learning, characterized in that: The steps include: S1, labeling missing values: select the IoT multi-sensor time series S with some missing data, import it into this method through the pytorch toolkit to generate the labeling matrix M; S2, data processing: normalize the IoT multi-sensor time series S to obtain a normalized data set X; S3, data sampling: use python tools to filter out the complete data in the normalized data set X and obtain the IoT data set X'; S4. Calculate the correlation coefficient: Use the Gaussian copula function to model the IoT dataset X', and calculate the correlation coefficients between the sensors to form a correlation matrix R, and use the correlation coefficient matrix R to complete the IoT dataset; S5, space-time coding: Input the IoT dataset X' and the label matrix M into the space-time encoder for space-time coding to obtain the hidden state matrix; S6, spatial decoding: Use the spatial decoder to decode the spatiotemporal coding sequence to obtain the algorithm prediction value at time t, use the algorithm prediction value to preliminarily fill the missing values ​​to obtain a data set, and then aggregate the data set to obtain a feature matrix; S7, bidirectional time prediction: in the forward and reverse order of time, steps S5 to S6 are performed respectively to obtain a forward hidden state matrix, a forward feature matrix, a reverse hidden state matrix and a reverse feature matrix, and the forward hidden state matrix, the forward feature matrix, the reverse hidden state matrix and the reverse feature matrix are combined by MLP to complete the missing values; The calculation of the correlation coefficient in step S4 includes the following steps: S41. Use Gaussian copula function to model the IoT dataset X': C(x ·1 ,x ·2 ,……,x ·d ;∑)=Φ d (F -1 (x ·1 ),F -1 (x ·2 ),……,Φ -1 (x ·d );0,R) (4-1) Among them, C(x ·1 ,x ·2 ,……,x ·d ; ∑) represents the copula function, Σ represents the covariance matrix of the copula function, x ·1 ,x ·2 ,……,x ·d is the set of sensor observations in the IoT dataset X', Φ d is a multivariate Gaussian distribution with a mean of 0, and R represents Φ d The correlation coefficient matrix of R is d×d, Φ -1 is the quantile function of the standard Gaussian distribution; S42. Derivative the established model to obtain the copula density function of C: Where detR represents the determinant of R, exp(·) represents the exponential function, and I d represents the d-dimensional identity matrix; S43. Calculate the probability density function: Where f(·) is the density function, c is the copula density function, and f i (x ·i ) represents each variable x ·i The distribution function of S44. Find the random function: S45. Calculate the correlation coefficient matrix: Substitute X' into formula (4-4), combine formula (4-2) and formula (4-4) into a system of equations, and calculate the correlation coefficient matrix R.

2. According to the method for completing missing values ​​of multi-sensor time series in the Internet of Things based on graph learning in claim 1, it is characterized in that: The elements in the labeling matrix M are either 0 or 1, where 0 indicates that the corresponding position data of the IoT multi-sensor time series S is missing, and 1 indicates that the corresponding position data of the IoT multi-sensor time series S is complete.

3. According to the method for missing value completion of multi-sensor time series in the Internet of Things based on graph learning in claim 1, it is characterized in that: The spatiotemporal coding in step S5 comprises the following steps: S5. Spatial encoding: Use the pytorch toolkit in python to perform spatial encoding. The formula is as follows: Among them, MPNN() represents the message passing function, x ki is the observation value obtained by sensor i at time k; x kj is the observation value obtained by sensor j at time k; R is the correlation coefficient matrix, γ, ρ are two learnable multi-layer perceptrons MLP; r ij represents the correlation coefficient between sensor i and sensor j; S52, using a gated recurrent unit to perform time encoding on formula (5-1), the formula is as follows: in, They are reset gate and update gate respectively. is the hidden state of IoT data at time point t, is the hidden state of IoT data at time point t-1, is the value of this time step, is the label matrix M and Corresponding elements; get the spatiotemporal coding sequence 4. According to the method for completing missing values ​​of multi-sensor time series in the Internet of Things based on graph learning in claim 1, it is characterized in that: The spatial decoding in step S6 comprises the following steps: S61. Decode the spatiotemporal coding sequence using a spatial decoder to obtain the algorithm prediction value at time t, as follows: in, is the predicted value for time t generated by the algorithm, h t-1 is the hidden state at time t-1 output by the spatiotemporal encoder, V is a learnable matrix, and b is a learnable bias vector; S62, preliminarily filling the original vacant values ​​with the predicted values: in, represents the data set at time t after initial filling, m tj , x tj is the value in the label matrix M and the normalized dataset X corresponding to time t, Indicates m tj Bitwise inversion; S63. Use the MPNN function to aggregate the data at time t, as shown in formula (6-3): in, Represents the feature matrix of each node after preliminary completion. MPNN is the same as formula (5-1), m ti , They respectively represent whether the data of the i-th sensor is missing at time t and its hidden state.

5. According to the method for completing missing values ​​of multi-sensor time series in the Internet of Things based on graph learning in claim 1, it is characterized in that: The two-way time prediction in step S7 is as follows: in, It represents the algorithm's completion value for the missing value of the i-th sensor at the t-th time point, MLP represents a multi-layer perceptron, Represent the forward feature matrix and reverse feature matrix of each node respectively, Represent the forward hidden state matrix and reverse hidden state matrix of each node respectively.

6. According to the method for completing missing values ​​of multi-sensor time series in the Internet of Things based on graph learning in claim 1, it is characterized in that: The data processing in step S2 is to bring the IoT time series S into formula (2-1), as follows: Among them, s ij represents the data observed by sensor j at the i-th moment in S, s ·j represents all the data observed by sensor j, They represent the maximum and minimum values ​​of all data observed by sensor j, respectively, and x ij Represents a normalized data, and X represents the set of all normalized data.