A traffic flow data completion method and system based on prior knowledge and Transformer
Through a priori knowledge and Transformer method, the cumulative error and computing resource consumption problems in traffic flow data completion are solved, efficient and accurate data interpolation are achieved, and the real-time and generalization capabilities of the model are improved.
Patent Information
- Application Number
- CN202510661111.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing time series data completion method has cumulative errors, high computing resource consumption and lack of prior knowledge guidance when processing traffic flow data, resulting in insufficient real-time and generalization capabilities of the model.
Using a priori knowledge and Transformer method, the transformation of high-dimensional time series data to low-dimensional space is achieved through space-time feature embedding, projection attention mechanism and embedded attention, combined with low-rank structure, and the interpolation of missing data is optimized by minimizing the loss function.
It improves the accuracy and efficiency of traffic flow data completion, reduces computing resource consumption, and enhances the model's ability to capture and generalizes the essential characteristics of data.
Smart Images

Figure CN120216932B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data completion, and particularly to a traffic flow data completion method and system based on prior knowledge and Transformer (converter). Background Art
[0002] Time series data completion generally refers to filling in missing values in time series data. These time series data may have a very high data missing rate due to bad weather, sensor failures, transmission losses, etc. For traffic flow prediction, the missing of time series data poses a huge challenge. If the missing values can be correctly imputed, the data utilization rate can be improved, which is helpful for traffic flow prediction. Therefore, time series data completion is of great value.
[0003] The mainstream time series completion methods can be divided into the following two categories: time series completion methods based on autoregressive models and time series completion methods based on low-rank models. Time series completion based on autoregression often fills in missing data by autoregressive filling of the missing data, but this method will cause the generation of cumulative errors, and the calculation speed and memory of the model are also restricted. The time series completion method based on the low-rank model, although providing an effective inductive bias, is limited by the low model expressiveness and simplified assumptions, may produce an "overly smooth" completion result, and also lacks the guidance of prior knowledge, may produce overfitting to the observed data, and lacks generality.
[0004] Publication No.: CN118760677A, Title: A Missing Data Completion Method Based on Sparse Topological Spatiotemporal Attention Mechanism. The method includes: performing random missing processing on time series data and marking the missing values in the time series data; performing normalization processing on the processed time series data to construct a data set containing missing values; constructing a sparse topological spatiotemporal attention model, where the sparse topological spatiotemporal attention model includes a time attention module and a sparse topological attention module; inputting the data set into the sparse topological spatiotemporal attention model, calculating the loss function, training the sparse topological spatiotemporal attention model based on the loss function to obtain a data completion model based on the sparse topological spatiotemporal attention mechanism; inputting the missing data into the data completion model to obtain complete data. This method can capture the complex dependencies between data at different time points and different spatial positions. By integrating topological structure information, it can accurately complete the missing data and improve the accuracy of the completion result. Although the missing data completion method based on the sparse topological spatiotemporal attention mechanism is innovative in capturing complex dependencies, it may run slowly in an environment with limited computing resources, affecting real-time performance; at the same time, when dealing with long time series data and complex spatiotemporal correlations, there may be cumulative errors and the risk of oversmoothing, and the integration of prior knowledge is insufficient, resulting in limited generalization ability. Summary of the Invention
[0005] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further detailed in the Detailed Implementation section. The Summary of the Invention section of the present invention does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.
[0006] To solve the problems of cumulative error caused by autoregressive models, limited model expressiveness of low-rank models, and lack of prior knowledge guidance, the present invention proposes a traffic flow data completion method and system based on prior knowledge and Transformer.
[0007] In a first aspect, a traffic flow data completion method based on prior knowledge and Transformer is provided, including:
[0008] Obtaining a traffic flow data set;
[0009] Performing missing coding on the traffic flow data set and randomly generating some artificial missing values to simulate data loss;
[0010] Embedding and integrating timestamp information and node information based on spatiotemporal feature embedding technology;
[0011] Implementing the conversion of high-dimensional time series data to a low-dimensional space based on the projection attention mechanism;
[0012] Apply the embedded attention to the spatial dimension and capture the spatial correlation between sensors through the embedded attention;
[0013] Output the predicted values of the missing data and evaluate the loss of the predicted values;
[0014] Based on the optimization result of minimizing the loss function, impute the missing data in the traffic flow dataset;
[0015] The embedding integration of the timestamp information and the node information based on the spatio-temporal feature embedding technology includes:
[0016] Convert the input time series data into a hidden state representation through a linear transformation;
[0017] The linear transformation formula is:
[0018] ;
[0019] Among them, is the hidden state representation, is the input time series data, is the time series length, is the feature dimension, is the weight matrix, is the target hidden dimension, is the bias vector;
[0020] Concatenate the hidden state representation with the node distance matrix;
[0021] Merge and embed the position encoding and the timestamp information to construct the final spatio-temporal input embedding;
[0022] The conversion from high-dimensional time series data to low-dimensional space based on the projection attention mechanism includes the following steps:
[0023] S4.1 Introduce a projection matrix based on the projection attention mechanism and map the high-dimensional time series data to a low-dimensional space through the projection matrix;
[0024] S4.2 Enhance the ability to capture the correlation between time series by applying the temporal self-attention mechanism in the low-dimensional space;
[0025] S4.3 Add S4.1 and S4.2 to the Transformer through a residual connection;
[0026] The traffic flow dataset includes: the number of sensors, the number of data points, the number of channels, the missing value situation, the average vehicle speed, the node distance matrix, the timestamp information, and the time series data.
[0027] Further, the missing coding of the traffic flow data set includes:
[0028] Create a binary matrix of the same size as the traffic flow data set;
[0029] Traverse the traffic flow data set. If there is no data at the current position, represent the missing data with 0 in the binary matrix; otherwise, represent the existing data with 1 in the binary matrix.
[0030] The randomly generating some artificial missing values includes:
[0031] Artificially introduce missing values at the positions in the binary matrix with a value of 1 through a random seed;
[0032] Record the artificially missing positions through an artificial mask.
[0033] Further, the splicing of the hidden state representation and the node distance matrix includes the following steps:
[0034] Extract the node distance vector from the node distance matrix and extend the node distance vector to a time length matching the input time series data;
[0035] The extension formula is:
[0036] ;
[0037] where is the node distance matrix, is the distance vector, is the node, is the number of sensors, is the starting time point, is the time series length;
[0038] Splice the extended node distance vector and the hidden state representation;
[0039] The splicing formula is:
[0040] ;
[0041] where is the splicing result, is the splicing operation in the last dimension, is the hidden state representation, is the extended node distance vector.
[0042] Further, the merging and embedding of the position encoding and the timestamp information includes:
[0043] Generate positional encodings for each time step through sine and cosine functions;
[0044] The positional encoding formula is:
[0045] ;
[0046] ;
[0047] where, is the position of the time step, is the hidden dimension of the model, is the dimension index, is the sine value of the even dimension, is the cosine value of the odd dimension;
[0048] Encode the timestamp information through one-hot encoding to generate a binary vector encoding corresponding to the timestamp information;
[0049] Merge and embed the positional encoding and the timestamp information;
[0050] The merge embedding formula is:
[0051] ;
[0052] where, is the merge embedding result, is the concatenation result, is the time step, is the time step generated corresponding positional encoding, is the time step timestamp information.
[0053] Furthermore, the construction of the final spatio-temporal input embedding includes:
[0054] Map the node embedding information into a low-dimensional representation form that changes dynamically over time based on the time window length;
[0055] The low-dimensional representation form is:
[0056] ;
[0057] where, is the low-dimensional representation form of node within the time window , is the length of the time series, is the dimension before mapping the node embedding information;
[0058] Concatenate the merged embedding result with the low-dimensional representation of the node embedding information to construct the final spatio-temporal input embedding;
[0059] The combination formula is:
[0060] ;
[0061] Among them, is the final spatio-temporal input embedding, is the concatenation operation in the last dimension, is the merged embedding result, is the node within the time window of the low-dimensional representation.
[0062] Furthermore, the formula for mapping high-dimensional time series data to a low-dimensional space through a projection matrix is:
[0063] ;
[0064] Among them, is the low-dimensional projection result, is the final spatio-temporal input embedding, is the projection matrix, is the dimension of the target low-dimensional space;
[0065] The application of the temporal self-attention mechanism in the low-dimensional space includes the following steps:
[0066] S4.4 Perform projection transformation on the temporal query matrix, temporal key matrix, and temporal value matrix through the projection matrix;
[0067] The projection transformation formula is:
[0068] ;
[0069] Among them, is the temporal query matrix, is the temporal key matrix, is the temporal value matrix, is the temporal query matrix after projection transformation, is the temporal key matrix after projection transformation, is the temporal value matrix after projection transformation;
[0070] S4.5 Calculate the self-attention for the projected temporal query matrix, temporal key matrix, and temporal value matrix;
[0071] The formula for calculating self-attention is:
[0072] ;
[0073] Among them, is the activation function, is the self-attention.
[0074] Furthermore, the formula for adding S4.1 and S4.2 to the Transformer through residual connection is:
[0075] ;
[0076] Among them, is the normalization operation, is the residual result.
[0077] Furthermore, applying the embedded attention to the spatial dimension includes:
[0078] Performing parameter transformation on the residual result through a learnable parameter matrix;
[0079] The parameter transformation formula is:
[0080] ;
[0081] ;
[0082] ;
[0083] Among them, is the residual result, is the learnable parameter matrix, is the spatial query matrix, is the spatial key matrix, is the spatial value matrix;
[0084] Calculating the spatial attention weight between sensors through dot product similarity;
[0085] The spatial attention weight calculation formula is:
[0086] ;
[0087] Among them, is the spatial attention weight, is the activation function, is to transpose, is the dimension of the key vector, is the embedding vector, is the embedding vector;
[0088] Updating the spatial value matrix through the spatial attention weight to obtain the updated node feature representation;
[0089] ;
[0090] Among them, is the updated node feature representation, is the number of sensors, is the embedding vector, is the learnable parameter matrix, is the spatial attention weight.
[0091] Furthermore, predicting the missing data value and evaluating the loss of the predicted value includes:
[0092] Further processing the hidden state through a feed-forward neural network and finally generating a predicted value in the output layer;
[0093] ;
[0094] Among them, is the updated node feature representation, is the predicted value, is the first-layer weight matrix, is the first-layer bias vector, is the second-layer weight matrix, is the second-layer bias vector, is the activation function, is the predicted value output function;
[0095] Calculating the loss value of the mean absolute error based on the artificial mask;
[0096] ;
[0097] Among them, is the loss value of the mean absolute error, is the predicted value, is the true value, is the value in the mask matrix at dimension and time under, is the feature dimension, is the time series length;
[0098] Adjust the projection matrix and parameter matrix through the loss value to minimize the mean absolute error.
[0099] Furthermore, imputing the missing data in the traffic flow dataset based on the optimization result of minimizing the loss function includes:
[0100] Using the predicted value after minimizing the mean absolute error for data imputation;
[0101] The imputation formula is:
[0102] ;
[0103] Among them, is the result after interpolation, is the original data, is the binary matrix, is the element-wise multiplication, is the predicted value.
[0104] In a second aspect, a traffic flow data completion system based on prior knowledge and Transformer is provided, including:
[0105] A data collection module, a data encoding module, an information fusion module, a spatio-temporal interaction module, and an evaluation and interpolation module:
[0106] The data collection module is used to collect traffic flow data sets;
[0107] The data encoding module is used to perform missing encoding and random artificial missing encoding on the traffic flow data set in the data collection module;
[0108] The information fusion module is used to fuse the missing encoding of the data encoding module and the traffic flow data set in the data collection module;
[0109] The spatio-temporal interaction module is used to apply temporal attention and spatial attention to the fusion result of the information fusion module to capture low-dimensional information and output a predicted value;
[0110] The evaluation and interpolation module is used to perform loss evaluation on the predicted value, and use the predicted value after minimizing the mean absolute error operation for data interpolation.
[0111] The beneficial effects of the present invention are:
[0112] 1. A traffic flow data completion method based on prior knowledge and Transformer provided by the present invention can effectively guide the model to learn more reasonable interpolation values by embedding the distance information of sensors as prior knowledge in the embedding layer. Especially for data sets with clear patterns in geographical distribution, it can not only improve the accuracy of interpolation, but also enable the model to better generalize to unseen data.
[0113] 2. A traffic flow data completion method based on prior knowledge and Transformer provided by the present invention, by adopting the Transformer model, not only effectively solves the cumulative error problem existing in the previous autoregressive models, but also significantly improves the calculation speed with its parallel computing ability and reduces the memory consumption.
[0114] 3. A traffic flow data completion method based on prior knowledge and Transformer provided by the present invention realizes the conversion of high-dimensional time series data to a low-dimensional space through a projection attention mechanism, introduces low-rankness, improves the model interpolation efficiency, and enhances the model's ability to capture the essential features of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0115] Figure 1 It is a schematic flowchart of a traffic flow data completion method based on prior knowledge and Transformer provided by an embodiment of the present invention.
[0116] Figure 2 It is a schematic structural diagram of a traffic flow data completion method provided by an embodiment of the present invention.
[0117] Figure 3 It is a schematic diagram of a traffic flow data completion system based on prior knowledge and Transformer provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0118] The exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully communicated to those skilled in the art.
[0119] With the rapid development of big data technology and Internet of Things devices, the scale and complexity of time series data are increasing day by day, which poses higher requirements for data completion technology. Currently, sensor networks are widely distributed in various fields, from industrial monitoring to environmental monitoring, and then to smart city projects. The time series data generated by these sensors is huge in volume and strong in real-time. However, due to factors such as hardware failures, communication interruptions, or harsh environmental conditions, the problem of data loss has become particularly prominent. This requires more efficient and accurate data completion methods to ensure data integrity and availability.
[0120] In view of this, how to effectively combine emerging technologies to improve the accuracy and efficiency of time series data completion has become a research hotspot. On the one hand, traditional completion methods are unable to cope when faced with massive and high-dimensional time series data, and there is an urgent need for innovative technical means to solve problems such as cumulative errors and over-smoothing. On the other hand, the development of technologies such as cloud computing, edge computing, and artificial intelligence has made it possible to achieve this goal. In particular, the progress of deep learning and machine learning algorithms enables the model to better understand and predict patterns and trends in time series. In addition, by closely integrating these advanced technologies with actual application scenarios, not only can the data completion process be optimized, but also the potential value of the data can be further explored, providing strong data support for decision-making support systems.
[0121] Therefore, by introducing prior knowledge and the Transformer model and combining with low-rank structures, the accuracy and efficiency of time series data completion can be effectively improved. Prior knowledge can help the model better understand the data background, reduce sensitivity to outliers, and provide reasonable estimates in the absence of sufficient observed data. The Transformer model, with its powerful parallel processing ability and ability to capture long-term dependencies, can more accurately predict missing values and overcome the cumulative error problem in traditional methods. In addition, the application of low-rank structures helps to reveal the hidden correlations and patterns within time series data, further enhancing the stability and generalization ability of the completion algorithm. Combining these technical advantages can not only significantly improve the quality of time series data, but also provide a more solid foundation for subsequent data analysis and decision support.
[0122] The specific implementation of the technical solution of the present invention includes the following content:
[0123] Reference Figure 1 , a traffic flow data completion method based on prior knowledge and Transformer, the method comprising:
[0124] S1, obtaining a traffic flow data set;
[0125] S2, performing missing coding on the traffic flow data set and randomly generating some artificial missing values to simulate data loss;
[0126] S3, embedding and integrating timestamp information and node information based on spatio-temporal feature embedding technology;
[0127] S4, realizing the conversion of high-dimensional time series data to a low-dimensional space based on the projection attention mechanism;
[0128] S5, applying embedded attention to the spatial dimension and capturing the spatial correlation between sensors through embedded attention;
[0129] S6. Output the predicted values of the missing data and evaluate the loss of the predicted values.
[0130] S7. Based on the optimization result of minimizing the loss function, impute the missing data in the traffic flow dataset.
[0131] In step S1, the traffic flow dataset includes: the number of sensors, the number of data points, the number of channels, the situation of missing values, the average vehicle speed, the node distance matrix, the timestamp information, and the time series data.
[0132] In an embodiment of the present invention, the traffic flow dataset contains: the number of sensors, the number of data points, the number of channels, the situation of missing values, the average vehicle speed, the node distance matrix, the timestamp information, and the time series data, which are combined into a set denoted as Data. The number of sensors refers to the total number of sensors participating in data collection, which helps to evaluate the data coverage and density; the number of data points refers to the specific number of data entries recorded in the entire dataset, which affects the granularity of data analysis; the number of channels represents the number of different types of data streams included in each data record, such as traffic flows in different directions or types; the situation of missing values describes the proportion of missing data points in the dataset and their distribution, which is of great significance for evaluating data integrity and formulating filling strategies; the average vehicle speed provides information about traffic fluency and helps to identify congested areas; the node distance matrix records the relative distances between various sensors, which is crucial for constructing an accurate spatial model; the timestamp information records the specific time points of data collection, which helps to analyze the time series characteristics and change trends of the data; the time series data records the data values collected at different time points.
[0133] For step S2, the missing encoding of the traffic flow dataset includes:
[0134] Create a binary matrix of the same size as the traffic flow dataset;
[0135] Traverse the traffic flow dataset. If there is no data at the current position, represent the missing data with 0 in the binary matrix; otherwise, represent the existing data with 1 in the binary matrix.
[0136] The random generation of some artificial missing values includes:
[0137] Artificially make the positions with a value of 1 in the binary matrix missing through a random seed;
[0138] Record the artificially missing positions through an artificial mask.
[0139] In another embodiment of the present invention, when performing missing encoding on the traffic flow dataset, first create a binary matrix of the same size as the dataset , which is used to mark the data existence status of each position (1 indicates the existence of data, and 0 indicates the absence of data). Traverse the entire traffic flow data set. For the position of each data point, if there is no data at the current position, the corresponding position in the binary matrix is represented by 0 to indicate the absence of data; if there is data, it is represented by 1. To distinguish and efficiently process block missing and point missing, when it is detected that the continuous missing data exceeds a certain threshold, for example, 100 consecutive missing data, it is regarded as block missing, and the sparse matrix is used to record the start and end positions and their indexes of these block missing regions to save storage space and facilitate subsequent analysis. On this basis, to simulate the data loss situation in the real world, a random seed is selected to ensure the reproducibility of the experiment, and the positions with a value of 1 in the binary matrix are randomly selected and marked as artificially missing (that is, changing their value from 1 to 0), and the block missing region can also be selectively extended. Finally, all artificially missing positions are recorded through an artificial mask matrix . Among them, the positions with a value of 0 represent the positions of the data points that are artificially set as missing, while the positions with a value of 1 represent the original data points that have not been modified.
[0140] For step S3, the embedding and integration of the timestamp information and node information based on the spatio-temporal feature embedding technology includes:
[0141] Convert the input time series data into a hidden state representation through a linear transformation;
[0142] The linear transformation formula is:
[0143] ;
[0144] Among them, is the hidden state representation, is the input time series data, is the time series length, is the feature dimension, is the weight matrix, is the target hidden dimension, is the bias vector;
[0145] Concatenate the hidden state representation with the node distance matrix;
[0146] Merge and embed the position encoding and timestamp information to construct the final spatio-temporal input embedding.
[0147] Furthermore, the concatenating of the hidden state representation with the node distance matrix includes the following steps:
[0148] Extract the node distance vector from the node distance matrix, and extend the node distance vector to a time length that matches the input time series data;
[0149] The extension formula is:
[0150] ;
[0151] Wherein, is the node distance matrix, is the distance vector, is the node, is the number of sensors, is the starting time point, is the time series length;
[0152] Concatenate the extended node distance vector with the hidden state representation;
[0153] The concatenation formula is:
[0154] ;
[0155] Wherein, is the concatenation result, is the concatenation operation in the last dimension, is the hidden state representation, is the extended node distance vector.
[0156] In an embodiment of the present invention, the time series data to be processed is extracted from the traffic flow dataset to form an input matrix , and a suitable weight matrix and bias vector are selected through random initialization. For each time point in the input matrix , a linear transformation is performed using the selected weight matrix and bias vector to obtain the hidden state representation . Align the node distance matrix with the hidden state . It is necessary to adjust the shape of the node distance matrix to create a three-dimensional tensor of size , and reuse the same node distance information so that each time point has corresponding node distance information. On the premise of keeping the time dimension consistent, the node distance matrix is flattened along the feature dimension into a form of size and concatenated with the hidden state to finally form a new comprehensive representation matrix .
[0157] Furthermore, the merging and embedding of the position encoding and the timestamp information includes:
[0158] Generate positional encodings for each time step using sine and cosine functions;
[0159] The positional encoding formula is:
[0160] ;
[0161] ;
[0162] where, is the position of the time step, is the hidden dimension of the model, is the dimension index, is the sine value of the even dimension, is the cosine value of the odd dimension;
[0163] Encode the timestamp information using one-hot encoding to generate a binary vector encoding corresponding to the timestamp information;
[0164] Merge and embed the positional encoding and the timestamp information;
[0165] The merge embedding formula is:
[0166] ;
[0167] where, is the merge embedding result, is the concatenation result, is the time step, is the time step generated corresponding positional encoding, is the time step timestamp information.
[0168] Furthermore, construct the final spatio-temporal input embedding, including:
[0169] Map the node embedding information to a low-dimensional representation form that changes dynamically over time based on the time window length;
[0170] The low-dimensional representation form is:
[0171] ;
[0172] where, is the low-dimensional representation form of node within the time window , is the length of the time series, is the dimension before mapping the node embedding information;
[0173] Concatenate the merged embedding result with the low-dimensional representation of the node embedding information to construct the final spatio-temporal input embedding;
[0174] The combination formula is:
[0175] ;
[0176] where, is the final spatio-temporal input embedding, is the concatenation operation in the last dimension, is the merged embedding result, is the node in the time window within the low-dimensional representation.
[0177] In another embodiment of the present invention, there is currently a traffic flow data set, which contains 5 time steps, each time step has 3 feature dimensions, the data set includes 4 nodes, and the target hidden dimension of the model is 8, and the original dimension of the node embedding is 16. First, generate position encoding for each time step. For each time step , from 0 to 4, the position encoding value can be calculated. For , the position encoding value of the even dimension is , and the position encoding value of the odd dimension is ; for , the position encoding value of the even dimension is , and the position encoding value of the odd dimension is . Through the position encoding, a position encoding matrix with a size of can be obtained. Next, perform one-hot encoding on the timestamp information. The time window contains 5 time points, and each time period has a unique timestamp. Create a one-hot encoding matrix with a size of , where each row represents the timestamp information of a time point. Merge and embed the position encoding matrix and the one-hot encoding matrix into the comprehensive representation that already contains the hidden state and node information. The size of the comprehensive representation should be . Based on the length of the time window, map the node embedding information into a low-dimensional representation that changes dynamically with time. The target low dimension is , and the low-dimensional representation of each node within the time window is obtained. Concatenate the comprehensive representation and the target low dimension along the last dimension, and the final spatio-temporal input embedding obtained has a size of .
[0178] For step S4, the conversion of high-dimensional time series data to a low-dimensional space based on the projection attention mechanism includes:
[0179] The conversion of high-dimensional time series data to a low-dimensional space based on the projection attention mechanism includes the following steps:
[0180] S4.1 Introduce a projection matrix based on the projection attention mechanism, and map the high-dimensional time series data to a low-dimensional space through the projection matrix;
[0181] S4.2 By applying the temporal self-attention mechanism in the low-dimensional space, enhance the ability to capture the correlation between time series;
[0182] S4.3 Add S4.1 and S4.2 to the Transformer through residual connection.
[0183] Furthermore, the formula for mapping the high-dimensional time series data to a low-dimensional space through the projection matrix is:
[0184] ;
[0185] Among them, is the low-dimensional projection result, is the final spatio-temporal input embedding, is the projection matrix, is the dimension of the target low-dimensional space;
[0186] The application of the temporal self-attention mechanism in the low-dimensional space includes the following steps:
[0187] S4.4 Perform projection transformation on the time query matrix, time key matrix, and time value matrix through the projection matrix;
[0188] The projection transformation formula is:
[0189] ;
[0190] Among them, is the time query matrix, is the time key matrix, is the time value matrix, is the time query matrix after projection transformation, is the time key matrix after projection transformation, is the time value matrix after projection transformation;
[0191] S4.5 Calculate the self-attention of the time query matrix, time key matrix, and time value matrix after projection transformation;
[0192] The formula for calculating self-attention is:
[0193] ;
[0194] Among them, is the activation function, is the self-attention.
[0195] Furthermore, the formula for adding S4.1 and S4.2 to the Transformer through residual connection is:
[0196] ;
[0197] wherein, is the normalization operation, is the residual result.
[0198] In an embodiment of the present invention, by introducing a randomly initialized projection matrix of size , the spatio-temporal input embedding is mapped from the original size to the target low-dimensional space dimension .
[0199] First, the spatio-temporal input embedding is transformed using the projection matrix to obtain the low-dimensional projection result . Then, a temporal self-attention mechanism is applied in the low-dimensional space to enhance the ability to capture the correlation between time series. For this purpose, a temporal query matrix , a temporal key matrix and a temporal value matrix are constructed, and these matrices are projected and transformed using the projection matrix to calculate the self-attention result. Finally, the low-dimensional projection result is added to the self-attention result through residual connection and undergoes layer normalization processing to obtain the final representation.
[0200] For step S5, applying the embedded attention to the spatial dimension and capturing the spatial correlation between sensors through the embedded attention includes:
[0201] Applying the embedded attention to the spatial dimension includes:
[0202] Performing parameter transformation on the residual result through a learnable parameter matrix;
[0203] The parameter transformation formula is:
[0204] ;
[0205] ;
[0206] ;
[0207] wherein, is the residual result, is a learnable parameter matrix, is a spatial query matrix, is a spatial key matrix, is a spatial value matrix;
[0208] Calculate the spatial attention weights between sensors through dot - product similarity;
[0209] The formula for calculating the spatial attention weights is:
[0210] ;
[0211] where, is the spatial attention weight, is the activation function, is to transpose, is the dimension of the key vector, is the embedding vector, is the embedding vector;
[0212] Update the spatial value matrix through the spatial attention weights to obtain the updated node feature representation;
[0213] ;
[0214] where, is the updated node feature representation, is the total number of nodes, is the embedding vector, is a learnable parameter matrix, is the spatial attention weight.
[0215] In an embodiment of the present invention, the time step is 5, the total number of nodes is 4, the dimension of the key vector is 8, and the residual result , its size is , in order to perform parameter transformation, three learnable parameter matrices are introduced, and the size of each matrix is . For the embedding vector of node at time step : ; For the embedding vector of node at time step : ; For the embedding vector of node at time step : ; For the embedding vector of node at time step : ; The embedded vectors are transformed by a learnable parameter matrix to obtain a spatial query matrix , a spatial key matrix , and a spatial value matrix . The spatial query vector of node : ; The spatial query vector of node : ; The spatial query vector of node : ; The spatial query vector of node : ; The corresponding vectors of the spatial key matrix and the spatial value matrix can also be obtained through parameter transformation. The spatial attention weights between sensors are calculated by dot - product similarity. The spatial attention weight between node and node : ; The spatial attention weight between node and node : ; The spatial attention weight between node and node : . Finally, the spatial value matrix is updated by the spatial attention weights to update the feature representation of node : .
[0216] For step S6, outputting the predicted value of the missing data and evaluating the loss of the predicted value includes:
[0217] The hidden state is further processed by a feed - forward neural network, and finally a predicted value is generated at the output layer;
[0218] ;
[0219] Wherein, is the updated node feature representation, is the predicted value, is the first - layer weight matrix, is the first - layer bias vector, is the second - layer weight matrix, is the second - layer bias vector, is the activation function, is the predicted - value output function;
[0220] Calculate the loss value of the mean absolute error based on the artificial mask;
[0221] ;
[0222] Wherein, is the loss value of the mean absolute error, is the predicted value, is the true value, is the value in the mask matrix at dimension and time ; is the feature dimension, is the length of the time series;
[0223] The projection matrix and the parameter matrix are adjusted by the loss value to minimize the mean absolute error.
[0224] In one embodiment of the present invention, by obtaining the updated node feature representation with a size of , the features are input into a feed-forward network, is the weight matrix of the first layer with a size of ; is the bias vector of the first layer with a size of ; is the weight matrix of the second layer with a size of ; is the bias vector of the first layer with a size of . , , , , the predicted value obtained after being processed by the feed-forward network, , Calculate the loss value of the mean absolute error for the predicted value, the true value, and the matrix mask: , and adjust the projection matrix and the learnable parameter matrix by the loss value to minimize the mean absolute error.
[0225] For step S7, imputing the missing data in the traffic flow dataset based on the optimization result of minimizing the loss function includes:
[0226] The predicted value after minimizing the mean absolute error operation is used for data imputation;
[0227] The imputation formula is:
[0228] ;
[0229] where, is the result after imputation, is the original data, is the binary matrix, is the element-wise multiplication, is the predicted value.
[0230] In another embodiment of the present invention, the original data , binary matrix , predicted value , through the interpolation formula: .
[0231] Reference Figure 2 , schematic structural diagram of a traffic flow data completion method, including:
[0232] Obtain a traffic flow data set;
[0233] Extract node information and timestamp information from the traffic flow data set;
[0234] Perform missing coding and artificial missing coding on the traffic flow data set, and then perform linear transformation;
[0235] Embed the final spatio-temporal input into spatio-temporal interaction, evaluate the obtained predicted value and use it for missing data interpolation.
[0236] Reference Figure 3 , schematic diagram of a traffic flow data completion system based on prior knowledge and Transformer, including: a data collection module, a data coding module, an information fusion module, a spatio-temporal interaction module, and an evaluation and interpolation module:
[0237] The data collection module is used to collect a traffic flow data set;
[0238] The data coding module is used to perform missing coding and random artificial missing coding on the traffic flow data set in the data collection module;
[0239] The information fusion module is used to fuse the missing coding of the data coding module and the traffic flow data set in the data collection module;
[0240] The spatio-temporal interaction module is used to capture low-dimensional information by applying temporal attention and spatial attention to the fusion result of the information fusion module, and output a predicted value;
[0241] The evaluation and interpolation module is used to evaluate the loss of the predicted value, and use the predicted value after minimizing the mean absolute error operation for data interpolation.
[0242] Those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments.
[0243] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A traffic flow data completion method based on prior knowledge and Transformer, characterized in that The method includes: Obtain a traffic flow dataset; Perform missing coding on the traffic flow dataset and randomly generate some artificial missing values to simulate data loss; Embed and integrate timestamp information and node information based on spatio-temporal feature embedding technology; Based on the projection attention mechanism, realize the conversion of high-dimensional time series data to a low-dimensional space; Apply embedded attention to the spatial dimension and capture the spatial correlation between sensors through embedded attention; Output the predicted values of the missing data and evaluate the loss of the predicted values; Based on the optimization result of minimizing the loss function, impute the missing data in the traffic flow dataset; The embedding and integration of timestamp information and node information based on spatio-temporal feature embedding technology includes: Convert the input time series data into a hidden state representation through linear transformation; The linear transformation formula is: ; Among them, is the hidden state representation, is the input time series data, is the time series length, is the feature dimension, is the weight matrix, is the target hidden dimension, is the bias vector; Concatenate the hidden state representation with the node distance matrix; Merge and embed the position encoding and timestamp information to construct the final spatio-temporal input embedding; The realization of the conversion of high-dimensional time series data to a low-dimensional space based on the projection attention mechanism includes the following steps: S4.1 Introduce a projection matrix based on the projection attention mechanism and map the high-dimensional time series data to a low-dimensional space through the projection matrix; S4.2 Apply the temporal self-attention mechanism in the low-dimensional space to enhance the ability to capture the correlation between time series; S4.3 Add S4.1 and S4.2 to the Transformer through residual connection; The traffic flow dataset includes: the number of sensors, the number of data points, the number of channels, the missing value situation, the average vehicle speed, the node distance matrix, the timestamp information, and the time series data.
2. A traffic flow data completion method based on prior knowledge and Transformer according to claim 1, characterized in that: The missing coding of the traffic flow dataset includes: Create a binary matrix of the same size as the traffic flow dataset; Traverse the traffic flow dataset. If there is no data at the current position, represent the missing data with 0 in the binary matrix. Otherwise, represent the existing data with 1 in the binary matrix; The random generation of some artificial missing values includes: Artificially miss the positions with a value of 1 in the binary matrix through a random seed; Record the artificially missing positions through an artificial mask.
3. A traffic flow data completion method based on prior knowledge and Transformer according to claim 1, characterized in that: The concatenation of the hidden state representation and the node distance matrix includes the following steps: Extract the node distance vector from the node distance matrix and extend the node distance vector to a time length matching the input time series data; The extension formula is: ; Among them, is the node distance matrix, is the distance vector, is the node, is the number of sensors, is the starting time point, is the length of the time series; Concatenate the extended node distance vector with the hidden state representation; The concatenation formula is: ; Among them, is the splicing result, is the splicing operation performed on the last dimension, is the hidden state representation, is the extended node distance vector.
4. A traffic flow data completion method based on prior knowledge and Transformer according to claim 1, characterized in that: The merging and embedding of the position encoding and timestamp information includes: Generate positional encodings for each time step using sine and cosine functions; The formula for positional encoding is: ; ; Among them, is the position of the time step, is the hidden dimension of the model, is the dimension index, is the sine value of the even dimension, is the cosine value of the odd dimension; Encode the timestamp information using one-hot encoding to generate a binary vector encoding corresponding to the timestamp information; Merge and embed the positional encoding and the timestamp information; The formula for merge embedding is: ; Among them, is the merged embedding result, is the concatenation result, is the time step, is the time step and is the corresponding position encoding generated, is the time step and is the timestamp information of it.
5. A traffic flow data completion method based on prior knowledge and Transformer as claimed in claim 1, characterized in that: The construction of the final spatio-temporal input embedding includes: Based on the time window length, map the node embedding information into a low-dimensional representation form that changes dynamically over time; The low-dimensional representation form is: ; Among them, is the node in the time window with a low-dimensional representation form, is the length of the time series, is the dimension before the mapping of the node embedding information; Concatenate the merge embedding result with the low-dimensional representation form of the node embedding information to construct the final spatio-temporal input embedding; The combination formula is: ; Among them, is the final spatio-temporal input embedding, is the concatenation operation in the last dimension, is the merged embedding result, is the node in the time window is the low-dimensional representation form within.
6. A traffic flow data completion method based on prior knowledge and Transformer as claimed in claim 1, characterized in that: The formula for mapping high-dimensional time series data to a low-dimensional space through a projection matrix is: ; Among them, is the low-dimensional projection result, is the final spatio-temporal input embedding, is the projection matrix, is the dimension of the target low-dimensional space; The application of the temporal self-attention mechanism in the low-dimensional space includes the following steps: S4.4 Perform projection transformation on the time query matrix, time key matrix, and time value matrix through the projection matrix; The formula for projection transformation is: ; Among them, is the time query matrix, is the time key matrix, is the time value matrix, is the time query matrix after projection transformation, is the time key matrix after projection transformation, is the time value matrix after projection transformation; S4.5 Calculate self-attention for the time query matrix, time key matrix, and time value matrix after projection transformation; The formula for calculating self-attention is: ; Among them, is an activation function, is self-attention.
7. A traffic flow data completion method based on prior knowledge and Transformer as claimed in claim 1, characterized in that: The formula for adding S4.1 and S4.2 to the Transformer through residual connection is: ; Among them, is the normalization operation, is the residual result.
8. A traffic flow data completion method based on prior knowledge and Transformer as claimed in claim 1, characterized in that: The application of the embedded attention to the spatial dimension includes: Perform parameter transformation on the residual result through a learnable parameter matrix; The formula for parameter transformation is: ; ; ; Among them, is the residual result, is the learnable parameter matrix, is the spatial query matrix, is the spatial key matrix, is the spatial value matrix; Calculate the spatial attention weights between sensors through dot product similarity; The formula for calculating the spatial attention weights is: ; Among them, is the spatial attention weight, is the activation function, is to transpose, is the dimension of the key vector, is the embedding vector, is the embedding vector; Update the spatial value matrix through the spatial attention weights to obtain the updated node feature representation; ; Among them, is the updated node feature representation, is the number of sensors, is the embedding vector, is the learnable parameter matrix, is the spatial attention weight.
9. A traffic flow data completion method based on prior knowledge and Transformer as claimed in claim 1, characterized in that: Output the predicted values of the missing data and evaluate the loss of the predicted values, including: Further process the hidden state through a feed-forward neural network and finally generate predicted values at the output layer; ; Among them, is the updated node feature representation, is the predicted value, is the weight matrix of the first layer, is the bias vector of the first layer, is the weight matrix of the second layer, is the bias vector of the second layer, is the activation function, is the predicted value output function; Calculate the loss value of the mean absolute error based on the artificial mask; ; Among them, is the loss value of the mean absolute error, is the predicted value, is the true value, is the value in the mask matrix at dimension and time ; is the feature dimension, is the length of the time series; Adjust the projection matrix and the parameter matrix through the loss value to minimize the mean absolute error.
10. A traffic flow data completion method based on prior knowledge and Transformer as claimed in claim 1, characterized in that: Based on the optimization result of minimizing the loss function, perform imputation on the missing data in the traffic flow dataset, including: Use the predicted values after minimizing the mean absolute error for data imputation; The imputation formula is: ; Among them, is the result after interpolation, is the original data, is a binary matrix, is element-wise multiplication, is the predicted value.
11. A traffic flow data completion system based on prior knowledge and Transformer, used to implement the method described in any one of claims 1 to 10, including a data collection module, a data encoding module, an information fusion module, a spatio-temporal interaction module, and an evaluation and imputation module, characterized in that: The data collection module is used to collect traffic flow data sets; The data encoding module is used to perform missing encoding and random artificial missing encoding on the traffic flow data set in the data collection module; The information fusion module is used to fuse the missing encoding of the data encoding module and the traffic flow data set in the data collection module; The spatio-temporal interaction module is used to capture low-dimensional information by applying temporal attention and spatial attention to the fusion result of the information fusion module, and output a predicted value; The evaluation and imputation module is used to evaluate the loss of the predicted value, and use the predicted value after minimizing the mean absolute error operation for data imputation.
Citation Information
Patent Citations
Missing data completion method based on sparse topology space-time attention mechanism
CN118760677A
Traffic state data interpolation method and device, electronic equipment and readable storage medium
CN119811094A