Transformer-Based Multi-Task Traffic Flow Prediction Method, Device, Terminal, and Storage Medium
By using a Transformer-based multi-tasking traffic flow prediction method in traffic flow and traffic speed prediction, the spatiotemporal characteristics of historical traffic data are extracted and feature enhancement between decoders is solved, and a more efficient traffic flow and traffic speed prediction is achieved.
Patent Information
- Application Number
- CN202210660535.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-06-13
AI Technical Summary
The prior art ignores the high correlation between these two tasks in traffic flow and traffic speed prediction, resulting in low accuracy of prediction results and reducing the efficiency of traffic management.
The multi-task traffic flow prediction method based on Transformer is adopted, and the spatiotemporal characteristics of historical traffic data are extracted through a shared encoder, and the traffic flow and traffic speed are predicted respectively in two independent decoders. A hierarchical feature extraction module is designed between each decoder to extract node-level and spatial-level features from one prediction task to enhance the features of another prediction task.
Through the feature enhancement mechanism, the accuracy of traffic flow and traffic speed prediction results is improved, the efficiency of traffic management is improved, and the two prediction tasks are characterized by mutual promotion.
Smart Images

Figure CN115271157B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic prediction, and particularly to a multi-task traffic flow prediction method, device, terminal and storage medium based on Transformer. Background Art
[0002] The prediction of traffic flow and speed plays an important role in urban traffic management. Accurate prediction of traffic flow and speed can help traffic departments conduct traffic management, provide efficient route planning and navigation for citizens' travel, and is an indispensable part of intelligent transportation systems.
[0003] Currently, when people predict traffic flow and traffic speed, they usually use machine learning models to process these two tasks separately and then combine these learning results. However, the above approach ignores the high correlation between traffic flow and traffic speed, which affects the accuracy of the prediction results of traffic flow and traffic speed, and also reduces the efficiency of traffic management.
[0004] Therefore, the existing technology needs to be improved. Summary of the Invention
[0005] The main object of the present invention is to provide a multi-task traffic flow prediction method, device, intelligent terminal and storage medium based on Transformer, which can enhance the features of the two prediction tasks respectively based on the correlation between traffic flow and traffic speed, and improve the accuracy of the prediction results.
[0006] To achieve the above object, in the first aspect of the present invention, a multi-task traffic flow prediction method based on Transformer is provided, and the method includes:
[0007] Input historical traffic data into an encoder for feature extraction to obtain historical spatio-temporal features;
[0008] Input real-time traffic flow data and real-time traffic speed data into a decoder respectively, and input the historical spatio-temporal features into the decoder respectively;
[0009] Extract first spatial features in each decoder and extract second spatial features from another decoder, and fuse the first spatial features and the second spatial features to obtain complementary post-spatial features of each decoder;
[0010] Fuse the complementary post-spatial features with node features extracted from the complementary post-spatial features of another decoder to obtain complementary post-node features output by each decoder;
[0011] Based on the complementary post-node features, obtain traffic flow prediction results and traffic speed prediction results.
[0012] Optionally, extracting the second spatial feature from the decoder includes:
[0013] Based on the real-time traffic flow data and the real-time traffic speed data, obtain a correlation coefficient matrix according to the Kendall rank correlation coefficient;
[0014] Obtain the temporal feature output by the decoding layer of the decoder;
[0015] Based on the correlation coefficient matrix and the temporal feature, obtain the second spatial feature according to the graph convolutional neural network.
[0016] Optionally, at least two layers of temporal trend-aware self-attention layers are connected in series in the decoding layer of the decoder, and the output result of the last temporal trend-aware self-attention layer is the temporal feature.
[0017] Optionally, the expression of the temporal trend-aware self-attention layer is:
[0018] SelfTrAttention(Q, K, V) = softmax(conv(Q) × conv(K)) ⊙ V,
[0019] where conv represents a 1x1 convolution in the temporal dimension, Q is the query vector, K is the key-value vector, and V is the value vector.
[0020] Optionally, fusing the complemented spatial feature with the node feature extracted from the complemented spatial feature of another decoder to obtain the complemented node feature output by each decoder includes:
[0021] Based on the real-time traffic flow data and the real-time traffic speed data, obtain a correlation coefficient matrix according to the Kendall rank correlation coefficient;
[0022] Based on the correlation coefficient matrix and the complemented spatial feature of another decoder, obtain the node feature by using a convolution operation;
[0023] Fuse the complemented spatial feature and the node feature to obtain the complemented node feature.
[0024] Optionally, the step of obtaining the node feature by using a convolution operation based on the correlation coefficient matrix and the complemented spatial feature of another decoder includes:
[0025] Perform a convolution operation on the correlation coefficient matrix using a 1x1 convolution kernel to obtain a first convolution result;
[0026] Perform a convolution operation on the complemented spatial feature of another decoder using a 1x1 convolution kernel to obtain a second convolution result;
[0027] Perform a vector multiplication on the first convolution result and the second convolution result to obtain the node feature.
[0028] Optionally, the inputting the historical traffic data into an encoder for feature extraction to obtain historical spatio-temporal features includes:
[0029] Extract time features based on a time trend-aware self-attention layer;
[0030] Extract spatial features based on a spatial dynamic graph convolutional network layer;
[0031] Obtain the historical spatio-temporal features based on the time features and the spatial features.
[0032] The second aspect of the present invention provides a multi-task traffic flow prediction device, wherein the device includes:
[0033] A historical spatio-temporal feature extraction module, configured to input historical traffic data into an encoder for feature extraction to obtain historical spatio-temporal features;
[0034] A data input module, configured to input real-time traffic flow data and real-time traffic speed data into a decoder respectively and input the historical spatio-temporal features into the decoder respectively;
[0035] A spatial feature module, configured to extract a first spatial feature in each decoder and extract a second spatial feature from another decoder, and fuse the first spatial feature and the second spatial feature to obtain a complementary post-spatial feature of each decoder;
[0036] A node feature module, configured to fuse the complementary post-spatial feature with a node feature extracted from the complementary post-spatial feature of another decoder to obtain a complementary post-node feature output by each decoder;
[0037] An output module, configured to obtain a traffic flow prediction result and a traffic speed prediction result based on the complementary post-node feature.
[0038] The third aspect of the present invention provides an intelligent terminal, the intelligent terminal includes a memory, a processor, and a Transformer-based multi-task traffic flow prediction program stored on the memory and executable on the processor. When the Transformer-based multi-task traffic flow prediction program is executed by the processor, the steps of any one of the above-mentioned Transformer-based multi-task traffic flow prediction methods are implemented.
[0039] A fourth aspect of the present invention provides a computer-readable storage medium, on which a Transformer-based multi-task traffic flow prediction program is stored. When the Transformer-based multi-task traffic flow prediction program is executed by a processor, the steps of any one of the above-mentioned Transformer-based multi-task traffic flow prediction methods are implemented.
[0040] As can be seen from the above, compared with the prior art, the solution of the present invention sets one encoder and two decoders, uses the two encoders to process the traffic flow task and the traffic speed task respectively, extracts the historical spatio-temporal features of the historical traffic data by the encoder and inputs them into each decoder to enhance the features of the real-time traffic data input in each decoder; moreover, each decoder hierarchically extracts spatial features and node features from the other decoder to fuse and complement with its own features, so as to enhance the features of the other prediction task through one prediction task, realize mutual promotion in features, and improve the accuracy of the prediction result. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 is a model block diagram of the Transformer-based multi-task traffic flow prediction method of the present invention;
[0043] Figure 2 is a schematic flowchart of the Transformer-based multi-task traffic flow prediction method provided by an embodiment of the present invention;
[0044] Figure 3 is Figure 1 a specific flowchart of extracting the second spatial feature in step S300 of the embodiment;
[0045] Figure 4 is Figure 1 a specific flowchart of step S400 of the embodiment;
[0046] Figure 5 is Figure 1 a specific flowchart of step S100 of the embodiment;
[0047] Figure 6 is a schematic structural diagram of the Transformer-based multi-task traffic flow prediction device provided by an embodiment of the present invention;
[0048] Figure 7 It is a block diagram of the internal structure principle of an intelligent terminal provided by an embodiment of the present invention. Specific embodiments
[0049] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0050] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0051] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0052] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0053] As used in this specification and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0054] Next, in conjunction with the accompanying drawings of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0055] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Persons skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0056] Traffic flow prediction and traffic speed prediction are important research directions in the fields of intelligent transportation and smart cities, and have received increasing attention in recent years. Accurate and rapid traffic flow prediction and traffic speed prediction are of great help for efficient traffic management and route planning.
[0057] However, the existing method of separately processing traffic flow prediction and traffic speed prediction as different prediction tasks, due to not considering the correlation between traffic flow and traffic speed, results in low accuracy of the prediction results and cannot improve the efficiency of traffic management.
[0058] The present invention proposes a new prediction algorithm for traffic flow and traffic speed prediction, specifically including using a shared encoder to extract spatio-temporal features from historical traffic data, and using two independent decoders to predict traffic flow and traffic speed respectively. And a hierarchical feature extraction module is designed between the two decoders to extract node-level and spatial-level features from one prediction task to enhance the features of the other prediction task, so as to achieve the effect of improving the other prediction task. This enables the two prediction tasks to promote each other in terms of features and jointly improve the prediction accuracy.
[0059] Exemplary method
[0060] An embodiment of the present invention provides a multi-task traffic flow prediction method based on Transformer, as Figure 1 shown, using a shared Transformer encoder and two independent Transformer decoders. The encoder is used to extract spatio-temporal features from historical traffic data, and the decoders are used to predict traffic flow and traffic speed respectively. In terms of time relationship and time features, a TemporalTrend-Aware Self-Attention module is adopted in both the encoder and the decoder to extract time features. In terms of spatial relationship and spatial features, a Spatial Dynamic Graph Convolution Network module is adopted to extract spatial features. These modules are stacked to form the basic structure for the encoder and decoder to extract spatio-temporal features of traffic data.
[0061] Specifically, as Figure 2As shown in the figure, the above prediction method includes the following steps:
[0062] Step S100: Input the historical traffic data into the encoder for feature extraction to obtain historical spatio-temporal features;
[0063] Specifically, the traffic data is sampled data at consecutive time moments. The traffic data of the previous T consecutive time moments is used as historical traffic data, and the traffic data after the T-th moment is real-time traffic data. The historical traffic data includes historical traffic flow data and historical traffic speed data. The historical traffic data is first normalized, and then mapped and feature-extracted through the input mapping layer, and after temporal position embedding and spatial position embedding, it is input into the encoder. The encoder includes multiple encoder layers, and spatio-temporal features are extracted in the encoder layer to obtain the historical spatio-temporal features output by the encoder. Since a shared encoder is used, it can make the features of the input historical traffic data interact fully.
[0064] In this embodiment, The historical traffic data represented is for road network nodes, where T is the number of historical time steps. Assuming that the data sampling interval of the node is once every 5 minutes, then the time step interval is also 5 minutes. One hour has 12 time steps. N is the number of nodes on the road network. C is the number of channels of the input signal. In the present invention, C = 2, that is, two channels of traffic flow and traffic speed. First, the input mapping layer maps and extracts features from the input historical traffic data d model is the dimension of the neural network framework model. Since time features and spatial features play important roles in traffic data, temporal position embedding and spatial position embedding are also used to enable the model to obtain time features and spatial features.
[0065] The method adopted for temporal position embedding is as follows:
[0066]
[0067]
[0068] where t is the position of the input data, and d is the d-th dimension of the vector. d model is the dimension of the neural network framework model.
[0069] For spatial position embedding, the method of is first initialized and then Laplace smoothing is performed to obtain the spatial position embedding of each node.
[0070] Step S200: Input the real-time traffic flow data and the real-time traffic speed data into a decoder respectively, and input the historical spatio-temporal features into the decoder respectively;
[0071] Specifically, first, the real-time traffic flow data and the real-time traffic speed data are normalized, and then, after being mapped and feature-extracted through the input mapping layer, and with temporal position embedding (Temporal Position Embed) and spatial position embedding (Spatial Position Embed), they are input into a decoder respectively.
[0072] In this embodiment, the decoder includes two serially-connected time trend-aware self-attention layers. The historical spatio-temporal features are input into the second time trend-aware self-attention layer, so that the decoder can better utilize the historical spatio-temporal features for feature enhancement.
[0073] Input the historical spatio-temporal features output by the encoder into the decoder, and the decoder can predict the traffic flow and speed at the next time point according to the features of the historical data and the real-time traffic flow data.
[0074] Step S300: Extract the first spatial feature in each decoder and extract the second spatial feature from the other decoder, and fuse the first spatial feature and the second spatial feature to obtain the complementary post-spatial feature of each decoder;
[0075] Step S400: Fuse the complementary post-spatial feature with the node feature extracted from the complementary post-spatial feature of the other decoder to obtain the complementary post-node feature output by each decoder;
[0076] Specifically, since there is a correlation between the traffic flow prediction task and the traffic speed prediction task, in order to learn the information representation of another prediction task in one prediction task, the present invention first fuses the spatial features extracted by each decoder with the spatial features extracted from the other decoder to achieve the effect of complementing and enhancing the extracted spatial features. Then, the complementary post-spatial features in each decoder are fused with the node features extracted from the complementary post-spatial features of the other decoder to obtain the output result of the decoder, that is, the complementary post-node feature. Among them, the extraction of spatial features can set a spatial dynamic graph convolution network (Spatial Dynamic Graph ConvolutionNetwork) layer in the decoder as in this embodiment, and the spatial features are extracted by using the spatial dynamic graph convolution network module, denoted as GCN(.); the extraction of node features from the decoder can perform a convolution operation on the complementary post-spatial features in the decoder to obtain the node features, denoted as conv(.).
[0077] In this embodiment, a feature extraction module including node level and spatial level is designed between two independent encoders. The feature extraction module extracts node features and spatial features from a prediction task to enhance the features of another prediction task, so as to improve the effect of the other prediction task. The two prediction tasks promote each other in terms of features, improving the overall prediction effect of the model.
[0078] Step S500: Obtain traffic flow prediction results and traffic speed prediction results based on the complementary node features.
[0079] Specifically, after obtaining the complementary node features output by the decoder, input them into a linear layer to obtain traffic flow prediction results and traffic speed prediction results.
[0080] As described above, the present invention extracts spatio-temporal features from historical traffic data by using an encoder, inputs the historical spatio-temporal features into two decoders for respectively predicting traffic flow and traffic speed, and constructs a hierarchical feature extraction module between the two decoders. By extracting node and spatial level features in one decoder to enhance the features of the other decoder, the two encoders can jointly improve the feature extraction effect, so as to accurately and quickly predict traffic flow and traffic speed.
[0081] In one embodiment, as Figure 3 shown, extracting the second spatial feature from the decoder in the above step S300 specifically includes the following steps:
[0082] Step S310: Based on real-time traffic flow data and real-time traffic speed data, obtain a correlation coefficient matrix according to the Kendall rank correlation coefficient;
[0083] Specifically, in order to enhance the effect of feature complementarity, this embodiment introduces the Kendall rank correlation coefficient (KRCC). The Kendall rank correlation coefficient is used to measure the order correlation relationship between two variables. Calculate the corresponding KRCC according to the input real-time traffic flow data and real-time traffic speed data to obtain a correlation coefficient matrix. Using the Kendall rank correlation coefficient to represent the correlation relationship between traffic flow and traffic speed, extracting complementary features under the guidance of the Kendall rank correlation coefficient can be more effective.
[0084] Among them, the calculation expression of the Kendall rank correlation coefficient is as follows:
[0085]
[0086] Among them, sgn(·) is the sign function, (x i , y i ), (x j , yj ) are two variables.
[0087] It should be noted that this expression is used to calculate the correlation coefficient on a node. For the N nodes of the traffic prediction task, each node needs to be calculated separately to construct a correlation coefficient matrix to represent the correlation coefficient P ∈ R of the entire traffic node. N×1 .
[0088] Step S320: Obtain the time features output by the decoding layer of the decoder;
[0089] Specifically, the decoder includes multiple decoding layers. When processing traffic flow data, the above decoding layers include a decoding layer for extracting time features. As Figure 1 shown, in this embodiment, two time trend-aware self-attention layers are connected in series in the decoder to extract features in time, that is, to learn the time dependence between different time steps. The output result of the last time trend-aware self-attention layer is used as the time feature. Of course, the number of time trend-aware self-attention layers is not limited and can be one or more than two.
[0090] In this embodiment, the self-attention mechanism adopted by the time trend-aware self-attention layer is also improved, and a 1x1 convolutional layer operation in the time dimension is adopted. The specific expression is:
[0091] SelfTrAttention(Q, K, V) = softmax(conv(Q) × conv(K)) ⊙ V,
[0092] where conv represents a 1x1 convolution in the time dimension, Q is the query vector, K is the key-value vector, and V is the value vector.
[0093] Due to the adoption of this convolution, when calculating the attention, the context in time is considered, so that the model has a certain perception of the trend in the time series.
[0094] Step S330: Based on the correlation coefficient matrix and the time features, obtain the second spatial feature according to the graph convolutional neural network.
[0095] Specifically, multiply the correlation coefficient matrix and the time features, and then obtain the second spatial feature according to the graph convolutional neural network.
[0096] In this embodiment, the two decoders respectively correspond to the feature extraction of traffic flow and traffic speed, including the extraction and modeling of the spatial relationship between traffic flow and traffic speed. When the decoder extracts its own spatial features, the time features output by the last temporal trend-aware self-attention layer are subjected to feature extraction according to the graph convolutional neural network to obtain the first spatial features; since the Kendall rank correlation coefficient is considered, when extracting the second spatial features from the decoder, the time features output by the temporal trend-aware self-attention layer need to be matrix-multiplied with the correlation coefficient matrix, and then feature extraction is performed according to the graph convolutional neural network to obtain the second spatial features.
[0097] Among them, for the traffic flow prediction task, the specific calculation method is as follows:
[0098]
[0099] Among them, S speed is the time feature output by the temporal trend-aware self-attention layer in the traffic speed prediction task, S flow is the time feature output by the temporal trend-aware self-attention layer in the traffic flow prediction task, GCN is the graph convolutional neural network, P is the correlation coefficient matrix, is the complementary post-spatial feature of the decoder corresponding to the traffic flow prediction task. GCN(S flow ) is the first spatial feature, GCN(S speed ⊙P) is the second spatial feature, and Norm[] is the normalization layer.
[0100] For the traffic speed prediction task, the specific calculation method is as follows:
[0101]
[0102] Among them, S speed is the time feature output by the temporal trend-aware self-attention layer in the traffic speed prediction task, S flow is the time feature output by the temporal trend-aware self-attention layer in the traffic flow prediction task, GCN is the graph convolutional neural network, P is the correlation coefficient matrix, is the complementary post-spatial feature of the decoder corresponding to the traffic speed prediction task. GCN(S speed ) is the first spatial feature, GCN(S flow ⊙P) is the second spatial feature, and Norm[] is the normalization layer.
[0103] As described above, under the guidance of KRCC, feature extraction at the spatial level is performed, so that effective features are extracted from one decoder to enhance the features of the other decoder.
[0104] In one embodiment, as Figure 4As shown, the above step S400 specifically includes the following steps:
[0105] Step S410: Based on the real-time traffic flow data and real-time traffic speed data, obtain the correlation coefficient matrix according to the Kendall rank correlation coefficient;
[0106] Specifically, referring to the content described in the above step S310, obtain the correlation coefficient matrix.
[0107] Step S420: Based on the correlation coefficient matrix and the complementary post-space features of another decoder, obtain the node features by using convolution operations;
[0108] Specifically, since the Kendall rank correlation coefficient is considered, when extracting the node features of another decoder, it is necessary to first perform convolution operations on the correlation coefficient matrix and the complementary post-space features of another decoder, and then perform matrix multiplication to obtain the node features.
[0109] In this embodiment, first perform a convolution operation on the correlation coefficient matrix using a 1x1 convolution kernel to obtain a convolution result; also perform a convolution operation on the complementary post-space features of another decoder using a 1x1 convolution kernel to obtain a convolution result; multiply the two convolution results to obtain the node features.
[0110] Step S430: Fuse the complementary post-space features and the node features to obtain the complementary post-node features.
[0111] Specifically, add the node features and the complementary post-space features as feature vectors to obtain the complementary post-node features.
[0112] In this embodiment, the two decoders respectively correspond to the feature extraction of traffic flow and traffic speed.
[0113] For the traffic flow prediction task, the specific calculation method is as follows:
[0114]
[0115] Among them, is the complementary post-space feature in the traffic speed prediction task, is the complementary post-space feature in the traffic flow prediction task, conv is a convolutional layer with a 1x1 convolution kernel, P is the correlation coefficient matrix, is the node feature, is the complementary post-node feature, Norm[] is the normalization layer.
[0116] For the traffic speed prediction task, the specific calculation method is as follows:
[0117]
[0118] Among them, is the complementary post - spatial feature in the traffic speed prediction task, is the complementary post - spatial feature in the traffic flow prediction task, conv is a convolutional layer with a 1x1 convolutional kernel, and P is the correlation coefficient matrix, is the node feature, is the complementary post - node feature, and Norm[] is the normalization layer.
[0119] As described above, node - level feature extraction is carried out under the guidance of KRCC. Effective features are extracted from one decoder to enhance the features of another decoder.
[0120] In one embodiment, as Figure 5 shown, the above - mentioned step S100 specifically includes the following steps:
[0121] Step S110: Extract time features based on the time - trend - aware self - attention layer;
[0122] Step S120: Extract spatial features based on the spatial dynamic graph convolutional network layer;
[0123] Step S130: Obtain historical spatio - temporal features based on the time features and spatial features.
[0124] Specifically, in extracting the time relationship and time features of traffic data, a time - trend - aware self - attention mechanism module is used for feature extraction. In extracting the spatial relationship and spatial features of traffic data, a spatial dynamic graph convolutional network module is used for feature extraction. In the encoder, the time features and spatial features are interacted to obtain historical spatio - temporal features. Among them, the spatial dynamic graph convolutional network layer uses the spatial dynamic graph convolutional network module (DynGCN), which models the representation learning on the dynamic graph as the aggregation of time and spatial information. The model combines the spatial convolution of the graph convolutional neural network (GCN) to extract the structural information on the graph, and the causal convolution of the temporal convolutional neural network (TCN) to extract the historical information in the time series. At the same time, an adaptive model update mechanism is added to the spatial convolutional layer, so that the model parameters are updated adaptively with the change of the graph structure, which is a common technical means in the field and will not be elaborated here.
[0125] As described above, the encoder encodes the input data, enables the input features to interact fully, and extracts the spatio - temporal features existing therein.
[0126] Exemplary device
[0127] As Figure 6As shown, corresponding to the above-mentioned Transformer-based multi-task traffic flow prediction method, an embodiment of the present invention further provides a Transformer-based multi-task traffic flow prediction device, and the above-mentioned Transformer-based multi-task traffic flow prediction device includes:
[0128] A historical spatio-temporal feature extraction module 600, configured to input historical traffic data into an encoder for feature extraction to obtain historical spatio-temporal features;
[0129] A data input module 610, configured to input real-time traffic flow data and real-time traffic speed data into one decoder respectively and input the historical spatio-temporal features into the decoder respectively;
[0130] A spatial feature module 620, configured to extract a first spatial feature in each decoder and extract a second spatial feature from another decoder, and fuse the first spatial feature and the second spatial feature to obtain a complementary post-spatial feature of each decoder;
[0131] A node feature module 630, configured to fuse the complementary post-spatial feature with node features extracted from the complementary post-spatial feature of another decoder to obtain a complementary post-node feature output by each decoder;
[0132] An output module 640, configured to obtain a traffic flow prediction result and a traffic speed prediction result based on the complementary post-node feature.
[0133] Specifically, in this embodiment, the specific functions of the modules of the above-mentioned Transformer-based multi-task traffic flow prediction device can refer to the corresponding descriptions in the above-mentioned Transformer-based multi-task traffic flow prediction method, which will not be elaborated here.
[0134] Based on the above embodiment, the present invention further provides an intelligent terminal, and its principle block diagram can be as Figure 7As shown in the figure. The above intelligent terminal includes a processor, a memory, a network interface, and a display screen connected by a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a multi-task traffic flow prediction program based on Transformer. The internal memory provides an environment for the operation of the operating system and the multi-task traffic flow prediction program based on Transformer in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the multi-task traffic flow prediction program based on Transformer is executed by the processor, it implements the steps of any of the above multi-task traffic flow prediction methods based on Transformer. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen.
[0135] Those skilled in the art can understand that Figure 7 The principle block diagram shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0136] In one embodiment, an intelligent terminal is provided. The above intelligent terminal includes a memory, a processor, and a multi-task traffic flow prediction program based on Transformer stored on the above memory and executable on the above processor. When the multi-task traffic flow prediction program based on Transformer is executed by the above processor, the following operation instructions are performed:
[0137] Input historical traffic data into the encoder for feature extraction to obtain historical spatio-temporal features;
[0138] Input real-time traffic flow data and real-time traffic speed data into one decoder each and input the historical spatio-temporal features into the decoders respectively;
[0139] Extract the first spatial feature in each decoder and extract the second spatial feature from the other decoder, and fuse the first spatial feature and the second spatial feature to obtain the complementary post-spatial feature of each decoder;
[0140] Fuse the complementary post-spatial feature with the node feature extracted from the complementary post-spatial feature of the other decoder to obtain the complementary post-node feature output by each decoder;
[0141] Based on the complementary post-node feature, obtain a traffic flow prediction result and a traffic speed prediction result.
[0142] Optionally, extracting the second spatial feature from the decoder includes:
[0143] Based on the real-time traffic flow data and the real-time traffic speed data, obtaining a correlation coefficient matrix according to the Kendall rank correlation coefficient;
[0144] Obtaining the temporal feature output by the decoding layer of the decoder;
[0145] Based on the correlation coefficient matrix and the temporal feature, obtaining the second spatial feature according to the graph convolutional neural network.
[0146] Optionally, at least two layers of temporal trend-aware self-attention layers are connected in series in the decoding layer of the decoder, and the output result of the last temporal trend-aware self-attention layer is the temporal feature.
[0147] Optionally, the expression of the temporal trend-aware self-attention layer is:
[0148] SelfTrAttention(Q, K, V) = softmax(conv(Q) × conv(K)) ⊙ V,
[0149] where conv represents a 1x1 convolution in the time dimension, Q is the query vector, K is the key-value vector, and V is the value vector.
[0150] Optionally, fusing the complementary post-spatial feature with the node feature extracted from the complementary post-spatial feature of another decoder to obtain the complementary post-node feature output by each decoder includes:
[0151] Based on the real-time traffic flow data and the real-time traffic speed data, obtaining a correlation coefficient matrix according to the Kendall rank correlation coefficient;
[0152] Based on the correlation coefficient matrix and the complementary post-spatial feature of another decoder, obtaining the node feature by using a convolution operation;
[0153] Fusing the complementary post-spatial feature and the node feature to obtain the complementary post-node feature.
[0154] Optionally, the step of obtaining the node feature by using a convolution operation based on the correlation coefficient matrix and the complementary post-spatial feature of another decoder includes:
[0155] Performing a convolution operation on the correlation coefficient matrix by using a 1x1 convolution kernel to obtain a first convolution result;
[0156] Performing a convolution operation on the complementary post-spatial feature of another decoder by using a 1x1 convolution kernel to obtain a second convolution result;
[0157] Multiply the first convolution result and the second convolution result vectorially to obtain the node features.
[0158] Optionally, the inputting the historical traffic data into the encoder for feature extraction to obtain historical spatio-temporal features includes:
[0159] Extract time features based on a time trend-aware self-attention layer;
[0160] Extract spatial features based on a spatial dynamic graph convolutional network layer;
[0161] Obtain the historical spatio-temporal features based on the time features and the spatial features.
[0162] An embodiment of the present invention also provides a computer-readable storage medium, on which a multi-task traffic flow prediction program based on Transformer is stored. When the multi-task traffic flow prediction program based on Transformer is executed by a processor, the steps of any one of the multi-task traffic flow prediction methods provided by the embodiments of the present invention are implemented.
[0163] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0164] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the above device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0165] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0166] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0167] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the above-mentioned division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0168] If the above-mentioned integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The above computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the above computer program includes computer program code, and the above computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The above computer-readable medium can include: any entity or device capable of carrying the above computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the above computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0169] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the present invention in essence, and should all be included in the protection scope of the present invention.
Claims
1. A multi-task traffic flow prediction method based on Transformer, including a traffic flow prediction task and a traffic speed prediction task, characterized in that, the method includes: Inputting historical traffic data into an encoder for feature extraction to obtain historical spatio-temporal features; Inputting real-time traffic flow data and real-time traffic speed data into a decoder respectively and inputting the historical spatio-temporal features into the decoder respectively; Extracting first spatial features in each decoder and extracting second spatial features from another decoder, and fusing the first spatial features and the second spatial features to obtain complementary post-spatial features of each decoder; Fusing the complementary post-spatial features with node features extracted from the complementary post-spatial features of another decoder to obtain complementary post-node features output by each decoder; Based on the complementary post-node features, obtaining a traffic flow prediction result and a traffic speed prediction result; The extracting second spatial features from another decoder includes: Based on the real-time traffic flow data and the real-time traffic speed data, obtaining a correlation coefficient matrix according to the Kendall rank correlation coefficient; Obtaining time features output by the decoding layer of the decoder; Based on the correlation coefficient matrix and the time features, obtaining the second spatial features according to the graph convolutional neural network; The fusing the complementary post-spatial features with node features extracted from the complementary post-spatial features of another decoder to obtain complementary post-node features output by each decoder includes: Based on the real-time traffic flow data and the real-time traffic speed data, obtaining a correlation coefficient matrix according to the Kendall rank correlation coefficient; Based on the correlation coefficient matrix and the complementary post-spatial features of another decoder, obtaining the node features by using a convolution operation; Fusing the complementary post-spatial features and the node features to obtain the complementary post-node features.
2. The multi-task traffic flow prediction method based on Transformer according to claim 1, characterized in that, At least two layers of time trend-aware self-attention layers are connected in series in the decoding layer of the decoder, and the output result of the last time trend-aware self-attention layer is the time feature.
3. The multi-task traffic flow prediction method based on Transformer according to claim 2, characterized in that, The expression of the time trend-aware self-attention layer is: , Among them, conv represents a 1x1 convolution in the time dimension. is the query vector. is the key-value vector. is the value vector.
4. The multi-task traffic flow prediction method based on Transformer according to claim 1, characterized in that, The obtaining the node features by using a convolution operation based on the correlation coefficient matrix and the complementary post-spatial features of another decoder includes: Performing a convolution operation on the correlation coefficient matrix by using a 1x1 convolution kernel to obtain a first convolution result; Performing a convolution operation on the complementary post-spatial features of another decoder by using a 1x1 convolution kernel to obtain a second convolution result; Performing a vector multiplication on the first convolution result and the second convolution result to obtain the node features.
5. The multi-task traffic flow prediction method based on Transformer according to claim 1, characterized in that, Inputting the historical traffic data into an encoder for feature extraction to obtain historical spatio-temporal features, including: Extracting temporal features based on a time-trend aware self-attention layer; Extracting spatial features based on a spatial dynamic graph convolutional network layer; Obtaining the historical spatio-temporal features based on the temporal features and the spatial features.
6. A multi-task traffic flow prediction device based on Transformer, characterized in that, the device is used to implement the steps of the multi-task traffic flow prediction method based on Transformer according to any one of claims 1-5, and the device includes: A historical spatio-temporal feature extraction module, configured to input historical traffic data into an encoder for feature extraction to obtain historical spatio-temporal features; A data input module, configured to input real-time traffic flow data and real-time traffic speed data into a decoder respectively and input the historical spatio-temporal features into the decoder respectively; A spatial feature module, configured to extract first spatial features in each decoder and extract second spatial features from the other decoder, and fuse the first spatial features and the second spatial features to obtain complementary post-spatial features of each decoder; A node feature module, configured to fuse the complementary post-spatial features with node features extracted from the complementary post-spatial features of the other decoder to obtain complementary post-node features output by each decoder; An output module, configured to obtain a traffic flow prediction result and a traffic speed prediction result based on the complementary post-node features.
7. An intelligent terminal, characterized in that, the intelligent terminal includes a memory, a processor, and a multi-task traffic flow prediction program based on Transformer stored on the memory and executable on the processor. When the multi-task traffic flow prediction program based on Transformer is executed by the processor, it implements the steps of the multi-task traffic flow prediction method based on Transformer according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, a multi-task traffic flow prediction program based on Transformer is stored on the computer-readable storage medium. When the multi-task traffic flow prediction program based on Transformer is executed by a processor, it implements the steps of the multi-task traffic flow prediction method based on Transformer according to any one of claims 1-5.
Citation Information
Patent Citations
Space-time prediction method based on knowledge distillation in industrial Internet of Things edge equipment
CN113988263A
A distributed network traffic data decomposition method
WO2021186158A1