A traffic prediction method based on multimodal data fusion and its application

By constructing a traffic prediction model based on multimodal input data and a traffic prediction model based on multimodal data fusion, combined with traffic data of multiple modes, the problem of using only single mode data in the prior art is solved, and more accurate traffic state prediction is achieved.

CN115293428BActive Publication Date: 2025-05-13UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210944879.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2025-05-13
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

Existing traffic prediction methods use only one type of traffic data and fail to fully utilize multiple types of traffic data, thus limiting prediction performance.

Method used

A traffic prediction method based on multimodal data fusion is proposed. Accurate prediction of traffic state is achieved by constructing multimodal input data and constructing a traffic prediction model based on multimodal data fusion, including input conversion module, spatiotemporal embedding module, cross-modal attention module, maximum pooling fusion layer, spatiotemporal attention module and output linear layer.

Benefits of technology

By combining traffic data of multiple modes, the accuracy of traffic state prediction is improved, the problem of insufficient information of single mode data is overcome, and traffic data of multiple modes is effectively utilized, which improves prediction performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115293428B_ABST
    Figure CN115293428B_ABST
Patent Text Reader

Abstract

The present invention discloses a traffic prediction method and application based on multimodal data fusion, the method steps include: 1. constructing multimodal input data, 2. processing the multimodal input data using an input conversion module, 3. generating a spatiotemporal embedding module using spatiotemporal information, 4. processing data between different modalities using a cross-modal attention module, 5. fusing multimodal data using a maximum pooling fusion layer, 6. further performing data conversion using a spatiotemporal attention module, 7. converting output prediction results using an output linear layer, 8. iteratively performing network training to obtain a trained model. The present invention can efficiently combine traffic data of multiple modalities to achieve accurate traffic status prediction, thereby effectively helping urban traffic managers to make overall arrangements in advance and reduce urban road congestion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of traffic prediction, and specifically relates to a traffic prediction method and application based on multimodal data fusion. Background Art

[0002] With the increase in the number of cars in cities, urban traffic congestion is becoming more and more serious. Using historical traffic data to predict urban traffic conditions in the future can help urban traffic managers take measures in advance to slow down or avoid traffic congestion, and can also help travelers make reasonable travel plans.

[0003] The existing mainstream related technologies all use deep neural networks to realize traffic condition prediction. However, the existing mainstream related technologies only use one type of traffic data when making traffic predictions, ignoring the fact that traffic sensors can generate multiple types of traffic data at the same time, and fail to make full use of the existing rich traffic data to improve prediction performance. Summary of the invention

[0004] The present invention aims to solve the deficiencies of the above-mentioned prior art and proposes a traffic prediction method based on multimodal data fusion, in order to efficiently combine traffic data of multiple modes to achieve accurate traffic status prediction, thereby effectively helping urban traffic managers to make overall arrangements in advance and reduce urban road congestion.

[0005] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme:

[0006] The traffic prediction method based on multimodal data fusion of the present invention is characterized in that it comprises the following steps:

[0007] Step 1: Construct multimodal input data X;

[0008] Step 1.1: Build a directed road network graph in, It is the set of all traffic sensors in the road network; It is the set of road segments between various traffic sensors; is the adjacency matrix, the adjacency matrix If the element value in is 1, it means that there is a road section connecting the two traffic sensors. If the element value is 0, it means that there is no road section connecting the two traffic sensors.

[0009] Step 1.2: Road network diagram The N traffic sensors in the system record the traffic status data of C modes every time step, and after normalizing the traffic status data of each mode, the traffic status data of N traffic sensors in L time steps are obtained.

[0010] Step 1.3, from X all Select C types of traffic state data of Y consecutive historical time steps and use them as multimodal input data Let the sub-input data of the cth mode be expressed as T <L;

[0011] Step 2: Construct a traffic prediction model based on multimodal data fusion, including: input conversion module, spatiotemporal embedding module, cross-modal attention module, maximum pooling fusion layer, spatiotemporal attention module and output linear layer;

[0012] The input conversion module includes: an input linear layer and a position embedding layer;

[0013] The spatiotemporal embedding module includes: a spatial embedding module and a temporal embedding module;

[0014] The cross-modal attention module includes: a first cross-modal attention layer, a first feed-forward neural network, a second cross-modal attention layer, and a second feed-forward neural network;

[0015] The spatiotemporal attention module includes: a temporal attention layer, a third feedforward neural network, a spatial attention layer, and a fourth feedforward neural network;

[0016] Step 3, processing of the input conversion module;

[0017] Step 3.1: The input linear layer converts the sub-input data X of the cth mode into c Perform the transformation process to obtain the transformation data of the cth mode containing the D-dimensional latent space

[0018] Step 3.2: The position embedding layer transforms the data Z of the cth modality c 0 Perform position embedding operation to obtain the data of the cth mode after embedding the position Thus, the data of C modes after the embedding position are obtained and connected to obtain the connected data

[0019] Step 4: Processing of the spatiotemporal embedding module;

[0020] Step 4.1: The spatial embedding module uses the node2vec method to transform the adjacency matrix Transformed into a spatial embedding matrix

[0021] Step 4.2, processing of the time embedding module;

[0022] Step 4.2.1: The time embedding module transforms the traffic state data X intoall Convert into a sampling signal in the frequency domain, and analyze the sampling signal in the frequency domain to obtain F time period information;

[0023] Step 4.2.2: Encode the F periodic information using one-hot encoding to obtain the F relative position vectors of the lth time step and connect them to obtain the periodic embedding vector V corresponding to the lth time step. l ;

[0024] Step 4.2.3: Connect the period embedding vectors of the selected T consecutive historical time steps and the period embedding vectors corresponding to the subsequent T' consecutive future time steps, and then obtain the time embedding matrix after processing through the fully connected layer. T′ <L;

[0025] Step 4.3: Add the spatial embedding matrix SE and the temporal embedding matrix TE to obtain the spatiotemporal embedding vector Among them, the spatiotemporal embedding subvector containing the historical time step information is expressed as The spatiotemporal embedding subvector containing the information of future time steps is expressed as

[0026] Step 5: Processing of the cross-modal attention module;

[0027] Step 5.1: Z 1 With E (T) After connection, we get the tensor And input into the first cross-modal attention layer, after being processed by the fully connected layer with ReLU as the activation function, the query, key, and value tensors corresponding to the hth attention head are obtained respectively. Then use formula (1) to get the tensor output by the first cross-modal attention layer

[0028]

[0029] In formula (1), || h∈H It means that H subspaces are concatenated in sequence; d represents the dimension of the subspace of each attention head; and H×d=D;

[0030] Step 5.4: Convert the tensor Z 2 Input into the first feedforward neural network, and obtain the tensor output by the first feedforward neural network through formula (2):

[0031] Z 3 =ReLU(Z 2 W1+b1)W2+b2 (2)

[0032] In formula (2), W1 and W2 are learnable weight parameters in the first feedforward neural network; b1 and b2 are learnable bias parameters in the first feedforward neural network;

[0033] Step 5.5: The tensor Z 3 After being processed by the second cross-modal attention layer and the second feedforward neural network, the tensor is obtained. And used as output data of the cross-modal attention module;

[0034] Step 6: Processing of the maximum pooling fusion layer;

[0035] According to the order of each mode, take out the tensor Z respectively 4 The tensor of one dimension in the C modalities is spliced ​​to obtain a spliced ​​tensor of one dimension, thereby obtaining a spliced ​​tensor of D dimensions in C modalities and splicing them into the final interleaved spliced ​​tensor, which is then input into the maximum pooling fusion layer for multimodal fusion to obtain the fused data.

[0036] Step 7: Processing of the spatiotemporal attention module;

[0037] Step 7.1: Z 5 and E (T′) After the connection, we get the tensor And input into the temporal attention layer, after being processed by the fully connected layer with ReLU as the activation function, the query, key, and value tensors corresponding to the hth attention head are obtained respectively Thus, using formula (3), we can get the tensor Z output by the temporal attention layer: 6 :

[0038]

[0039] In formula (3), represents the attention score matrix corresponding to the hth attention head in the temporal attention layer, and is obtained by formula (4):

[0040]

[0041] In formula (4), is the attention score matrix , where is the attention score between the y-th time step and the z-th time step on the x-th traffic sensor; represents the correlation between the xth traffic sensor corresponding to the hth attention head at the yth time step and the zth time step, and is obtained by formula (5):

[0042]

[0043] In formula (5), yes where represents the vector of the x-th traffic sensor and the y-th time step, yes where represents the vector of the x-th traffic sensor and the z-th time step;

[0044] Step 7.2: The tensor Z output by the temporal attention layer 6 Input into the third feedforward neural network for processing to obtain a tensor

[0045] Step 7.3: Z 7 and E (T′) After the connection, we get the tensor And input into the spatial attention layer, after being processed by the fully connected layer with ReLU as the activation function, the query, key, and value tensors corresponding to the hth attention head are obtained respectively Thus, the tensor Z output by the temporal attention layer is obtained using formula (6): 8 :

[0046]

[0047] In formula (6), represents the attention score matrix corresponding to the hth attention head in the spatial attention layer, and is obtained by formula (7):

[0048]

[0049] In formula (6), is the attention score matrix is the attention score between the βth traffic sensor and the γth traffic sensor at the ath time step, represents the correlation between the βth traffic sensor and the γth traffic sensor at the ath time step corresponding to the hth attention head, and is obtained by formula (8);

[0050]

[0051] In formula (8), yes where represents the vector of the ath time step and the βth traffic sensor, yes where represents the vector of the ath time step and the γth traffic sensor;

[0052] Step 7.4: The output tensor Z of the spatial attention layer 8Input to the fourth feedforward neural network for further processing, and get the tensor output by the feedforward neural network

[0053] Step 8: The tensor Z 9 After the transformation of the output linear layer, the prediction result of the multimodal input data X is obtained

[0054] Step 9: Network training;

[0055] Step 9.1: Use formula (7) to construct the loss function

[0056]

[0057] In formula (7), is the prediction result of the nth future time step, Y n The label value of the nth future time step; Θ is all parameters of the traffic prediction model based on multimodal data fusion; T' is the total number of prediction steps in the future time;

[0058] Step 8.2: Use back propagation and gradient descent method to train the traffic prediction model based on multimodal data fusion, and calculate the loss value. When the number of iterations reaches the threshold ξ or the loss value does not decrease for a certain number of consecutive rounds, stop the training to obtain the optimal model after training and its optimal parameters Θ.

[0059] The present invention discloses an electronic device, comprising a memory and a processor, wherein the memory is used to store a program supporting the processor to execute the traffic prediction method, and the processor is configured to execute the program stored in the memory.

[0060] The present invention provides a computer-readable storage medium on which a computer program is stored. The computer-readable storage medium is characterized in that when the computer program is run by a processor, the steps of the traffic prediction method are executed.

[0061] Compared with the prior art, the present invention has the following beneficial effects:

[0062] 1. The present invention combines traffic data of multiple modes to predict traffic status, which can overcome the problem of insufficient information of single modal data, thereby improving the prediction accuracy of target modal data;

[0063] 2. The present invention realizes the learning and processing of different modal data through the cross-modal attention module, and captures important information between different modalities through maximum pooling fusion, thereby effectively utilizing traffic data of multiple modalities and improving prediction accuracy;

[0064] 3. The present invention deeply mines spatiotemporal information through the spatiotemporal embedding mechanism, providing more spatiotemporal information for the attention mechanism, thereby helping to improve the model learning efficiency and prediction accuracy;

[0065] 4. The present invention realizes the learning of important spatiotemporal information through the spatiotemporal attention module, thereby helping to generate more accurate prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 This is a model framework diagram of the present invention;

[0067] Figure 2 It is a structural diagram of the spatiotemporal embedding module of the present invention;

[0068] Figure 3 It is a schematic diagram of the interleaved splicing operation in the maximum pooling fusion module of the present invention, wherein the number in each unit represents the modal sequence number to which the element in the corresponding tensor belongs. DETAILED DESCRIPTION

[0069] In this embodiment, a traffic prediction method based on multimodal data fusion includes the following steps:

[0070] Step 1: Construct multimodal input data X;

[0071] Step 1.1: Build a directed road network graph in, It is the set of all traffic sensors in the road network; It is the set of road segments between various traffic sensors; is the adjacency matrix, the adjacency matrix If the element value in is 1, it means that there is a road section connecting the two traffic sensors. If the element value is 0, it means that there is no road section connecting the two traffic sensors.

[0072] Step 1.2: Road network diagram The N traffic sensors in the system record the traffic status data of C modes every time step (for example, every 5 minutes), and after normalizing the traffic status data of each mode, the traffic status data of N traffic sensors in L time steps are obtained.

[0073] Step 1.3, from X all Select C types of traffic state data of T consecutive historical time steps and use them as multimodal input data In this embodiment, C=3, that is, there are three modes of traffic status data. In addition, let the sub-input data of the cth mode be expressed as T <L;

[0074] Step 2: Construct a traffic prediction model based on multimodal data fusion, including: input conversion module, spatiotemporal embedding module, cross-modal attention module, maximum pooling fusion layer, spatiotemporal attention module and output linear layer;

[0075] The input transformation module includes: input linear layer and position embedding layer;

[0076] The spatiotemporal embedding module includes: a spatial embedding module and a temporal embedding module;

[0077] The cross-modal attention module includes: a first cross-modal attention layer, a first feed-forward neural network, a second cross-modal attention layer, and a second feed-forward neural network;

[0078] The spatiotemporal attention module includes: a temporal attention layer, a third feedforward neural network, a spatial attention layer, and a fourth feedforward neural network;

[0079] Step 3: Processing of the input conversion module, the purpose is to process and convert the input data of each mode separately to obtain data suitable for subsequent reading and processing by each module;

[0080] Step 3.1: Figure 1 As shown in the input linear layer in , the input linear layer converts the sub-input data X of the cth mode into c Perform the transformation process to obtain the transformation data of the cth mode containing the D-dimensional latent space

[0081] Step 3.2: Figure 1 As shown in the position embedding layer in , the position embedding layer transforms the data Z of the cth modality c 0 Perform position embedding operation to obtain the data of the cth mode after embedding the position Thus, the data of C modes after the embedding position are obtained and connected to obtain the connected data

[0082] Step 4: Processing of the spatiotemporal embedding module. The structure of the spatiotemporal embedding module is as follows: Figure 2 As shown, it contains spatial embedding module and temporal embedding module, which are used to mine deep spatiotemporal information, thereby providing more information for subsequent attention modules;

[0083] Step 4.1: The spatial embedding module uses the node2vec method to transform the adjacency matrix Transformed into a spatial embedding matrix

[0084] Step 4.2, processing of the time embedding module;

[0085] Step 4.2.1: The time embedding module uses discrete Fourier transform to transform the traffic state data Xall Convert into a sampling signal in the frequency domain, and analyze the sampling signal in the frequency domain to obtain F time period information;

[0086] Step 4.2.2: Encode the F periodic information using one-hot encoding to obtain the F relative position vectors of the lth time step and connect them to obtain the periodic embedding vector V corresponding to the lth time step. l ; For example, assume that the data set starts at 00:00 on July 8, and there are a total of 5 time period information, representing 1 week, 1 day, 12 hours, 8 hours and 6 hours. Next, assume that the time step of 13:00 on July 10 is one-hot encoded, then the encoding positions in the 5 corresponding encoding vectors are 3 (the third day in a week), 157 (the 157th time step in a day), 2 (in the second 12-hour cycle in 24 hours), 2 (in the second 8-hour cycle in 24 hours), 3 (in the third 6-hour cycle in 24 hours), and then connect these 5 relative position vectors to get the period embedding vector corresponding to the time step.

[0087] Step 4.2.3: Connect the period embedding vectors of the selected T consecutive historical time steps and the period embedding vectors corresponding to the subsequent T' consecutive future time steps, and then obtain the time embedding matrix after processing through the fully connected layer. T′ <L;

[0088] Step 4.3: Add the spatial embedding matrix SE and the temporal embedding matrix TE to obtain the spatiotemporal embedding vector Among them, the spatiotemporal embedding subvector containing historical time step information is expressed as The spatiotemporal embedding subvector containing the information of future time steps is expressed as

[0089] Step 5: Figure 1 As shown in the cross-modal attention module in Figure 2, the data is processed by the cross-modal attention module to learn and mine information between different modalities;

[0090] Step 5.1: Z 1 With E (T) After connection, we get the tensor And input into the first cross-modal attention layer, after being processed by the fully connected layer with ReLU as the activation function, the query, key, and value tensors corresponding to the h-th attention head are obtained respectively. Formula (1) is then used to enhance the intra-modal and inter-modal feature representation of the data, and the tensor output by the first cross-modal attention layer is obtained:

[0091]

[0092] In formula (1), || h∈H It means that H subspaces are concatenated in sequence; d represents the dimension of the subspace of each attention head; and H×d=D;

[0093] Step 5.4: Convert the tensor Z 2 Input to Figure 1 The first feedforward neural network is further processed in the first feedforward neural network, and the tensor output by the first feedforward neural network is obtained by formula (2):

[0094] Z 3 =ReLU(Z 2 W1+b1)W2+b2 (2)

[0095] In formula (2), W1 and W2 are learnable weight parameters in the first feedforward neural network; b1 and b2 are learnable bias parameters in the first feedforward neural network;

[0096] Step 5.5, tensor Z 3 Then go through Figure 1 After the second cross-modal attention layer and the second feed-forward neural network processing, the data representation is further enhanced to obtain the tensor And serve as the output data of the cross-modal attention module;

[0097] Step 6: Figure 1 As shown in , the data enters the maximum pooling fusion layer for processing, and the data of the three modes are interlaced and spliced ​​and then processed by maximum pooling. After fusion, the most significant feature representation is obtained and used as the output data; as shown in Figure 3 As shown, the tensor Z is taken out in the order of the three modes. 4 The tensor of one dimension in the data is concatenated to obtain a concatenated tensor of one dimension in three modes, thereby obtaining a concatenated tensor of D dimensions in three modes and concatenating them into the final interleaved concatenated tensor, which is then input into the maximum pooling fusion layer for multimodal fusion to obtain the fused data.

[0098] Step 7: Figure 1 As shown, the fused data Z 5 Enter the spatiotemporal attention module for processing, and the fused data Z 5 The feature representation is enhanced in time and space dimensions;

[0099] Step 7.1: Z 5 and E (T′) After the connection, we get the tensor and enter into Figure 1In the mid-time attention layer, after being processed by the fully connected layer with ReLU as the activation function, the query, key, and value tensors corresponding to the h-th attention head are obtained respectively. Thus, the feature representation of the time dimension of the data is enhanced using formula (3), and the tensor Z output by the time attention layer is obtained. 6 :

[0100]

[0101] In formula (3), represents the attention score matrix corresponding to the hth attention head, and is obtained by formula (4):

[0102]

[0103] In formula (4), is the attention score matrix , where is the attention score between the y-th time step and the z-th time step on the x-th traffic sensor; represents the correlation between the xth traffic sensor corresponding to the hth attention head at the yth time step and the zth time step, and is obtained by formula (5):

[0104]

[0105] In formula (5), yes where represents the vector of the x-th traffic sensor and the y-th time step, yes where represents the vector of the x-th traffic sensor and the z-th time step;

[0106] Step 7.2: Figure 1 The tensor Z output by the mid-time attention layer 6 Input to Figure 1 The tensor is processed in the third feedforward neural network

[0107] Step 7.3: Z 7 and E (T′) After the connection, we get the tensor and enter into Figure 1 In the spatial attention layer, after being processed by the fully connected layer with ReLU as the activation function, the query, key, and value tensors corresponding to the hth attention head are obtained respectively. Thus, the feature representation of the spatial dimension of the data is enhanced using formula (6) to obtain the tensor Z output by the temporal attention layer: 8 :

[0108]

[0109] In formula (5), represents the attention score matrix corresponding to the h-th attention head, and is obtained by formula (7):

[0110]

[0111] In formula (6), is the attention score matrix is the attention score between the βth traffic sensor and the γth traffic sensor at the ath time step, represents the correlation between the βth traffic sensor and the γth traffic sensor at the ath time step corresponding to the hth attention head, and is obtained by formula (8);

[0112]

[0113] In formula (8), yes where represents the vector of the ath time step and the βth traffic sensor, yes where represents the vector of the ath time step and the γth traffic sensor;

[0114] Step 7.4: Figure 1 The output tensor Z of the spatial attention layer 8 Input to Figure 1 The fourth feedforward neural network is further processed and the tensor output by the fourth feedforward neural network is obtained

[0115] Step 8. Tensor Z 9 go through Figure 1 After the transformation of the output linear layer, the prediction result of the multimodal input data X is obtained

[0116] Step 9: Network training, obtaining the trained model through continuous iteration;

[0117] Step 9.1: Use formula (7) to construct the loss function Sum the absolute errors at each time step:

[0118]

[0119] In formula (7), is the prediction result of the nth future time step, Y n The label value of the nth future time step; Θ is all parameters of the traffic prediction model based on multimodal data fusion;

[0120] Step 8.2: Use back propagation and gradient descent method to train the traffic prediction model based on multimodal data fusion and calculate the loss value. When the number of iterations reaches the threshold ξ or the loss value does not decrease for a certain number of consecutive rounds, stop training to obtain the optimal parameters Θ of the model and the trained model.

[0121] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute a traffic prediction method, and the processor is configured to execute the program stored in the memory.

[0122] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium, and the steps of the traffic prediction method are performed when the computer program is executed by a processor.

Claims

1. A traffic prediction method based on multimodal data fusion, characterized in that: The steps include: Step 1: Construct multimodal input data X; Step 1.1: Build a directed road network graph in, is the set of all traffic sensors in the road network; ε is the set of road sections between each traffic sensor; is the adjacency matrix, the adjacency matrix If the element value in is 1, it means that there is a road section connecting the two traffic sensors. If the element value is 0, it means that there is no road section connecting the two traffic sensors. Step 1.2: Road network diagram The N traffic sensors in the system record the traffic status data of C modes every time step, and after normalizing the traffic status data of each mode, the traffic status data of N traffic sensors in L time steps are obtained. Step 1.3, from X all Select C types of traffic state data of T consecutive historical time steps and use them as multimodal input data Let the sub-input data of the cth mode be expressed as T <L; Step 2: Construct a traffic prediction model based on multimodal data fusion, including: input conversion module, spatiotemporal embedding module, cross-modal attention module, maximum pooling fusion layer, spatiotemporal attention module and output linear layer; The input conversion module includes: an input linear layer and a position embedding layer; The spatiotemporal embedding module includes: a spatial embedding module and a temporal embedding module; The cross-modal attention module includes: a first cross-modal attention layer, a first feed-forward neural network, a second cross-modal attention layer, and a second feed-forward neural network; The spatiotemporal attention module includes: a temporal attention layer, a third feedforward neural network, a spatial attention layer, and a fourth feedforward neural network; Step 3, processing of the input conversion module; Step 3.1: The input linear layer converts the sub-input data X of the cth mode into c Perform the transformation process to obtain the transformation data of the cth mode containing the D-dimensional latent space Step 3.2: The position embedding layer transforms the data Z of the cth modality c 0 Perform position embedding operation to obtain the data of the cth mode after embedding the position Thus, the data of C modes after the embedding position are obtained and connected to obtain the connected data Step 4: Processing of the spatiotemporal embedding module; Step 4.1: The spatial embedding module uses the node2vec method to transform the adjacency matrix Transformed into a spatial embedding matrix Step 4.2, processing of the time embedding module; Step 4.2.1: The time embedding module transforms the traffic state data X into all Convert into a sampling signal in the frequency domain, and analyze the sampling signal in the frequency domain to obtain F time period information; Step 4.2.2: Encode the F periodic information using one-hot encoding to obtain the F relative position vectors of the lth time step and connect them to obtain the periodic embedding vector V corresponding to the lth time step. l ; Step 4.2.3: Connect the period embedding vectors of the selected T consecutive historical time steps and the period embedding vectors corresponding to the subsequent T' consecutive future time steps, and then obtain the time embedding matrix after processing through the fully connected layer. T′ <L; Step 4.3: Add the spatial embedding matrix SE and the temporal embedding matrix TE to obtain the spatiotemporal embedding vector Among them, the spatiotemporal embedding subvector containing historical time step information is expressed as The spatiotemporal embedding subvector containing the information of future time steps is expressed as Step 5: Processing of the cross-modal attention module; Step 5.1: Z 1 With E (T) After connection, we get the tensor And input into the first cross-modal attention layer, after being processed by the fully connected layer with ReLU as the activation function, the query, key, and value tensors corresponding to the hth attention head are obtained respectively. Then use formula (1) to get the tensor output by the first cross-modal attention layer In formula (1), || h∈H It means that H subspaces are concatenated in sequence; d represents the dimension of the subspace of each attention head; and H×d=D; Step 5.4: Convert the tensor Z 2 Input into the first feedforward neural network, and obtain the tensor output by the first feedforward neural network through formula (2): WITH 3 =ReLU(Z 2 W1+b1)W2+b2 (2) In formula (2), W1 and W2 are learnable weight parameters in the first feedforward neural network; b1 and b2 are learnable bias parameters in the first feedforward neural network; Step 5.5: The tensor Z 3 After being processed by the second cross-modal attention layer and the second feedforward neural network, the tensor is obtained. And used as output data of the cross-modal attention module; Step 6: Processing of the maximum pooling fusion layer; According to the order of each mode, take out the tensor Z respectively 4 The tensor of one dimension in the C modalities is spliced ​​to obtain a spliced ​​tensor of one dimension, thereby obtaining a spliced ​​tensor of D dimensions in C modalities and splicing them into the final interleaved spliced ​​tensor, which is then input into the maximum pooling fusion layer for multimodal fusion to obtain the fused data. Step 7: Processing of the spatiotemporal attention module; Step 7.1: Z 5 and E (T′) After the connection, we get the tensor And input into the temporal attention layer, after being processed by the fully connected layer with ReLU as the activation function, the query, key, and value tensors corresponding to the hth attention head are obtained respectively Thus, using formula (3), we can get the tensor Z output by the temporal attention layer: 6 : In formula (3), represents the attention score matrix corresponding to the hth attention head in the temporal attention layer, and is obtained by formula (4): In formula (4), is the attention score matrix , where is the attention score between the y-th time step and the z-th time step on the x-th traffic sensor; represents the correlation between the xth traffic sensor corresponding to the hth attention head at the yth time step and the zth time step, and is obtained by formula (5): In formula (5), yes where represents the vector of the x-th traffic sensor and the y-th time step, yes where represents the vector of the x-th traffic sensor and the z-th time step; Step 7.2: The tensor Z output by the temporal attention layer 6 Input into the third feedforward neural network for processing to obtain a tensor Step 7.3: Z 7 and E (T′) After the connection, we get the tensor And input into the spatial attention layer, after being processed by the fully connected layer with ReLU as the activation function, the query, key, and value tensors corresponding to the hth attention head are obtained respectively Thus, the tensor Z output by the temporal attention layer is obtained using formula (6): 8 : In formula (6), represents the attention score matrix corresponding to the hth attention head in the spatial attention layer, and is obtained by formula (7): In formula (6), is the attention score matrix is the attention score between the βth traffic sensor and the γth traffic sensor at the ath time step, represents the correlation between the βth traffic sensor and the γth traffic sensor at the ath time step corresponding to the hth attention head, and is obtained by formula (8); In formula (8), yes where represents the vector of the ath time step and the βth traffic sensor, yes where represents the vector of the ath time step and the γth traffic sensor; Step 7.4: The output tensor Z of the spatial attention layer 8 Input to the fourth feedforward neural network for further processing, and get the tensor output by the feedforward neural network Step 8: The tensor Z 9 After the transformation of the output linear layer, the prediction result of the multimodal input data X is obtained Step 9: Network training; Step 9.1: Use formula (7) to construct the loss function In formula (7), is the prediction result of the nth future time step, Y n The label value of the nth future time step; Θ is all parameters of the traffic prediction model based on multimodal data fusion; T' is the total number of prediction steps in the future time; Step 8.2: Use back propagation and gradient descent method to train the traffic prediction model based on multimodal data fusion, and calculate the loss value. When the number of iterations reaches the threshold ξ or the loss value does not decrease for a certain number of consecutive rounds, stop training to obtain the optimal model after training and its optimal parameters Θ.

2. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports a processor to execute the method according to claim 1, and the processor is configured to execute the program stored in the memory.

3. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program executes the steps of the method according to claim 1 when executed by a processor.

Citation Information

Patent Citations

  • Semi-supervised symbol network embedding method and system based on improved graph convolutional network

    CN111401514A

  • Course field multi-modal document classification method based on cross-modal attention convolutional neural network

    CN111985369A