Real-time traffic jam prediction method and system based on multi-integration deep learning

Through multi-integrated deep learning methods, combined with real-time monitoring of video data and Multi-AM-OL-LSTM network, the problem of low accuracy of traffic congestion prediction is solved, more accurate traffic flow prediction is achieved, and path planning errors are reduced.

CN120356334APending Publication Date: 2025-07-22SICHUAN POLICE COLLEGE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510634037.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing traffic congestion prediction methods have low accuracy and are easily affected by the original data, resulting in large errors in traffic management.

Method used

The multi-integrated deep learning method is adopted, combined with real-time monitoring of video data, traffic flow, vehicle speed variance, front distance variance, front time distance variance and conflict rate are obtained, and the Multi-AM-OL-LSTM network is used for prediction, and time series features are extracted through the multi-head attention layer and LSTM layer to achieve real-time traffic flow prediction.

Benefits of technology

It improves the accuracy of traffic congestion prediction, can capture short-term and long-term dependencies effectively in real time, dynamically adjust new data, and reduces errors in driving path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356334A_ABST
    Figure CN120356334A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time traffic jam prediction method and system based on multi-integration deep learning, relates to the technical field of traffic safety, and aims to solve the problem of low accuracy of existing traffic jam prediction, comprehensively considers a vehicle speed variance, an interval variance and a time interval variance, and avoids the problem that a clustering result is easily influenced by original data. And the prediction accuracy of the traffic jam is improved. The network model combines multi-head attention, online learning and an LSTM network, can effectively predict the traffic flow in real time by capturing short-term and long-term dependency relationships, dynamically adjusts new data through online learning, and improves the accuracy of congestion prediction by using an integration method, thereby avoiding errors in specific driving path planning, and improving the traffic congestion prediction accuracy. And the accuracy of driving path planning is not influenced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic safety, and specifically provides a real-time traffic congestion prediction method and system based on multi-integrated deep learning. Background Art

[0002] Traffic congestion hinders travel in our daily life, leading to various problems such as time waste, economic losses, pollution, and energy consumption. Achieving real-time traffic congestion prediction can save a large amount of response time for route guidance, control, and law enforcement, thus realizing proactive traffic management.

[0003] Existing traffic congestion prediction methods comprehensively consider three parameters: traffic flow, average speed, and density. The three-dimensional parameters are converted into a one-dimensional time series through a clustering method, and a BP neural network is used to predict the short-term traffic state. However, the clustering results are easily affected by the original data. Therefore, it leads to the problem of low accuracy in traffic congestion prediction. Summary of the Invention

[0004] The purpose of the present invention is to provide a real-time traffic congestion prediction method and system based on multi-integrated deep learning for the problem of low accuracy in existing traffic congestion prediction.

[0005] The technical solution adopted by the present invention to solve the above technical problems is as follows:

[0006] A real-time traffic congestion prediction method based on multi-integrated deep learning includes the following steps:

[0007] Based on real-time monitoring videos, traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate are obtained, and traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate are input into a prediction network to obtain the total delay time of the output traffic flow. Finally, traffic congestion prediction is performed according to the total delay time of the traffic flow.

[0008] The prediction network is Multi-AM-OL-LSTM. The Multi-AM-OL-LSTM includes an input layer, a multi-head attention layer, an LSTM layer, and an output layer. The input layer receives input data and converts the input data into a three-dimensional tensor. The three-dimensional tensor is the number of samples × time steps × number of features. The three-dimensional tensor is input into the LSTM layer, and the LSTM layer extracts the time series features h” i of the three-dimensional tensor. The time series features h” i are input into the multi-head attention layer. The multi-head attention layer calculates its attention weights according to the correlation of each time step in the sequence, and then obtains weighted features. The output layer maps the weighted features generated by the attention layer into the final prediction value.

[0009] Further, the specific steps for obtaining traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate are as follows:

[0010] Step 1: Obtain historical surveillance video data, and use the Video Capture function in the OpenCV library to read and process the surveillance video data frame by frame to obtain consecutive frame images, and perform grayscale processing on each frame image in the consecutive frame images;

[0011] Step 2: Use the YOLO-v8 network to identify each vehicle in each frame image, and obtain the center point pixel coordinates of each monitoring bounding box. Then, use the Lucas-Kanade method to track each center point pixel to obtain vehicle trajectory information;

[0012] Step 3: Obtain the vehicle ID, head direction, vehicle coordinates, and vehicle passing time difference based on the vehicle trajectory information, obtain the traffic flow based on the vehicle ID, obtain the vehicle speed based on the vehicle passing time difference, obtain the vehicle speed variance based on the vehicle speeds of all vehicles, and obtain the headway spacing between adjacent vehicles based on the vehicle coordinates and head directions of adjacent vehicles;

[0013] Step 4: Repeat Step 3 to obtain all headway spacings, and use all headway spacings to obtain the headway spacing variance;

[0014] Step 5: Obtain the headway spacing between adjacent vehicles and the speed of the rear vehicle in the adjacent vehicles, and divide the headway spacing between adjacent vehicles by the speed of the rear vehicle to obtain the headway time;

[0015] Step 6: Obtain the headway time variance based on all headway times;

[0016] Step 7: Divide the headway spacing between adjacent vehicles by the speed difference between the adjacent vehicles to obtain the TTC. Determine those with TTC greater than the threshold as collisions, otherwise as non-collisions. Finally, divide the number of collisions by the traffic volume to obtain the conflict rate.

[0017] Further, the total delay time of the traffic flow is obtained through the following steps:

[0018] Step 1: Obtain the length of the monitored section in the monitoring area and the expected speed of this section, that is, the speed limit value of this section, and obtain the expected passing time based on the length of the monitored section and the expected speed of this section;

[0019] Step 2: Obtain the time when the vehicle enters and exits the monitored section according to the surveillance video to obtain the actual passing time;

[0020] Step 3: Subtract the actual passing time from the expected passing time to obtain the total delay time of the traffic flow.

[0021] Furthermore, the time series feature h” i is expressed as:

[0022] h” i = o t *tanh(C t )

[0023]

[0024]

[0025] where W f is the weight of the forget gate, h t-1 is the input of the previous neuron, x t is the data input, b f is the bias value of the forget gate, f t is the output of the forget gate, W i is the weight of the input gate, h t-1 is the input of the previous neuron, x t is the data input, b t is the bias value of the input gate, i t is the output of the input gate, W C is the weight of the candidate value, h t-1 is the input of the previous neuron, x t is the data input, b C is the bias value of the candidate value, is the output of the candidate value, C t-1 is the output of the previous candidate value, C t is the update value of the candidate value, W o is the weight of the output gate, h t-1 is the input of the previous neuron, x t is the data input, b o is the bias value of the output gate, o t is the output of the output gate, h t is the value passed to the next neuron.

[0026] Furthermore, the multi-head attention layer specifically performs the following steps:

[0027] Step 1: Perform a linear transformation on the input time series feature to generate a query vector q i , a key vector k i and a value vector v i , expressed as:

[0028]

[0029] where W q , W k, W v is a learnable weight matrix, and x i is the time step in the time series features;

[0030] Step 2: For each query vector q i , each attention head calculates its similarity with all key vectors k j through the dot product function to obtain the attention scores, and the attention scores are expressed as:

[0031] score h (q i , k j ) = q T · k j

[0032] Step 3: Use the Softmax function to normalize the attention scores of each attention head into a probability distribution, which is expressed as:

[0033]

[0034] where α ij represents the attention weight of the i-th query to the j-th key, n represents all time steps, and l represents the l-th time step;

[0035] Step 4: Use the attention weight α ij to perform a weighted sum on the value vector v j to obtain the output vector y i of each attention head, and y i is expressed as;

[0036]

[0037] Step 5: Concatenate the outputs of all attention heads and generate the final output through a linear transformation, which is expressed as:

[0038]

[0039] where W o is a learnable weight matrix.

[0040] A real-time traffic congestion prediction system based on multi-integration deep learning, and the system specifically executes the following steps:

[0041] Based on real-time monitoring videos, obtain traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate, and input the traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate into the prediction network to obtain the total delay time of the output traffic flow, and finally perform traffic congestion prediction according to the total delay time of the traffic flow;

[0042] The prediction network is Multi-AM-OL-LSTM. The Multi-AM-OL-LSTM includes an input layer, a multi-head attention layer, an LSTM layer, and an output layer. The input layer receives input data and converts the input data into a three-dimensional tensor. The three-dimensional tensor is the number of samples × the number of time steps × the number of features. The three-dimensional tensor is input into the LSTM layer, and the LSTM layer extracts the time series feature h” of the three-dimensional tensor. i The time series feature h” i is input into the multi-head attention layer. The multi-head attention layer calculates its attention weights according to the relevance of each time step in the sequence, and then obtains the weighted feature. The output layer maps the weighted feature generated by the attention layer to the final predicted value.

[0043] Further, the specific steps for obtaining traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate are as follows:

[0044] Step 1: Obtain historical surveillance video data, and use the Video Capture function in the OpenCV library to read and process the surveillance video data frame by frame to obtain consecutive frame images, and perform grayscale processing on each frame image in the consecutive frame images.

[0045] Step 2: Use the YOLO-v8 network to identify each vehicle in each frame image, and obtain the center point pixel coordinates of each monitoring bounding box. Then, use the Lucas-Kanade method to track each center point pixel to obtain vehicle trajectory information.

[0046] Step 3: Obtain the vehicle ID, vehicle head direction, vehicle coordinates, and vehicle passing time difference according to the vehicle trajectory information, obtain the traffic flow according to the vehicle ID, obtain the vehicle speed according to the vehicle passing time difference, obtain the vehicle speed variance according to the vehicle speeds of all vehicles, and obtain the headway spacing between adjacent vehicles according to the vehicle coordinates and vehicle head directions of adjacent vehicles.

[0047] Step 4: Repeat Step 3 to obtain all headway spacings, and use all headway spacings to obtain the headway spacing variance.

[0048] Step 5: Obtain the headway spacing between adjacent vehicles and the speed of the rear vehicle in the adjacent vehicles, and divide the headway spacing between adjacent vehicles by the speed of the rear vehicle to obtain the headway time.

[0049] Step 6: Obtain the headway time variance according to all headway times.

[0050] Step 7: Divide the headway spacing between adjacent vehicles by the speed difference between the adjacent vehicles to obtain the TTC. Those with a TTC greater than the threshold are judged as collisions, otherwise as non-collisions. Finally, divide the number of collisions by the traffic volume to obtain the conflict rate.

[0051] Further, the total delay time of the traffic flow is obtained through the following steps:

[0052] Step 1: Obtain the length of the monitored section within the monitoring area and the desired speed of this section, i.e., the speed limit value of this section, and obtain the desired travel time based on the length of the monitored section and the desired speed of this section;

[0053] Step 2: According to the monitoring video, obtain the time when the vehicle enters and exits the monitored section to obtain the actual travel time;

[0054] Step 3: Subtract the desired travel time from the actual travel time to obtain the total delay time of the traffic flow.

[0055] Further, the time series feature h” i is expressed as:

[0056] h” i = o t *tanh(C t )

[0057]

[0058]

[0059] where W f is the weight of the forget gate, h t-1 is the input of the previous neuron, x t is the data input, b f is the bias value of the forget gate, f t is the output of the forget gate, W i is the weight of the input gate, h t-1 is the input of the previous neuron, x t is the data input, b t is the bias value of the input gate, i t is the output of the input gate, W C is the weight of the candidate value, h t-1 is the input of the previous neuron, x t is the data input, b C is the bias value of the candidate value, is the output of the candidate value, C t-1 is the output of the previous candidate value, C t is the update value of the candidate value, W o is the weight of the output gate, h t-1 is the input of the previous neuron, x t is the data input, b o is the bias value of the output gate, o tis the output of the output gate, h t is the value passed to the next neuron.

[0060] Furthermore, the multi-head attention layer specifically performs the following steps:

[0061] Step 1: Perform a linear transformation on the input time series features to generate a query vector q i , a key vector k i and a value vector v i , expressed as:

[0062]

[0063] where W q , W k , W v are learnable weight matrices, and x i is the time step in the time series features;

[0064] Step 2: For each query vector q i , each attention head calculates its similarity with all key vectors k j through a dot product function to obtain attention scores, which are expressed as:

[0065] score h (q i , k j ) = q T ·k j

[0066] Step 3: Use the Softmax function to normalize the attention scores of each attention head into a probability distribution, expressed as:

[0067]

[0068] where α ij represents the attention weight of the i-th query to the j-th key, n represents all time steps, and l represents the l-th time step;

[0069] Step 4: Use the attention weight α ij to perform a weighted sum on the value vector v j to obtain the output vector y i of each attention head, and y i is expressed as;

[0070]

[0071] Step 5: Concatenate the outputs of all attention heads and generate the final output through a linear transformation, expressed as:

[0072]

[0073] Among them, W o is a learnable weight matrix.

[0074] The beneficial effects of the present invention are as follows:

[0075] This application comprehensively considers the vehicle speed variance, the headway spacing variance, and the headway time variance, avoiding the problem that the clustering results are easily affected by the original data. Furthermore, it improves the prediction accuracy of traffic congestion.

[0076] The network model of this application combines multi-head attention, online learning, and the LSTM network, which can effectively predict traffic flow in real time by capturing short-term and long-term dependencies, dynamically adjust new data through online learning, and improve the accuracy of congestion prediction using the ensemble method. Furthermore, it avoids the problem of errors in the specific driving route planning and does not affect the accuracy rate of the driving route planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 is the online learning flowchart of this application;

[0078] Figure 2 is the schematic diagram of the model structure of this application;

[0079] Figure 3 is the schematic diagram of the observation point location;

[0080] Figure 4 is the model prediction result Figure 1 ;

[0081] Figure 5 is the model prediction result Figure 2 ;

[0082] Figure 6 is the model prediction result Figure 3 ;

[0083] Figure 7 is the model prediction result Figure 4 ;

[0084] Figure 8 is the model prediction result Figure 5 ;

[0085] Figure 9 is the model prediction result Figure 6 . DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] It should be specifically noted that, without conflict, the various embodiments disclosed in this application can be combined with each other.

[0087] Specific Embodiment 1: A real-time traffic congestion prediction method based on multi-integrated deep learning described in this embodiment includes the following steps:

[0088] Based on real-time monitoring videos, obtain traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate, and input the traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate into the prediction network to obtain the total delay time of the output traffic flow. Finally, perform traffic congestion prediction based on the total delay time of the traffic flow;

[0089] The prediction network is Multi-AM-OL-LSTM. The Multi-AM-OL-LSTM includes an input layer, a multi-head attention layer, an LSTM layer, and an output layer. The input layer receives input data and converts the input data into a three-dimensional tensor. The three-dimensional tensor is the number of samples × time steps × number of features. The three-dimensional tensor is input into the LSTM layer, and the LSTM layer extracts the time series features h” i , the time series features h” i are input into the multi-head attention layer. The multi-head attention layer calculates its attention weights according to the correlation of each time step in the sequence, and then obtains the weighted features. The output layer maps the weighted features generated by the attention layer to the final prediction value.

[0090] Specific Embodiment 2: This embodiment is a further description of Specific Embodiment 1. The difference between this embodiment and Specific Embodiment 1 is that the specific steps for obtaining the traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate are as follows:

[0091] Step 1: Obtain historical monitoring video data, and use the Video Capture function in the OpenCV library to read and process the monitoring video data frame by frame to obtain consecutive frame images, and perform grayscale processing on each frame image in the consecutive frame images;

[0092] Step 2: Use the YOLO-v8 network to identify each vehicle in each frame image, and obtain the center point pixel coordinates of each monitoring bounding box. Then, use the Lucas-Kanade method to track each center point pixel to obtain vehicle trajectory information;

[0093] Step 3: Obtain the vehicle ID, head direction, vehicle coordinates, and vehicle passing time difference according to the vehicle trajectory information, obtain the traffic flow according to the vehicle ID, obtain the vehicle speed according to the vehicle passing time difference, obtain the vehicle speed variance according to the vehicle speeds of all vehicles, and obtain the headway spacing between adjacent vehicles according to the vehicle coordinates and head directions of adjacent vehicles;

[0094] Step 4: Repeat Step 3 to obtain all headway distances, and use all headway distances to obtain the variance of headway distances;

[0095] Step 5: Obtain the headway distance between adjacent vehicles and the speed of the rear vehicle in the adjacent vehicles, and divide the headway distance between adjacent vehicles by the speed of the rear vehicle to obtain the headway time;

[0096] Step 6: Based on all headway times, further obtain the variance of headway times;

[0097] Step 7: Divide the headway distance between adjacent vehicles by the speed difference of the adjacent vehicles to obtain the TTC, and judge those with TTC greater than the threshold as collisions, otherwise as non-collisions. Finally, divide the number of collisions by the traffic volume to obtain the conflict rate.

[0098] Specific Embodiment 3: This embodiment is a further description of Specific Embodiment 2. The difference between this embodiment and Specific Embodiment 2 is that the total delay time of the traffic flow is obtained through the following steps:

[0099] Step 1: Obtain the length of the monitored section in the monitoring area and the expected speed of this section, that is, the speed limit value of this section, and obtain the expected travel time according to the length of the monitored section and the expected speed of this section;

[0100] Step 2: According to the monitoring video, obtain the time when the vehicle enters and exits the monitored section to obtain the actual travel time;

[0101] Step 3: Subtract the expected travel time from the actual travel time to obtain the total delay time of the traffic flow.

[0102] Specific Embodiment 4: This embodiment is a further description of Specific Embodiment 3. The difference between this embodiment and Specific Embodiment 3 is that the time series feature h” i is expressed as:

[0103] h” i =o t *tanh(C t )

[0104]

[0105]

[0106] where, W f is the weight of the forgetting gate, h t-1 is the input of the previous neuron, x t is the data input, b f is the bias value of the forgetting gate, f t is the output of the forgetting gate, W iis the weight of the input gate, h t-1 is the input of the previous neuron, x t is the data input, b t is the bias value of the input gate, i t is the output of the input gate, W C is the weight of the candidate value, h t-1 is the input of the previous neuron, x t is the data input, b C is the bias value of the candidate value, is the output of the candidate value, C t-1 is the output of the previous candidate value, C t is the updated value of the candidate value, W o is the weight of the output gate, h t-1 is the input of the previous neuron, x t is the data input, b o is the bias value of the output gate, o t is the output of the output gate, h t is the value passed to the next neuron.

[0107] Specific Embodiment 5: This embodiment is a further description of Specific Embodiment 4. The difference between this embodiment and Specific Embodiment 4 is that the multi-head attention layer specifically performs the following steps:

[0108] Step 1: Perform a linear transformation on the input time series features to generate a query vector q i 、key vector k i and value vector v i , expressed as:

[0109]

[0110] Where, W q , W k , W v are learnable weight matrices, x i is the time step in the time series features;

[0111] Step 2: For each query vector q i , each attention head calculates its similarity with all key vectors k j through the dot product function to obtain attention scores, and the attention scores are expressed as:

[0112] score h (q i ,k j ) = q T ·k j

[0113] Step 3: Use the Softmax function to normalize the attention scores of each attention head into a probability distribution, expressed as:

[0114]

[0115] where α ij represents the attention weight of the i-th query to the j-th key, n represents all time steps, and l represents the l-th time step;

[0116] Step 4: Use the attention weight α ij to perform a weighted sum on the value vector v j to obtain the output vector y i of each attention head, y i is expressed as;

[0117]

[0118] Step 5: Concatenate the outputs of all attention heads and generate the final output through a linear transformation, expressed as:

[0119]

[0120] where W o is a learnable weight matrix.

[0121] Online learning

[0122] As Figure 1 shown. The core idea of online learning is to minimize the prediction error L by iteratively updating the model parameters Θ. Given a continuously arriving data stream The goals of online learning are as follows:

[0123] 1) Update: At time step t, the model f t uses the sample (x t , y t ) for update

[0124] 2) Prediction: After the model is updated, it should be able to make more accurate predictions for future samples x t+1 .

[0125] LSTM

[0126] The core of LSTM is a neural network unit (cell) that can store and control memory. At each time step, the LSTM unit dynamically adjusts the selective memory and forgetting of information through three gating mechanisms (input gate, forget gate, and output gate), so as to effectively model short-term dependencies and long-term dependencies.

[0127] 1) Forget gate: The forget gate is used to control the memory C at the previous momentt-1 Which information in it needs to be forgotten. The value of the forget gate is:

[0128] f t = σ(W f · [h t-1 , x t + b f )

[0129] Where x t is the current input, h t-1 is the hidden state at the previous moment, W f and b f are the weight matrix and bias respectively, and σ is the Sigmoid activation function. f t ∈ [0, 1] represents the forgetting ratio of the memory. When it is close to 1, more information is retained, and when it is close to 0, more information is forgotten.

[0130] 2) Input gate: The input gate determines the degree of update of the new information to the memory cell at the current time step. It includes two steps: information filtering and candidate memory generation:

[0131] i t = σ(W i · [h t-1 , x t + b i )

[0132]

[0133] Where i t is the weight of the input gate, determining the degree of introduction of new information;

[0134] 3) Output gate: The output gate controls which information in the memory cell contributes to the output state h t at the current moment. Its expression and output state are as follows:

[0135] o t = σ(W o · [h t-1 , x t + b o )

[0136] h t = o t · tanh(C t ) #(1)

[0137] Where o t is the weight of the output gate, determining which information flows out of the memory cell; h t is the hidden state at the current moment; C t is the memory cell state at the current moment.

[0138] 4) Memory cell update: The LSTM dynamically updates the state of the memory cell through the forget gate and the input gate:

[0139]

[0140] This update mechanism enables the LSTM to retain long-term information while introducing new short-term information as needed.

[0141] Multi-Head Attention Mechanism

[0142] The multi-head attention mechanism is an extension of the basic attention mechanism, aiming to enhance the model's ability to simultaneously focus on different parts of the input sequence, thereby improving its performance in complex tasks. Different from the standard attention mechanism, which calculates a single attention weight for each element, the multi-head attention mechanism performs multiple attention calculations in parallel, with each calculation using a different learned weight matrix. This enables the model to capture various relationships in the data by simultaneously focusing on different aspects of the input. The core steps of multi-head attention are as follows:

[0143] 1) Input representation: The input sequence X = (x1, x2,..., x n ) generates multiple sets of query (Q), key (K), and value (V) vectors through linear transformation, with each attention head corresponding to a set of vectors:

[0144]

[0145] where, is the learnable weight matrix for each attention head h, and H is the total number of attention heads.

[0146] 2) Calculate attention scores: For each query vector q i , calculate its similarity with all key vectors k j , and calculate it for each attention head through the dot product function:

[0147] score h (q i , k j ) = q T ·k j

[0148] 3) Normalization: Normalize the attention scores of each attention head into a probability distribution through the Softmax function:

[0149]

[0150] 4) Output calculation: Use the normalized attention weights of each attention head to calculate the value vector v jThe weighted sum is obtained to get the output of each head:

[0151] 5) Concatenation and linear transformation: Concatenate the outputs of all attention heads and generate the final output through a linear transformation:

[0152]

[0153] where, W o is a learnable weight matrix used to combine the outputs of all attention heads.

[0154] Multi-AM-OL-LSTM

[0155] The structure of the Multi-AM-OL-LSTM model studied in this application is as Figure 2 shown. The model mainly consists of an input layer, a multi-head attention layer, an LSTM layer, an output layer, and an online update module.

[0156] In the input layer, the model receives the normalized and preprocessed time series traffic flow data, which contains four features: flow rate, vehicle speed variance, vehicle distance variance, and time interval variance. The input data is organized as a three-dimensional tensor with a shape of (number of samples × time steps × number of features), so as to capture the dynamic spatio-temporal characteristics of the traffic flow and provide high-quality input for the subsequent layers.

[0157] The multi-head attention layer aims to enhance the model's ability to simultaneously focus on different parts of the time series data. This layer receives the time series features output by the LSTM layer and applies multiple attention heads in parallel. Each attention head calculates its attention weights according to the relevance of each time step in the sequence, enabling the model to capture different aspects of the input data. Subsequently, the attention mechanism is applied to the weighted features, enhancing the representation of traffic flow dynamics.

[0158] The LSTM layer can extract the dynamic features in the time series data layer by layer. It can learn long-term dependencies and sequence patterns from the traffic flow data, improving the model's ability to predict future traffic conditions.

[0159] The online learning module is integrated into the model training process and dynamically updated under the sliding window mechanism of the time series data. Each time during training, the model uses the latest 5 samples for training and predicts the value of the next time step. This enables the model to continuously adapt to new data without the need for a large amount of historical data.

[0160] The output layer consists of a fully connected (Dense) layer, which maps the weighted features generated by the attention mechanism to the final predicted value. The model outputs the total delay time of the traffic flow.

[0161] In summary, the Multi-AM-OL-LSTM model combines multi-head attention, online learning, and the LSTM network. It can effectively predict traffic flow in real time by capturing short-term and long-term dependencies, dynamically adjust new data through online learning, and improve prediction accuracy using an ensemble method.

[0162] The experimental data studied in this application is sourced from a monitoring station located on the Chengyu Ring Expressway (Chongqing section). This section of the expressway connects Jiangjin District, Chongqing, with Hejiang County, Sichuan Province, and is an important economic corridor and transportation hub. The specific location of the monitoring station is as Figure 3 shown.

[0163] This study is based on the YOLOv8 algorithm and focuses on processing surveillance video data from the Jiangjin to Hejiang section of the Chengyu Ring Expressway (Chongqing section). Specifically, the algorithm is used to extract six different sets of traffic flow data from the videos captured at a predetermined time interval. The data collection time period is set from 6:00 to 18:00 every day.

[0164] The main objective of this application study is to solve the problem of short-term traffic flow prediction. Once a series of traffic flow data is obtained, it is first aggregated into 5-minute data streams. In these data streams, the input features are defined as traffic volume, vehicle speed variance, vehicle distance variance, and time interval variance, while the output feature is the average delay per kilometer.

[0165] Parameter Model Settings

[0166] During the training process, to optimize the model's performance on the training set and test set, it is necessary to appropriately configure the model parameters. For the proposed Multi–AM–OL–LSTM model, three layers of LSTM are adopted, with each layer containing 128 units. In addition, the model introduces a multi-head attention mechanism, with each attention head having a dimension of 64 and a total of 8 attention heads being used. After that, a Flatten operation is applied to convert the data into a one-dimensional vector.

[0167] The model uses the mean squared error (MSE) as the loss function and the Adam optimizer for parameter updates. Through parameter tuning during the training process, the batch size is finally determined to be 5, the number of training epochs is 100, and the learning rate is 0.001. These settings result in the model having a low loss value and good data fitting effect. The expression for the mean squared error (MSE) is as follows:

[0168]

[0169] where y i is the true value, is the predicted value, and m is the number of samples.

[0170] Selection of Evaluation Metrics

[0171] In this application, the Mean Relative Error (MRE), Root Mean Square Error (RMSE), and Theil's U coefficient (U) are used as evaluation indicators to determine the prediction performance and quality of the model. The expressions are respectively

[0172]

[0173]

[0174]

[0175] where: m is the total number of data sets; is the actual value; is the predicted value; is the mean of the true values.

[0176] Result Analysis

[0177] The predicted values generated by the Multi–AM–OL–LSTM model were compared with the actual observed values, and the results are as Figures 4 - 9 shown. In addition, to verify the prediction effect of the proposed model, a horizontal comparison was also conducted, and three other models were used: OL-RNN, OL-GRU, and OL-LSTM. The results of this comparative analysis are summarized in Table 1:

[0178] Table 1. Comparison of results of different models.

[0179]

[0180] The proposed Multi–AM–OL–LSTM model in this application shows significant advantages in multiple evaluation indicators. As shown in Table 1, the average performance analysis of six scenario data sets shows that the Multi–AM–OL–LSTM model has reached the optimal values in indicators such as root mean square error (RMSE), mean relative error ratio (MRE), and Theil's U coefficient. Specifically, the mean relative error (MRE) - a key indicator to measure the relative size of prediction errors - reaches the lowest value in the proposed model which is reduced by 0.0188, 0.0289, and 0.0410 compared with OL-RNN, OL-GRU, and OL-LSTM respectively.

[0181] The root mean square error (RMSE), which measures the absolute difference between the predicted value and the actual value and effectively captures the impact of large errors on the overall performance, is minimized in the proposed model and is reduced by 11.127, 12.178, and 15.316 compared with OL-RNN, OL-GRU, and OL-LSTM respectively.

[0182] In addition, the Theil's U coefficient, which is used to evaluate the uniformity of the error distribution (the lower the value, the more uniform the distribution), shows that the Multi–AM–OL–LSTM model outperforms other models in this metric. and right behind it, while the OL-LSTM performs the worst.

[0183] These experimental results strongly validate the effectiveness of the proposed model in traffic congestion prediction tasks. By introducing the multi-head attention mechanism (MHA), the model significantly enhances the adaptive allocation of spatio-temporal features and can dynamically focus on the key spatio-temporal nodes in the traffic flow. This ability enables the model to capture global temporal dependencies and spatial correlations more precisely.

[0184] Compared with traditional RNN, GRU, and basic LSTM models, the Multi–AM–OL–LSTM model utilizes the gating mechanism of LSTM and the multi-dimensional feature fusion ability of MHA to demonstrate stronger prediction robustness in complex scenarios.

[0185] Notably, the improvements in the RMSE and Theil's U coefficient of the model are particularly significant, reducing by 50.2% and 24.0% respectively compared to the traditional LSTM model. These results highlight the sensitivity of the model to extreme events (such as traffic accidents or sudden congestion) and its ability to stabilize the distribution of prediction errors. In addition, the parallel multi-head structure of MHA promotes the collaborative optimization of multi-scale spatio-temporal features, further enhancing the adaptability and generalization ability of the model in dynamic traffic systems. This integrated method provides an effective prediction solution that can balance accuracy and interpretability in congestion prediction.

[0186] This application proposes a new online learning method based on the integration of multi-head attention mechanism and LSTM for real-time prediction of traffic congestion. This method takes traffic flow, vehicle speed variance, vehicle distance variance, and time interval variance as inputs, fully considering safety factors. In addition, the average delay time per kilometer is used as the output to achieve the prediction of traffic congestion. Then, the proposed method is tested using six datasets (each dataset lasts for 12 hours) from a highway. The results show that the proposed method outperforms the baseline methods (OL-RNN, OL-GRU, and OL-LSTM) in terms of prediction performance.

[0187] However, there are still some limitations in this study: 1) Limited data diversity: The dataset used for evaluation comes from a single highway section and the collection time is within a 12-hour period. The limitations in this geographical and time scope may not fully capture the variability of traffic patterns under different regions, road types, or seasonal conditions. 2) Generalizability of the model: Given that this study focuses on a specific highway, there is still uncertainty about the performance of the model when applied to other traffic networks or atypical congestion scenarios. To verify the robustness and generalizability of the model, further validation is required on a more diverse dataset.

[0188] It should be noted that the specific implementation manners are only explanations and illustrations of the technical solutions of the present invention, and the scope of the right protection cannot be limited thereby. Those that are only partial changes made according to the claims and specifications of the present invention should still fall within the protection scope of the present invention.

Claims

1. A real-time traffic congestion prediction method based on multi-integrated deep learning, characterized in that It includes the following steps: Based on the real-time monitoring video, obtain the traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate, and input the traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate into the prediction network to obtain the total delay time of the output traffic flow. Finally, conduct traffic congestion prediction based on the total delay time of the traffic flow; The prediction network is Multi-AM-OL-LSTM, and the Multi-AM-OL-LSTM includes an input layer, a multi-head attention layer, an LSTM layer, and an output layer. The input layer receives input data and converts the input data into a three-dimensional tensor. The three-dimensional tensor is the number of samples × the number of time steps × the number of features. The three-dimensional tensor is input into the LSTM layer, and the LSTM layer extracts the time series feature h” of the three-dimensional tensor. i , the time series feature h” i is input into the multi-head attention layer. The multi-head attention layer calculates its attention weights according to the relevance of each time step in the sequence, and then obtains the weighted feature. The output layer maps the weighted feature generated by the attention layer into the final predicted value.

2. The real-time traffic congestion prediction method based on multi-integrated deep learning according to claim 1, characterized in that The specific steps for obtaining the traffic flow, vehicle speed variance, headway spacing variance, headway time variance, and conflict rate are as follows: Step 1: Obtain historical monitoring video data, and use the Video Capture function in the OpenCV library to read and process the monitoring video data frame by frame to obtain consecutive frame images, and perform grayscale processing on each frame image in the consecutive frame images; Step 2: Use the YOLO-v8 network to identify each vehicle in each frame image, and obtain the center point pixel coordinates of each monitoring bounding box. Then, use the Lucas-Kanade method to track each center point pixel to obtain vehicle trajectory information; Step 3: Obtain the vehicle ID, head direction, vehicle coordinates, and vehicle passing time difference based on the vehicle trajectory information, obtain the traffic flow based on the vehicle ID, obtain the vehicle speed based on the vehicle passing time difference, obtain the vehicle speed variance based on the vehicle speeds of all vehicles, and obtain the headway spacing between adjacent vehicles based on the vehicle coordinates and head directions of adjacent vehicles; Step 4: Repeat Step 3 to obtain all headway spacings, and use all headway spacings to obtain the headway spacing variance; Step 5: Obtain the headway spacing between adjacent vehicles and the speed of the rear vehicle in the adjacent vehicles, and divide the headway spacing between adjacent vehicles by the speed of the rear vehicle to obtain the headway time; Step 6: Based on all headway times, further obtain the headway time variance; Step 7: Divide the headway spacing between adjacent vehicles by the speed difference between the adjacent vehicles to obtain the TTC. Determine those with TTC greater than the threshold as collisions, otherwise as non-collisions. Finally, divide the number of collisions by the traffic volume to obtain the conflict rate.

3. A real-time traffic congestion prediction method based on multi-integrated deep learning according to claim 2, characterized in that The total delay time of the traffic flow is obtained through the following steps: Step 1: Obtain the length of the monitored section in the monitoring area and the expected speed of this section, that is, the speed limit value of this section, and obtain the expected passing time based on the length of the monitored section and the expected speed of this section; Step 2: Based on the monitoring video, obtain the time when the vehicle enters and exits the monitored section to obtain the actual passing time; Step 3: Subtract the actual passing time from the expected passing time to obtain the total delay time of the traffic flow.

4. A real-time traffic congestion prediction method based on multi-integrated deep learning according to claim 1, characterized in that The time series feature h” i is expressed as: h” i = o t *tanh(C t ) Among them, W f is the weight of the forget gate, h t-1 is the input of the previous neuron, x t is the data input, b f is the bias value of the forget gate, f t is the output of the forget gate, W i is the weight of the input gate, h t-1 is the input of the previous neuron, x t is the data input, b t is the bias value of the input gate, i t is the output of the input gate, W C is the weight of the candidate value, h t-1 is the input of the previous neuron, x t is the data input, b C is the bias value of the candidate value, is the output of the candidate value, C t-1 is the output of the previous candidate value, C t is the updated value of the candidate value, W o is the weight of the output gate, h t-1 is the input of the previous neuron, x t is the data input, b o is the bias value of the output gate, o t is the output of the output gate, h t is the value passed to the next neuron.

5. A real-time traffic congestion prediction method based on multi-integrated deep learning according to claim 4, characterized in that The multi-head attention layer specifically performs the following steps: Step 1: Perform a linear transformation on the input time series features to generate a query vector q i , a key vector k i and a value vector v i , expressed as: Among them, W q , W k , W v are learnable weight matrices, and x i is the time step in the time series feature; Step 2: For each query vector q i , each attention head calculates its similarity with all key vectors k j through a dot product function to obtain attention scores, which are expressed as: score h (q i ,k j )=q T ·k j Step 3: Use the Softmax function to normalize the attention scores of each attention head into a probability distribution, expressed as: Among them, α ij represents the attention weight of the i-th query to the j-th key, n represents all time steps, and l represents the l-th time step; Step 4: Use the attention weight α ij to perform weighted summation on the value vector v j to obtain the output vector y of each attention head i , y i is expressed as; Step 5: Concatenate the outputs of all attention heads and generate the final output through a linear transformation, expressed as: Among them, W o is a learnable weight matrix.

6. A real-time traffic congestion prediction system based on multi-integrated deep learning, characterized in that The system specifically performs the following steps: Based on real-time monitoring videos, traffic flow, speed variance, headway spacing variance, headway time variance, and conflict rate are obtained. Then, the traffic flow, speed variance, headway spacing variance, headway time variance, and conflict rate are input into a prediction network to obtain the total delay time of the output traffic flow. Finally, traffic congestion prediction is carried out according to the total delay time of the traffic flow; The prediction network is Multi-AM-OL-LSTM, and the Multi-AM-OL-LSTM includes an input layer, a multi-head attention layer, an LSTM layer, and an output layer. The input layer receives input data and converts the input data into a three-dimensional tensor, where the three-dimensional tensor is the number of samples × time steps × number of features. The three-dimensional tensor is input into the LSTM layer, and the LSTM layer extracts the time series feature h” of the three-dimensional tensor. i , the time series feature h” i is input into the multi-head attention layer. The multi-head attention layer calculates its attention weights according to the correlation of each time step in the sequence, and then obtains the weighted feature. The output layer maps the weighted feature generated by the attention layer to the final predicted value.

7. The real-time traffic congestion prediction system based on multi-integrated deep learning according to claim 6, characterized in that The specific steps for obtaining the traffic flow, speed variance, headway spacing variance, headway time variance, and conflict rate are as follows: Step 1: Obtain historical monitoring video data, and use the Video Capture function in the OpenCV library to read and process the monitoring video data frame by frame to obtain consecutive frame images, and perform grayscale processing on each frame image in the consecutive frame images; Step 2: Use the YOLO-v8 network to identify each vehicle in each frame image, and obtain the center point pixel coordinates of each monitoring bounding box. Then, use the Lucas-Kanade method to track each center point pixel to obtain vehicle trajectory information; Step 3: Obtain the vehicle ID, head direction, vehicle coordinates, and vehicle passing time difference according to the vehicle trajectory information. Obtain the traffic flow according to the vehicle ID, obtain the vehicle speed according to the vehicle passing time difference, obtain the speed variance according to the vehicle speeds of all vehicles, and obtain the headway spacing between adjacent vehicles according to the vehicle coordinates and head directions of adjacent vehicles; Step 4: Repeat Step 3 to obtain all headway spacings, and use all headway spacings to obtain the headway spacing variance; Step 5: Obtain the headway spacing between adjacent vehicles and the speed of the following vehicle in the adjacent vehicles, and divide the headway spacing between adjacent vehicles by the speed of the following vehicle to obtain the headway time; Step 6: Obtain the headway time variance based on all headway times; Step 7: Divide the headway spacing between adjacent vehicles by the speed difference between the adjacent vehicles to obtain the TTC. Those with TTC greater than the threshold are judged as collisions, otherwise as non-collisions. Finally, divide the number of collisions by the traffic volume to obtain the conflict rate.

8. A real-time traffic congestion prediction system based on multi-integrated deep learning according to claim 7, characterized in that The total delay time of the traffic flow is obtained through the following steps: Step 1: Obtain the length of the monitored section in the monitored area and the expected speed of this section, that is, the speed limit value of this section, and obtain the expected passing time according to the length of the monitored section and the expected speed of this section; Step 2: Obtain the time when the vehicle enters and exits the monitored section according to the monitoring video to obtain the actual passing time; Step 3: Subtract the expected passing time from the actual passing time to obtain the total delay time of the traffic flow.

9. The real-time traffic congestion prediction system based on multi-integrated deep learning according to claim 6, characterized in that The time series feature h” i is expressed as: h” i = o t *tanh(C t ) Among them, W f is the weight of the forget gate, h t-1 is the input of the previous neuron, x t is the data input, b f is the bias value of the forget gate, f t is the output of the forget gate, W i is the weight of the input gate, h t-1 is the input of the previous neuron, x t is the data input, b t is the bias value of the input gate, i t is the output of the input gate, W C is the weight of the candidate value, h t-1 is the input of the previous neuron, x t is the data input, b C is the bias value of the candidate value, is the output of the candidate value, C t-1 is the output of the previous candidate value, C t is the updated value of the candidate value, W o is the weight of the output gate, h t-1 is the input of the previous neuron, x t is the data input, b o is the bias value of the output gate, o t is the output of the output gate, h t is the value passed to the next neuron.

10. A real-time traffic congestion prediction system based on multi-integrated deep learning according to claim 9, characterized in that The multi-head attention layer specifically performs the following steps: Step 1: Perform a linear transformation on the input time series features to generate a query vector q i , a key vector k i and a value vector v i , expressed as: Among them, W q , W k , W v are learnable weight matrices, and x i is the time step in the time series feature; Step 2: For each query vector q i , each attention head calculates its similarity with all key vectors k j through a dot product function to obtain attention scores, which are expressed as: score h (q i ,k j )=q T ·k j Step 3: Use the Softmax function to normalize the attention scores of each attention head into a probability distribution, expressed as: where α ij represents the attention weight of the i-th query to the j-th key, n represents all time steps, and l represents the l-th time step; Step 4: Utilize the attention weight α ij to perform weighted summation on the value vector v j to obtain the output vector y of each attention head i , y i is expressed as; Step 5: Concatenate the outputs of all attention heads and generate the final output through a linear transformation, expressed as: Among them, W o is a learnable weight matrix.

Citation Information

Cited By

  • Traffic flow prediction method based on time-frequency domain joint modeling

    CN121354346A