Double-flow learning prediction model based on space-time fusion
Through a dual-flow learning prediction model based on space-time fusion, combined with TF online monitoring system, Coord-GCN and LSTM, the problem of lack of timeliness of traffic prediction methods is solved, and accurate prediction of the spatio-temporal and spatial characteristic parameters of highway toll stations is achieved, which improves traffic efficiency and avoids paralysis of toll stations.
Patent Information
- Application Number
- CN202510055235.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
AI Technical Summary
The existing vehicle flow prediction methods lack timeliness and it is difficult to predict the spatio-temporal and spatial characteristics of highway toll stations in real time, resulting in a decrease in vehicle traffic efficiency and may even cause paralysis of toll stations.
The dual-stream learning prediction model based on space-time fusion is adopted, combined with the TF online monitoring system, coordinate global convolutional network (Coord-GCN) and long-term memory network (LSTM), to achieve seamless fusion of space-time information, capture important models and long-term dependencies in vehicle traffic data, and make real-time predictions.
Accurate and timely prediction of the spatial and temporal characteristics of traffic flow is achieved, the traffic efficiency of highway toll stations is improved, and the risk of paralysis of toll stations is avoided.
Smart Images

Figure CN120012827A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of highway toll collection, and in particular to a dual-stream learning prediction model based on spatiotemporal fusion. Background Art
[0002] Highway toll stations are necessary facilities for collecting tolls from vehicles. H-TS must be set up at toll roads or toll intersections. The congestion state of H-TS depends on the smoothness of the traffic flow (TF). Too much TF will lead to the decline of H-TS performance. After the efficiency of H-TS degrades to a certain extent, it will face paralysis and damage. If the H-TS with serious inefficiency cannot clear the TF in time, it will seriously affect the driving performance and safety performance of the vehicle. Therefore, real-time analysis and prediction of the spatiotemporal performance of TF is crucial. At present, preliminary research on TF mainly focuses on the prediction of space and time. Generally speaking, when the spatial density rises to 70% or the processing time increases to 0.2s, we believe that H-TS reaches the level of paralysis. However, in actual operation, it is difficult to directly obtain the spatial density and temporal clues of traffic flow. Directly measuring parameters such as traffic density requires field measurement and analysis, which will cause huge cost waste. Therefore, it is very necessary to obtain the spatial information and temporal clues of TF in real time and accurately through simple and easy-to-measure characteristic parameters.
[0003] In existing research, there are three main methods for TF prediction: statistics, machine learning, and deep learning. For sections where TF is relatively stable, statistical methods such as ARIMA, Kalman filtering, and linear regression can be used. However, since linear models have obvious nonlinear change characteristics, they are no longer suitable for highway lane-level TF prediction. Machine learning models train a large amount of data through supervised learning to establish a mapping relationship between input and output, which is suitable for nonlinear TF prediction.
[0004] Although machine learning methods can effectively capture the nonlinear change characteristics of TF, their structure is simple and it is challenging to learn deeper complex dynamics. In recent years, with the continuous acquisition of multi-source big data, deep learning methods have gradually played a leading role in short-term TF prediction and have broad application prospects. Recurrent neural networks (RNNs) can process sequences as the main structure of deep learning models. Compared with BPNNs, the loop feedback mechanism enables RNNs to remember the information of the last moment and capture the temporal characteristics of TF. However, when the network depth is too deep and it cannot learn the long-term dependence on TF, problems such as gradient disappearance and explosion are prone to occur. The long short-term memory neural network (LSTM) uses gate units to transmit effective historical information to achieve long-term and short-term memory of TF, thereby obtaining better prediction results. In addition, some variants based on the LSTM model have shown excellent capabilities in TF prediction. These models have achieved good results in time series prediction, but due to the large number of parameters, they cannot meet the needs of real-time prediction.
[0005] With the development of artificial intelligence, many scholars are not satisfied with the improvement of accuracy, but are more concerned about how to combine real-time and accuracy to meet the deployment requirements of real life. Although many scholars are committed to the accurate and real-time prediction and analysis of TF, many previous studies have focused on the spatial attention of traffic flow. Specifically, when constructing the corresponding theoretical or data model, the spatial mechanism and time of TF are calculated, and then the relevant parameters are imported into the model to obtain the corresponding TF prediction value. However, although these methods can achieve high-precision predictions through advanced machine learning and artificial intelligence, they often cannot predict TF performance parameters in real time.
[0006] The development of digital twin technology makes it possible to predict TF in real time. Digital twin is a real-time simulation process that combines multidisciplinary theories, multi-physics field knowledge, multi-scale information and multi-probability thinking. It makes full use of physical entity characteristics, sensor measurement parameters and TF historical data to build virtual-real mapping, digital twin display, monitor and adjust the entire process of TF. With the advancement of computer technology and the enhancement of computing power, digital twin technology based on information-physical systems has once again been favored by researchers. Due to the lack of timeliness of TF parameter prediction methods, it is difficult to predict TF in time. What is difficult to predict in time is the spatiotemporal characteristic parameters. At the same time, in the actual operation of the vehicle, it is not convenient to obtain the real-time TF data of the entire H-TS. Therefore, this paper proposes a dual-stream learning prediction model based on spatiotemporal fusion. Summary of the invention
[0007] The present invention discloses a dual-stream learning prediction model based on spatiotemporal fusion, aiming to solve the technical problem raised in the background technology that the spatiotemporal characteristic parameters of TF are difficult to predict in time due to the lack of timeliness of the prediction method of TF parameters.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A dual-stream learning prediction model based on spatiotemporal fusion includes a TF online monitoring system, a coordinate global convolutional network (Coord-GCN) and a long short-term memory network (LSTM). The TF online monitoring system takes TF data as input, completes the spatial position curve through the coordinate-global convolutional network, learns the complete TF spatial feature representation, and uses LSTM to capture the time clues of TF data. The time-space curve is imported into the ST-Fusion module trained with historical cycle data for reshaping and processing, so as to achieve seamless fusion of spatiotemporal information, capture the most important patterns and long-term dependencies in TF data, and then predict the final TF. The LSTM core is a storage unit composed of an input gate, a forget gate and an output gate. The forget gate selectively forgets unimportant information in the previous unit through a Sigmoid function and retains relevant important clues; the input gate determines whether relevant signals are added to the current process; the output gate determines which information is regarded as output in the current process.
[0010] In a preferred solution, the specific calculation formula of the LSTM is as follows:
[0011]
[0012] where x t is the input signal; h t is the output signal; t is the number of LSTM; f t is the output of the forget gate; i t and are the output stage and current stage of the input gate respectively; o t and C t is the output state and current state of the output gate; W and b are the weights and offsets of the corresponding groups; σ is the Sigmoid activation function; tanh is the tanh activation function.
[0013] In a preferred solution, the Coord-GCN is constructed using coordinate convolution - Coord-Conv. Coord-Conv uses the extracted horizontal clues i and vertical clues j as additional features to cascade with the original mapping, and uses parallel methods to add the i coordinate horizontal information and j coordinate vertical information obtained through linear transformation, and inputs the two feature maps into k×1 Conv and 1×k Conv to obtain the horizontal spatial features and vertical spatial features of TF. In order to make the spatial information obtained in each path not limited to a single path, the two features are then converted by 1×k Conv and k×1 Conv to obtain the overall spatial information of the feature map, and finally the obtained spatial map is fused pixel by pixel. The overall calculation is shown in the following formula:
[0014]
[0015] in For cascade; out up is the output of the on-road feature map; out low is the output of the lower feature map; out spatial is the TF space feature of the module.
[0016] In a preferred solution, after obtaining the temporal and spatial information of TF respectively, ST-Fusion is used to effectively fuse the temporal information and spatial information to accurately decode complex multidimensional data. ST-Fusion converts x into a matrix by aligning and connecting tensor slices at each time step t. The calculation is shown in the following formula:
[0017]
[0018] The ST-Fusion Layer processes the extended sequence data with lower computational complexity by using a simplified hardware-aware parallel algorithm. The output y is obtained from the ST-Fusion Layer. in , ST-Fusion is constructed as shown in the following formula:
[0019]
[0020] Residual connections are used to allow gradients to propagate between layers without being reduced, and layer normalization and Dropout are applied to stabilize the learning process and prevent overfitting of the model pre-training data;
[0021] The calculation of application layer normalization is as follows:
[0022]
[0023] Among them, μ and σ 2 are the mean and variance of the feature dimension d, ρ and ζ are the learnable scale and displacement, which facilitate the adjustment of the normalization effect.
[0024] From the above, it can be seen that the dual-stream learning prediction model based on spatiotemporal fusion includes a TF online monitoring system, a coordinate global convolutional network-Coord-GCN and a long short-term memory network-LSTM. The TF online monitoring system uses TF data as input, completes the spatial position curve through the coordinate-global convolutional network, learns the complete TF spatial feature representation, and uses LSTM to capture the time clues of TF data, imports the time-space curve into the ST-Fusion module trained with historical cycle data for reshaping and processing, realizes the seamless fusion of spatiotemporal information, captures the most important patterns and long-term dependencies in TF data, and then predicts the final TF. The LSTM core is a storage unit composed of an input gate, a forget gate and an output gate. The forget gate selectively forgets unimportant information in the previous unit through a Sigmoid function and retains relevant important clues; the input gate determines whether to add relevant signals to the current process; the output gate determines which information is regarded as output in the current process. The dual-stream learning prediction model based on spatiotemporal fusion provided by the present invention has the following technical effects:
[0025] (1) Combining advanced technologies such as big data, cloud computing, and the Internet of Things, a digital twin framework based on spatiotemporal learning prediction models (SP-STMs) is constructed to predict and complete the spatiotemporal information curve of TF according to the actual H-TS needs.
[0026] (2) Coord-GCN and LSTM are used to obtain the location node information and time clues of TF respectively to provide real-time feedback on the online status of TF.
[0027] (3) The designed spatiotemporal fusion module seamlessly integrates spatial-temporal information, eliminating the need for a separate data processing stage and reducing computational overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is a digital twin framework diagram based on the ST-SPMs model of the dual-stream learning prediction model based on spatiotemporal fusion proposed in the present invention.
[0029] Figure 2 This is the LSTM architecture diagram of the dual-stream learning prediction model based on spatiotemporal fusion proposed in the present invention.
[0030] Figure 3 This is the Coord-GCN architecture diagram of the dual-stream learning prediction model based on spatiotemporal fusion proposed in this invention.
[0031] Figure 4 This is an example diagram of the mutual fusion of time information and space information of the dual-stream learning prediction model based on time-space fusion proposed in the present invention. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0033] Reference Figure 1-Figure 4 , based on the dual-stream learning prediction model of spatiotemporal fusion, a digital twin structure is used to build a TF online monitoring system, using TF data as input, completing the spatial position curve through the coordinate-global convolutional network, learning the complete TF spatial feature representation, and using LSTM to capture the time clues of TF data. Then, the spatiotemporal curve is imported into the ST-Fusion module trained with historical cycle data for reshaping and processing, achieving seamless fusion of spatiotemporal information, effectively capturing the most important patterns and long-term dependencies in TF data, and then predicting the final TF. This comprehensive design enables the ST-SPMs model to accurately provide precise and reliable TF dynamic predictions and effectively capture the complex interactions between spatiotemporal factors.
[0034] In a preferred embodiment, LSTM is an advanced version of RNN, which can solve the problems of long-term dependency, gradient vanishing and gradient exploding. Therefore, LSTM has achieved remarkable results in predicting time features. For LSTM, the core is a storage unit composed of input gate, forget gate and output gate.
[26] The structure of LSTM is as follows Figure 2 As shown in the figure, the forget gate selectively forgets the unimportant information in the previous unit through the Sigmoid function and retains the relevant important clues; the input gate decides to add the relevant signal to the current process; the output gate decides which information is regarded as the output of the current process. The specific calculation is shown in the formula:
[0035]
[0036] where x t is the input signal; h t is the output signal; t is the number of LSTM; f t is the output of the forget gate; i t and are the output stage and current stage of the input gate respectively; o t and C t is the output state and current state of the output gate; W and b are the weights and offsets of the corresponding groups; σ is the Sigmoid activation function; tanh is the tanh activation function.
[0037] In a preferred embodiment, "Coord-Conv" introduces coordinate information on the basis of standard convolution, and extracts horizontal clues i and vertical clues j as additional features and cascades with the original mapping. This method can effectively solve the spatial position nodes that conventional convolution cannot take into account, and enhance the perception of object boundary information. The disadvantage is that it has a large number of parameters similar to conventional convolution. GCN can effectively avoid this problem by using nonlinear convolution. Therefore, this paper designs a Coord-GCN module to organically combine "Coord-Conv" with GCN. The purpose is to capture the spatial position of the object under the premise of reducing the number of parameters. Its structure is as follows: Figure 3 shown.
[0038] In a preferred embodiment, the method uses a parallel method to add the i-coordinate horizontal information and j-coordinate vertical information obtained through linear transformation, and inputs the two feature maps into k×1 Conv and 1×k Conv to obtain the horizontal spatial features and vertical spatial features of TF. In order to make the spatial information obtained in each path not limited to a single path, the two features are transformed by 1×k Conv and k×1 Conv to obtain the overall spatial information of the feature map, and finally the obtained spatial map is fused pixel by pixel. The overall calculation is shown in the formula:
[0039]
[0040] in For cascade; out up is the output of the on-road feature map; out low is the output of the lower feature map; out spatial is the TF space feature of the module.
[0041] In a preferred embodiment, after acquiring the time-space information of TF respectively, ST-Fusion is used to effectively fuse the time information and the space information to accurately decode the complex multi-dimensional data, such as Figure 1 As shown. ST-Fusion converts x into a matrix by aligning and concatenating tensor slices at each time step t. Its calculation is shown in formula (3). By reshaping the tensor, we T×d Get a new embedding T encapsulates the total length of time T×S, i.e., the time dimension, simplifies the composite dimension, effectively unifies the spatiotemporal information, and captures the complex patterns in the data:
[0042]
[0043] The fused spatiotemporal information is Figure 4As shown, as the input of ST-Fusion Layer, long sequence data is effectively processed through a simplified hardware-aware parallel algorithm. Unlike traditional Transformer, ST-Fusion Layer processes extended sequences with lower computational complexity, which is crucial for tasks that require long-term dependency modeling, such as TF.
[0044] In a preferred embodiment, once we obtain the output y from the ST-Fusion Layer in , ST-Fusion is constructed as shown in the formula:
[0045]
[0046] Here, we use residual connections to allow gradients to propagate between layers without decreasing, and apply layer normalization and Dropout to stabilize the learning process and prevent overfitting of the model pre-training data. The calculation of applying layer normalization is shown in the following formula:
[0047]
[0048] Among them, μ and σ 2 is the mean and variance of the feature dimension d, ρ and ζ are learnable scales and shifts, which facilitate the adjustment of the normalization effect. This mechanism helps to alleviate internal covariate shift, accelerate training convergence, and improve the overall performance of deep neural models.
[0049] Experimental results show that the ST-SPMs model has achieved significant improvements in prediction accuracy and real-time performance, opening up a new perspective for traffic flow prediction.
[0050] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A dual-stream learning prediction model based on spatiotemporal fusion, including TF online monitoring system, coordinate global convolutional network-Coord-GCN and long short-term memory network-LSTM, characterized by: The TF online monitoring system takes TF data as input, completes the spatial position curve through the coordinate-global convolutional network, learns the complete TF spatial feature representation, and uses LSTM to capture the time clues of TF data. The time-space curve is imported into the ST-Fusion module trained with historical cycle data for reshaping and processing, so as to achieve seamless fusion of time and space information, capture the most important patterns and long-term dependencies in TF data, and then predict the final TF. The LSTM core is a storage unit composed of an input gate, a forget gate and an output gate. The forget gate selectively forgets unimportant information in the previous unit through a Sigmoid function and retains relevant important clues; the input gate decides to add relevant signals to the current process; the output gate decides which information is regarded as output in the current process.
2. The dual-stream learning prediction model based on spatiotemporal fusion according to claim 1 is characterized in that: The specific calculation formula of the LSTM is as follows: where x t is the input signal; h t is the output signal; t is the number of LSTM; f t is the output of the forget gate; i t and are the output stage and current stage of the input gate respectively; o t and C t is the output state and current state of the output gate; W and b are the weights and offsets of the corresponding groups; σ is the Sigmoid activation function; tanh is the tanh activation function.
3. The dual-stream learning prediction model based on spatiotemporal fusion according to claim 1 is characterized in that: The Coord-GCN is constructed using coordinate convolution - Coord-Conv, which concatenates the extracted horizontal clues i and vertical clues j with the original mapping as additional features.
4. The dual-stream learning prediction model based on spatiotemporal fusion according to claim 3 is characterized in that: The horizontal information of the i coordinate and the vertical information of the j coordinate obtained by linear transformation are added in parallel, and the two feature maps are input into k×1Conv and 1×kConv to obtain the horizontal and vertical spatial features of TF. In order to make the spatial information obtained in each path not limited to a single path, the two features are transformed by 1×kConv and k×1Conv to obtain the overall spatial information of the feature map, and finally the obtained spatial map is fused pixel by pixel. The overall calculation is shown in the following formula: in For cascade; out up is the output of the on-road feature map; out low is the output of the lower feature map; out spatial is the TF space feature of the module.
5. The dual-stream learning prediction model based on spatiotemporal fusion according to claim 4 is characterized in that: After acquiring the temporal and spatial information of TF respectively, ST-Fusion is used to effectively fuse temporal and spatial information and accurately decode complex multidimensional data. ST-Fusion converts x into a matrix by aligning and connecting tensor slices at each time step t. The calculation is shown in the following formula:
6. The dual-stream learning prediction model based on spatiotemporal fusion according to claim 5 is characterized in that: As the input of ST-Fusion Layer, the long sequence data is effectively processed through a simplified hardware-aware parallel algorithm. ST-Fusion Layer processes the extended sequence with lower computational complexity and obtains the output y from ST-Fusion Layer. in , ST-Fusion is constructed as shown in the following formula: Residual connections are used to allow gradients to propagate between layers without being reduced, and layer normalization and Dropout are applied to stabilize the learning process and prevent overfitting of the model pre-training data.
7. The dual-stream learning prediction model based on spatiotemporal fusion according to claim 6, characterized in that: The calculation of application layer normalization is as follows: Among them, μ and σ 2 are the mean and variance of the feature dimension d, ρ and ζ are the learnable scale and displacement, which facilitate the adjustment of the normalization effect.