Vehicle flow prediction method and system based on bimodal optimization embedded learning model

Through the dual-mode optimization embedded learning model, the problem of insufficient vehicle flow prediction accuracy in the existing technology is solved, high-precision prediction of complex traffic scenarios is achieved, and decision-making support capabilities for traffic management are improved.

CN120356348AActive Publication Date: 2025-07-22FUJIAN NORMAL UNIV

Patent Information

Application Number
CN202510831040.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-22
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing vehicle flow prediction methods have insufficient prediction accuracy in complex traffic scenarios, insufficient multi-source data fusion capabilities, limited robustness and generalization capabilities, making it difficult to deal with emergencies and nonlinear dynamics.

Method used

The dual-modal optimization embedded learning model is adopted, multi-source data is integrated, and the cross-modal feature fusion mechanism and robust model architecture are designed. The pre-transformer with layer normalization and customized traffic event prompt templates are used to extract and predict features using attention alignment module and sequence prediction module.

Benefits of technology

It improves the accuracy and reliability of vehicle flow forecasting, can deal with emergencies in complex traffic scenarios, and provides accurate traffic management support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356348A_ABST
    Figure CN120356348A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle flow prediction method and system based on a bimodal optimization embedded learning model, and the system comprises a data preprocessing module, a bimodal coding module, an attention alignment module and a sequence prediction module, is suitable for the field of intelligent transportation, and aims at achieving the prediction of the vehicle flow through multi-source data fusion and deep model architecture innovation. The problem that a traditional method is insufficient in prediction precision in a complex traffic scene is solved, and efficient decision support is provided for urban traffic management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation, and particularly to a vehicle flow prediction method and system based on a dual-modal optimized embedding learning model. Background Art

[0002] With the acceleration of the global urbanization process, the number of urban motor vehicles has shown an explosive growth. Problems such as traffic congestion and low commuting efficiency have become severe challenges in modern urban governance. Accurate vehicle flow prediction, as the core link of the intelligent transportation system, can provide key decision-making basis for traffic planning, signal timing optimization, dynamic route guidance, etc., and is of great significance for alleviating congestion and improving the operation efficiency of the traffic system. However, the limitations of existing prediction methods in complex traffic scenarios are becoming increasingly prominent, and a technological breakthrough is urgently needed.

[0003] Traditional statistical prediction methods, such as the moving average method and the exponential smoothing method, mainly rely on the time series law of historical data and capture trends through simple weighting or smoothing processing. The principles of these methods are simple and the calculation cost is low, but the significant defect is that they can only describe linear relationships and cannot model the complex non-linear dynamics in the traffic system. For example, during sudden gathering activities such as large-scale sports events and concerts, the vehicle flow in local areas of the city will show non-periodic violent fluctuations. Due to the lack of the ability to perceive external factors, the prediction error of traditional methods can reach more than 50% during peak hours, making it difficult to meet the sudden changes in traffic conditions.

[0004] Prediction models relying on a single data source also have significant deficiencies. Although the road sensor network can collect key indicators such as vehicle flow and vehicle speed in real time, there are problems such as limited coverage, data loss caused by equipment failures, and only reflecting the status of local sections; traffic monitoring video data is limited by camera perspectives, weather conditions, and target recognition accuracy, and it is difficult to stably provide global traffic information. More importantly, a single data source cannot integrate multi-dimensional influencing factors such as the behavior of traffic participants and the distribution of urban functional areas, resulting in insufficient adaptability of the prediction model to complex scenarios. For example, when extreme weather causes some sensor failures, the model relying only on sensor data will fail due to incomplete data, and without the support of supplementary data sources such as social media, it is difficult to obtain real-time road conditions through other channels, ultimately leading to mistakes in traffic guidance decisions.

[0005] In recent years, the development of machine learning and deep learning technologies has provided a new path for traffic prediction. Methods such as support vector machines and artificial neural networks have improved the fitting ability of complex patterns through non-linear mapping, but still face two core challenges: (1)Insufficient ability to fuse multi-source heterogeneous data: The traffic system involves modal information such as numerical time-series data, text semantic data, and visual image data. Existing models are difficult to efficiently integrate the feature associations of different modalities under a unified framework, resulting in insufficient information utilization. (2)Limited robustness and generalization ability: Data defects such as sensor noise and image recognition errors will significantly affect the model performance. Moreover, most models rely on the traffic patterns of specific cities for training, and the prediction accuracy will drop drastically in scenarios with different road network structures and travel habits. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a vehicle flow prediction method and system based on a dual-modal optimized embedding learning model. By integrating multi-source data, designing a cross-modal feature fusion mechanism and a robust model architecture, it breaks through the bottleneck of traditional methods and provides a high-precision and strong-generalization prediction solution for urban traffic management.

[0007] To achieve the above purpose, the present invention adopts the following technical solutions: A vehicle flow prediction method based on a dual-modal optimized embedding learning model, including the following steps: S1: Preprocess the time-series data and text data; S2: Input the preprocessed data into the dual-modal optimized embedding learning model for training. The dual-modal optimized embedding learning model includes a dual-modal encoding module, an attention alignment module, and a sequence prediction module; In the dual-modal encoding module, when encoding the preprocessed time-series data, layer normalization pre-Transformer is used for feature extraction to obtain time-series features; when encoding the preprocessed text data, a customized traffic event prompt template is used to guide GPT-2 to convert the prompt input into tokens, and after being processed by the pre-trained encoder, a semantic vector representation including event types and impact degrees, that is, text features, is obtained. The attention alignment module uses the attention mechanism to align and integrate the time-series features and text features; The sequence prediction module integrates the time-series features and text features to predict the vehicle flow in future periods.

[0008] In a preferred embodiment, the time-series data includes vehicle flow, vehicle speed, and lane occupancy, and the text data includes vehicle traffic event descriptions, congestion complaint information, and road construction reminder content.

[0009] In a preferred embodiment, for the data preprocessing in S1, a filtering algorithm is used to remove the noise in the data, an outlier detection algorithm is used to identify and process abnormal data points, and then a normalization formula is used to transform the data so that the data from different data sources have the same scale.

[0010] In a preferred embodiment, the sequence prediction module in S2 performs vehicle flow sequence prediction based on the integrated features, adopts a layer normalized front-end Transformer architecture, and has a hidden unit number range of [200, 800]. The number of hidden units is adjusted according to the amount of data and the length of the time series. The number of training iterations ranges from [1000, 5000], and the learning rate ranges from [0.001, 0.01]. The learning rate and number of iterations are adjusted according to the convergence speed and performance during the training process.

[0011] In a preferred embodiment, the customized traffic event prompt template in S2 refers to a text template containing the structured fields of "[event type], [road section name] and [impact level]" to guide GPT-2 to process traffic-related text, so that the semantic vector output by the model contains key information such as the type of traffic event, location of occurrence and impact level.

[0012] In a preferred embodiment, the future time period division in S2 refers to the demand for real-time regulation, management decision-making and strategic planning in the field of traffic forecasting, including short-term of 5-15 minutes, medium-term of 1-3 hours and long-term of 1-7 days; the accuracy of the prediction results is evaluated by the mean square error and mean absolute error indicators.

[0013] In a preferred embodiment, the attention alignment module in S2 generates query Q, key K and value V respectively by linear transformation of temporal features and text features, and uses the scaled dot product attention formula Calculate inter-modality correlation score , and through the dynamic weight formula Adjust the text feature weights by the scaling factor Weighted fusion is achieved after balancing the feature weights, and T represents transposition.

[0014] In a preferred embodiment, the sequence prediction module in S2 processes the embedded sequence through a multi-head self-attention mechanism, which projects the embedded sequence into multiple subspaces through a multi-head self-attention mechanism and calculates the attention weight matrix of each subspace in parallel. ,in i Represents the first i Attention head, Q i is the query matrix, K i is the key matrix, V i is a value matrix used to calculate the attention weights in the subspace; then concatenated and passed through a feedforward neural network Enhance the non - linear expression, where, x represents the feature vector input to the feed - forward neural network, W 1 and W 2 are weight matrices of different layers, b 1 and b 2 are the corresponding bias terms. The non - linear transformation of features is realized through matrix operations and activation functions; after residual connection and layer normalization, the prediction sequence is generated by the linear function y = xW + b, where W is the weight matrix of the prediction layer and b is the bias; subsequently, the non - linear expression ability of features is enhanced by the feed - forward neural network, and the training stability is optimized through residual connection and layer normalization; finally, the processed features are mapped to the target space by the linear prediction function to generate the prediction sequence of future multiple time steps.

[0015] The present invention also provides a vehicle flow prediction system based on a dual - mode optimized embedding learning model, which runs the vehicle flow prediction method based on the dual - mode optimized embedding learning model, including a data pre - processing module, a dual - mode encoding module, an attention alignment module and a sequence prediction module.

[0016] Compared with the prior art, the present invention has the following beneficial effects: by multi - source data fusion and deep model optimization, the prediction accuracy and reliability are improved, providing refined support for urban traffic management. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic structural diagram of a preferred embodiment of the present invention; Figure 2 is a schematic flow diagram of a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The present invention will be further described below with reference to the drawings and embodiments.

[0019] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0020] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.

[0021] A vehicle flow prediction method based on a dual-modal optimized embedding learning model, referring to Figure 1-2 , includes the following steps: First step, multi-source data collection and integration: Connect with the road sensor system of the urban traffic management department to obtain raw data covering multiple dimensions of the traffic system from multiple channels, including: real-time collecting core time-series data such as vehicle flow, vehicle speed, and lane occupancy through the urban road sensor network at a frequency of not less than every 5 minutes to accurately reflect the immediate traffic status of the road; connecting with mainstream map navigation software to obtain macroscopic traffic situation data such as road congestion status, traffic condition changes, and estimated travel time in real time; using web crawler technology to grab traffic-related text information on social media platforms hourly, and screening through keywords to retain effective content such as traffic event descriptions, congestion complaints, and construction reminders posted by users to supplement the emergency information not covered by official data sources in a timely manner, forming a dual-modal raw data set containing numerical time-series data and text semantic data.

[0022] Second step, data cleaning and standardized preprocessing: (1) In the data cleaning link, for multi-source heterogeneous traffic-related data, a combination of filtering algorithms and outlier detection algorithms is used to improve data quality: for time-series data such as vehicle flow and vehicle speed collected by road sensors, the median filtering algorithm is used to effectively remove noise caused by environmental interference, equipment failures and other factors, and statistical methods such as the 3σ principle are used to identify and process abnormal peak and valley values to ensure the reliability of time-series data; for unstructured data such as social media texts and map navigation software, through keyword filtering, text deduplication and logical verification, garbage information, duplicate content and invalid data irrelevant to traffic are removed, and key information such as effective traffic event descriptions and congestion status is retained.

[0023] (2)In the data standardization process, to eliminate the scale differences of different data sources, a normalization method adapted to the characteristics of bimodal data is adopted: for numerical data such as traffic flow and vehicle speed, the Z-score normalization formula is used to transform it into a standard normal distribution to unify the dimension of continuous variables; for traffic condition classification data, it is transformed into digital features recognizable by the model through numerical mapping or one-hot encoding; for social media text data, after preprocessing such as word segmentation and stop word removal, it is converted into semantic vectors of a fixed dimension with the help of a word vector model to align the feature scales of text information and numerical time series data, providing standardized and unified input data for subsequent bimodal model training.

[0024] The third step, bimodal model training and feature fusion: Input the preprocessed data into a three-layer architecture model including bimodal encoding, attention alignment, and sequence prediction: (1)Bimodal encoding module: For time series data, a pre-normalized Transformer architecture is used. The training process is stabilized through layer normalization, and the multi-head self-attention mechanism is used to capture time step dependencies, and a feed-forward network is combined to extract deep time series features; for text data, a customized traffic event prompt template (such as "[Event type] occurred at [Time] on [Road section name], causing a traffic impact of [Impact degree]") is used to guide GPT-2 to convert the prompt input into tokens, and after being processed by the pre-trained encoder, a 768-dimensional semantic vector representation containing the event type (such as accident / construction), affected road section and level is obtained.

[0025] (2)Attention alignment module: The attention alignment module generates query Q, key K, and value V from time series features and text features respectively through linear transformation, and calculates the correlation score between modalities using the scaled dot-product attention formula where the feature weights are balanced through a scaling factor ( d k is the dimension of the key vector) to achieve weighted fusion. For example, taking the time series features as the query, and the text features as the key and value, higher weights are assigned to text features with high correlation.

[0026] (3)Sequence prediction module: The sequence prediction module adopts a pre-normalized Transformer variant architecture. The embedded sequence is projected into multiple subspaces through the multi-head self-attention mechanism, and the attention weight matrices of each subspace are calculated in parallel and then concatenated and enhanced with non-linear expression through a feed-forward neural network (where x is the input feature vector, W 1 and W 2 are weight matrices, b1 and b 2 is the bias term), and after residual connection and layer normalization, a prediction sequence is generated by the linear function y = xW + b.

[0027] Fourth step, multi-time-span traffic flow prediction and application: (1) Traffic flow prediction: The fully trained model can predict the vehicle traffic flow in future time periods. In practical applications, the time span of prediction is set according to different requirements. For example, in the traffic signal control scenario, short-term vehicle traffic flow prediction is required to adjust the signal timing in real time. In the traffic planning scenario, medium-term or long-term prediction may be needed to provide a basis for traffic facility construction and optimization. At a busy intersection, the trained model is used to predict the change in vehicle traffic flow within the next 15 minutes. The model comprehensively considers the current road sensor data, the vehicle queuing situation in traffic monitoring videos, the surrounding road congestion information provided by map navigation software, and traffic event reports on social media, and gives a relatively accurate vehicle traffic flow prediction result.

[0028] In Figure 1 RV(t) represents the vehicle traffic flow fluctuation value, X T represents the input feature vector of the vehicle traffic flow time series data, represents the hidden layer feature after the time series data is processed by the RV encoder, represents the output feature of the (i + 1)-th layer of the RV encoder, P S represents the output of the prompt template before the text data is processed by GPT-2, represents the semantic feature vector after the text data is encoded by GPT-2, represents the normalized semantic weight output by the text encoder, H C represents the fused feature vector after the time series feature and the text feature are aligned by attention, X M represents the final prediction output.

[0029] (2) Application examples: The traffic management department adjusts traffic management strategies in a timely manner according to the prediction results of the model. When it is predicted that the traffic flow on a certain road will increase significantly in the future and congestion may occur, measures are taken in advance, such as adjusting the signal timing, increasing the green light duration in the direction of this road; issuing traffic guidance information to guide vehicles to avoid congested sections; arranging traffic police to the scene for dredging, etc. During a large-scale event in a certain city, through the vehicle flow prediction method of the present invention, the traffic flow changes around the event were accurately predicted. The traffic management department formulated a detailed traffic dredging plan in advance, effectively avoiding traffic congestion, ensuring the smooth progress of the event and the normal travel of surrounding residents. Verified by actual traffic data, this method significantly reduces the prediction error, can effectively integrate dynamic factors such as sudden accidents and holiday effects, and provides accurate support for traffic signal timing optimization, route guidance and resource planning. In a concert scenario, the social media text "The concert at XX Stadium will end at 20:00" is mapped to "Event type = large-scale event ending, Road section = around XX Stadium, Impact degree = severe" through the prompt template. The weight of the "ending" dimension in the semantic vector generated by GPT-2 is increased by 40%. The model predicts the traffic flow peak 3 hours in advance to be 2.8 times that of daily (predicted value 1800 vehicles / hour, actual 1750 vehicles / hour). The traffic department will assist in optimizing the dredging strategy to improve traffic efficiency. The present invention breaks through the limitations of traditional methods, realizes accurate modeling of the complex dynamics of urban traffic, and has important theoretical significance and engineering application value.

Claims

1. A vehicle flow prediction method based on a dual-modal optimized embedding learning model, characterized in that, It includes the following steps: S1: Preprocess the time series data and text data; S2: Input the preprocessed data into the dual-modal optimization embedding learning model for training. The dual-modal optimization embedding learning model includes a dual-modal encoding module, an attention alignment module, and a sequence prediction module; In the dual-modal encoding module, when encoding the preprocessed time series data, use layer normalization pre-Transformer for feature extraction to obtain time series features; When encoding the preprocessed text data, use a customized traffic event prompt template to guide GPT-2 to convert the prompt input into tokens, and after being processed by the pre-trained encoder, obtain a semantic vector representation including the event type and impact degree, that is, text features; The attention alignment module uses the attention mechanism to align and integrate the time series features and text features; The sequence prediction module integrates the time series features and text features to predict the vehicle flow in the future period.

2. The vehicle flow prediction method based on a dual-modal optimized embedding learning model according to claim 1, characterized in that The time series data includes vehicle flow, vehicle speed, and lane occupancy, and the text data includes vehicle traffic event descriptions, congestion complaint information, and road construction reminder content.

3. A vehicle traffic flow prediction method based on a dual-modal optimized embedding learning model according to claim 1, characterized in that, For the data preprocessing in S1, use a filtering algorithm to remove the noise in the data, use an outlier detection algorithm to identify and process the outlier data points, and then use a normalization formula to transform the data to make the data from different data sources have the same scale.

4. A vehicle flow prediction method based on a dual-modal optimized embedding learning model according to claim 1, characterized in that, The sequence prediction module in S2 performs vehicle flow sequence prediction based on the integrated features, uses a layer normalization pre-Transformer architecture, the number of its hidden units ranges from [200, 800], adjusts the number of hidden units according to the data volume and the length of the time series, the number of training iterations ranges from [1000, 5000], and the learning rate ranges from [0.001, 0.01]. Adjust the learning rate and the number of iterations according to the convergence speed and performance during the training process.

5. The vehicle flow prediction method based on a dual-modal optimized embedding learning model according to claim 1, characterized in that The customized traffic event prompt template in S2 refers to a text template containing structured fields of "[event type], [section name], and [impact degree]" to guide GPT-2 to process traffic-related text, so that the semantic vector output by the model contains key information such as traffic event type, occurrence location, and impact level.

6. The vehicle traffic flow prediction method based on a dual-modal optimized embedding learning model according to claim 1, wherein The division of the future period in S2 refers to the requirements for real-time regulation, management decision-making, and strategic planning in the field of traffic prediction, including short-term of 5-15 minutes, medium-term of 1-3 hours, and long-term of 1-7 days; the accuracy of the prediction results is evaluated by mean square error and mean absolute error metrics.

7. A vehicle flow prediction method based on a dual-modal optimized embedding learning model according to claim 1, characterized in that, The attention alignment module in S2 generates query Q, key K, and value V for the time series features and text features respectively through linear transformation, using the scaled dot product attention formula Calculate the correlation score between modes , and through the dynamic weight formula Adjust the text feature weights, where the weighted fusion is achieved after balancing the feature weights by a scaling factor and T represents the transpose.

8. A vehicle traffic flow prediction method based on a dual-modal optimized embedding learning model according to claim 1, characterized in that, The sequence prediction module in S2 processes the embedded sequence through a multi-head self-attention mechanism, which projects the embedded sequence into multiple subspaces through the multi-head self-attention mechanism and calculates the attention weight matrices of each subspace in parallel. , where i represents the i th attention head in the multi-head self-attention mechanism, Q i is the query matrix, K i is the key matrix, V i is the value matrix, which is used to calculate the attention weights within the subspace. Then concatenate, and then pass through a feed-forward neural network Enhance the non - linear expression, where x represents the feature vector input to the feed - forward neural network, W 1 and W 2 are weight matrices of different layers, b 1 and b 2 are the corresponding bias terms. The non - linear transformation of features is realized through matrix operations and activation functions; after residual connection and layer normalization, the prediction sequence is generated by the linear function y = xW + b, where W is the weight matrix of the prediction layer and b is the bias; subsequently, the feed - forward neural network is used to enhance the non - linear expression ability of features, and the training stability is optimized through residual connection and layer normalization; finally, the processed features are mapped to the target space through the linear prediction function to generate the prediction sequence for multiple future time steps.

9. A vehicle traffic flow prediction system based on a dual-modal optimized embedding learning model, characterized in that Run a vehicle flow prediction method based on the dual-modal optimization embedding learning model described in any one of claims 1-8 above, including a data preprocessing module, a dual-modal encoding module, an attention alignment module, and a sequence prediction module.

Citation Information

Patent Citations

  • CEEMDAN-RF-LSTM-based traffic flow time sequence data prediction method and system

    CN117574080A

  • Urban road traffic flow state prediction method and system based on big data

    CN118447687A

  • Traffic flow prediction system based on deep learning and dynamic network analysis and application method thereof

    CN118675324A

  • Multi-modal traffic flow prediction method based on multi-source data feature fusion

    CN119323879A

  • Traffic flow prediction optimization method based on intelligent optimization algorithm

    CN119600812A

Cited By

  • Multi-modal traffic flow prediction method based on artificial intelligence and related equipment

    CN121236923A

  • Deep learning-based road traffic flow short-term prediction method and system

    CN121725643A