Intelligent transportation management multi-modal digital monitoring method and system based on deep learning

By collecting multi-source traffic data and using deep learning methods to establish spatiotemporal correlations, standardized data sequences are generated for traffic situation assessment and flow field abrupt change identification. This solves the problem of insufficient multi-source data processing capabilities and achieves efficient traffic monitoring and management.

CN121583116BActive Publication Date: 2026-04-28SICHUAN VOCATIONAL & TECHN COLLEGE OF COMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN VOCATIONAL & TECHN COLLEGE OF COMM
Filing Date
2026-01-23
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies lack the ability to process multi-source traffic data in intelligent transportation management, making it difficult to establish effective correlation mapping between different types of data and fully extract the spatiotemporal features of the data. This results in low quality of standardized data sequence generation, a lack of accurate data support for subsequent situation assessment, and an inability to accurately reflect the real state of traffic operation. Furthermore, the lack of ability to identify and finely monitor sudden changes in traffic flow field affects the timeliness and accuracy of traffic management decisions.

Method used

By deploying heterogeneous traffic sensing devices to collect multi-source traffic data, establishing spatiotemporal correlations using a multi-source traffic data fusion engine, generating standardized traffic data sequences, encoding and decoding them using a traffic situation VAE dynamic model, capturing spatiotemporal dependencies by combining a flow field mutation Transformer model, and using heterogeneous node traffic sensing algorithms for data correlation analysis and collaborative verification to generate high-quality, refined monitoring data, and finally constructing a multi-dimensional monitoring view through a multimodal digital monitoring module.

Benefits of technology

It has achieved high-quality traffic situation assessment and flow field change identification, generated high-quality and refined monitoring data, improved the real-time, accuracy and comprehensiveness of traffic monitoring, provided a scientific basis for transportation management decisions, and met the needs of refined management under complex road networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583116B_ABST
    Figure CN121583116B_ABST
Patent Text Reader

Abstract

The application discloses a deep learning-based intelligent transportation management multi-modal digital monitoring method and system, multi-source traffic data is collected through heterogeneous traffic perception equipment, transmitted to a multi-source traffic data fusion engine for feature extraction and correlation mapping, and a standardized traffic data sequence is generated; the sequence is input into a traffic situation dynamic model for encoding and decoding, and a preliminary traffic situation assessment result is output; the result is input into a flow field mutation recognition model to capture a space-time dependence relationship and identify a traffic flow field mutation area; heterogeneous node traffic perception algorithms are used for correlation analysis and collaborative verification of the mutation area data to generate refined monitoring data; and finally, a multi-dimensional view is constructed through a multi-modal digital monitoring module to realize real-time monitoring of road network traffic operation states. The application improves multi-source data processing efficiency and situation analysis accuracy, quickly identifies flow field mutations, generates high-quality monitoring data, and meets the refined management requirements of complex road networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation management technology, and in particular to a multimodal digital monitoring method and system for intelligent transportation management based on deep learning. Background Technology

[0002] With the acceleration of urbanization, the continuous expansion of road networks, and the ever-increasing number of vehicles, traffic operation is exhibiting complex and ever-changing characteristics, significantly increasing the demands for real-time and accurate traffic management. Current traffic management needs to integrate multiple types of data, including vehicle trajectories, road cross-sectional flow, traffic signal phases, road surface conditions, and meteorological influences. Through dynamic perception of traffic conditions and timely identification of sudden changes in flow fields, comprehensive monitoring of the road network's operational status can be achieved. Traditional monitoring methods rely on single-device data collection, making it difficult to handle the collaborative processing needs of multi-source data and lacking in-depth mining of traffic spatiotemporal correlation characteristics. This fails to meet the practical application scenarios of refined management under complex road networks. Therefore, a deep learning-based multimodal digital monitoring method is urgently needed, integrating technologies such as multi-source data fusion, dynamic model analysis, and heterogeneous node perception to improve the effectiveness and reliability of traffic monitoring.

[0003] Existing technologies for intelligent transportation management have two significant drawbacks: First, they lack the ability to process multi-source traffic data, making it difficult to establish effective correlation mappings between different types of data and fully extract the spatiotemporal features from the data. This results in low-quality standardized data sequences, a lack of accurate data support for subsequent situation assessments, and an inability to accurately reflect the true state of traffic operations. Second, they lack the ability to identify and monitor traffic flow field abrupt changes in detail. They lack depth in the dynamic analysis of traffic conditions, cannot effectively capture the spatiotemporal dependencies in traffic data, and struggle to quickly identify areas of abrupt changes in traffic flow. Furthermore, the correlation analysis and collaborative verification of heterogeneous node data in abrupt change areas are inadequate, failing to generate high-quality, detailed monitoring data. This, in turn, affects the timeliness and accuracy of traffic management decisions. Summary of the Invention

[0004] To overcome the shortcomings and deficiencies of existing technologies, this invention provides a multimodal digital monitoring method for intelligent transportation management based on deep learning.

[0005] A deep learning-based multimodal digital monitoring method for intelligent transportation management includes:

[0006] Step S1: Collect multi-source traffic data by using heterogeneous traffic sensing devices deployed at road network calibration nodes. The multi-source traffic data includes vehicle trajectory data, road cross-section flow data, traffic signal phase data, road surface condition data, and meteorological impact data. Transmit the collected multi-source traffic data to the multi-source traffic data fusion engine.

[0007] Step S2: Use a multi-source traffic data fusion engine to extract features and perform association mapping on the received multi-source traffic data, establish spatiotemporal associations between different types of traffic data, and generate standardized traffic data sequences.

[0008] Step S3: Input the standardized traffic data sequence into the traffic situation VAE dynamic model. The encoder of the model performs hierarchical feature encoding on the standardized traffic data sequence to obtain the traffic situation latent vector. Then, the decoder of the model reconstructs the traffic situation latent vector and outputs the preliminary traffic situation assessment results.

[0009] Step S4: Input the preliminary traffic situation assessment results into the flow field abrupt change Transformer model, capture the spatiotemporal dependencies in the preliminary traffic situation assessment results through the model's multi-head attention mechanism, generate a traffic flow field spatiotemporal feature matrix in combination with the location encoding module, and identify the traffic flow field abrupt change region based on the traffic flow field spatiotemporal feature matrix.

[0010] Step S5: Use the heterogeneous node traffic perception algorithm to perform heterogeneous node data association analysis on the multi-source traffic data corresponding to the traffic flow field change area. Through the node feature matching module and data collaborative verification module of the algorithm, generate refined monitoring data of the traffic flow field change area.

[0011] Step S6: Based on the refined monitoring data of traffic flow field change areas, construct a multi-dimensional monitoring view through the intelligent transportation management multi-modal digital monitoring module to monitor the traffic operation status of the road network in real time.

[0012] Furthermore, the expression for the traffic situation VAE dynamic model is: ,in, The loss function of the traffic situation VAE dynamic model is... Let be the posterior probability distribution of the model encoder. For encoder network parameters, This is the latent vector for traffic situation. To standardize traffic data sequences, Let be the likelihood probability distribution of the model decoder. For decoder network parameters, Let be the prior probability distribution of the latent vector. The KL divergence function is used; the model also incorporates a traffic situation time-series weighting factor. Construct the time-series optimization loss function: ,in, The length of the time series. For the first Latent vector of traffic situation at any given time. For the first Standardized traffic data sequences for each time period.

[0013] Furthermore, the expression for generating the spatiotemporal feature matrix of the traffic flow field in the Transformer model of the flow field mutation is as follows: ,in, The traffic flow field spatiotemporal characteristic matrix, For multi-head attention computation function, For querying the matrix, The key matrix, Let be a value matrix, and These are the weight parameters for the query, key, and value matrices, respectively. This is a preliminary traffic situation assessment result. For position encoding functions, For time dimension parameters, The spatial dimension parameter is used; the expression for identifying abrupt changes in the flow field in the model is as follows: ,in, The results of traffic flow field abrupt change identification. For Sigmoid activation function, Conv1d This is a one-dimensional convolution operation. Weight matrix for mutation identification, Bias parameters for mutation identification.

[0014] Furthermore, the heterogeneous node data association analysis expression of the heterogeneous node traffic perception algorithm is as follows: ,in, The correlation coefficient for heterogeneous node data. This represents the number of heterogeneous sensing nodes. For the number of traffic data types, For the first The sensing node and the first Association weights of traffic-like data This is the function for calculating feature similarity. For the first The feature vector of each sensing node For the first The feature vector of traffic data; the expression for generating refined monitoring data by the algorithm is: ,in, To refine the monitoring data, For feature concatenation function, Optimize the weight matrix for the data. For residual connection functions, This is the data co-verification error vector.

[0015] Furthermore, the feature extraction and association mapping expression of the multi-source traffic data fusion engine is as follows: ,in, For the fused traffic feature vector, LayerNorm ( () represents the layer normalization operation. It is a feedforward neural network. For the feature concatenate operation, This is the feature vector of the vehicle's driving trajectory data. This is the feature vector of road cross-section traffic flow data. For traffic signal phase data feature vectors, For road surface condition data feature vectors, The feature vector represents the meteorological impact data; the standardized traffic data sequence generation expression of the engine is: ,in, To standardize traffic data sequences, To fuse the mean of the feature vectors, To fuse the standard deviation of the feature vectors, Adjust the weight matrix for standardization.

[0016] Furthermore, the expression for constructing the multi-dimensional monitoring view of the intelligent transportation management multimodal digital monitoring module is as follows: ,in, For multi-dimensional monitoring view, For view projection functions, These are the weighting parameters for traffic flow field characteristics. To refine the weighting parameters of the monitoring data, To standardize the data weight parameters, the real-time monitoring output expression of the module is: ,in, To monitor the output results in real time, The Softmax activation function is used. For linear transformation operations, This is the output layer weight matrix. These are the output layer bias parameters.

[0017] Further, step S3 includes the following sub-steps: S31, extracting standardized traffic data sequences from the multi-source traffic data fusion engine, dividing the sequences into spatiotemporal dimensions, determining the spatial region range corresponding to each time segment, and forming spatiotemporal block data; S32, inputting the spatiotemporal block data into the encoder of the traffic situation VAE dynamic model, extracting local features from the spatiotemporal block data through the convolutional layer of the encoder, and then mapping the local features into high-dimensional feature vectors through a fully connected layer; S33, calculating the mean and variance of the high-dimensional feature vectors using the probability distribution layer of the encoder, constructing a probability distribution of the traffic situation latent vector based on the calculated mean and variance, and sampling the traffic situation latent vector from this probability distribution; S34, inputting the traffic situation latent vector into the decoder of the traffic situation VAE dynamic model, increasing the dimension of the traffic situation latent vector through the deconvolutional layer of the decoder, and then generating a preliminary traffic situation assessment result with the same dimension as the original standardized traffic data sequence through the output layer.

[0018] Further, step S4 includes the following sub-steps: S41, receiving the preliminary traffic situation assessment results output by the traffic situation VAE dynamic model, converting the data format of the results into a matrix format recognizable by the flow field mutation Transformer model, and obtaining the initial input matrix; S42, performing position encoding processing on the initial input matrix, calculating the corresponding position encoding vector according to the temporal and spatial positions of each element in the matrix, adding the position encoding vector to the elements of the initial input matrix, and obtaining the input matrix with position information; S43, inputting the input matrix with position information into the multi-head attention layer of the flow field mutation Transformer model, calculating the query, key, and value matrices of the input matrix through multiple parallel attention heads, and then concatenating the output results of each attention head to obtain the multi-head attention output; S44, inputting the multi-head attention output into the feedforward neural network layer of the flow field mutation Transformer model, performing nonlinear transformation on the multi-head attention output through the feedforward neural network to enhance the model's ability to fit complex features, outputting a traffic flow field spatiotemporal feature matrix, and determining the location and range of the traffic flow field mutation region based on this matrix.

[0019] Further, step S5 includes the following sub-steps: S51, acquiring traffic flow field changeover region information identified by the flow field changeover Transformer model, determining heterogeneous traffic sensing nodes related to the changeover region based on this information, and collecting the raw traffic data collected by these nodes; S52, performing feature extraction on the collected raw traffic data, extracting vehicle density, driving speed, and number of stops from the raw data of each node to form a node feature vector set; S53, inputting the node feature vector set into the node feature matching module of the heterogeneous node traffic sensing algorithm, determining the node combination with similar features by calculating the cosine similarity between the feature vectors of different nodes, and establishing the data association relationship between nodes; S54, using the data collaborative verification module of the heterogeneous node traffic sensing algorithm to perform cross-verification on the data of associated nodes, removing abnormal data points, completing missing data, and generating refined monitoring data of the traffic flow field changeover region.

[0020] A deep learning-based intelligent transportation management multimodal digital monitoring system includes:

[0021] The multi-source traffic data acquisition and transmission unit is connected to the heterogeneous traffic sensing devices deployed at the road network calibration nodes. It receives vehicle trajectory data, road cross-sectional flow data, traffic signal phase data, road surface condition data and meteorological impact data collected by the heterogeneous traffic sensing devices, and transmits the collected data to the multi-source traffic data fusion processing unit.

[0022] The multi-source traffic data fusion processing unit is connected to the multi-source traffic data acquisition and transmission unit. It receives the transmitted multi-source traffic data, performs feature extraction and correlation mapping on the data, generates a standardized traffic data sequence, and transmits the standardized traffic data sequence to the traffic situation VAE dynamic analysis unit.

[0023] The traffic situation VAE dynamic analysis unit is connected to the multi-source traffic data fusion processing unit. It receives standardized traffic data sequences, encodes and decodes the data through the traffic situation VAE dynamic model, outputs preliminary traffic situation assessment results, and transmits the preliminary traffic situation assessment results to the flow field change Transformer identification unit.

[0024] The Transformer flow field change identification unit is connected to the traffic situation VAE dynamic analysis unit. It receives the preliminary traffic situation assessment results, captures the spatiotemporal dependencies in the data through the Transformer flow field change model, generates a traffic flow field spatiotemporal feature matrix, identifies traffic flow field change areas, and transmits the change area information to the heterogeneous node traffic perception and analysis unit.

[0025] The heterogeneous node traffic perception and analysis unit is connected to the flow field change Transformer identification unit. It receives information on traffic flow field change areas, uses the heterogeneous node traffic perception algorithm to perform correlation analysis and collaborative verification on the relevant data of the change areas, generates refined monitoring data, and transmits the refined monitoring data to the multimodal digital monitoring output unit.

[0026] The multimodal digital monitoring output unit is connected to the heterogeneous node traffic perception and analysis unit. It receives refined monitoring data, constructs a multi-dimensional monitoring view, monitors the traffic operation status of the road network in real time, and outputs the monitoring results.

[0027] Beneficial effects:

[0028] This invention proposes a deep learning-based multimodal digital monitoring method for intelligent transportation management. It collects various types of data by deploying heterogeneous traffic sensing devices and establishes spatiotemporal correlations between different data sources using a multi-source traffic data fusion engine. This fully extracts the spatiotemporal features of the data, generating high-quality standardized traffic data sequences. This completely solves the problems of insufficient multi-source traffic data processing capabilities and inaccurate data support in existing technologies, providing a reliable data foundation for subsequent situation assessment. The standardized data is then encoded and decoded using a traffic situation VAE dynamic model to output accurate preliminary traffic situation assessment results. Combined with a flow field mutation Transformer model to capture the spatiotemporal dependencies of the data, it quickly identifies traffic flow field mutation regions. Simultaneously, heterogeneous node traffic sensing algorithms are used to perform correlation analysis and collaborative verification on the data in mutation regions, generating high-quality, refined monitoring data. This effectively compensates for the shortcomings of existing technologies in traffic flow field mutation identification and refined monitoring capabilities. Finally, a multi-dimensional monitoring view is constructed through a multimodal digital monitoring module, enabling real-time monitoring of road network traffic operation status. This significantly improves the real-time performance, accuracy, and comprehensiveness of traffic monitoring, providing a scientific basis for transportation management decisions and meeting the needs of refined management under complex road networks. Attached Figure Description

[0029] Figure 1 This is a flowchart of the method steps of the present invention;

[0030] Figure 2 This is a diagram showing the unit composition for implementing the method of the present invention. Detailed Implementation

[0031] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] like Figure 1 As shown, the deep learning-based multimodal digital monitoring method for intelligent transportation management includes:

[0033] Step S1: Collect multi-source traffic data by using heterogeneous traffic sensing devices deployed at road network calibration nodes. The multi-source traffic data includes vehicle trajectory data, road cross-section flow data, traffic signal phase data, road surface condition data, and meteorological impact data. Transmit the collected multi-source traffic data to the multi-source traffic data fusion engine.

[0034] Specifically, the implementation of step S1 requires the deployment of heterogeneous traffic sensing equipment at key nodes of the road network. The equipment types include high-definition video detectors, microwave radar detectors, coil detectors, and meteorological sensors. The high-definition video detectors are deployed at intervals of 500 to 800 meters, the microwave radar detectors are deployed at intervals of 1,000 to 1,500 meters, the coil detectors are embedded in the lanes at intervals of 200 to 300 meters, and the meteorological sensors are installed on the guardrails on both sides of the road or on top of traffic signal poles. The multi-source traffic data collected by these devices specifically includes: vehicle trajectory data, collected 10 times per second, recording the longitude, latitude, instantaneous speed, and direction of travel; road cross-section flow data, calculating the total number of vehicles passing through the cross-section every 5 minutes, distinguishing between small, medium, and large vehicles; traffic signal phase data, collecting the duration of red, green, and yellow lights in real time with millisecond-level accuracy; road surface condition data, detecting the dryness, wetness, water accumulation, or icing status of the road surface every 30 minutes; and meteorological impact data, recording rainfall, visibility, and wind speed every 15 minutes. All collected data is transmitted to the multi-source traffic data fusion engine via 4G or 5G wireless communication modules at a transmission rate of no less than 10Mbps, with a data packet loss rate controlled within 0.1%, ensuring that the data is delivered to subsequent processing stages in real time and completely.

[0035] Step S2: Use a multi-source traffic data fusion engine to extract features and perform association mapping on the received multi-source traffic data, establish spatiotemporal associations between different types of traffic data, and generate standardized traffic data sequences.

[0036] Specifically, in step S2, after receiving the data, the multi-source traffic data fusion engine first activates the feature extraction module, employing corresponding extraction methods for different types of data: for vehicle trajectory data, it extracts the average speed, speed fluctuation amplitude, and trajectory deviation of each vehicle over a continuous 5-minute period; for road cross-section flow data, it extracts the peak flow, flow rate, and proportion of different vehicle types within a 5-minute time period; for traffic signal phase data, it extracts the time proportion of each phase and phase switching interval within a signal cycle; for road surface condition data, it extracts the duration of road surface condition and the frequency of condition switching; for meteorological impact data, it extracts the cumulative rainfall, visibility change trend, and duration of wind force level. After feature extraction is completed, the association mapping module is activated to establish association relationships between different data based on spatiotemporal coordinates. For example, it associates vehicle trajectory data and road cross-section flow data for the same time period and road segment, and associates road surface condition data with meteorological impact data. Simultaneously, it combines the road network topology to establish spatial associations for traffic data of adjacent road segments and temporal associations for data of the same type over consecutive time periods. Through the above processing, standardized traffic data sequences are generated, with a uniform time interval of 5 minutes. Each sequence contains 28 feature dimensions, covering the core features of various types of data. The data format adopts JSON, which facilitates subsequent model reading and processing.

[0037] Step S3: Input the standardized traffic data sequence into the traffic situation VAE dynamic model. The encoder of the model performs hierarchical feature encoding on the standardized traffic data sequence to obtain the traffic situation latent vector. Then, the decoder of the model reconstructs the traffic situation latent vector and outputs the preliminary traffic situation assessment results.

[0038] Specifically, step S3 first divides the standardized traffic data sequence into a training segment and a real-time processing segment according to time sequence. The training segment contains sequence data from the past 72 hours, while the real-time processing segment contains sequence data from the current hour and the next hour. The standardized traffic data sequence from the real-time processing segment is input into the traffic situation VAE dynamic model. The model first performs hierarchical feature encoding through an encoder, which contains three convolutional layers and two fully connected layers. The first convolutional layer uses 3×3 convolutional kernels (32 kernels) with a stride of 1 to capture local features of the input sequence. The second convolutional layer uses 3×3 convolutional kernels (64 kernels) with a stride of 1 to further extract deep local features. The third convolutional layer uses 5×5 convolutional kernels (128 kernels) with a stride of 2 to expand the feature receptive field. The first fully connected layer then maps the features output by the convolutional layers into a 256-dimensional vector, and the second fully connected layer maps them into a 64-dimensional traffic situation latent vector. At the same time, the mean and variance of the latent vector are calculated to construct the probability distribution of the latent vector. After encoding, the model reconstructs the traffic situation latent vector through a decoder. The decoder contains two fully connected layers and three deconvolutional layers. The first fully connected layer maps the 64-dimensional latent vector to a 256-dimensional vector, and the second fully connected layer maps it to a vector that matches the dimension of the encoder input features. The first deconvolutional layer uses 5×5 deconvolution kernels, with 128 kernels and a stride of 2 to increase the feature dimension. The second deconvolutional layer uses 3×3 deconvolution kernels, with 64 kernels and a stride of 1. The third deconvolutional layer uses 3×3 deconvolution kernels, with 32 kernels and a stride of 1. Finally, it outputs preliminary traffic situation assessment results consistent with the dimensions of the standardized traffic data sequence. The assessment results include three core indicators: traffic flow level, speed level, and congestion risk level. Each indicator is divided into 5 levels, achieving preliminary quantification of the traffic situation.

[0039] Step S4: Input the preliminary traffic situation assessment results into the flow field abrupt change Transformer model, capture the spatiotemporal dependencies in the preliminary traffic situation assessment results through the model's multi-head attention mechanism, generate a traffic flow field spatiotemporal feature matrix in combination with the location encoding module, and identify the traffic flow field abrupt change region based on the traffic flow field spatiotemporal feature matrix.

[0040] Specifically, step S4 first performs data preprocessing on the preliminary traffic situation assessment results, converting various level indicators in the assessment results into numerical data. Among them, traffic flow level is divided into 1 to 5, corresponding to below 200 vehicles / hour, 200-400 vehicles / hour, 400-600 vehicles / hour, 600-800 vehicles / hour, and above 800 vehicles / hour, respectively; speed level is divided into 1 to 5, corresponding to above 60 km / h, 40-60 km / h, 30-40 km / h, 20-30 km / h, and below 20 km / h, respectively; and congestion risk level is divided into 1 to 5, corresponding to very low, low, medium, high, and very high, respectively. The converted data forms an initial input matrix, with the number of rows in the matrix being the number of sequences in the real-time processing segment, and the number of columns being 3 (corresponding to the three types of indicators). The initial input matrix is ​​fed into the Transformer model for flow field mutation. The model first adds positional information to the matrix through a positional encoding module. The positional encoding is generated using sine and cosine functions, with the time dimension encoding period ranging from 100 to 10000, and the spatial dimension encoding based on road segment numbers, ensuring the uniqueness of the positional information for each element. Next, the model initiates a multi-head attention mechanism with eight attention heads. Each attention head calculates the query, key, and value matrices of the input matrix. The query and key matrices are multiplied by a dot product to calculate the attention weights. After normalization using the Softmax function, the weights are multiplied by the value matrix to obtain the output of a single attention head. The outputs of the eight attention heads are concatenated and subjected to a linear transformation to obtain the multi-head attention output. This multi-head attention output is then fed into a feedforward neural network layer. The feedforward neural network contains two linear transformation layers, with a ReLU activation function in between. The first linear transformation layer maps the input dimension from 24 dimensions (8 attention heads × 3 indicators) to 128 dimensions, and the second linear transformation layer maps it back to 24 dimensions, enhancing the model's ability to fit complex features. After the above processing, the model outputs a traffic flow field spatiotemporal feature matrix. The number of rows in the matrix is ​​the same as the initial input matrix, and the number of columns is 24. Based on this matrix, the feature difference values ​​of adjacent time periods and adjacent road segments are calculated. When the difference value exceeds the preset threshold (traffic flow difference threshold is 200 vehicles / hour, speed difference threshold is 15 km / h, and congestion risk difference threshold is level 2), it is determined to be a traffic flow field change area. At the same time, the road segment number, change start time, and change type (flow change, speed change, or congestion risk change) of the change area are recorded.

[0041] Step S5: Use the heterogeneous node traffic perception algorithm to perform heterogeneous node data association analysis on the multi-source traffic data corresponding to the traffic flow field change area. Through the node feature matching module and data collaborative verification module of the algorithm, generate refined monitoring data of the traffic flow field change area.

[0042] Specifically, step S5 first identifies the traffic flow field abrupt change area information based on the Transformer model, and then determines the heterogeneous traffic sensing devices involved in that area. These devices include high-definition video detectors, microwave radar detectors, loop detectors, and weather sensors, typically numbering 3 to 5. The exact number depends on the length of the road segment in the abrupt change area, with one core sensing device corresponding to every 500 meters of road segment. The raw traffic data collected by these devices during the abrupt change period (30 minutes before and after the start of the abrupt change) is retrieved using the device numbers. This raw data includes the raw coordinate data of vehicle trajectories, the raw pulse data from the loop detectors, the raw echo data from the microwave radar, and the raw measurement data from the weather sensors. A heterogeneous node traffic perception algorithm is used to process these raw data. The algorithm first activates the node feature matching module to extract features from the raw data of each sensing device: for high-definition video detectors, vehicle contour features, vehicle count features, and vehicle trajectory features are extracted; for microwave radar detectors, vehicle distance features, speed features, and radar reflection intensity features are extracted; for coil detectors, pulse frequency features, pulse duration features, and vehicle passage time interval features are extracted; for meteorological sensors, raw rainfall values, raw visibility values, and raw wind force values ​​are extracted. After feature extraction, the similarity of features between different devices is calculated using the cosine similarity calculation method, with a similarity threshold set to 0.7. When the feature similarity between two devices exceeds the threshold, a data association relationship is established between the two, forming an associated device group. Next, the algorithm activates the data collaborative verification module to cross-compare data of the same type within the associated device group. For example, it compares the vehicle counts from the high-definition video detector with those from the loop detector. When the difference exceeds 5%, abnormal data is identified. The abnormal data is corrected using a weighted average method (weights are determined according to device accuracy, with a weight of 0.4 for the high-definition video detector and 0.6 for the loop detector). For missing data, interpolation of synchronous data from adjacent devices is used to complete the data, with an interpolation interval of 1 minute. After feature matching and collaborative verification, refined monitoring data for traffic flow change areas is generated. The data includes vehicle flow, average speed, vehicle density, road surface conditions, and weather conditions every minute during the change period. The data accuracy is improved by more than 30% compared to the original data, providing accurate data support for the subsequent construction of monitoring views.

[0043] Step S6: Based on the refined monitoring data of traffic flow field change areas, construct a multi-dimensional monitoring view through the intelligent transportation management multi-modal digital monitoring module to monitor the traffic operation status of the road network in real time.

[0044] Specifically, step S6 first collects refined monitoring data of traffic flow field abrupt change areas. The data time range is 30 minutes before and after the start time of the abrupt change, with a 1-minute time interval. The data dimensions include vehicle flow, average speed, vehicle density, road surface conditions (dry, wet, waterlogged, icy), and meteorological conditions (rainfall, visibility, wind speed), totaling 8 core data dimensions. Each data point includes a corresponding timestamp and data value (road surface conditions and meteorological conditions have been converted into numerical codes). This refined monitoring data is then input into the intelligent transportation management multimodal digital monitoring module. The module includes four core components: a data import interface, a data processing unit, a view generation unit, and an output unit. The data import interface uses an API interface, supporting real-time data streaming import with an import rate of no less than 5Mbps, ensuring that data enters the processing stage in real time. The data processing unit classifies and processes the imported refined monitoring data: for numerical data such as vehicle flow, average speed, and vehicle density, time series smoothing is performed using the moving average method with a window size of 5 minutes to eliminate data fluctuation noise; for categorized data such as road surface condition and meteorological conditions, the duration of the condition and the frequency of condition switching are statistically analyzed to determine the changing trends of road surface condition and meteorological conditions in areas of sudden change. After processing, the view generation unit constructs multi-dimensional monitoring views, including time-series trend views, spatial distribution views, and multi-factor correlation views. The time-series trend view uses time as the horizontal axis and vehicle flow, average speed, and vehicle density as the vertical axes, generating three trend curves in red, blue, and green, clearly showing the data's trend over time. The spatial distribution view uses a road network map of abrupt change areas as a basis, marking the congestion level of different road segments with different colors (green, yellow, orange, and red), with darker colors indicating more severe congestion. It also marks the location of sensing devices and data collection points. The multi-factor correlation view uses average vehicle speed as the vertical axis and rainfall and road surface condition as the horizontal axes (road surface condition converted to numerical values), generating a scatter plot to show the relationship between speed, rainfall, and road surface condition. All views are displayed through the output unit, which supports large-screen display, computer terminal display, and mobile terminal display. The large-screen display resolution is no less than 1920×1080, with a refresh rate of 1 minute / time. Computer terminals and mobile terminals access the data via web pages or apps, with data transmission latency controlled within 2 seconds. By displaying multi-dimensional monitoring views in real time, the traffic operation status of the road network can be monitored in real time. Managers can intuitively grasp the traffic conditions in areas of sudden change, formulate traffic diversion strategies in a timely manner, and improve traffic management efficiency.

[0045] Preferably, the expression for the traffic situation VAE dynamic model is: ,in, The loss function of the traffic situation VAE dynamic model is... Let be the posterior probability distribution of the model encoder. For encoder network parameters, This is the latent vector for traffic situation. To standardize traffic data sequences, Let be the likelihood probability distribution of the model decoder. For decoder network parameters, Let be the prior probability distribution of the latent vector. The KL divergence function is used; the model also incorporates a traffic situation time-series weighting factor. Construct the time-series optimization loss function: ,in, The length of the time series. For the first Latent vector of traffic situation at any given time. For the first Standardized traffic data sequences for each time period.

[0046] Specifically, when implementing the traffic situation VAE dynamic model, a basic loss function is first constructed. This function contains two core parts: one is a loss term based on data reconstruction probability, used to measure the difference between the model decoder output and the original standardized traffic data sequence, ensuring that the reconstructed data can accurately reflect the characteristics of the original data; the other is a loss term based on probability distribution difference, used to control the deviation between the traffic situation latent vector probability distribution output by the model encoder and the preset prior probability distribution, avoiding the latent vector distribution being too discrete and affecting the accuracy of situation assessment. Based on the basic loss function, a traffic situation time-series weight factor is introduced. This factor is dynamically adjusted according to the importance of traffic data in different time segments. For example, the weight factor for the morning peak (7:00-9:00) and evening peak (17:00-19:00) periods is set to 1.5, for off-peak periods (9:00-17:00, 19:00-22:00) it is set to 1.0, and for nighttime periods (22:00-7:00 the next day) it is set to 0.8. A time-series optimized loss function is constructed through the time-series weight factor. During model training, a 72-hour standardized traffic data sequence was used as the training sample. The batch size was set to 32, and the number of iterations was 100. After each training round, the temporal optimization loss function value was calculated. Training was stopped when the change in the loss function value was less than 0.001 for five consecutive rounds, and the final model parameters were determined. This model improves its adaptability to traffic conditions at different times through loss function optimization, making the output traffic condition latent vector more accurate. This, in turn, improves the reliability of the preliminary traffic condition assessment results and provides a high-quality data foundation for subsequent flow field abrupt change identification.

[0047] Preferably, the expression for generating the spatiotemporal feature matrix of the traffic flow field in the Transformer model of the flow field mutation is: ,in, The traffic flow field spatiotemporal characteristic matrix, For multi-head attention computation function, For querying the matrix, The key matrix, Let be a value matrix, and These are the weight parameters for the query, key, and value matrices, respectively. This is a preliminary traffic situation assessment result. For position encoding functions, For time dimension parameters, The spatial dimension parameter is used; the expression for identifying abrupt changes in the flow field in the model is as follows: ,in, The results of traffic flow field abrupt change identification. For Sigmoid activation function, Conv1d This is a one-dimensional convolution operation. Weight matrix for mutation identification, Bias parameters for mutation identification.

[0048] Specifically, the first step in implementing the Transformer model for traffic flow abrupt changes is to generate a spatiotemporal feature matrix of the traffic flow field. The preliminary traffic situation assessment results are first converted into three types of matrices: query, key, and value. The weight parameters of these three types of matrices are determined through iterative optimization using training data. During training, historical data containing traffic flow abrupt change events from the past 30 days are used, with a learning rate set to 0.001. Gradient descent is used to adjust the weight parameters to ensure the matrix effectively reflects traffic situation characteristics. The location encoding module generates encoding vectors based on both the time and spatial dimensions. The time dimension encoding period is set according to the traffic data update frequency, with each 5-minute interval corresponding to a period of 100 to 10,000 time units. The spatial dimension encoding is generated based on road segment numbers, with each road segment assigned a unique encoding value to ensure accurate differentiation of traffic data from different spatiotemporal locations. After generating the feature matrix, the model enters the flow field abrupt change identification stage. First, a one-dimensional convolution operation is used to enhance the feature matrix. The convolution kernel size is set to 3, the number of kernels is 64, and the stride is 1. Then, an activation function is used to convert the convolution output into a value between 0 and 1 to determine whether a flow field abrupt change exists. The weight matrix and bias parameters for mutation identification are determined through training. The training samples include normal traffic flow data and mutated traffic flow data, with the ratio of the two types of data set to 3:1. Training is stopped when the model's recognition accuracy stabilizes above 95%. Finally, the mutation region is determined by judging whether the output value is greater than 0.5, which greatly improves the accuracy and efficiency of flow field mutation identification.

[0049] Preferably, the heterogeneous node data association analysis expression of the heterogeneous node traffic perception algorithm is: ,in, The correlation coefficient for heterogeneous node data. This represents the number of heterogeneous sensing nodes. For the number of traffic data types, For the first The sensing node and the first Association weights of traffic-like data This is the function for calculating feature similarity. For the first The feature vector of each sensing node For the first The feature vector of traffic data; the expression for generating refined monitoring data by the algorithm is: ,in, To refine the monitoring data, For feature concatenation function, Optimize the weight matrix for the data. For residual connection functions, This is the data co-verification error vector.

[0050] Specifically, the heterogeneous node traffic perception algorithm first calculates the correlation coefficient of heterogeneous node data. This coefficient calculation requires statistics on the number of heterogeneous sensing nodes and the number of traffic data types. Typically, 3 to 5 sensing nodes are deployed in a traffic flow field abrupt change area, covering 4 to 5 types of traffic data. The correlation weight is determined based on the accuracy of the node equipment. The correlation weight for high-definition video detectors and microwave radar detectors is set to 0.3, for coil detectors to 0.25, and for meteorological sensors to 0.15, ensuring that high-precision equipment data accounts for a higher proportion in the correlation analysis. The feature similarity calculation uses the cosine similarity method. During calculation, the feature vectors of each node are first standardized to eliminate the influence of dimensions, and then the cosine value of the angle between the vectors is calculated. The similarity threshold is set to 0.7. When the similarity between two nodes exceeds the threshold, it is determined that their data have a strong correlation and they are grouped into the same associated device group. When generating refined monitoring data, the feature vectors of nodes within the associated device group are first weighted. The weight parameters are optimized through training data. The training samples include traffic data from different weather conditions and different time periods. The training objective is to ensure that the error between the output data and the actual traffic conditions is less than 5%. Simultaneously, a residual connection mechanism is introduced to process the error vector generated during the data collaborative verification process, supplement the missing information in the feature vector, and finally integrate all processed features through feature splicing operation to generate refined monitoring data. The data time interval is set to 1 minute to ensure that the data can reflect the changes in traffic conditions in the sudden change area in real time.

[0051] Preferably, the feature extraction and association mapping expression of the multi-source traffic data fusion engine is as follows: ,in, For the fused traffic feature vector, LayerNorm ( () represents the layer normalization operation. It is a feedforward neural network. For the feature concatenate operation, This is the feature vector of the vehicle's driving trajectory data. This is the feature vector of road cross-section traffic flow data. For traffic signal phase data feature vectors, For road surface condition data feature vectors, The feature vector represents the meteorological impact data; the standardized traffic data sequence generation expression of the engine is: ,in, To standardize traffic data sequences, To fuse the mean of the feature vectors, To fuse the standard deviation of the feature vectors, Adjust the weight matrix for standardization.

[0052] Specifically, the multi-source traffic data fusion engine first performs feature extraction and association mapping. When extracting features from various types of traffic data, the feature dimensions for vehicle trajectory data are set to 6 (average speed, speed fluctuation amplitude, etc.), road cross-section flow data features are set to 5 (flow peak, vehicle type ratio, etc.), traffic signal phase data features are set to 4 (phase time ratio, switching interval, etc.), road surface condition data features are set to 3 (state duration, switching frequency, etc.), and meteorological impact data features are set to 3 (cumulative rainfall, visibility trend, etc.), for a total of 21 basic feature dimensions. After feature concatenation, the data is input into a feedforward neural network for feature fusion. The neural network contains two hidden layers: the first layer has 128 neurons, and the second layer has 64 neurons. The activation function is ReLU. The neural network maps the 21 basic features into a 28-dimensional fused feature vector. Then, layer normalization is used to eliminate the dimensional differences between features, ensuring that the values ​​of each dimension of the fused feature vector are distributed between 0 and 1. When generating standardized traffic data sequences, the mean and standard deviation of the fused feature vectors are first calculated. The mean is determined by averaging all fused feature vectors over the past 24 hours, and the standard deviation is calculated using a rolling standard deviation method with a window size of 12 time units (one unit every 5 minutes, for a total of 60 minutes) to ensure that the mean and standard deviation are updated in real time. The standardized adjustment weight matrix is ​​determined based on the traffic flow of different road segments, with a weight of 1.2 for main roads, 1.0 for secondary roads, and 0.8 for local roads. The resulting standardized data sequence effectively integrates multi-source data information, providing a unified data format for subsequent model analysis.

[0053] Preferably, the multi-dimensional monitoring view construction expression of the intelligent transportation management multimodal digital monitoring module is as follows: ,in, For multi-dimensional monitoring view, For view projection functions, These are the weighting parameters for traffic flow field characteristics. To refine the weighting parameters of the monitoring data, To standardize the data weight parameters, the real-time monitoring output expression of the module is: ,in, To monitor the output results in real time, The Softmax activation function is used. For linear transformation operations, This is the output layer weight matrix. These are the output layer bias parameters.

[0054] Specifically, the multimodal digital monitoring module for intelligent transportation management first constructs a multi-dimensional monitoring view. The view projection function uses orthogonal projection to project three types of data—the traffic flow spatiotemporal feature matrix, refined monitoring data, and standardized traffic data sequences—to the same dimensional space. The projection dimension is set to 24 to ensure that all three types of data can be displayed in the same view. The weight parameters for the three types of data are determined based on their importance: the traffic flow spatiotemporal feature matrix has a weight of 0.4, the refined monitoring data has a weight of 0.35, and the standardized traffic data sequence has a weight of 0.25, ensuring that key data has a higher proportion in the view. When generating real-time monitoring output results, the feature vectors of the multi-dimensional monitoring view are first mapped to 10-dimensional vectors through a linear transformation operation. The weight matrix and bias parameters of the linear transformation are determined through training. The training samples include monitoring data for different traffic states (smooth, slow, congested). Training stops when the model's accuracy in classifying traffic states reaches over 92%. Then, an activation function converts the linear transformation output into a probability distribution, where each probability value corresponds to a traffic state, with the highest probability value corresponding to the current traffic state. The output layer weight matrix is ​​initialized using the Xavier method to ensure gradient stability during training. The initial value of the bias parameter is set to 0. Through iterative optimization using the backpropagation algorithm, the final output real-time monitoring results can accurately reflect the traffic operation status of the road network and provide direct basis for traffic management decisions.

[0055] Preferably, step S3 includes the following sub-steps: S31, extracting a standardized traffic data sequence from the multi-source traffic data fusion engine, dividing the sequence into spatiotemporal dimensions, determining the spatial region range corresponding to each time segment, and forming spatiotemporal block data; S32, inputting the spatiotemporal block data into the encoder of the traffic situation VAE dynamic model, extracting local features from the spatiotemporal block data through the convolutional layer of the encoder, and then mapping the local features into high-dimensional feature vectors through a fully connected layer; S33, calculating the mean and variance of the high-dimensional feature vectors using the probability distribution layer of the encoder, constructing a probability distribution of the traffic situation latent vector based on the calculated mean and variance, and sampling the traffic situation latent vector from this probability distribution; S34, inputting the traffic situation latent vector into the decoder of the traffic situation VAE dynamic model, increasing the dimension of the traffic situation latent vector through the deconvolutional layer of the decoder, and then generating a preliminary traffic situation assessment result with the same dimension as the original standardized traffic data sequence through the output layer.

[0056] Specifically, in step S3, S31 is executed first, extracting standardized traffic data sequences from the multi-source traffic data fusion engine. This data is divided into time segments of 5 minutes each, and spatially into spatial regions of 2-kilometer road segments based on road network topology, forming spatiotemporal block data. Each block contains all traffic feature information for the corresponding time period and region. Next, in step S32, the spatiotemporal block data is input into the encoder of the traffic situation VAE dynamic model. The encoder's first convolutional layer uses a 3×3 kernel, 32 kernels, and a stride of 1 to extract local features from the block data and filter redundant information. The second convolutional layer also uses a 3×3 kernel, 64 kernels, and a stride of 1 to further deepen feature extraction. The third convolutional layer uses a 5×5 kernel, 128 kernels, and a stride of 2 to expand the receptive field and capture broader spatiotemporal correlations. Finally, the first fully connected layer maps the features output from the convolutional layers to a 256-dimensional high-dimensional feature vector. The first layer compresses the data into a 64-dimensional feature vector using a second fully connected layer. Then, in step S33, the mean and variance of the 64-dimensional feature vector are calculated using the probability distribution layer of the encoder. Based on the mean and variance, a traffic situation latent vector probability distribution conforming to a normal distribution is constructed. Traffic situation latent vectors are obtained from this distribution through random sampling to ensure that the latent vectors are both representative and random. Finally, in step S34, the traffic situation latent vectors are input into the decoder. The first deconvolution layer of the decoder uses a 5×5 deconvolution kernel, 128 samples, and a stride of 2 to increase the feature dimension. The second deconvolution layer uses a 3×3 deconvolution kernel, 64 samples, and a stride of 1. The third deconvolution layer uses a 3×3 deconvolution kernel, 32 samples, and a stride of 1 to gradually restore the feature dimension. Then, the output layer generates a preliminary traffic situation assessment result with the same dimension as the original standardized traffic data sequence. This result includes five levels of flow rate, speed, and congestion risk indicators, providing an accurate situational basis for subsequent flow field abrupt change identification.

[0057] Preferably, step S4 includes the following sub-steps: S41, receiving the preliminary traffic situation assessment result output by the traffic situation VAE dynamic model, converting the result into a matrix format recognizable by the flow field mutation Transformer model to obtain an initial input matrix; S42, performing position encoding processing on the initial input matrix, calculating the corresponding position encoding vector based on the temporal and spatial positions of each element in the matrix, adding the position encoding vector to the elements of the initial input matrix to obtain an input matrix with position information; S43, inputting the input matrix with position information into the multi-head attention layer of the flow field mutation Transformer model, calculating the query, key, and value matrices of the input matrix through multiple parallel attention heads, and then concatenating the output results of each attention head to obtain the multi-head attention output; S44, inputting the multi-head attention output into the feedforward neural network layer of the flow field mutation Transformer model, performing nonlinear transformation on the multi-head attention output through the feedforward neural network to enhance the model's ability to fit complex features, outputting a traffic flow field spatiotemporal feature matrix, and determining the location and range of the traffic flow field mutation region based on this matrix.

[0058] Specifically, in step S4, S41 is first performed to receive the preliminary traffic situation assessment results output by the traffic situation VAE dynamic model. The data format conversion module converts the level indicators in the assessment results into numerical data, forming an initial input matrix with the same number of rows as the number of real-time processing time-segment sequences and 3 columns (corresponding to flow rate, speed, and congestion risk). This ensures that the matrix format meets the input requirements of the flow field abrupt change Transformer model. Next, S42 is executed to perform position encoding processing on the initial input matrix. Temporal position encoding calculates the encoding vector using a sine and cosine function with a period range of 100 to 10000 based on the time sequence of each element. Spatial position encoding generates a unique encoding vector based on the number of the road segment corresponding to each element. The two types of encoding vectors are superimposed on the corresponding elements of the initial input matrix to obtain an input matrix with position information, enabling the model to recognize the spatiotemporal attributes of the data. Then, S43 is performed to input the input matrix with position information into the multi-head attention layer of the model. This layer is configured with 8 parallel attention heads. Each attention head calculates its query, key, and value matrices. Attention weights are calculated by the dot product of the query and key matrices, normalized using the Softmax function, and multiplied by the value matrix to obtain the output of a single attention head. The outputs of the eight attention heads are concatenated column-wise and then compressed into a fixed-dimensional multi-head attention output through a linear transformation, enhancing the model's ability to capture spatiotemporal dependencies. Finally, S44 is executed, inputting the multi-head attention output into a feedforward neural network layer. This layer contains two linear transformation layers. The first layer maps the input dimension from 24 dimensions (8 attention heads × 3 indicators) to 128 dimensions, introducing nonlinear features through the ReLU activation function. The second layer maps the 128-dimensional features back to 24 dimensions, outputting a traffic flow spatiotemporal feature matrix. Based on this matrix, the feature difference values ​​of adjacent time periods and road segments are calculated. When the difference value exceeds a preset threshold (flow rate of 200 vehicles / hour, speed of 15 km / h, congestion risk level 2), the location and range of the traffic flow abrupt change area are determined, providing accurate information on the abrupt change area for subsequent refined monitoring.

[0059] Preferably, step S5 includes the following sub-steps: S51, acquiring traffic flow field changeover region information identified by the flow field changeover Transformer model, determining heterogeneous traffic sensing nodes related to the changeover region based on this information, and collecting the raw traffic data collected by these nodes; S52, performing feature extraction on the collected raw traffic data, extracting vehicle density, driving speed, and number of stops from the raw data of each node to form a node feature vector set; S53, inputting the node feature vector set into the node feature matching module of the heterogeneous node traffic sensing algorithm, determining the node combination with similar features by calculating the cosine similarity between the feature vectors of different nodes, and establishing the data association relationship between nodes; S54, using the data collaborative verification module of the heterogeneous node traffic sensing algorithm to perform cross-verification on the data of associated nodes, removing abnormal data points, completing missing data, and generating refined monitoring data of the traffic flow field changeover region.

[0060] Specifically, in step S5, S51 is initiated first to acquire information on traffic flow field abrupt change areas identified by the Transformer model, including the abrupt change road segment number and the start time of the abrupt change. Based on the road segment number, the heterogeneous traffic sensing devices deployed in that area are determined. Typically, one core device corresponds to every 500 meters of road segment, for a total of 3 to 5 devices. The device types include high-definition video detectors, microwave radar detectors, coil detectors, and meteorological sensors. Raw traffic data collected by these devices within 30 minutes before and after the start time of the abrupt change is collected, including vehicle raw trajectory coordinates, coil pulse signals, radar echo data, and raw meteorological measurements. Next, S52 is executed to extract features from the collected raw data. Specific features are extracted for different device types: high-definition video detectors extract vehicle contour, count, and motion trajectory features; microwave radar detectors extract vehicle distance, speed, and reflection intensity features; coil detectors extract pulse frequency, duration, and vehicle passing interval features; and meteorological sensors extract raw values ​​of rainfall, visibility, and wind force. A set of node feature vectors containing multi-dimensional features is formed. Then, in step S53, the set of node feature vectors is input into the node feature matching module of the heterogeneous node traffic perception algorithm. The cosine similarity calculation method is used to calculate the similarity between feature vectors of different nodes. The similarity threshold is set to 0.7. When the similarity between two nodes exceeds the threshold, it is determined that there is a strong correlation between their data. The data association relationship between nodes is established, forming a group of associated devices to ensure the reliability of data sources. Finally, step S54 is executed. The data collaboration verification module of the algorithm is used to perform cross-verification on the same type of data in the group of associated devices. For example, the vehicle counts of video and loop detectors are compared. When the difference exceeds 5%, it is determined to be abnormal data. Weighted correction is performed according to the device accuracy (0.4 for video detectors and 0.6 for loop detectors). Missing data is filled by interpolation at 1-minute intervals of data from adjacent devices. Refined monitoring data of traffic flow change areas is generated. The data accuracy is improved by more than 30% compared with the original data, providing high-quality data support for the construction of multimodal monitoring views.

[0061] The traffic situation VAE dynamic model in this invention is a deep learning model used to hierarchically encode and decode standardized traffic data sequences and output preliminary traffic situation assessment results. Its implementation requires first extracting standardized data from a multi-source traffic data fusion engine, dividing it into spatiotemporal blocks according to 5-minute time segments and 2-kilometer spatial regions; then inputting the block data into an encoder containing three convolutional layers (using 3×3 / 3×3 / 5×5 convolutional kernels, with a number of 32 / 64 / 128 and a stride of 1 / 1 / 2) and two fully connected layers (mapped to 256-dimensional and 64-dimensional), calculating the mean and variance of the feature vectors to construct a normal distribution, and sampling to obtain the traffic situation latent vector; finally, through a decoder containing two fully connected layers and three deconvolutional layers (5×5 / 3×3 / 3×3 deconvolutional kernels, with a number of 128 / 64 / 32 and a stride of 2 / 1 / 1), the assessment results containing five levels of flow rate, speed, and congestion risk indicators are reconstructed and generated. This model aims to accurately extract the spatiotemporal characteristics of traffic data and generate a reliable preliminary situation assessment. It addresses the problem of traditional models struggling to capture dynamic changes in traffic conditions, providing a high-quality data foundation for subsequent flow field abrupt changes and improving the accuracy and adaptability of traffic situation analysis.

[0062] The Transformer model for traffic flow abrupt changes is a core model used to capture the spatiotemporal dependencies of traffic data and identify abrupt change regions in traffic flow. Its implementation requires first receiving preliminary traffic situation assessment results, converting the level indicators into a numerical initial input matrix (3 columns, corresponding to flow rate, speed, and congestion risk); then, the matrix is ​​positionally encoded (time encoding uses sine and cosine functions with periods of 100-10000, spatial encoding is based on road segment numbers), resulting in an input matrix with spatiotemporal attributes; subsequently, it is input into a multi-head attention layer containing 8 attention heads, calculating the query, key, and value matrix, and obtaining the output through dot product and Softmax, which is then linearly transformed and input into a feedforward neural network (2 linear transformation layers, dimension 24→128→24, ReLU activation), outputting a spatiotemporal feature matrix of the traffic flow field; finally, it calculates the feature difference values ​​between adjacent time periods / road segments, and those exceeding the threshold (flow rate 200 vehicles / hour, speed 15 km / h, congestion risk level 2) are identified as abrupt change regions. The model's function is to efficiently capture the spatiotemporal correlation of traffic flow and accurately locate the regions and extent of abrupt change in the flow field. This method overcomes the limitations of traditional methods in identifying sudden changes in traffic flow in real time, providing accurate data for timely response to traffic anomalies and improving the real-time and targeted nature of traffic management.

[0063] The heterogeneous node traffic perception algorithm is a key algorithm used to perform correlation analysis and collaborative verification of heterogeneous device data in traffic flow field abrupt change areas, generating refined monitoring data. Its implementation requires first acquiring information about the abrupt change area (road segment number, start time), identifying 3-5 heterogeneous devices (one per 500 meters, including video / radar / loop detectors and meteorological sensors), and collecting raw data for 30 minutes before and after the abrupt change. Then, it extracts device-specific features (vehicle outlines and counts from video, distance and speed from radar, pulse features from loop detectors, and rainfall and visibility from meteorological sensors) to form a node feature vector set. Next, it matches associated nodes using cosine similarity calculation (threshold 0.7) to form associated device groups. Finally, it cross-validates the data within each group (vehicle count differences exceeding 5% are corrected with weights of 0.4 / 0.6), and fills in missing data using interpolation at 1-minute intervals, generating refined data with an accuracy improvement of over 30%. The algorithm's function is to integrate heterogeneous device data, remove anomalies, fill in missing data, and generate high-precision monitoring data. This addresses the challenges of data collaboration and inconsistent quality from heterogeneous devices, providing reliable data support for the construction of multimodal monitoring views and ensuring the accuracy of traffic monitoring results.

[0064] The multi-source traffic data fusion engine algorithm is a fundamental algorithm used to integrate multiple types of raw traffic data and generate standardized traffic data sequences. Its implementation requires first collecting vehicle trajectory, cross-sectional flow, signal phase, road surface condition, and meteorological data through heterogeneous devices. This data is then transmitted to the engine, where various data features are extracted (6-dimensional trajectory, 5-dimensional flow, 4-dimensional phase, 3-dimensional road surface, and 3-dimensional meteorological data, totaling 21 dimensions). These features are then concatenated and input into a feedforward neural network (2 hidden layers, 128 / 64 neurons, ReLU activation), and layer normalization generates a 28-dimensional fused feature vector. Subsequently, the mean (statistics over the past 24 hours) and standard deviation (12 time unit rolling window) of the fused features are calculated, combined with road segment weights (1.2 for main roads, 1.0 for secondary roads, and 0.8 for local roads) to generate a standardized data sequence. The algorithm's function is to break down the barriers between multi-source data formats, extract core features, and unify data standards. This addresses the problem of low efficiency in subsequent analysis caused by the fragmented and inconsistent formats of traditional data processing, providing consistent data input for traffic situation VAE models and flow field abrupt change Transformer models, ensuring the efficiency and orderliness of the entire monitoring method's data processing stage.

[0065] like Figure 2As shown, the deep learning-based intelligent transportation management multimodal digital monitoring system includes: a multi-source traffic data acquisition and transmission unit, which is connected to heterogeneous traffic sensing devices deployed at road network calibration nodes. This unit receives vehicle trajectory data, road cross-sectional flow data, traffic signal phase data, road surface condition data, and meteorological impact data collected by the heterogeneous traffic sensing devices and transmits the collected data to a multi-source traffic data fusion processing unit; a multi-source traffic data fusion processing unit, connected to the multi-source traffic data acquisition and transmission unit, receives the transmitted multi-source traffic data, performs feature extraction and correlation mapping on the data, generates a standardized traffic data sequence, and transmits the standardized traffic data sequence to a traffic situation VAE dynamic analysis unit; and a traffic situation VAE dynamic analysis unit, connected to the multi-source traffic data fusion processing unit, receives the standardized traffic data sequence, encodes and decodes the data using a traffic situation VAE dynamic model, outputs a preliminary traffic situation assessment result, and transmits the preliminary traffic situation assessment result to... The system comprises the following components: a flow field abrupt change Transformer identification unit; a flow field abrupt change Transformer identification unit connected to the traffic situation VAE dynamic analysis unit; a flow field abrupt change Transformer identification unit receiving preliminary traffic situation assessment results; a flow field abrupt change Transformer model capturing spatiotemporal dependencies in the data to generate a traffic flow field spatiotemporal feature matrix; identification of traffic flow field abrupt change regions; transmission of abrupt change region information to the heterogeneous node traffic perception analysis unit; a heterogeneous node traffic perception analysis unit connected to the flow field abrupt change Transformer identification unit receiving traffic flow field abrupt change region information; a heterogeneous node traffic perception algorithm performing correlation analysis and collaborative verification on the abrupt change region-related data to generate refined monitoring data; transmission of refined monitoring data to the multimodal digital monitoring output unit; and a multimodal digital monitoring output unit connected to the heterogeneous node traffic perception analysis unit receiving refined monitoring data, constructing a multi-dimensional monitoring view, real-time monitoring of road network traffic operation status, and outputting monitoring results.

[0066] The deep learning-based multimodal digital monitoring method for intelligent transportation management offers several advantages: First, it boasts superior multi-source data processing capabilities, comprehensively collecting various data such as vehicle trajectories and road cross-sectional flow through heterogeneous traffic sensing devices. A multi-source traffic data fusion engine then extracts and maps data features, constructs spatiotemporal relationships between data, and generates standardized data sequences, providing a high-quality data foundation for subsequent analysis. Second, it offers high accuracy in traffic situation analysis, relying on specialized dynamic models to encode and decode standardized data, effectively mining situational features and outputting reliable preliminary assessment results. Third, it efficiently identifies abrupt changes in the flow field, using specific models to capture spatiotemporal dependencies and combining location encoding to generate feature matrices for rapid location of abrupt flow field changes. Fourth, it provides highly refined monitoring data, using heterogeneous node traffic sensing algorithms to perform correlation analysis and collaborative verification of data in abrupt change areas, eliminating abnormal data and supplementing missing data. Fifth, it presents multi-dimensional monitoring, constructing a multi-dimensional view through multimodal digital monitoring modules to achieve real-time and comprehensive monitoring of the road network's operational status, providing rich references for management decisions.

[0067] This method addresses the shortcomings of existing technologies in multi-source data processing capabilities. It ensures comprehensive data collection by deploying various types of heterogeneous sensing devices, and then utilizes the feature extraction and association mapping functions of a multi-source data fusion engine to break down barriers between different types of data, establish effective correlations, and generate high-quality standardized sequences, thus solving the problem of insufficient data support. Addressing the deficiencies in existing technologies for identifying and finely monitoring flow field mutations, it first uses a dynamic model to accurately analyze traffic conditions, then uses a specialized model to capture spatiotemporal dependencies to quickly identify mutation regions, and finally uses a heterogeneous node algorithm to deeply process the mutation region data, completing anomaly removal and data completion to generate finely monitored data. Simultaneously, it combines a multimodal monitoring module to achieve comprehensive presentation, completely overcoming the shortcomings of existing technologies in mutation identification and fine-grained monitoring, and meeting the needs of complex road network management.

[0068] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0069] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multimodal digital monitoring method for intelligent transportation management based on deep learning, characterized in that, include: Step S1: Collect multi-source traffic data by using heterogeneous traffic sensing devices deployed at road network calibration nodes. The multi-source traffic data includes vehicle trajectory data, road cross-section flow data, traffic signal phase data, road surface condition data, and meteorological impact data. Transmit the collected multi-source traffic data to the multi-source traffic data fusion engine. Step S2: Use a multi-source traffic data fusion engine to extract features and perform association mapping on the received multi-source traffic data, establish spatiotemporal associations between different types of traffic data, and generate standardized traffic data sequences. Step S3: Input the standardized traffic data sequence into the traffic situation VAE dynamic model. The encoder of the model performs hierarchical feature encoding on the standardized traffic data sequence to obtain the traffic situation latent vector. Then, the decoder of the model reconstructs the traffic situation latent vector and outputs the preliminary traffic situation assessment results. Step S4: Input the preliminary traffic situation assessment results into the flow field abrupt change Transformer model, capture the spatiotemporal dependencies in the preliminary traffic situation assessment results through the model's multi-head attention mechanism, generate a traffic flow field spatiotemporal feature matrix in combination with the location encoding module, and identify the traffic flow field abrupt change region based on the traffic flow field spatiotemporal feature matrix. Step S5: Use the heterogeneous node traffic perception algorithm to perform heterogeneous node data association analysis on the multi-source traffic data corresponding to the traffic flow field change area. Through the node feature matching module and data collaborative verification module of the algorithm, generate refined monitoring data of the traffic flow field change area. Step S6: Based on the refined monitoring data of traffic flow field change areas, construct a multi-dimensional monitoring view through the intelligent transportation management multi-modal digital monitoring module to monitor the traffic operation status of the road network in real time; The expression for the traffic situation VAE dynamic model is: ,in, The loss function of the traffic situation VAE dynamic model is... Let be the posterior probability distribution of the model encoder. For encoder network parameters, This is the latent vector for traffic situation. To standardize traffic data sequences, Let be the likelihood probability distribution of the model decoder. For decoder network parameters, Let be the prior probability distribution of the latent vector. The KL divergence function is used; the model also incorporates a traffic situation time-series weighting factor. Construct the time-series optimization loss function: ,in, The length of the time series. For the first Latent vector of traffic situation at any given time. For the first Standardized traffic data sequences for each time period.

2. The deep learning-based multimodal digital monitoring method for intelligent transportation management according to claim 1, characterized in that, The expression for generating the spatiotemporal feature matrix of the traffic flow field in the Transformer model of the flow field mutation is as follows: ,in, The traffic flow field spatiotemporal characteristic matrix, For multi-head attention computation function, For querying the matrix, The key matrix, Let be a value matrix, and These are the weight parameters for the query, key, and value matrices, respectively. This is a preliminary traffic situation assessment result. For position encoding functions, For time dimension parameters, The spatial dimension parameter is used; the expression for identifying abrupt changes in the flow field in the model is as follows: ,in, The results of traffic flow field abrupt change identification. For Sigmoid activation function, Conv1d This is a one-dimensional convolution operation. Weight matrix for mutation identification, Bias parameters for mutation identification.

3. The deep learning-based intelligent transportation management multimodal digital monitoring method according to claim 1, characterized in that, The heterogeneous node traffic perception algorithm's heterogeneous node data association analysis expression is as follows: ,in, The correlation coefficient for heterogeneous node data. This represents the number of heterogeneous sensing nodes. For the number of traffic data types, For the first The sensing node and the first Association weights of traffic-like data This is the function for calculating feature similarity. For the first The feature vector of each sensing node For the first The feature vector of traffic data; the expression for generating refined monitoring data by the algorithm is: ,in, To refine the monitoring data, For feature concatenation function, Optimize the weight matrix for the data. For residual connection functions, This is the data co-verification error vector.

4. The deep learning-based multimodal digital monitoring method for intelligent transportation management according to claim 1, characterized in that, The feature extraction and association mapping expression of the multi-source traffic data fusion engine is as follows: ,in, For the fused traffic feature vector, LayerNorm ( () represents the layer normalization operation. It is a feedforward neural network. For the feature concatenate operation, This is the feature vector of the vehicle's driving trajectory data. This is the feature vector of road cross-section traffic flow data. For traffic signal phase data feature vectors, For road surface condition data feature vectors, The feature vector represents the meteorological impact data; the standardized traffic data sequence generation expression of the engine is: ,in, To standardize traffic data sequences, To fuse the mean of the feature vectors, To fuse the standard deviation of the feature vectors, Adjust the weight matrix for standardization.

5. The deep learning-based multimodal digital monitoring method for intelligent transportation management according to claim 1, characterized in that, The expression for constructing the multi-dimensional monitoring view of the intelligent transportation management multimodal digital monitoring module is as follows: ,in, For multi-dimensional monitoring view, For view projection functions, These are the weighting parameters for traffic flow field characteristics. To refine the weighting parameters of the monitoring data, To standardize the data weight parameters, the real-time monitoring output expression of the module is: ,in, To monitor the output results in real time, The Softmax activation function is used. For linear transformation operations, This is the output layer weight matrix. These are the output layer bias parameters.

6. The deep learning-based multimodal digital monitoring method for intelligent transportation management according to claim 1, characterized in that, Step S3 includes the following sub-steps: S31, extracting standardized traffic data sequences from the multi-source traffic data fusion engine, dividing the sequences into spatiotemporal dimensions, determining the spatial region range corresponding to each time segment, and forming spatiotemporal block data; S32, inputting the spatiotemporal block data into the encoder of the traffic situation VAE dynamic model, extracting local features from the spatiotemporal block data through the convolutional layer of the encoder, and then mapping the local features into high-dimensional feature vectors through a fully connected layer; S33, calculating the mean and variance of the high-dimensional feature vectors using the probability distribution layer of the encoder, constructing a probability distribution of the traffic situation latent vector based on the calculated mean and variance, and sampling the traffic situation latent vector from this probability distribution; S34, inputting the traffic situation latent vector into the decoder of the traffic situation VAE dynamic model, increasing the dimension of the traffic situation latent vector through the deconvolutional layer of the decoder, and then generating a preliminary traffic situation assessment result with the same dimension as the original standardized traffic data sequence through the output layer.

7. The deep learning-based multimodal digital monitoring method for intelligent transportation management according to claim 1, characterized in that, Step S4 includes the following sub-steps: S41. Receive the preliminary traffic situation assessment results output by the traffic situation VAE dynamic model, convert the data format of the results into a matrix format recognizable by the flow field mutation Transformer model, and obtain the initial input matrix; S42. Perform position encoding processing on the initial input matrix. Calculate the corresponding position encoding vector based on the temporal and spatial positions of each element in the matrix, and add the position encoding vector to the elements of the initial input matrix to obtain an input matrix with position information; S43. Input the input matrix with position information into the multi-head attention layer of the flow field mutation Transformer model. Calculate the query, key, and value matrices of the input matrix through multiple parallel attention heads, and then concatenate the outputs of each attention head to obtain the multi-head attention output; S44. Input the multi-head attention output into the feedforward neural network layer of the flow field mutation Transformer model. Perform nonlinear transformation on the multi-head attention output through the feedforward neural network to enhance the model's ability to fit complex features, outputting a traffic flow field spatiotemporal feature matrix. Determine the location and range of the traffic flow field mutation region based on this matrix.

8. The deep learning-based multimodal digital monitoring method for intelligent transportation management according to claim 1, characterized in that, Step S5 includes the following sub-steps: S51, obtaining traffic flow field change region information identified by the flow field change Transformer model, determining heterogeneous traffic sensing nodes related to the change region based on the information, and collecting the raw traffic data collected by these nodes; S52, performing feature extraction on the collected raw traffic data, extracting vehicle density, driving speed, and number of stops from the raw data of each node to form a node feature vector set. S53. Input the node feature vector set into the node feature matching module of the heterogeneous node traffic perception algorithm. By calculating the cosine similarity between the feature vectors of different nodes, determine the combination of nodes with similar features and establish the data association between nodes. S54. Use the data collaborative verification module of the heterogeneous node traffic perception algorithm to cross-verify the data of the associated nodes, remove abnormal data points, fill in the missing data, and generate refined monitoring data of traffic flow field change areas.

9. A deep learning-based intelligent transportation management multimodal digital monitoring system, employing the deep learning-based intelligent transportation management multimodal digital monitoring method as described in any one of claims 1-8, characterized in that, include: The multi-source traffic data acquisition and transmission unit is connected to the heterogeneous traffic sensing devices deployed at the road network calibration nodes. It receives vehicle trajectory data, road cross-sectional flow data, traffic signal phase data, road surface condition data and meteorological impact data collected by the heterogeneous traffic sensing devices, and transmits the collected data to the multi-source traffic data fusion processing unit. A multi-source traffic data fusion processing unit is connected to the multi-source traffic data acquisition and transmission unit. It receives the transmitted multi-source traffic data, performs feature extraction and correlation mapping on the data, generates a standardized traffic data sequence, and transmits the standardized traffic data sequence to the traffic situation VAE dynamic analysis unit. The traffic situation VAE dynamic analysis unit is connected to the multi-source traffic data fusion processing unit. It receives standardized traffic data sequences, encodes and decodes the data through the traffic situation VAE dynamic model, outputs preliminary traffic situation assessment results, and transmits the preliminary traffic situation assessment results to the flow field change Transformer identification unit. The Transformer flow field change identification unit is connected to the traffic situation VAE dynamic analysis unit. It receives the preliminary traffic situation assessment results, captures the spatiotemporal dependencies in the data through the Transformer flow field change model, generates a traffic flow field spatiotemporal feature matrix, identifies traffic flow field change areas, and transmits the change area information to the heterogeneous node traffic perception and analysis unit. The heterogeneous node traffic perception and analysis unit is connected to the flow field change Transformer identification unit. It receives information on traffic flow field change areas, uses the heterogeneous node traffic perception algorithm to perform correlation analysis and collaborative verification on the relevant data of the change areas, generates refined monitoring data, and transmits the refined monitoring data to the multimodal digital monitoring output unit. The multimodal digital monitoring output unit is connected to the heterogeneous node traffic perception and analysis unit. It receives refined monitoring data, constructs a multi-dimensional monitoring view, monitors the traffic operation status of the road network in real time, and outputs the monitoring results.

Citation Information

Patent Citations

  • Transform-based highway traffic state estimation method

    CN117576897A

  • Highway event detection algorithm based on data situation analysis

    CN121122007A