A traffic warning method and device based on a spatio-temporal dynamic network
By acquiring multi-source traffic data, performing time and space registration, feature encoding and graph convolution processing, and creating explicit physical graphs and implicit semantic graphs, it solves the problem of low prediction accuracy in the Internet of Vehicles system and realizes high-precision traffic warning and real-time prediction.
Patent Information
- Application Number
- CN202510546427.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In the Internet of Vehicles system, due to factors such as environmental interference and vehicle movement, the traffic prediction accuracy is low.
By acquiring multi-source traffic data, registering, feature coding and normalizing the time dimension and spatial dimension, creating explicit physical maps and implicit semantic maps, using spatial graph convolution, time convolution and causal expansion convolution to extract timing-dependent features, and using spatial-temporal position coding and sparse self-attention mechanism optimization, feature fusion is performed to determine traffic warning information.
The traffic prediction accuracy is improved, especially in high-speed and high-density scenarios, the vehicle movement trajectory prediction error is reduced by 30%, the traffic flow prediction accuracy is improved by 25%, the communication link quality prediction error is ≤5%, and real-time early warning with low latency is achieved.
Smart Images

Figure CN120071632B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of traffic control systems, and particularly to a traffic warning method and device based on a spatio-temporal dynamic network. Background Art
[0002] Currently, C-V2X (Cellular Vehicle-to-Everything), as a vehicle wireless communication technology, can achieve all-round communication between vehicles and the surrounding environment (including other vehicles, technical facilities, pedestrians, and the network), and has a wide range of applications in the field of vehicle communication technology. At the same time, the GNN (Graph Neural Network) based on the vehicle networking system can obtain vehicle and surrounding information through C-V2X, so as to predict traffic flow. However, in the actual use process, due to the interference of factors such as environmental interference and vehicle movement, the prediction accuracy of the vehicle networking system is often low. Summary of the Invention
[0003] The purpose of the embodiments of this application is to provide a traffic warning method and device based on a spatio-temporal dynamic network to solve the problem of low prediction accuracy in the vehicle networking system. The specific technical solutions are as follows:
[0004] In the first aspect of the embodiments of this application, first, a traffic warning method based on a spatio-temporal dynamic network is provided. The method includes:
[0005] Obtain multi-source traffic data, where the multi-source traffic data includes vehicle-side data, roadside data, and network data;
[0006] Perform registration on the multi-source traffic data in the time dimension and the space dimension to obtain registered traffic data;
[0007] Perform feature encoding on the registered traffic data to obtain encoded data, where the encoded data includes visual data, point cloud data, and communication data;
[0008] Perform normalization processing on the encoded data to obtain normalized visual data, point cloud data, and communication data;
[0009] Create an explicit physical graph according to the normalized visual data, point cloud data, and communication data, where the explicit physical graph is used to characterize communication quality and physical distance; mine potential information according to the normalized visual data, point cloud data, and communication data, and create an implicit semantic graph, where the potential information includes at least one of potential vehicle direction change, potential braking, and potential collision information;
[0010] According to the explicit physical graph and the implicit semantic graph, temporal dependence features are extracted through spatial graph convolution, temporal convolution, and causal dilated convolution, and are optimized through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features;
[0011] The temporal dependence features and the semantic features are subjected to feature fusion to obtain fused features;
[0012] Traffic warning information is determined according to the fused features, where the traffic warning information includes whether a collision occurs and / or whether the traffic flow is greater than a preset threshold.
[0013] In a possible implementation manner, the extracting temporal dependence features through spatial graph convolution, temporal convolution, and causal dilated convolution according to the explicit physical graph and the implicit semantic graph includes:
[0014] According to the explicit physical graph and the implicit semantic graph, spatial graph convolution is used to extract and aggregate image spatial features to obtain spatial graph features;
[0015] According to the extracted spatial graph features, temporal convolution is performed to obtain graph features that satisfy the temporal order;
[0016] According to the obtained graph features that satisfy the temporal order, causal dilated convolution is performed to obtain the temporal dependence features that satisfy the causality law.
[0017] In a possible implementation manner, the optimizing through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features includes:
[0018] According to the explicit physical graph and the implicit semantic graph, coordinate normalization is performed to obtain normalized coordinate features;
[0019] The normalized coordinate features are subjected to sine time encoding to obtain sine time-encoded features;
[0020] The sine time-encoded features are divided into multiple neighborhood grids; for each node in the grid, feature calculation is performed according to the node and its multiple adjacent local nodes to obtain local window features; according to the calculated local window features, multiple-granularity feature extraction and fusion are performed through multi-level memory nodes to obtain the semantic features.
[0021] In a possible implementation manner, the extracting temporal dependence features through spatial graph convolution, temporal convolution, and causal dilated convolution according to the explicit physical graph and the implicit semantic graph, and optimizing through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features includes:
[0022] According to the explicit physical graph and the implicit semantic graph, temporal dependence features are extracted through spatial graph convolution, temporal convolution, and causal dilated convolution, and are optimized through spatio-temporal position encoding and a sparse self-attention mechanism to obtain semantic features;
[0023] A first image feature is extracted through a local spatio-temporal model, the explicit physical graph, and the implicit semantic graph; a second image feature is extracted through a global semantic model, the explicit physical graph, and the implicit semantic graph;
[0024] Through cross-attention, using the local spatio-temporal model, the temporal dependence features are calculated based on the first image feature and the second image feature; through cross-attention, using the global semantic model, the semantic features are calculated based on the first image feature and the second image feature.
[0025] In a possible implementation manner, after determining the traffic warning information according to the fusion features, the traffic warning method based on the spatio-temporal dynamic network further includes:
[0026] According to the traffic warning information, calculate the contribution degrees of the attention heads in the global semantic model;
[0027] According to the calculated contribution degrees, for the attention heads with corresponding contribution degrees lower than the preset threshold, prune the corresponding convolution channels.
[0028] In a second aspect of the embodiments of the present application, a traffic warning device based on a spatio-temporal dynamic network is provided, and the device includes:
[0029] A data acquisition module, configured to acquire multi-source traffic data, where the multi-source traffic data includes vehicle-side data, roadside data, and network data;
[0030] A data registration module, configured to perform registration on the multi-source traffic data in the time dimension and the space dimension to obtain registered traffic data;
[0031] A data encoding module, configured to perform feature encoding on the registered traffic data to obtain encoded data, where the encoded data includes visual data, point cloud data, and communication data;
[0032] A normalization module, configured to perform normalization processing on the encoded data to obtain normalized visual data, point cloud data, and communication data;
[0033] A map creation module, configured to create an explicit physical map according to the normalized visual data, point cloud data, and communication data, where the explicit physical map is used to characterize communication quality and physical distance; and to mine potential information and create an implicit semantic map according to the normalized visual data, point cloud data, and communication data, where the potential information includes at least one of potential vehicle direction change, potential braking, and potential collision information.
[0034] A feature extraction module, configured to extract temporal dependence features according to the explicit physical map and the implicit semantic map through spatial graph convolution, temporal convolution, and causal dilated convolution, and to optimize them through spatio-temporal position encoding and a sparse self-attention mechanism to obtain semantic features.
[0035] A feature fusion module, configured to perform feature fusion on the temporal dependence features and the semantic features to obtain fused features.
[0036] A warning determination module, configured to determine traffic warning information according to the fused features, where the traffic warning information includes whether a collision occurs and / or whether the traffic flow is greater than a preset threshold.
[0037] In a possible implementation, the feature extraction module includes:
[0038] A graph feature acquisition sub-module, configured to extract and aggregate image spatial features through spatial graph convolution according to the explicit physical map and the implicit semantic map to obtain spatial graph features.
[0039] A temporal convolution sub-module, configured to perform temporal convolution on the extracted spatial graph features to obtain graph features that satisfy the temporal order.
[0040] A causal dilated convolution sub-module, configured to perform causal dilated convolution on the obtained graph features that satisfy the temporal order to obtain the temporal dependence features that satisfy the causality law.
[0041] In a possible implementation, the feature extraction module includes:
[0042] A coordinate normalization sub-module, configured to perform coordinate normalization according to the explicit physical map and the implicit semantic map to obtain coordinate-normalized features.
[0043] A temporal encoding sub-module, configured to perform sine temporal encoding on the coordinate-normalized features to obtain sine-temporal-encoded features.
[0044] A feature extraction sub-module, configured to divide the features encoded by the sine time into multiple neighborhood grids; for each node in the grid, calculate features based on the node and multiple adjacent local nodes of the node to obtain local window features; according to the calculated local window features, extract and fuse features of multiple granularities through multi-level memory nodes to obtain the semantic features.
[0045] In a possible implementation manner, the feature extraction module is specifically configured to extract temporal dependence features according to the explicit physical graph and the implicit semantic graph through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimize them through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features; extract first image features through a local spatio-temporal model, the explicit physical graph, and the implicit semantic graph; extract second image features through a global semantic model, the explicit physical graph, and the implicit semantic graph; calculate the temporal dependence features through cross-attention, using the local spatio-temporal model, based on the first image features and the second image features; calculate the semantic features through cross-attention, using the global semantic model, based on the first image features and the second image features.
[0046] In a possible implementation manner, the traffic warning device based on the spatio-temporal dynamic network further includes:
[0047] A pruning module, configured to calculate the contribution degree of each attention head in the global semantic model according to the traffic warning information; prune the corresponding convolution channels of the attention heads whose corresponding contribution degrees are lower than a preset threshold according to the calculated contribution degrees.
[0048] Another aspect of the embodiments of the present application further provides an electronic device, including:
[0049] A memory, configured to store a computer program;
[0050] A processor, configured to implement the traffic warning method based on the spatio-temporal dynamic network described in any one of the above when executing the program stored in the memory.
[0051] Another aspect of the embodiments of the present application further provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and the computer program, when executed by a processor, implements the traffic warning method based on the spatio-temporal dynamic network described in any one of the above.
[0052] Another aspect of the embodiments of the present application further provides a computer program product containing instructions, which, when running on a computer, causes the computer to execute the traffic warning method based on the spatio-temporal dynamic network described in any one of the above.
[0053] Advantages of the embodiments of the present application:
[0054] A traffic warning method and device based on a spatio-temporal dynamic network provided by an embodiment of the present application. The method includes: obtaining multi-source traffic data, where the multi-source traffic data includes vehicle-side data, roadside data, and network data; registering the multi-source traffic data in the time dimension and the space dimension to obtain registered traffic data; performing feature encoding on the registered traffic data to obtain encoded data, where the encoded data includes visual data, point cloud data, and communication data; performing normalization processing on the encoded data to obtain normalized visual data, point cloud data, and communication data; creating an explicit physical map according to the normalized visual data, point cloud data, and communication data, where the explicit physical map is used to characterize communication quality and physical distance; mining potential information according to the normalized visual data, point cloud data, and communication data, and creating an implicit semantic map, where the potential information includes at least one of potential vehicle direction change, or potential braking and potential collision information; extracting temporal dependence features according to the explicit physical map and the implicit semantic map through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimizing through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features; performing feature fusion on the temporal dependence features and the semantic features to obtain fusion features; determining traffic warning information according to the fusion features, where the traffic warning information includes whether a collision occurs, and / or whether the traffic flow is greater than a preset threshold. Through the solution of the embodiment of the present application, multi-source traffic data can be obtained, and a dynamic spatio-temporal map including an explicit physical map and an implicit semantic map can be created according to the multi-source data, so as to extract and fuse temporal dependence features and semantic features according to the dynamic spatio-temporal map of the explicit physical map and the implicit semantic map to obtain fusion features, and thus determine traffic warning information according to the fusion features, realizing the improvement of prediction accuracy through the acquisition and processing of various data.
[0055] Of course, it is not necessary for any product or method implementing the present application to achieve all the above advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments according to these drawings.
[0057] Figure 1 It is a schematic flowchart of a traffic warning method based on a spatio-temporal dynamic network provided by an embodiment of the present application;
[0058] Figure 2aAnother flowchart of the traffic warning method based on the spatio-temporal dynamic network provided by the embodiment of the present application;
[0059] Figure 2b A flowchart of creating a dynamic representation modeling flowchart provided by the embodiment of the present application;
[0060] Figure 2c A flowchart of dynamic representation modeling provided by the embodiment of the present application;
[0061] Figure 3 A flowchart of obtaining temporal dependence features provided by the embodiment of the present application;
[0062] Figure 4 A flowchart of obtaining semantic features provided by the embodiment of the present application;
[0063] Figure 5 A structural diagram of a traffic warning device based on the spatio-temporal dynamic network provided by the embodiment of the present application;
[0064] Figure 6 A structural diagram of an electronic device provided by the embodiment of the present application. Detailed implementation manners
[0065] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0066] In the first aspect of the embodiment of the present application, first, a traffic warning method based on the spatio-temporal dynamic network is provided. Refer to Figure 1 , Figure 1 A flowchart of the traffic warning method based on the spatio-temporal dynamic network provided by the embodiment of the present application. The method includes:
[0067] Step S11, obtaining multi-source traffic data, where the multi-source traffic data includes vehicle-side data, roadside data, and network data;
[0068] Step S12, registering the multi-source traffic data in the time dimension and the space dimension to obtain the registered traffic data;
[0069] Step S13, performing feature encoding on the registered traffic data to obtain the encoded data, where the encoded data includes visual data, point cloud data, and communication data;
[0070] Step S14, perform normalization processing on the encoded data to obtain normalized visual data, point cloud data, and communication data;
[0071] Step S15, create an explicit physical map according to the normalized visual data, point cloud data, and communication data, where the explicit physical map is used to characterize communication quality and physical distance; mine potential information according to the normalized visual data, point cloud data, and communication data, and create an implicit semantic map, where the potential information includes at least one of potential vehicle direction change, or potential braking and potential collision information;
[0072] Step S16, extract temporal dependence features according to the explicit physical map and the implicit semantic map through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimize through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features;
[0073] Step S17, perform feature fusion on the temporal dependence features and the semantic features to obtain fused features;
[0074] Step S18, determine traffic warning information according to the fused features, where the traffic warning information includes whether a collision occurs, and / or whether the traffic flow is greater than a preset threshold.
[0075] Corresponding to the above step S11, where the multi-source traffic data includes vehicle-side data, roadside data, and network data. When obtaining the multi-source traffic data, refer to Figure 2a , the multi-source traffic data may include: obtaining vehicle-mounted sensors, roadside units, and cellular network data. Specifically, the vehicle-side data may include information such as the vehicle's speed, lane, position, as well as the vehicle's visual data and point cloud data; the roadside data may include information such as the road speed limit, road surface conditions, such as whether it is slippery, etc.; the network data may include information such as cellular network, mobile network, C-V2X (Cellular Vehicle-to-Everything, vehicle networking).
[0076] It should be noted that the method of the embodiment of the present application can be implemented through a network model, and the network model can be deployed in an intelligent terminal device on the roadside. Through this intelligent terminal device, the road traffic conditions and warnings can be carried out.
[0077] Corresponding to the above step S12, when performing registration in the time dimension and space dimension on the multi-source traffic data, it is mainly to perform spatio-temporal alignment on the multi-source traffic data, that is, time synchronization and space registration. Since the time of data from different sources may vary, through time synchronization, the data from different sources can be unified to the same time dimension, eliminating the time deviation between multiple devices / sensors, and ensuring that data collection or operations are based on a unified time benchmark. When performing space registration, the data from different sources can be unified to the same space coordinate system, solving the problem of coordinate deviation between different sensors, and ensuring that the data corresponds precisely in the unified space coordinate system, thus facilitating subsequent processing. Specifically, calculations can be performed through methods such as cubic spline interpolation.
[0078] Corresponding to the above step S13, among them, the encoded data includes visual data, point cloud data, and communication data. Performing feature encoding on the registered traffic data can achieve classification of different data types, and distinguish the registered traffic data into visual data, point cloud data, and communication data. Specifically, encoding processing can be performed through methods such as label encoding, one-hot encoding, frequency encoding, and target encoding.
[0079] Corresponding to the above step S14, performing normalization processing on the encoded data can eliminate the dimensional difference between different features and improve the efficiency and effect of model training. Specifically, through normalization, all features can be in a similar numerical range (such as [0, 1] or [-1, 1]), ensuring the balance of feature contribution degrees.
[0080] Corresponding to the above step S15, this step mainly constructs a dynamic graph, which includes an explicit physical graph and an implicit semantic graph. Based on the normalized visual data, point cloud data, and communication data, an explicit physical graph and an implicit semantic graph can be created. Among them, the explicit physical graph is used to characterize communication quality and physical distance. Among them, the potential information includes at least one of the potential direction change of the vehicle, potential braking, and potential collision information. The explicit graph generates an adjacency matrix and edge weights based on communication quality and physical distance. Specifically, the communication quality and the physical distance between vehicles can be determined through visual data, point cloud data, and communication data, and an explicit physical graph can be created according to the preset edge weights. Based on the normalized visual data, point cloud data, and communication data, potential information can be mined. The intention can be predicted through LSTM (Long Short-Term Memory, a special recurrent neural network) and potential relationships can be mined through GAT (Graph Attention Network). The potential information includes at least one of the potential direction change of the vehicle, potential braking, and potential collision information, such as whether the vehicle will change lanes, whether the vehicle will brake, and whether there is a possibility of collision. In the actual use process, a dual-graph coupling process with spatio-temporal consistency constraints can also be performed on the dynamic graph including the explicit physical graph and the implicit semantic graph. In this application, aiming at the high dynamics of vehicles, roadside devices, and communication links in the vehicle network, the limitations of traditional static graph models are broken through. Through the dynamic spatio-temporal graph generation and adaptive update mechanism, network topology changes (such as vehicle movement, communication link fluctuations) are captured in real time, and the joint representation ability of traffic state and communication quality is improved. See Figure 2b , the construction process of the dynamic spatio-temporal graph can include: after obtaining multi-source data, data preprocessing is performed. The specific preprocessing process can include: spatio-temporal alignment, then interpolation synchronization, then multi-modal feature encoding, and then for the encoded features, through visual features: CNN (Convolutional Neural Networks) network, point cloud features: 3D (three-dimensional) convolution, and communication features: GNN (Graph Neural Network), feature acquisition is performed, and then feature normalization is carried out. Finally, a dynamic spatio-temporal graph is constructed based on the normalized features. The specific construction process includes: explicit graph construction, dynamic adjacency matrix generation, edge weight calculation, real-time topology update, to obtain the processing result of the first branch, and through implicit semantic graph construction, vehicle intention prediction, semantic association modeling, potential relationship extraction is performed to obtain the processing result of the second path, and then dual-graph coupling and spatio-temporal consistency constraints are performed according to the processing results of the two paths, and finally feature extraction and fusion are carried out.
[0081] Corresponding to the above step S16, when performing spatial graph convolution, temporal convolution, and causal dilated convolution based on the explicit physical graph and the implicit semantic graph, spatial convolution can update the node representation by aggregating the information of the node and its neighbors, thereby capturing the spatial topological relationship; causal dilated convolution can combine the composite operation of causal constraints and dilated convolution, aiming to efficiently model long-term temporal dependencies. Specifically, the methods of spatial graph convolution, temporal convolution, and causal dilated convolution can refer to the prior art. Through spatio-temporal position encoding and sparse self-attention mechanism optimization, when obtaining semantic features, spatio-temporal position encoding is used to explicitly model the spatial topological relationship and the position information of the time series in the model, and sparse self-attention improves the model efficiency and alleviates the memory bottleneck in long sequence modeling by reducing the scope and complexity of attention calculation. In one example, the method of this step can be implemented through the STGNN (Spatio-Temporal Graph Neural Network)-Transformer (a deep learning model architecture) collaborative framework. Among them, the STGNN branch (local spatio-temporal model) performs spatial graph convolution. Specifically, it can aggregate the features of neighboring nodes by improving GATv2 (Graph Attention Networks v2, an improved network of graph attention networks), introduce edge attributes such as relative position and communication delay for spatial convolution; then perform temporal convolution through a temporal convolution network. Specifically, it can perform temporal convolution through TCN (Temporal Convolutional Network); finally, perform causal dilated convolution. Specifically, it can process the historical 5 frames of data through Dilation = 2 (a dilated convolution network with a dilation factor of 2) to extract temporal dependence features. The semantic features can be extracted through the Transformer branch, that is, the global semantic model. Specifically, spatio-temporal position encoding can be performed, such as UTM (a coordinate system including longitude zone, latitude zone, distance east, and distance north) coordinate normalization and sine time encoding, and then sparse self-attention optimization can be performed, such as combining the local window (8-neighborhood) with the global memory node. Among them, spatio-temporal position encoding can encode the position information of time and space into the model when processing spatio-temporal data. The sparse self-attention mechanism optimization can reduce the computational amount in the attention mechanism, for example, by restricting the number of key-value pairs that each query focuses on, thereby reducing the complexity of the model and the consumption of computing resources. Through the deep fusion of STGNN and Transformer, the present invention proposes spatio-temporal cross-attention fusion, dynamically fuses the spatial adjacency matrix and the temporal dependence weight, realizes the unified representation of multi-source network states, solves the core bottleneck of traditional models in spatio-temporal dependence modeling, and provides a theoretical breakthrough and technical tool for multi-source fusion in the vehicle network.Moreover, the hybrid architecture of STGNN and Transformer can achieve multi-granularity modeling of local spatio-temporal patterns (such as vehicle following and sudden lane changes) and global long-term dependencies (such as traffic flow periodicity and communication interference propagation), overcoming the defect of insufficient spatio-temporal correlation modeling of a single model.
[0082] For the STGNN-Transformer network, a cross-attention module can be designed, such as a gate control algorithm, so that features of different modalities can complement and enhance each other. Image features and point cloud features can interact through the attention mechanism to capture their spatio-temporal correlations and achieve feature acquisition between the two branches. In one possible implementation, according to the explicit physical graph and the implicit semantic graph, temporal dependence features are extracted through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimized through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features, including: according to the explicit physical graph and the implicit semantic graph, temporal dependence features are extracted through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimized through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features; first image features are extracted through a local spatio-temporal model and the explicit physical graph and the implicit semantic graph; second image features are extracted through a global semantic model and the explicit physical graph and the implicit semantic graph; through cross-attention, using the local spatio-temporal model, the temporal dependence features are calculated according to the first image feature and the second image feature; through cross-attention, using the global semantic model, the semantic features are calculated according to the first image feature and the second image feature.
[0083] See Figure 2c , the dynamic representation modeling process may include: multi-source data input, preprocessing and feature extraction, and then dynamic graph construction; then local spatio-temporal features are obtained through the STGNN branch, and global friendship features are obtained through the Transformer branch; then through spatio-temporal cross-attention, dynamic representation fusion is performed to achieve spatio-temporal cross-attention fusion, and finally the fused features are output.
[0084] Corresponding to the above step S17, for the temporal dependence features and the semantic features, feature fusion can be performed through feature fusion methods such as weighted summation. Specifically, other feature fusion methods can also be referred to the prior art for feature fusion. In this application, heterogeneous data such as vehicle sensor data (position, speed), C-V2X communication status (channel quality, delay), and environmental information (road conditions, weather) are uniformly encoded as spatio-temporal graph node and edge features, and cross-modal feature alignment and semantic enhancement are achieved through an adaptive graph attention mechanism, which can improve the perception robustness of complex scenarios.
[0085] Corresponding to the above step S18, traffic warning information is determined according to the fusion features, where the traffic warning information includes whether a collision occurs and / or whether the traffic flow is greater than a preset threshold. Specifically, the traffic flow can be predicted through the fusion features. When the traffic flow is greater than the preset threshold, a warning message is sent, or the driving trajectory of the vehicle is predicted, so that when it is determined that a collision may occur, a warning message is sent.
[0086] In one example, see Figure 2a , including: obtaining in-vehicle sensors, roadside units, and cellular network data to obtain multi-source heterogeneous data as data input; then, after spatio-temporal alignment, data cleaning, and feature encoding processing, it is input into the edge and processing layer for processing; then, according to the processing results, a dynamic graph is constructed. Through node definition: vehicle / roadside unit, edge definition: spatio-temporal relationship, dynamic update mechanism, and real-time topology adjustment, a dynamic graph is constructed; then, according to the constructed dynamic graph, feature extraction is performed through the STGNN module and the Transformer module. Then, the extracted features are fused through the feature fusion core, and after the fusion result is compressed by the model, it is processed by the edge large model. During the processing, an adaptive learning and dynamic topology update mechanism is also introduced. Then, the processing result is input into the edge-cloud collaborative optimization for model update and resource scheduling; and through the results of traffic flow prediction, collision warning, and path planning, it is output through intelligent decision-making.
[0087] It can be seen that through the solution of the embodiment of the present application, multi-source traffic data can be obtained, and a dynamic spatio-temporal graph including an explicit physical graph and an implicit semantic graph can be created according to the multi-source data, so as to extract and fuse temporal dependence features and semantic features according to the dynamic spatio-temporal graph of the explicit physical graph and the implicit semantic graph, obtain fusion features, and thus determine traffic warning information according to the fusion features, realizing the improvement of prediction accuracy through the acquisition and processing of various data.
[0088] In a possible implementation manner, see Figure 3 , the extraction of temporal dependence features according to the explicit physical graph and the implicit semantic graph through spatial graph convolution, temporal convolution, and causal dilated convolution includes:
[0089] Step 31, according to the explicit physical graph and the implicit semantic graph, through spatial graph convolution, extract and aggregate image spatial features to obtain spatial graph features;
[0090] Step S32, perform temporal convolution according to the extracted spatial graph features to obtain graph features that satisfy the time order;
[0091] Step S33, perform causal dilated convolution according to the obtained graph features that satisfy the time order to obtain the temporal dependence features that satisfy the causality law.
[0092] Among them, spatial graph convolution can model spatial relationships (such as traffic road network nodes, sensor network topologies) through graph structures, define the interaction weights between nodes using the adjacency matrix, and extract spatial features by combining graph neural networks. Specifically, it can include graph construction based on topology, such as mapping physical connections (such as roads, power lines) as edges of the graph, and aggregating neighborhood information through graph convolutional layers; then performing multi-modal feature fusion, such as integrating node attributes (such as traffic flow, meteorological data) and topological relationships, and dynamically adjusting neighborhood weights using the attention mechanism. Temporal convolution can slide a one-dimensional convolutional kernel along the time axis to extract dynamic patterns in the sequence, support parallel computing, and avoid the problem of gradient disappearance. Specifically, it can include causal constraints, such as ensuring that the output at the current moment only depends on historical inputs through zero-padding to avoid leakage of future information, and dilated convolution, such as inserting holes between convolutional kernel elements to exponentially expand the receptive field. Finally, residual connections are performed, such as introducing a residual structure when stacking multiple layers of dilated convolution to alleviate the degradation problem of deep networks. Causal dilated convolution can combine causal constraints and dilated convolution to efficiently model long-range dependencies while ensuring temporal causality. Specifically, it can be achieved through a staged dilation strategy, such as exponentially increasing the Dilation value layer by layer (such as 1, 2, 4, 8) to cover features at different time scales, and then performing dynamic kernel selection, such as adaptively adjusting the convolutional kernel weights through a gating mechanism to optimize the ability to capture different time patterns.
[0093] In a possible implementation, referring to Figure 4 , the semantic features obtained by optimizing through spatio-temporal position encoding and sparse self-attention mechanism include:
[0094] Step 41: Perform coordinate normalization on the basis of the explicit physical graph and the implicit semantic graph to obtain features with normalized coordinates;
[0095] Step 42: Perform sinusoidal time encoding on the features with normalized coordinates to obtain features with sinusoidal time encoding;
[0096] Step 43: Divide the features with sinusoidal time encoding into multiple neighborhood grids; for each node in the grid, calculate features based on this node and its multiple adjacent local nodes to obtain local window features; based on the calculated local window features, extract and fuse features of multiple granularities through multi-level memory nodes to obtain the semantic features.
[0097] Among them, the spatio-temporal position encoding can include coordinate normalization and sinusoidal time encoding. Coordinate normalization can unify features into the same coordinate system. Specifically, UTM coordinate normalization can be performed, which specifically includes: regional scaling, which can divide geographical regions and independently perform magnification and reduction normalization on the coordinates within each region to eliminate cross-regional scale differences, and dynamic parameter update, which periodically updates the magnification / reduction values for real-time data streams to adapt to changes in the coordinate range. The sinusoidal time encoding can include frequency function embedding, which maps the time step t to a multi-dimensional vector and captures the periodic features of the time series through combinations of sine / cosine functions with different frequencies, and time scale expansion, which generates composite encodings by combining multiple time granularities such as hours, days, and weeks to enhance the periodic modeling ability.
[0098] The optimization of the sparse self-attention mechanism can include probabilistic sparse screening, structured sparse design, and dynamic sparse patterns. Among them, probabilistic sparse screening can reduce the computational complexity by only retaining the key-value pairs that have the greatest impact on the current query through probabilistic methods (such as maximum entropy screening). The structured sparse design can limit the global attention range to a local sliding window (such as a fixed neighborhood or block), only calculate the similarity of the elements within the window, then divide the long sequence into sub-blocks, perform inter-block interaction at the coarse-grained layer, and perform intra-block refinement at the fine-grained layer, taking into account both local and global dependencies. That is, the features of the sinusoidal time encoding are divided into multiple neighborhood grids; for each node in the grid, feature calculations are performed based on the node and its multiple adjacent local nodes to obtain local window features. The dynamic sparse pattern can dynamically generate a sparse pattern (such as locality-sensitive hashing) based on the similarity between the query and the key, adaptively focus on the highly relevant regions, and then combine the global sparse head (capturing key features) with the local dense head (retaining details) to balance efficiency and accuracy. Thus, not only can the computational complexity be reduced and the computational efficiency be improved, but also the computational accuracy can be taken into account, and multi-granularity feature extraction and fusion are performed through multi-level memory nodes to obtain the semantic features.
[0099] In a possible implementation manner, after determining the traffic warning information according to the fusion features, see Figure 4 , the traffic warning method based on the spatio-temporal dynamic network further includes:
[0100] Step S41, calculating the contribution degrees of the attention heads in the global semantic model according to the traffic warning information;
[0101] Step S42, pruning the corresponding convolutional channels for the attention heads whose calculated contribution degrees are lower than the preset threshold.
[0102] By calculating the contribution degrees of the attention heads in the global semantic model, for the attention heads with corresponding contribution degrees lower than the preset threshold, pruning is performed on the corresponding convolutional channels. Specifically, the attention heads with low contribution degrees in the Transformer in the above embodiments can be removed, and redundant convolutional channels can be trimmed. Specifically, the determination and pruning of attention heads and channels can be performed through channel importance ranking and a greedy algorithm. In the actual use process, hardware adaptation and acceleration, dynamic offloading strategies, and protocol adaptation can also be performed to achieve a closed-loop of "training - compression - deployment - monitoring". At the same time, incremental learning is introduced to regularly update the model parameters for new scenario data (such as road construction, sudden weather). Through the embodiments of the present application, lightweight deployment and optimization can be achieved. Through model compression and hardware adaptation optimization, the model is marginalized to adapt to the architecture of edge servers and meet the real-time requirements.
[0103] The inventors have found through research that through the solution of the embodiments of the present application, by combining dynamic spatio-temporal graph modeling, spatio-temporal cross-attention mechanism, and lightweight edge deployment, the problems of "spatio-temporal fragmentation, modality conflict, and resource inefficiency" in dynamic network representation are solved, providing a feasible technical paradigm for the deep integration of intelligent transportation systems and communication. The specific core technical effects are as follows:
[0104] 1. Improvement in dynamic spatio-temporal modeling ability. Dynamic graph adaptive update: By adjusting the graph topology (node connection relationship and edge weight) in real time, the prediction error of vehicle movement trajectories is reduced by 30% (compared with traditional solutions), especially showing significant advantages in high-speed (>80 km / h) and high-density (>100 vehicles / km²) scenarios. Spatio-temporal coupled feature extraction: By combining local spatio-temporal convolution and the global attention mechanism of the Transformer, the prediction accuracy of traffic flow is improved by 25%, and the prediction error of communication link quality ≤ 5%.
[0105] 2. Optimization of multi-modal data fusion. Cross-modal feature alignment: Based on the graph attention mechanism, the semantic alignment efficiency of multi-source heterogeneous data (sensors, communication, environment) is improved by 40%, reducing feature conflicts caused by sampling frequency differences. Enhanced robustness: In complex environments such as rain, snow, and tunnel occlusion, the fluctuation range of the accuracy of collaborative perception tasks is reduced from ±15% in traditional methods to ±5%.
[0106] 3. Real-time performance and resource efficiency. Low-latency inference: Through model lightweighting (pruning + quantization), the inference latency at the edge ≤ 10 ms, supporting high-throughput processing of 1000 nodes / second. Saving of communication resources: The dynamic channel allocation strategy reduces redundant resource occupancy, the spectrum utilization rate is increased by 20%, and the energy consumption of the base station is reduced by 15%.
[0107] In the second aspect of the embodiments of the present application, a traffic warning device based on a spatio-temporal dynamic network is provided. See Figure 5, the device includes:
[0108] A data acquisition module 501 for acquiring multi-source traffic data, where the multi-source traffic data includes vehicle-side data, roadside data, and network data;
[0109] A data registration module 502 for registering the multi-source traffic data in the time dimension and the space dimension to obtain registered traffic data;
[0110] A data encoding module 503 for performing feature encoding on the registered traffic data to obtain encoded data, where the encoded data includes visual data, point cloud data, and communication data;
[0111] A normalization module 504 for performing normalization processing on the encoded data to obtain normalized visual data, point cloud data, and communication data;
[0112] A graph creation module 505 for creating an explicit physical graph according to the normalized visual data, point cloud data, and communication data, where the explicit physical graph is used to characterize communication quality and physical distance; mining potential information according to the normalized visual data, point cloud data, and communication data, and creating an implicit semantic graph, where the potential information includes at least one of potential vehicle direction change, or potential braking and potential collision information;
[0113] A feature extraction module 506 for extracting temporal dependence features according to the explicit physical graph and the implicit semantic graph through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimizing through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features;
[0114] A feature fusion module 507 for performing feature fusion on the temporal dependence features and the semantic features to obtain fusion features;
[0115] A warning determination module 508 for determining traffic warning information according to the fusion features, where the traffic warning information includes whether a collision occurs, and / or whether the traffic flow is greater than a preset threshold.
[0116] In a possible implementation manner, the feature extraction module includes:
[0117] A graph feature acquisition sub-module for extracting and aggregating image spatial features through spatial graph convolution according to the explicit physical graph and the implicit semantic graph to obtain spatial graph features;
[0118] A temporal convolution sub-module for performing temporal convolution on the extracted spatial graph features to obtain graph features that satisfy the temporal order;
[0119] The causal dilated convolution sub-module is used to perform causal dilated convolution on the obtained graph features that satisfy the time order, so as to obtain the time series dependence features that satisfy the causality law.
[0120] In a possible implementation manner, the feature extraction module includes:
[0121] The coordinate normalization sub-module is used to perform coordinate normalization on the explicit physical graph and the implicit semantic graph to obtain the features with normalized coordinates.
[0122] The time encoding sub-module is used to perform sine time encoding on the features with normalized coordinates to obtain the features with sine time encoding.
[0123] The feature extraction sub-module is used to divide the features with sine time encoding into multiple neighborhood grids; for each node in the grid, feature calculation is performed according to the node and its multiple adjacent local nodes to obtain local window features; according to the calculated local window features, multiple granularity features are extracted and fused through multi-level memory nodes to obtain the semantic features.
[0124] In a possible implementation manner, the feature extraction module is specifically used to extract time series dependence features according to the explicit physical graph and the implicit semantic graph through spatial graph convolution, time convolution and causal dilated convolution, and optimize them through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features; extract the first image features through the local spatio-temporal model and the explicit physical graph and the implicit semantic graph; extract the second image features through the global semantic model and the explicit physical graph and the implicit semantic graph; calculate the time series dependence features through cross-attention, using the local spatio-temporal model according to the first image features and the second image features; calculate the semantic features through cross-attention, using the global semantic model according to the first image features and the second image features.
[0125] In a possible implementation manner, the traffic warning device based on the spatio-temporal dynamic network further includes:
[0126] The pruning module is used to calculate the contribution degree of each attention head in the global semantic model according to the traffic warning information; according to the calculated contribution degree, prune the corresponding convolution channels of the attention heads whose corresponding contribution degrees are lower than the preset threshold.
[0127] It can be seen that through the solution of the embodiments of the present application, multi-source traffic data can be obtained, and a dynamic spatio-temporal graph including an explicit physical graph and an implicit semantic graph can be created based on the multi-source data, so as to extract and fuse temporal dependence features and semantic features according to the dynamic spatio-temporal graph of the explicit physical graph and the implicit semantic graph, obtain fused features, and thus determine traffic warning information according to the fused features, realizing the improvement of prediction accuracy through the acquisition and processing of various data.
[0128] The embodiments of the present application also provide an electronic device, as Figure 6 shown, including:
[0129] A memory 601 for storing a computer program;
[0130] A processor 602, when executing the program stored on the memory 601, implements the following steps:
[0131] Obtain multi-source traffic data, where the multi-source traffic data includes vehicle-side data, roadside data, and network data;
[0132] Perform registration on the multi-source traffic data in the time dimension and the space dimension to obtain registered traffic data;
[0133] Perform feature encoding on the registered traffic data to obtain encoded data, where the encoded data includes visual data, point cloud data, and communication data;
[0134] Perform normalization processing on the encoded data to obtain normalized visual data, point cloud data, and communication data;
[0135] According to the normalized visual data, point cloud data, and communication data, create an explicit physical graph, where the explicit physical graph is used to characterize communication quality and physical distance; according to the normalized visual data, point cloud data, and communication data, mine potential information, and create an implicit semantic graph, where the potential information includes at least one of potential vehicle direction change, or potential braking and potential collision information;
[0136] According to the explicit physical graph and the implicit semantic graph, extract temporal dependence features through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimize through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features;
[0137] Perform feature fusion on the temporal dependence features and the semantic features to obtain fused features;
[0138] Determine traffic warning information according to the fused features, where the traffic warning information includes whether a collision occurs, and / or whether the traffic flow is greater than a preset threshold.
[0139] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0140] The communication interface is used for communication between the above electronic device and other devices.
[0141] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.
[0142] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0143] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above traffic warning methods based on a spatio-temporal dynamic network are implemented.
[0144] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any one of the traffic warning methods based on a spatio-temporal dynamic network in the above embodiments.
[0145] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.
[0146] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0147] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device, electronic device, and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.
[0148] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A traffic warning method based on a spatio-temporal dynamic network, characterized in that, The method includes: Obtaining multi-source traffic data, where the multi-source traffic data includes vehicle-side data, roadside data, and network data; Registering the multi-source traffic data in the time dimension and the space dimension to obtain registered traffic data; Performing feature encoding on the registered traffic data to obtain encoded data, where the encoded data includes visual data, point cloud data, and communication data; Performing normalization processing on the encoded data to obtain normalized visual data, point cloud data, and communication data; Creating an explicit physical map based on the normalized visual data, point cloud data, and communication data, where the explicit physical map is used to characterize communication quality and physical distance; mining potential information based on the normalized visual data, point cloud data, and communication data, and creating an implicit semantic map, where the potential information includes at least one of potential vehicle direction change, or potential braking and potential collision information; Extracting temporal dependence features based on the explicit physical map and the implicit semantic map through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimizing through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features; Performing feature fusion on the temporal dependence features and the semantic features to obtain fusion features; Determining traffic warning information based on the fusion features, where the traffic warning information includes whether a collision occurs, and / or whether the traffic flow is greater than a preset threshold; The extracting temporal dependence features based on the explicit physical map and the implicit semantic map through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimizing through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features includes: extracting a first image feature through a local spatio-temporal model and the explicit physical map and the implicit semantic map; extracting a second image feature through a global semantic model and the explicit physical map and the implicit semantic map; calculating the temporal dependence features through cross-attention using the local spatio-temporal model based on the first image feature and the second image feature; calculating the semantic features through cross-attention using the global semantic model based on the first image feature and the second image feature.
2. The traffic warning method based on the spatio-temporal dynamic network according to claim 1, wherein The extracting temporal dependence features based on the explicit physical map and the implicit semantic map through spatial graph convolution, temporal convolution, and causal dilated convolution includes: Based on the explicit physical map and the implicit semantic map, performing extraction and aggregation of image spatial features through spatial graph convolution to obtain spatial graph features; Performing temporal convolution on the extracted spatial graph features to obtain graph features that satisfy the time order; Performing causal dilated convolution on the obtained graph features that satisfy the time order to obtain the temporal dependence features that satisfy the causality law.
3. The traffic warning method based on the spatio-temporal dynamic network according to claim 1, wherein, The optimizing through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features includes: Performing coordinate normalization based on the explicit physical map and the implicit semantic map to obtain coordinate-normalized features; Performing sine time encoding on the coordinate-normalized features to obtain sine time-encoded features; Divide the features of the sine time encoding into multiple neighborhood grids; for each node in the grid, calculate features based on the node and its multiple adjacent local nodes to obtain local window features; according to the calculated local window features, extract and fuse features of multiple granularities through multi-level memory nodes to obtain the semantic features.
4. The traffic warning method based on a spatio-temporal dynamic network according to claim 1, wherein After determining the traffic warning information according to the fusion features, the traffic warning method based on the spatio-temporal dynamic network further includes: Calculate the contribution degrees of the attention heads in the global semantic model according to the traffic warning information; According to the calculated contribution degrees, prune the corresponding convolutional channels of the attention heads with contribution degrees lower than the preset threshold.
5. A traffic warning device based on a spatio-temporal dynamic network, characterized in that, The device includes: A data acquisition module, configured to acquire multi-source traffic data, where the multi-source traffic data includes vehicle-side data, roadside data, and network data; A data registration module, configured to register the multi-source traffic data in the time dimension and the space dimension to obtain registered traffic data; A data encoding module, configured to perform feature encoding on the registered traffic data to obtain encoded data, where the encoded data includes visual data, point cloud data, and communication data; A normalization module, configured to perform normalization processing on the encoded data to obtain normalized visual data, point cloud data, and communication data; A graph creation module, configured to create an explicit physical graph according to the normalized visual data, point cloud data, and communication data, where the explicit physical graph is used to represent communication quality and physical distance; mine potential information according to the normalized visual data, point cloud data, and communication data, and create an implicit semantic graph, where the potential information includes at least one of potential direction change of a vehicle, or potential braking and potential collision information; A feature extraction module, configured to extract temporal dependence features according to the explicit physical graph and the implicit semantic graph through spatial graph convolution, temporal convolution, and causal dilated convolution, and optimize through spatio-temporal position encoding and sparse self-attention mechanism to obtain semantic features; A feature fusion module, configured to perform feature fusion on the temporal dependence features and the semantic features to obtain fusion features; A warning determination module, configured to determine traffic warning information according to the fusion features, where the traffic warning information includes whether a collision occurs, and / or whether the traffic flow is greater than a preset threshold; The feature extraction module is specifically configured to extract a first image feature through a local spatio-temporal model and the explicit physical graph and the implicit semantic graph; extract a second image feature through a global semantic model and the explicit physical graph and the implicit semantic graph; calculate the temporal dependence feature through cross-attention, using the local spatio-temporal model, according to the first image feature and the second image feature; calculate the semantic feature through cross-attention, using the global semantic model, according to the first image feature and the second image feature.
6. The traffic warning device based on the spatio-temporal dynamic network according to claim 5, characterized in that, The feature extraction module includes: The graph feature acquisition sub-module is used to extract and aggregate the image spatial features through spatial graph convolution based on the explicit physical graph and the implicit semantic graph, so as to obtain spatial graph features; The temporal convolution sub-module is used to perform temporal convolution on the extracted spatial graph features to obtain graph features that satisfy the temporal order; The causal dilated convolution sub-module is used to perform causal dilated convolution on the obtained graph features that satisfy the temporal order to obtain the temporal dependence features that satisfy the causality law.
7. The traffic warning device based on the spatio-temporal dynamic network according to claim 5, characterized in that The feature extraction module includes: The coordinate normalization sub-module is used to perform coordinate normalization based on the explicit physical graph and the implicit semantic graph to obtain features with normalized coordinates; The temporal encoding sub-module is used to perform sinusoidal temporal encoding on the features with normalized coordinates to obtain features with sinusoidal temporal encoding; The feature extraction sub-module is used to divide the features with sinusoidal temporal encoding into multiple neighborhood grids; for each node in the grid, feature calculation is performed according to the node and its multiple adjacent local nodes to obtain local window features; according to the calculated local window features, multiple-granularity feature extraction and fusion are performed through multi-level memory nodes to obtain the semantic features.
8. An electronic device, characterized in that, It includes: A memory for storing computer programs; A processor, when executing the programs stored on the memory, implements the traffic warning method based on the spatio-temporal dynamic network according to any one of claims 1-4.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the traffic warning method based on the spatio-temporal dynamic network according to any one of claims 1-4.
Citation Information
Patent Citations
Road network safety early warning method and device based on hologram and storage medium
CN119296322A