Highway traffic incident video detection method

By using multimodal data processing and Bayesian decision theory, the problem of insufficient modeling of time series and spatial context in highway traffic incident video detection is solved, achieving more efficient and accurate traffic incident detection that can adapt to complex environments.

CN120853121APending Publication Date: 2025-10-28SHAANXI COMM ELECTRONIC ENG TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511360243.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing methods for detecting traffic incidents on highways lack in-depth modeling of time series and spatial context, cannot adapt to complex weather conditions, have a high false alarm rate, insufficient real-time performance, and are weak in handling the correlation of events in long video sequences.

Method used

By acquiring samples of highway traffic incidents, a multimodal preprocessing and asymmetric spatiotemporal encoder is established to extract multimodal spatiotemporal joint feature data. Combined with Bayesian decision theory, the confidence of video and radar modes is analyzed to verify the probability of event type occurrence and complete the detection.

Benefits of technology

It improves detection speed and accuracy, enhances adaptability and real-time performance in complex environments, and reduces false alarm rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853121A_ABST
    Figure CN120853121A_ABST
Patent Text Reader

Abstract

The invention discloses an expressway traffic event video detection method, which relates to the technical field of video analysis, and comprises the following steps: acquiring a plurality of types of traffic event samples of an expressway, performing data calibration on an initial expressway traffic event training original data set according to multi-modal preprocessing, and obtaining an expressway traffic event training original data set; generating a highway traffic event multi-mode original training data set; performing spatio-temporal joint feature extraction on the multi-modal original training data set of the highway traffic events to obtain multi-modal spatio-temporal joint feature data, establishing a highway traffic event detection classification model, and evaluating probability distribution vectors of abnormal scores and event categories of the highway traffic events; and verifying the occurrence probability of the traffic event type under the given video stream and radar point cloud by using the probability distribution vector based on the abnormal score of the highway traffic event and the event type. The method has the advantages that the detection speed of the highway traffic incident is increased, and the detection accuracy of the traffic incident is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video analytics, specifically to a method for detecting traffic incidents on highways via video. Background Technology

[0002] Highway traffic incident video detection refers to the process of using video surveillance technology and computer vision methods to monitor, identify, analyze, and classify traffic incidents occurring on highways in real time.

[0003] Existing methods mainly rely on traditional computer vision techniques, such as background modeling, target tracking, or simple feature extraction based on deep learning. However, they may lack deep modeling of time series and spatial context, have insufficient adaptability to complex weather conditions, high false alarm rate, insufficient real-time performance, and weak processing of event correlation in long video sequences. Summary of the Invention

[0004] To address the aforementioned technical problems, a video detection method for highway traffic incidents is provided. This technical solution solves the problems of existing methods that mainly rely on traditional computer vision techniques, such as background modeling, target tracking, or simple feature extraction based on deep learning, but may lack deep modeling of time series and spatial context, have insufficient adaptability to complex weather conditions, high false alarm rate, insufficient real-time performance, and weak processing of event correlation in long video sequences.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for detecting traffic incidents on highways via video, comprising: S1. Obtain traffic event samples of several types on highways and build an initial highway traffic event training database; the traffic event samples include: video stream data and radar point cloud data. S2. Perform data calibration on the initial original training dataset of highway traffic events according to multimodal preprocessing to generate a multimodal original training dataset of highway traffic events. S3. Establish an asymmetric spatiotemporal encoder and extract spatiotemporal joint features from the original training data set of multimodal highway traffic events to obtain multimodal spatiotemporal joint feature data of highway traffic events. S4. Based on multimodal spatiotemporal joint feature data of highway traffic incidents, establish a highway traffic incident detection and classification model, and evaluate the anomaly score and probability distribution vector of the event category of highway traffic incidents. S5. Based on Bayesian decision theory, using the probability distribution vector of anomaly scores and event categories based on highway traffic events, analyze video modal confidence and radar modal confidence, verify the occurrence probability of traffic event types under given video streams and radar point clouds, and complete highway traffic event detection.

[0006] Preferably, step S2 specifically includes: For radar point cloud data in the original training dataset of highway traffic incidents, the polar coordinate system is converted to the Cartesian coordinate system to calibrate the spatial coordinate alignment between radar point cloud data and video stream data. For video stream data in the training dataset of highway traffic incidents, time synchronization between the timestamps of the video stream data and the timestamps of the radar point cloud data is determined by time interpolation through a sliding window. Based on radar point cloud data in the original training dataset of highway traffic incidents, a voxel network is established, voxel features are statistically analyzed, and then substituted into a 3D sparse convolutional neural network to extract multi-layer 3D voxel features from the radar point cloud data.

[0007] Preferably, step S2 further includes: By taking the maximum value along the Z-axis and compressing the multi-layer 3D voxel features of the radar point cloud data, a multi-layer 2D bird's-eye view feature map of the radar point cloud data is obtained. By using the camera's intrinsic parameter matrix, each pixel in the multi-layer 2D bird's-eye view feature map of the radar point cloud data is projected onto the pixel coordinates of the original image to obtain the radar point cloud data feature map. Based on the video stream data in the original training dataset of highway traffic incidents, the EfficientNet network is substituted into it. The video stream data is used as input and the visual feature map of the video stream data is used as output. The CBAM module and FPN structure are used to change the visual feature map to enhance spatial attention and fuse low-level details and high-level semantics. Based on visual feature maps from video stream data and feature maps from radar point cloud data, a matrix of multimodal data was stitched together to obtain a set of original multimodal training data for highway traffic events.

[0008] Preferably, step S3 specifically includes: Based on deformable convolution DCNv3, multi-scale features of feature maps in the original training dataset of multimodal highway traffic events are extracted to obtain spatial branch multi-scale feature parameters. Based on the time series prediction bidirectional gated recurrent BiGRU, the temporal dependency features in the original training dataset of multimodal highway traffic events are captured, and the temporal dependency feature parameters of the time branch are obtained. Based on spatial branch multi-scale feature parameters and temporal branch temporal dependent feature parameters, an asymmetric spatiotemporal encoder is established, and multimodal spatiotemporal joint feature data of highway traffic events is generated according to the spatiotemporal hybrid attention mechanism.

[0009] Preferably, step S4 specifically includes: A serial highway traffic incident detection and classification model is established based on SVM support vector machine. Based on the cascade highway traffic incident detection and classification model, the multimodal spatiotemporal joint feature data of highway traffic incidents is used as input, and the abnormal scores of highway traffic incidents quantified by the first detector of the model are used as output. The multimodal spatiotemporal joint feature data of highway traffic incidents corresponding to the abnormal scores of positively correlated traffic incidents are selected as input to the second detector to generate the classification probability of highway traffic incident categories. Combining the abnormal scores of highway traffic incidents and the classification probability of highway traffic incident categories, the probability distribution vector of abnormal scores and event categories of highway traffic incidents is output. Specifically, the series highway traffic incident detection and classification model is as follows:

[0010] In the formula, Anomaly scoring for highway traffic incidents. Classifying highway traffic incidents This is an abnormal rating matrix. For category classification matrix, For classification bias, Softmax is the normalization function. For the number of categories, 't' represents the index of the time step and the index of the spatial location, respectively. This provides multimodal spatiotemporal joint feature data of highway traffic events with time step t and spatial location s.

[0011] Preferably, step S5 specifically includes: Based on the abnormal scores of highway traffic events and the probability distribution vector of event categories, calculate the modal confidence of highway video streams and the modal confidence of highway radar point clouds. Based on the modal confidence of highway video streams and the modal confidence of highway radar point clouds, a Bayesian decision theory function is established to analyze the visual likelihood of highway video streams and the likelihood of radar point clouds, and to jointly evaluate the probability of traffic event types occurring under a given video stream and radar point cloud. Highway traffic event detection is achieved by classifying highway traffic events based on the probability of occurrence of traffic event types in a given video stream and radar point cloud.

[0012] Preferably, step S5 further includes: The calculation of the modal confidence scores of the highway video stream and the highway radar point cloud is specifically as follows:

[0013] In the formula, Modal confidence of highway video stream For the modal confidence of radar point cloud on highways Transpose the classification vector for highway traffic incidents. This is the weight matrix. This is the radar data feature map corresponding to time step t; Specifically, the Bayesian decision theory function is as follows:

[0014] In the formula, Given a video stream and the probability of traffic event types occurring in radar point clouds, Let visual likelihood function be the conditional probability of visual evidence given an event type. Given the event type, the likelihood function of the radar point cloud represents the conditional probability of radar evidence. The unconditional probability of occurrence of an event type. The event type to be determined is 'event', which represents all possible event types for the highway.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a video detection scheme for highway traffic incidents. It collects samples of several types of highway traffic incidents for data calibration to generate a multimodal raw training dataset. An asymmetric spatiotemporal encoder is used to extract multimodal spatiotemporal joint feature data of highway traffic incidents. A detection and classification model is then established based on this multimodal spatiotemporal joint feature data to evaluate incident anomaly scores and category probability distributions. By using Bayesian decision theory and combining the confidence levels of video and radar modalities, the probability of incident occurrence is analyzed, thereby improving the detection speed and accuracy of highway traffic incidents. Attached Figure Description

[0016] Figure 1 This is a flowchart of a video detection method for highway traffic incidents. Detailed Implementation

[0017] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0018] Reference Figure 1 As shown, a video detection method for highway traffic incidents includes: S1. Obtain traffic event samples of several types on highways and build an initial highway traffic event training database; the traffic event samples include: video stream data and radar point cloud data. S2. Perform data calibration on the initial original training dataset of highway traffic events according to multimodal preprocessing to generate a multimodal original training dataset of highway traffic events. Step S2 specifically includes: For radar point cloud data in the original training dataset of highway traffic incidents, the polar coordinate system is converted to the Cartesian coordinate system to calibrate the spatial coordinate alignment between radar point cloud data and video stream data. For video stream data in the training dataset of highway traffic incidents, time synchronization between the timestamps of the video stream data and the timestamps of the radar point cloud data is determined by time interpolation through a sliding window. Based on radar point cloud data in the original training dataset of highway traffic incidents, a voxel network is established, voxel features are statistically analyzed, and then substituted into a 3D sparse convolutional neural network to extract multi-layer 3D voxel features from the radar point cloud data.

[0019] As a further development, the aforementioned 3D sparse convolutional neural network can be based on the VoxelNet network structure or the 3D backbone network in networks such as CenterPoint, and the corresponding internal 3D sparse convolutional layers, submanifold convolutional layers and sparse pooling layers can be constructed according to the typical encoder structure, which will not be described in detail here. Step S2 also includes: By taking the maximum value along the Z-axis and compressing the multi-layer 3D voxel features of the radar point cloud data, a multi-layer 2D bird's-eye view feature map of the radar point cloud data is obtained. By using the camera's intrinsic parameter matrix, each pixel in the multi-layer 2D bird's-eye view feature map of the radar point cloud data is projected onto the pixel coordinates of the original image to obtain the radar point cloud data feature map. Based on the video stream data in the original training dataset of highway traffic incidents, the EfficientNet network is substituted into it. The video stream data is used as input and the visual feature map of the video stream data is used as output. The CBAM module and FPN structure are used to change the visual feature map to enhance spatial attention and fuse low-level details and high-level semantics. Based on visual feature maps from video stream data and feature maps from radar point cloud data, a matrix of multimodal data was stitched together to obtain a set of original multimodal training data for highway traffic events.

[0020] When using it, please refer to the steps outlined above: As a further step, the alignment problem of heterogeneous data is solved through strict spatiotemporal synchronization (coordinate transformation and time interpolation) and deep feature-level fusion. First, 3D sparse convolution and EfficientNet+CBAM+FPN are used to efficiently extract precise geometric features from radar and rich semantic features from video, respectively. Then, the radar features are projected onto the image coordinate system, and finally, pixel-level stitching is performed. This achieves complementary advantages between modalities (the combination of radar ranging and visual recognition), generating more comprehensive and less redundant fused features, thereby greatly improving the accuracy, robustness, and adaptability to harsh environments in subsequent event detection.

[0021] The specific implementation process for step S2 is as follows: Scenario: Detecting a vehicle breakdown event in the right lane of a highway.

[0022] Data input: Video stream: A 1080p@25FPS video clip from a fixed surveillance camera.

[0023] Radar point cloud: 77GHz millimeter-wave radar point cloud data from the same time period and area as the video above, with an output frequency of 20Hz.

[0024] Implementation process: Spatiotemporal alignment: The system reads the radar point cloud frame (polar coordinates) at second t.

[0025] Spatial alignment: Using the radar's installation position and angle parameters, all point cloud data are converted into a Cartesian coordinate system (X_radar, Y_radar, Z_radar) with the radar as the origin.

[0026] Time synchronization: The system detects that the timestamp of the current radar frame is t + 0.12s. By performing a sliding window search and interpolation in the video stream buffer, it finds the video frame (the k-th frame) whose timestamp is closest to t + 0.12s and uses it as the matching frame.

[0027] Radar Feature Extraction and Projection: The converted radar point cloud is voxelized (e.g., voxel size is 0.1m x 0.1m x 0.2m).

[0028] The voxel mesh is fed into a 3D sparse convolutional encoder similar to Second or CenterPoint. The network outputs a multi-layer 3D feature volume.

[0029] The volume is compressed by taking the maximum value along the Z-axis (height axis) to obtain a 2D radar feature map (BEV). Each pixel in this feature map contains the maximum feature response at that (X, Y) location, encoding information such as the height and density of the object at that location.

[0030] Using pre-calibrated intrinsic parameters (focal length, optical center) and extrinsic parameters (relative position and rotation between the camera and radar), a projection matrix is ​​constructed. Each pixel in this BEV radar feature map is projected onto the corresponding pixel coordinates of the k-th frame video image, generating a radar feature map that corresponds one-to-one with the video frame space.

[0031] Video feature extraction: Input the k-th frame of video image into the EfficientNet-B3 backbone.

[0032] CBAM modules are inserted at different stages of the backbone. For example, in deeper layers of the network, CBAM might focus spatially on road regions in the image and channel-wise on feature channels related to vehicle color and texture.

[0033] Using the FPN structure, the outputs of the last three layers of EfficientNet are fused to generate a video visual feature map that contains both high-level semantics (this is a car) and fine-grained location details (the edges of the car are very clear).

[0034] Feature fusion: Now, we have two feature maps with the same spatial dimensions (possibly adjusted by upsampling / downsampling) and perfectly aligned spatial locations: Feature map A (video): [C_v, H, W], which contains appearance, texture, and semantic information.

[0035] Feature map B (radar projection): [C_r, H, W], which contains distance, height, and point cloud density information.

[0036] The two feature maps are concatenated along the channel dimension to generate the final multimodal fusion feature map [C_v + C_r, H, W].

[0037] For example, in the image region (x, y) where the broken-down vehicle is located, the fused features may include: (from video) red, tire-treaded, car-shaped features; (from radar) features with a velocity of 0 (estimated from multi-frame point clouds), a height of 1.5 meters, and moderate point cloud echo intensity.

[0038] Output: A single sample from this multimodal training dataset of highway traffic incidents is no longer the original image and point cloud, but a fused feature map that has undergone deep processing, has highly condensed information, and whose bimodal information has been precisely aligned. This feature map will be directly fed into the S3 spatiotemporal encoder for further processing.

[0039] S3. Establish an asymmetric spatiotemporal encoder and extract spatiotemporal joint features from the original training data set of multimodal highway traffic events to obtain multimodal spatiotemporal joint feature data of highway traffic events. Step S3 specifically includes: Based on deformable convolution DCNv3, multi-scale features of feature maps in the original training dataset of multimodal highway traffic events are extracted to obtain spatial branch multi-scale feature parameters. Based on the time series prediction bidirectional gated recurrent BiGRU, the temporal dependency features in the original training dataset of multimodal highway traffic events are captured, and the temporal dependency feature parameters of the time branch are obtained. Based on spatial branch multi-scale feature parameters and temporal branch temporal dependent feature parameters, an asymmetric spatiotemporal encoder is established, and multimodal spatiotemporal joint feature data of highway traffic events is generated according to the spatiotemporal hybrid attention mechanism.

[0040] When using it, please refer to the steps outlined above: As a further development, an asymmetric spatiotemporal encoder was established. Deformable convolution (DCNv3) was used to adaptively capture multi-scale spatial features to address irregular variations in target shape and scale. Bidirectional gated recurrent units (BiGRU) were employed to model long-term temporal dependencies to understand the dynamic evolution of events. Finally, a spatiotemporal hybrid attention mechanism was used to dynamically fuse features from spatial and temporal branches, generating highly discriminative spatiotemporal joint features. The beneficial effect is a significant improvement in the representation ability of spatiotemporal patterns of complex traffic events, providing a crucial information foundation for subsequent accurate classification.

[0041] The specific implementation process for step S3 is as follows: Scenario: Continue detecting vehicle breakdown events in the right lane of the highway. S2 has output a sequence of 8 consecutive frames (approximately 0.3 seconds) of feature maps [F1, F2, ..., F8] that have fused radar and video information.

[0042] Implementation process: Spatial feature extraction (parallel processing of each frame): The fused feature map F_t of each frame is input into the spatial branch (a network consisting of multiple DCNv3 layers).

[0043] For frame 5 (F5), the DCNv3 network learns an offset so that the sampling points of its convolutional kernels are focused on key parts of the broken-down vehicle in the image (such as the vehicle's outline, stationary tires, and possibly hazard lights), rather than sampling the entire image region uniformly.

[0044] After processing, each frame is converted into a highly refined spatial feature vector S_t, which represents the most important spatial information within that frame. The final result is the spatial feature sequence [S1, S2, ..., S8].

[0045] Temporal feature extraction (processing the entire sequence): Input the spatial feature vector sequence [S1, S2, ..., S8] corresponding to the 8 frames into the temporal branch (BiGRU) in chronological order.

[0046] The BiGRU unit scans the sequence from left to right (forward) and from right to left (backward).

[0047] Forward scan: Captures the process of a target appearing -> moving at a speed significantly lower than the traffic flow -> coming to a complete stop.

[0048] Backward scan: reinforces the fact that it has now stopped in this state and backtracks to confirm that its motion trajectory before stopping was abnormal.

[0049] At each time step t, BiGRU outputs a temporal feature vector T_t that integrates past and future context information. We typically take the output T8 of the last time step as a summary of the entire sequence, or aggregate the outputs of all time steps.

[0050] Spatiotemporal hybrid attention fusion: The feature S8 representing the final spatial state and the feature T8 (or aggregated feature) representing the entire temporal process are input together into an attention module.

[0051] This module performs calculations: the current task is to determine if a vehicle has broken down. S8 provides strong spatial evidence that there is a stationary vehicle there, while T8 provides strong temporal evidence that the vehicle's movement was an anomalous process from motion to rest. Both are extremely important.

[0052] The attention mechanism is calculated to assign high weights to both (e.g., 0.55 for T8 and 0.45 for S8), and then they are weighted and fused together.

[0053] Final output: Generates a fixed-length, highly condensed multimodal spatiotemporal joint feature data of a highway traffic event. This feature vector encodes the complete spatiotemporal proposition that there is an object at position (x,y) that underwent an abnormal deceleration process until it came to a stop.

[0054] Output: This combined feature data will be directly fed into the S4 classification model. The classifier receives clear and discriminative features, classifies it as a vehicle breakdown event, and assigns a high anomaly score.

[0055] S4. Based on multimodal spatiotemporal joint feature data of highway traffic incidents, establish a highway traffic incident detection and classification model, and evaluate the anomaly score and probability distribution vector of the event category of highway traffic incidents. Step S4 specifically includes: A serial highway traffic incident detection and classification model is established based on SVM support vector machine. Based on the cascade highway traffic incident detection and classification model, the multimodal spatiotemporal joint feature data of highway traffic incidents is used as input, and the abnormal scores of highway traffic incidents quantified by the first detector of the model are used as output. The multimodal spatiotemporal joint feature data of highway traffic incidents corresponding to the abnormal scores of positively correlated traffic incidents are selected as input to the second detector to generate the classification probability of highway traffic incident categories. Combining the abnormal scores of highway traffic incidents and the classification probability of highway traffic incident categories, the probability distribution vector of abnormal scores and event categories of highway traffic incidents is output. Specifically, the series highway traffic incident detection and classification model is as follows:

[0056] In the formula, Anomaly scoring for highway traffic incidents. Classifying highway traffic incidents This is an abnormal rating matrix. For category classification matrix, For classification bias, Softmax is the normalization function. For the number of categories, 't' represents the index of the time step and the index of the spatial location, respectively. This provides multimodal spatiotemporal joint feature data of highway traffic events with time step t and spatial location s.

[0057] When using it, please refer to the steps outlined above: As a further development, a cascaded SVM classification model is employed. First, a first-level detector (such as a single-class SVM) quantifies the anomaly level of the input features, efficiently filtering out a large number of normal samples to address the imbalance between positive and negative samples. Then, only features corresponding to high anomaly scores are fed into a second-level multi-class SVM for accurate event classification, generating a probability distribution. Finally, the anomaly score is combined with the classification probability for output. The beneficial effects are a significant improvement in system processing efficiency and real-time performance, while the rich probabilistic output provides quantifiable confidence levels for subsequent decision-making.

[0058] The specific implementation process for step S4 is as follows: Scenario: Continuing with the previous vehicle breakdown detection. S3 outputs a spatiotemporal joint feature vector representing the suspected event.

[0059] Implementation process: Level 1 Detection - Anomaly Scoring: The feature vector is then input into the first detector (anomaly scoring SVM).

[0060] Because the characteristics of a stranded vehicle (a stationary obstacle) are drastically different from those of a continuously flowing traffic flow, the deviation is extremely large. The first detector outputs a very high anomaly score, such as 92 out of 100.

[0061] Because the score of 92 is much higher than the preset threshold (e.g., 50), the system determines that the scene is highly abnormal and passes the feature vector and its abnormality score to the second detector.

[0062] Level 2 Detection - Event Classification: Input filtering: Since the anomaly score of 92 is positively correlated, this data is allowed to enter the second level.

[0063] The feature vector is input into the second detector (a multi-class SVM classifier). This classifier is trained on various labeled samples collected, such as those showing breakdowns, accidents, congestion, and normal traffic.

[0064] The classifier analyzes features that contain both strong spatial information (the shape of a stationary object) and temporal information (the process from motion to rest). These features best match the support vectors of the anchor category.

[0065] The second detector outputs a probability distribution vector, for example: [anchorage: 0.88, traffic accident: 0.09, congestion: 0.03].

[0066] Combined output: The system ultimately outputs a joint result: Anomaly Score: 0.92; Event Category Probability Distribution: {'Breakdown': 0.88, 'Accident': 0.09, 'Congestion': 0.03} This result can be interpreted as follows: the system is 92% confident that the current scenario is abnormal, and among the abnormal events, there is an 88% chance that the vehicle has broken down, a 9% chance that it is a traffic accident, and a 3% chance that it is congestion.

[0067] Output: This vector, containing anomaly scores and probability distributions, provides all the necessary data input for the final Bayesian decision verification in step S5. S5 can then further integrate modality confidence scores to make the final decision.

[0068] S5. Based on Bayesian decision theory, using the probability distribution vector of anomaly scores and event categories based on highway traffic events, analyze video modal confidence and radar modal confidence, verify the occurrence probability of traffic event types under given video streams and radar point clouds, and complete highway traffic event detection. Step S5 specifically includes: Based on the abnormal scores of highway traffic events and the probability distribution vector of event categories, calculate the modal confidence of highway video streams and the modal confidence of highway radar point clouds. Based on the modal confidence of highway video streams and the modal confidence of highway radar point clouds, a Bayesian decision theory function is established to analyze the visual likelihood of highway video streams and the likelihood of radar point clouds, and to jointly evaluate the probability of traffic event types occurring under a given video stream and radar point cloud. Highway traffic event detection is achieved by classifying highway traffic events based on the probability of occurrence of traffic event types in a given video stream and radar point cloud.

[0069] Step S5 also includes: The calculation of the modal confidence scores of the highway video stream and the highway radar point cloud is specifically as follows:

[0070] In the formula, Modal confidence of highway video stream For the modal confidence of radar point cloud on highways Transpose the classification vector for highway traffic incidents. This is the weight matrix. This is the radar data feature map corresponding to time step t; Specifically, the Bayesian decision theory function is as follows:

[0071] In the formula, Given a video stream and the probability of traffic event types occurring in radar point clouds, Let visual likelihood function be the conditional probability of visual evidence given an event type. Given the event type, the likelihood function of the radar point cloud represents the conditional probability of radar evidence. The unconditional probability of occurrence of an event type. The event type to be determined is 'event', which represents all possible event types for the highway.

[0072] When using it, please refer to the steps outlined above: As a further development, based on Bayesian decision theory, the independent confidence levels of video and radar modes are first derived from the anomaly scores and classification probabilities output by S4. Then, using these confidence levels as weighting factors, a weighted likelihood function is constructed to adaptively fuse the dual-modal classification probabilities. The maximum posterior probability, given all observed evidence, is calculated as the final event determination criterion. Its beneficial effect is that it achieves dynamic and reliable cross-modal fusion through a theoretically optimal statistical inference framework, significantly improving the system's decision robustness and result interpretability in sensor-constrained scenarios.

[0073] The specific implementation process for step S5 is as follows: Scenario: Continue vehicle breakdown detection. From stage S4, we obtained the following results: Video branch: Anomaly score P_anomaly_v = 0.92 The event classification probability vector G_event_v = [Breakdown: 0.88, Accident: 0.09, Congestion: 0.03] (Here, G_event is the probability distribution vector). Radar branch: We have extracted a high-level feature vector H_t^radar from the radar data feature map Z_t^radar.

[0074] The event classification probability vector G_event_r = [Breakdown: 0.92, Accident: 0.07, Congestion: 0.01] a) Calculate the video modal confidence score E^u: According to the formula E^u = Sigmoid(P_anomaly · G_event^T) b) P_anomaly is a scalar of 0.92.

[0075] G_event is a row vector [0.88, 0.09, 0.03], and its transpose G_event^T is a column vector.

[0076] The inner product P_anomaly · G_event^T represents a comprehensive index obtained by weighting or modulating the classification probability vector G_event with the anomaly score P_anomaly. The final result is a scalar value.

[0077] A reasonable calculation is to take the product of the maximum probability in G_event and P_anomaly as the input.

[0078] Input = P_anomaly_v max(G_event_v) = 0.92 0.88 = 0.8096 Calculate E^u: E^u = Sigmoid(0.8096) ≈ 0.692 (because Sigmoid(0.8)≈0.69) b) Calculate the radar modal confidence level E^r: According to the formula E^r = AvgPool(M · H_t^radar), M ∈ R^(N×K) H_t^radar is a high-level feature vector (with dimensions K) of the radar data at time step t.

[0079] M is a learnable weight matrix (dimension N×K) used to map radar features to another space and compute attention or importance.

[0080] The result of the M · H_t^radar operation is an N-dimensional vector.

[0081] The AvgPool operation performs global average pooling on this N-dimensional vector, compressing it into a scalar that represents the confidence level of the radar mode.

[0082] The final output is a value E^r = 0.75.

[0083] 2. Bayesian Decision Fusion Input: Video modal confidence E^u = 0.692, Radar modal confidence E^r = 0.75 The event classification probabilities (i.e., likelihood P(E|event)) of the two modalities in phase S4: P_v(E^u | event): That is G_event_v = [Breakdown: 0.88, Accident: 0.09,Congestion: 0.03] P_r(E^r | event): That is G_event_r = [Breakdown: 0.92, Accident: 0.07,Congestion: 0.01] Applying Bayes' theorem: F(event|E^u,E^r) = [ P(E^u|event) P(E^r|event) P(event) ] / { Σ_{event'} [ P(E^u|event') P(E^r|event') P(event') ]} We calculate the posterior probability F(Breakdown|E^u, E^r) of the event Breakdown.

[0084] The prior probability P(event) is uniformly distributed, i.e.: P(Breakdown) = P(Accident) = P(Congestion) = 1 / 3 ≈ 0.333 Molecular calculations: P(E^u | Breakdown) P(E^r | Breakdown) P(Breakdown) = 0.88 0.92 0.333 ≈ 0.269; Denominator calculation (summing across all event types): Breakdown: 0.88 0.92 0.333 ≈ 0.269 Accident: 0.09 0.07 0.333 ≈ 0.0021 Congestion: 0.03 0.01 0.333 ≈ 0.0001 The sum of the denominators ≈ 0.269 + 0.0021 + 0.0001 = 0.2712 Calculate the posterior probability: F(Breakdown|E^u, E^r) = 0.269 / 0.2712 ≈ 0.992 3. Final decision and testing completed. The calculated posterior probability indicates that, given the video and radar evidence and their confidence levels, the probability that the event is a vehicle breakdown is 99.2%, which is much higher than the preset alarm threshold (e.g., 80%).

[0085] The system's final output: A vehicle breakdown event was detected, with a confidence level of 99.2% based on Bayesian decision-making.

[0086] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A method for detecting traffic incidents on highways via video, characterized in that, include: S1. Obtain samples of several types of traffic incidents on highways and build an initial training database of highway traffic incidents. The traffic incident samples include: video stream data and radar point cloud data; S2. Perform data calibration on the initial original training dataset of highway traffic events according to multimodal preprocessing to generate a multimodal original training dataset of highway traffic events. S3. Establish an asymmetric spatiotemporal encoder and extract spatiotemporal joint features from the original training data set of multimodal highway traffic events to obtain multimodal spatiotemporal joint feature data of highway traffic events. S4. Based on multimodal spatiotemporal joint feature data of highway traffic incidents, establish a highway traffic incident detection and classification model, and evaluate the anomaly score and probability distribution vector of the event category of highway traffic incidents. S5. Based on Bayesian decision theory, using the probability distribution vector of anomaly scores and event categories based on highway traffic events, analyze video modal confidence and radar modal confidence, verify the occurrence probability of traffic event types under given video streams and radar point clouds, and complete highway traffic event detection.

2. The method for detecting highway traffic incidents via video according to claim 1, characterized in that, Step S2 specifically includes: For radar point cloud data in the original training dataset of highway traffic incidents, the polar coordinate system is converted to the Cartesian coordinate system to calibrate the spatial coordinate alignment between radar point cloud data and video stream data. For video stream data in the training dataset of highway traffic incidents, time synchronization between the timestamps of the video stream data and the timestamps of the radar point cloud data is determined by time interpolation through a sliding window. Based on radar point cloud data in the original training dataset of highway traffic incidents, a voxel network is established, voxel features are statistically analyzed, and then substituted into a 3D sparse convolutional neural network to extract multi-layer 3D voxel features from the radar point cloud data.

3. The method for detecting highway traffic incidents via video according to claim 2, characterized in that, Step S2 also includes: By taking the maximum value along the Z-axis and compressing the multi-layer 3D voxel features of the radar point cloud data, a multi-layer 2D bird's-eye view feature map of the radar point cloud data is obtained. By using the camera's intrinsic parameter matrix, each pixel in the multi-layer 2D bird's-eye view feature map of the radar point cloud data is projected onto the pixel coordinates of the original image to obtain the radar point cloud data feature map. Based on the video stream data in the original training dataset of highway traffic incidents, the EfficientNet network is substituted into it. The video stream data is used as input and the visual feature map of the video stream data is used as output. The CBAM module and FPN structure are used to change the visual feature map to enhance spatial attention and fuse low-level details and high-level semantics. Based on visual feature maps from video stream data and feature maps from radar point cloud data, a matrix of multimodal data was stitched together to obtain a set of original multimodal training data for highway traffic events.

4. The method for video detection of highway traffic incidents according to claim 3, characterized in that, Step S3 specifically includes: Based on deformable convolution DCNv3, multi-scale features of feature maps in the original training dataset of multimodal highway traffic events are extracted to obtain spatial branch multi-scale feature parameters. Based on the time series prediction bidirectional gated recurrent BiGRU, the temporal dependency features in the original training dataset of multimodal highway traffic events are captured, and the temporal dependency feature parameters of the time branch are obtained. Based on spatial branch multi-scale feature parameters and temporal branch temporal dependent feature parameters, an asymmetric spatiotemporal encoder is established, and multimodal spatiotemporal joint feature data of highway traffic events is generated according to the spatiotemporal hybrid attention mechanism.

5. The method for video detection of highway traffic incidents according to claim 4, characterized in that, Step S4 specifically includes: A serial highway traffic incident detection and classification model is established based on SVM support vector machine. Based on the cascade highway traffic incident detection and classification model, the multimodal spatiotemporal joint feature data of highway traffic incidents is used as input, and the abnormal scores of highway traffic incidents quantified by the first detector of the model are used as output. The multimodal spatiotemporal joint feature data of highway traffic incidents corresponding to the abnormal scores of positively correlated traffic incidents are selected as input to the second detector to generate the classification probability of highway traffic incident categories. Combining the abnormal scores of highway traffic incidents and the classification probability of highway traffic incident categories, the probability distribution vector of abnormal scores and event categories of highway traffic incidents is output. Specifically, the series highway traffic incident detection and classification model is as follows: ; In the formula, Anomaly scoring for highway traffic incidents. Classifying highway traffic incidents This is an abnormal rating matrix. For category classification matrix, For classification bias, Softmax is the normalization function. For the number of categories, 't' represents the index of the time step and the index of the spatial location, respectively. Multimodal spatiotemporal joint feature data of highway traffic events at time step t and spatial location s.

6. The method for video detection of highway traffic incidents according to claim 5, characterized in that, Step S5 specifically includes: Based on the abnormal scores of highway traffic events and the probability distribution vector of event categories, calculate the modal confidence of highway video streams and the modal confidence of highway radar point clouds. Based on the modal confidence of highway video streams and the modal confidence of highway radar point clouds, a Bayesian decision theory function is established to analyze the visual likelihood of highway video streams and the likelihood of radar point clouds, and to jointly evaluate the probability of traffic event types occurring under a given video stream and radar point cloud. Highway traffic event detection is achieved by classifying highway traffic events based on the probability of occurrence of traffic event types in a given video stream and radar point cloud.

7. The method for video detection of highway traffic incidents according to claim 6, characterized in that, Step S5 also includes: The calculation of the modal confidence scores of the highway video stream and the highway radar point cloud is specifically as follows: ; In the formula, Modal confidence of highway video stream For the modal confidence of radar point cloud on highways Transpose the classification vector for highway traffic incidents. This is the weight matrix. This is the radar data feature map corresponding to time step t; Specifically, the Bayesian decision theory function is as follows: ; In the formula, Given a video stream and the probability of traffic event types occurring in radar point clouds, Let visual likelihood function be the conditional probability of visual evidence given an event type. Given the event type, the likelihood function of the radar point cloud represents the conditional probability of radar evidence. The unconditional probability of occurrence of an event type. The event type to be determined is 'event', which represents all possible event types for the highway.

Citation Information

Cited By

  • Highway traffic situation intelligent early warning method

    CN121122029A