Traffic event detection method based on deep space-time memory and interactive network

By applying deep spatiotemporal memory and interactive network methods in traffic video processing, combining spatiotemporal information perception network and timing feature learning network, the problem of traffic event identification and classification in traffic videos by drivers in the first perspective is solved, and more accurate traffic event positioning and classification are achieved.

CN120032291AActive Publication Date: 2025-05-23SHENZHEN DAJI TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510067732.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-23
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and classify traffic events in traffic videos from the driver's first perspective, especially when dealing with videos in dynamic and complex contexts captured by dash recorders.

Method used

A traffic event detection method based on deep spatiotemporal memory and interactive network is designed. Through the combination of spatiotemporal information perception network and timing feature learning network, the spatial and temporal information characteristics in video frames are extracted to achieve accurate positioning and classification of traffic events.

Benefits of technology

This method can more effectively extract the spatial and temporal information characteristics in the video frame, improve the accuracy of identification and classification of traffic events, especially when processing dash recorder video, it can more accurately locate the start time of traffic events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032291A_ABST
    Figure CN120032291A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic incident detection method based on deep space-time memory and an interactive network. Firstly, a space-time information sensing network is designed, core semantics from a small space area to a large space area and short-term features in a time domain can be effectively captured from input video frames, the extraction capability of video frame features is improved, and traffic events are classified more accurately. Secondly, a time sequence feature learning network is provided, and the problem that long-term time information cannot be fully captured when a spatio-temporal information sensing network processes video frames is solved. According to the method, more complete context time information features can be obtained by strengthening learning of the time features, and accurate positioning of exception occurrence and end moments is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent video processing, and in particular relates to a traffic event detection method based on deep spatiotemporal memory and interactive network. Background Art

[0002] Autonomous vehicles still face significant challenges in accident detection and prediction, and this issue has attracted widespread attention in industry and academia. In recent years, the popularity of dashboard cameras has made them an important tool for recording accident scenes. This unique advantage not only improves the ability to record accident scene data, but also provides richer sample data for the training of deep learning models. Models based on dashcams will become more reliable, significantly improving the decision-making ability of autonomous vehicles in accident scenarios.

[0003] However, the video captured by dashcams has significantly different characteristics from the data from highway or traffic light surveillance cameras. Dashcams record traffic videos in a horizontal view, and the camera and surrounding objects are in motion in the video, which increases the complexity of traffic event detection, especially when classifying objects close to the dashcam and traffic participants. Therefore, it is particularly important to identify traffic events (including abnormal situations of vehicles and pedestrians) in natural driving videos captured by on-board cameras. This task is not only applicable to road safety warning, autonomous driving, and traffic flow control, but also can effectively promote pedestrian protection. Although many researchers have devoted themselves to anomaly detection methods under static monitoring in recent years, these methods often find it difficult to achieve ideal traffic event detection performance due to the dynamic and complex background of driving videos.

[0004] Models applied to traffic incident detection in dashcams usually adopt a hybrid deep learning framework, combining a convolutional neural network (CNN) with a recurrent neural network (RNN) to identify potential hazards in dashcam videos. These methods use CNN to extract features and locate objects in images, and the extracted spatial features are then input into the RNN to capture temporal patterns and understand the evolution of these spatial features over time. For example, some studies use Mask R-CNN for semantic segmentation, and then combine CNN-LSTM networks to classify dangerous lane change detection based on mask-based frame coverage. However, unguided deep learning models may produce spurious patterns, and CNN models are also limited in extracting unobservable contextual relationships.

[0005] The current research on combining the spatiotemporal information of traffic videos needs further study. Therefore, designing a model that can simultaneously consider the differences in the motion changes of the dashcam and the characteristics of the target scale changes in the video, combining the spatial features and long-term and short-term time information in the video frame, and quickly detecting and accurately locating the start time of traffic events is still an urgent problem to be solved. Summary of the invention

[0006] Purpose of the invention: The purpose of the present invention is to provide a traffic incident detection method based on deep spatiotemporal memory and interactive networks. First, a spatiotemporal information perception network is designed, which can effectively capture the core semantics from small spatial areas to large spatial areas, as well as short-term features in the time domain from the input video frames, improve the ability to extract video frame features, and classify traffic events more accurately. Secondly, a temporal feature learning network is proposed to solve the problem that the spatiotemporal information perception network fails to fully capture long-term temporal information when processing video frames. By strengthening the learning of temporal features, this method can obtain more complete contextual temporal information features and improve the accurate positioning of the occurrence and end times of anomalies.

[0007] Technical solution: The present invention is a traffic incident detection method based on deep spatiotemporal memory and interactive network, which constructs a detection model that combines a spatiotemporal information perception network and a temporal feature learning network, takes video frames as input, and realizes the location of traffic accidents and the prediction of the start time of occurrence, including the following steps:

[0008] Step 1: Build a spatiotemporal information perception network, take video frames as input, and introduce MarphFC in the spatial dimension s (Fully connected layer in spatial domain) mines spatial semantic information and introduces MorphFC on the temporal channel t (Time Domain Fully Connected Layer) captures the relationship between input video features and time, and remaps the features of the time dimension to the spatial dimension, thereby extracting the spatial and temporal information features in the video frame;

[0009] Step 2: Build a temporal feature learning network to enhance the learning ability of temporal information features in video frames and comprehensively model the temporal information in video frames;

[0010] Step 3: Train the detection model using weighted cross entropy loss.

[0011] Furthermore, step 1 is as follows: Introduce MarphFC in the MFCMLP (temporal and spatial fully connected multilayer perceptron) module s (spatial domain fully connected layer) to hierarchically expand the receptive field of FC and make it run from a small area to a large area; the MarphFC cited s(Fully connected layer in spatial domain) processes each frame of video independently in the horizontal and vertical paths; taking the horizontal path as an example, given an input video frame X∈R HW×C , which is projected into a token sequence, first splits X horizontally, setting the block length to L, thus obtaining

[0012] X i ∈R L×C (1)

[0013] Among them, i∈{1,…,HW / L}, in order to reduce the computational cost, each X is transformed along the channel dimension i Divide into multiple groups, each of which has D channels, and get segmented blocks, each of which is

[0014] X i k ∈R LD (2)

[0015] Where k∈{1,…,C / D}, each data block is flattened into a 1D vector and the FC weight matrix W∈R is applied LD ×LD Each data block is transformed, that is, Y is

[0016]

[0017] After feature transformation, all blocks Y i k Reshape to the original dimension Y∈R H×W×C , the same is true for the vertical direction, except that the tag sequence is split vertically; the FC layer is applied to process each tag separately so that the groups can communicate along the channel dimension; finally, the horizontal, vertical and channel features are summed element by element to get the output result; as the network goes deeper, the block length L will increase hierarchically, so that the FC filter gradually discovers more core semantics from small to large spatial regions;

[0018] When processing traffic videos, it is necessary not only to process the information in the horizontal and vertical paths of continuous video frames, but also to process the information in the time channel. Since the MLP in VST usually operates independently on a position-by-position basis, that is, the feature vector at each position is independently linearly transformed through the fully connected layer, it is impossible to directly capture the correlation between different positions and the dependency on the time dimension. Therefore, the present invention also introduces MorphFC in MFCMLP t (Time Domain Fully Connected Layer), this module transforms the input feature map in the time dimension, flattens it, fully connects it, and then restores it to a high-dimensional space; this processing method can be used to capture the relationship between the input video features and time, and remap these time dimension features to the spatial dimension.

[0019] Furthermore, the introduction of MorphFC in MFCMLP t (Time Domain Fully Connected Layer), this module transforms the input feature map in the time dimension, flattens it, fully connects it, and then restores it to a high-dimensional space. Specifically:

[0020] Given an input video segment labeled X∈R H×W×T×C , first divide X into several groups along the channel dimension to reduce the computational cost, with D channels in each group, and get

[0021] X k ∈R H×W×T×D (4)

[0022] Among them, k∈{1,…,C / D}, for each spatial position s, the features of all frames are concatenated into a block where s∈{1,…,HW};

[0023] Reference FC matrix W∈R TD×TD Transform the time features to obtain:

[0024]

[0025] All blocks Y s k ∈R TD Resize back to the original tokens dimension and output Y∈R H×W×T×C In this way, the FC filter can simply aggregate the token relationships within the block along the time dimension, providing VST with stronger time dimension modeling capabilities. s A residual connection is used between the original input and output of the (spatial domain fully connected layer) to increase the stability of model training.

[0026] Furthermore, step 2 is specifically as follows: in the input time series, assume that the input time series is Where T represents the time step, after a 1×1 convolution transformation, we get:

[0027]

[0028] Among them, W 1 represents the weight parameter, which is the weight matrix of the above linear transformation, b 1 represents the bias in the above linear transformation.

[0029] The parallel processing part of the dilated convolution mainly includes two parallel convolution processing blocks. Each convolution block contains two dilated convolution layers. For each dilated convolution layer, the output of the lth layer is expressed as:

[0030] X (l) =DilatedConv(X (l-1) ,dilation=d l ) (7)

[0031] Among them, d l refers to the expansion rate of the lth layer;

[0032] Suppose the output of the first dilated convolutional block is X 1 , the output of the second dilated convolutional block is X 2 , the specific processing process of each block is as follows:

[0033] X 1 =DilatedConv 1 (W 1 Z+b 1 ) (8)

[0034] X 2 =DilatedConv 2 (W 2 Z+b 2 ) (9)

[0035] Among them, W 2 is the weight parameter in the dilated convolution operation, b 2 represents the bias in the second dilated convolutional layer.

[0036] The information obtained after residual connection and subsequent processing is:

[0037] X=X 1 +X 2 (10)

[0038] Then after a series of nonlinear activation and regularization processing, we get:

[0039] X norm =WeightNorm(X) (11)

[0040] X relu =Relu(X) (12)

[0041] X dropout = Dropout(X) (13)

[0042] The final output is:

[0043] Y=W 3 Dropout(Relu(WeightNorm(X 1 +X 2 )))+b 3(14)

[0044] Among them, W 3 represents the weight matrix of the final output layer, b 3 It represents the bias of the final output layer.

[0045] Furthermore, step 3 is specifically as follows:

[0046] For each video frame F[t], the final output anomaly score S[t]∈[0,1] is obtained, where 0 indicates no anomaly and 1 indicates an anomaly in the frame. A higher weight is given to the anomaly class that reflects the data distribution, and a weighted cross entropy loss is selected to train the model. For each weight ω i , the formula used is as follows:

[0047] ω i =e / e i (15)

[0048] Where e is the total number of examples in the dataset, e i is the number of examples of class i.

[0049] The present invention also discloses a computer device, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method of the present invention.

[0050] The present invention also discloses a computer-readable storage medium on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the method of the present invention are implemented.

[0051] The present invention also discloses a computer program product, comprising a computer program / instruction, which implements the steps of the method of the present invention when executed by a processor.

[0052] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0053] The present invention designs a deep spatiotemporal memory and interaction network, which extracts local to global spatiotemporal information in traffic videos through two different stages, and obtains a richer real-time description of traffic events in traffic videos with the driver as the first-person perspective without speculating on the future. Specifically: In the spatiotemporal information perception network, by introducing MarphFC in the spatial dimension s (spatial domain fully connected layer) to mine spatial semantic information, and use MorphFC on the temporal channel t(Fully connected layer in time domain) to capture the changes of video features over time. These temporal features will be remapped to spatial dimensions, solving the tedious operations of expanding, reshaping and folding video frames in the original sliding window in the video swin transformer (VST). s The residual connection method is used between the original input and output of the (spatial domain fully connected layer) to increase the stability of model training, so as to more effectively extract the spatial information and short-term temporal information in the video frame. In the temporal feature learning network, since the first stage mainly extracts short-term temporal information, in order to better locate the start time of the traffic accident, a dual-branch residual structure of the dilated convolution is used to extract the contextual long-term temporal information features. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is the overall framework diagram of the deep spatiotemporal memory and interactive network of the present invention.

[0055] Figure 2 This is the architecture diagram of the spatiotemporal information perception network.

[0056] Figure 3 is the MarphFC in the spatial dimension s (Fully connected layer in spatial domain)Fig.

[0057] Figure 4 is MarphFC acting on the time dimension t picture.

[0058] Figure 5 Learning networks for temporal features.

[0059] Figure 6 This is the principle diagram of the hole convolution.

[0060] Figure 7 A visual comparison of traffic incident detection examples on the DoTA dataset. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] The present invention proposes a traffic incident detection method based on deep spatiotemporal memory and interactive network, comprising the following steps:

[0063] The invention aims to solve the problem that in the traffic video from the driver's first-person perspective, obstacles appear in different sizes due to distance differences, which causes the size of the traffic event target to be identified to change accordingly, making it impossible to better identify traffic events and better classify and locate them. A traffic event detection method based on deep spatiotemporal memory and interactive network is proposed. The overall block diagram is as follows Figure 1 As shown in Figure 2, the model mainly consists of two parts: a spatiotemporal information perception network and a temporal feature learning network. In the spatiotemporal information perception network, MarphFC is introduced in the spatial dimension. s (Fully connected layer in spatial domain) mines spatial semantic information and introduces MorphFC on the temporal channel t (Fully connected layer in time domain) captures the relationship between input video features and time changes, and remaps these time dimension features to the spatial dimension, so as to better extract spatial information and short-term time information features in the video frame. In the temporal feature learning network, since the short-term time information features are extracted in the first stage of processing, in order to better locate the time when the traffic accident occurred, the dual-branch residual structure of the dilated convolution is used in this network structure to better extract long-term time information features. Finally, based on the local and global spatiotemporal information features extracted by the deep spatiotemporal memory and the interactive network, the time of the traffic incident and the category of the traffic incident can be more accurately located.

[0064] In order to better process the spatiotemporal features in traffic videos, the present invention proposes a spatiotemporal information perception network, which can effectively capture the spatial and temporal domain features in video frames. Specifically, after preprocessing the input video frames, the network processes different scales by dividing the obtained token sequence into blocks, and gradually captures spatial and temporal information from local to global by gradually expanding the length of the blocks. This processing method can expand from local information to global information in the spatial domain, and can also gradually capture information features in the temporal domain. Through this multi-scale spatiotemporal perception, the network can more comprehensively extract key features in video frames, so that in subsequent traffic video anomaly detection, abnormal events can be classified and abnormal time can be accurately located, so as to achieve better detection effects. The overall framework is as follows: Figure 2 ((a) is the overall architecture of the spatiotemporal information perception network, and (b) is the schematic diagram of the spatiotemporal fully connected multilayer perceptron in the spatiotemporal information perception network).

[0065] In the process of processing traffic videos, mining the core spatiotemporal semantic information is crucial for the recognition and classification of traffic accidents in videos. The current typical CNN and MLP only focus on modeling local or global information, so they cannot effectively capture the local to global information in the video sequence. In order to solve this problem, the present invention designs the MFCMLP module, in which the MarphFC is first introduced. s (Fully connected layer in spatial domain), which can hierarchically expand the receptive field of FC and make it run from small area to large area. s (Fully connected layer in spatial domain) processes each frame of video independently in horizontal and vertical paths. As shown in the following figure, the present invention mainly takes a horizontal block as an example (such as Figure 3 medium green block).

[0066] Specifically, given an input video frame X∈R HW×C , which is projected into a token sequence, the present invention first divides X horizontally and sets the block length to L, so that

[0067] X i ∈R L×C (1)

[0068] Where i∈{1,…,HW / L}. In order to reduce the computational cost, the present invention also transforms each X i is divided into multiple groups, each of which has D channels. Therefore, we can get segmented blocks, each of which is

[0069] X i k ∈R LD (2)

[0070] Next, each data block is flattened into a 1D vector and the FC weight matrix W∈R is applied. LD×LD Each data block is transformed, that is, Y is

[0071]

[0072] After feature transformation, all blocks Y i k Reshape to the original dimension Y∈R H×W×C . Vertical direction ( Figure 3The same is true for the yellow blocks in , except that the tag sequence is split vertically. In order to allow communication between groups along the channel dimension, the present invention also applies FC layers to process each tag separately. Finally, the horizontal, vertical and channel features are summed element by element to obtain the output result. As the network goes deeper, the block length L will increase hierarchically, allowing the FC filter to gradually discover more core semantics from small to large spatial regions.

[0073] When processing traffic videos, it is necessary not only to process the information in the horizontal and vertical paths of continuous video frames, but also to process the information in the time channel. Since the MLP in VST usually operates independently on a position-by-position basis, that is, the feature vector at each position is independently linearly transformed through the fully connected layer, it is impossible to directly capture the correlation between different positions and the dependency in the time dimension. Therefore, the present invention also introduces MorphFC in MFCMLP. t (Time Domain Fully Connected Layer), this module transforms the input feature map in the time dimension, flattens it, fully connects it, and then restores it to a high-dimensional space. This processing method can be used to capture the relationship between the input video features and time, and remap these time dimension features to the spatial dimension. Its principle and network architecture are as follows Figure 4 shown.

[0074] Assume that, given an input video segment labeled X∈R H×W×T×C , first divide X into several groups (each group has D channels) along the channel dimension to reduce the computational cost, and get

[0075] X k ∈R H×W×T×D (4)

[0076] Where k∈{1,…,C / D}. For each spatial position s, the present invention concatenates the features of all frames into a block where s∈{1,…,HW}.

[0077] Next, the present invention refers to the FC matrix W∈R TD×TD Transform the time features to obtain:

[0078]

[0079] Finally, the present invention converts all blocks Y s k ∈R TD Resize back to the original tokens dimension and output Y∈R H ×W×T×CIn this way, the FC filter can simply aggregate the token relationships within the block along the time dimension, providing VST with stronger time dimension modeling capabilities. s A residual connection is used between the original input and output of the (spatial domain fully connected layer) to increase the stability of model training.

[0080] Since the spatiotemporal information perception network can learn the dependency of the input frames by aggregating the time tags at each spatial position. However, there is a lot of local and global time information between video frames that needs to be processed to obtain more complete time features, so as to more accurately locate the time when the anomaly occurs. The present invention proposes a time series feature learning network (as shown below): Figure 5 As shown in Figure 3, the learning ability of temporal features in the processed feature graph is improved.

[0081] exist Figure 5 In the 1×1 convolution operation, only the temporal information of adjacent frames is considered. This operation can effectively capture the short-term dependencies between video frames and extract local temporal features. Secondly, the present invention uses parallel processing of dilated convolutions with a dilation rate of 3 in the time domain. This parameter enables the convolution kernel of the same size to obtain a larger receptive field without increasing the computational complexity. Figure 6 (Where (a) is the initial input convolution kernel that sets the receptive field through the expansion coefficient d; (b) is the feature extraction and expansion of the receptive field and the generation of the intermediate layer output through the dilation coefficient d; (c) is the further extraction of high-level features and the use of the final output for subsequent tasks.) As shown. At the same time, the present invention designs two serial structures to form a time series feature learning network in order to supplement the time information of the intervals of the input time series that are omitted under the dilation rate of 3 dilation convolution. Through such a design, even with a smaller convolution kernel, features of a longer time range can be captured, thereby processing global time information. This method can aggregate information from more distant time frames without losing local details.

[0082] exist Figure 5 In the input time series, assume that the input time series is Where T represents the time step, after a 1×1 convolution transformation, we get:

[0083]

[0084] The parallel processing part of the dilated convolution mainly includes two parallel convolution processing blocks. Each convolution block contains two dilated convolution layers. For each dilated convolution layer, the output of the lth layer can be expressed as:

[0085] X (l) =DilatedConv(X(l-1) ,dilation=d l ) (7)

[0086] Among them, d l Refers to the expansion rate of the lth layer.

[0087] Assume that the output of the first dilated convolutional block is X 1 , the output of the second dilated convolutional block is X 2 , the specific processing process of each block is as follows:

[0088] X 1 =DilatedConv 1 (W 1 Z+b 1 ) (8)

[0089] X 2 =DilatedConv 2 (W 2 Z+b 2 ) (9)

[0090] The information obtained after residual connection and subsequent processing is:

[0091] X=X 1 +X 2 (10)

[0092] Then after a series of nonlinear activation and regularization processing, we get:

[0093] X norm =WeightNorm(X) (11)

[0094] X relu =Relu(X) (12)

[0095] X dropout = Dropout(X) (13)

[0096] The final output is:

[0097] Y=W 3 Dropout(Relu(WeightNorm(X 1 +X 2 )))+b 3 (14)

[0098] The temporal feature learning network proposed in the present invention processes local and global temporal information in the video through dilated convolution and causal convolution, and realizes comprehensive modeling of temporal information in video frames through multi-branch and multi-scale feature fusion.

[0099] For each video frame F[t], the final output anomaly score S[t]∈[0,1], where 0 indicates no anomaly and 1 indicates an anomaly in the frame. In order to give higher weights to the anomaly classes that reflect the data distribution, the present invention selects weighted cross entropy loss to train the model. For each weight ω i , the formula used in the present invention is as follows:

[0100] ω i =e / e i (15)

[0101] Where e is the total number of examples in the dataset, e i is the number of examples of class i.

[0102] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

[0103] Furthermore, simulation experiments are performed to verify the effectiveness of the method proposed in the present invention.

[0104] Simulation experiment 1:

[0105] The proposed model is trained and evaluated on a server running Nvidia GeForce RTX 3090, and experiments are performed using Pytorch. Secondly, in order to make the training of the model more stable, the present invention uses the SDG optimizer for training, with a batch size of 4, a learning rate of 0.00005, a momentum of 0.9, and a video segment length of 8. To train the network, the size of the input video frame is set to 640×480, and the dataset is set to 100 epochs on the DoTA dataset. The present invention uses weighted cross entropy loss for training to solve the problem of data imbalance in the DoTA dataset, and ω n =0.3 and ω a = 0.7 are assigned to normal and abnormal classes respectively. And follow ω i =e / e i equation.

[0106] Table 1 AUC (%) of different methods on DoTA dataset

[0107]

[0108] The proposed traffic incident detection method based on deep spatiotemporal memory and interactive network is compared with existing related anomaly detection methods, including LSTM, ConvAE, ConvLSTMAE, TAD, TAD+ML, FC, Encoder-Decoder, AnoPred, FOL-Ensemble, Movad, TRN and STFE. Among them, FC is a video-level detection method, TAD, TAD+ML and Ensemble are segment-level detection methods based on positioning; ConvAE, ConvLSTMAE, AnoPred, LSTM, Encoder-Decoder, Movad, TRN and STFE are frame-level detection methods. Table 1 shows the AUC results of different methods on the DoTA dataset.

[0109] It can be observed from Table 1 that the AUC of the proposed method reaches 80.9%, which is better than all other methods. Compared with the video-level detection method FC, the AUC of FC is 61.7%, which is less accurate. Because this method processes the entire video rather than fine-grained segments or frames, it has limitations in capturing details. The AUCs of segment-level methods such as TAD and TAD+ML based on positioning are 69.2% and 69.7%, respectively, which are slightly better than the video-level methods, because segment-level detection can more accurately locate the time period when the event occurs. However, since these methods are still processed at the segment level, some important details may be missed. Frame-level detection methods include LSTM, ConvAE, ConvLSTMAE, AnoPred, Encoder-Decoder, Movad, TRN and STFE, etc. They capture more detailed temporal information by frame-by-frame analysis, so they generally outperform video-level and segment-level methods in performance, among which STFE performs well in frame-level detection with an AUC of 79.3%. Therefore, the method proposed in this invention is significantly better than other methods in terms of AUC performance. The model uses two different networks to extract local to global spatiotemporal information in traffic videos, thereby more accurately detecting the category of traffic events and locating the start time of traffic accidents. Simulation Experiment 2:

[0110] As can be seen from Table 2, there are obvious differences in the AUC performance of different methods in various traffic accident classification tasks. Compared with other methods, the method of the present invention achieves the best AUC scores in the vast majority of traffic accident categories. Methods such as AnoPred, FOL-STD, and FOL-Ensemble show fluctuations in performance across different categories. Especially in complex traffic accident types (such as the OO category), the AUC scores of these methods are generally low. However, the method proposed by the present invention also achieves high performance in such complex accidents, further demonstrating the robustness of the model. The method proposed by the present invention is more powerful than STFE in the vast majority of scenarios (such as ST, LA, OC, TC, VP, VO, OO, and UK), especially showing significant advantages in the two traffic event detection tasks of ST and OO. The detection performance in the AH category is slightly lower, probably because the target obstacles in the driving video are too small or there are some suddenly flying debris and other situations, and the method proposed by the present invention has not yet fully handled the detection of such traffic events, but it has more advantages in overall performance. In Table 2, the method proposed by the present invention has the best performance in categories such as ST, AH, LA, OC, TC, VP, etc., especially leading significantly in categories such as ST, LA, OC, and VP. As shown in Table 3, in the classification of traffic accidents with non-self anomalies, the method proposed by the present invention performs excellently in categories such as *ST, *AH, *LA, *OC, *TC, *VP, *VO, *OO, and *UK. Generally speaking, based on the deep spatio-temporal memory and interaction network, in almost all traffic accident categories, the method proposed by the present invention has achieved the highest AUC, especially in key categories such as ST, LA, OC, TC, VP, VO, OO, and UK, demonstrating strong classification accuracy. The method proposed by the present invention has obtained the highest AUC value in most categories and has good robustness and generalization ability.

[0111] Table 2 Comparison of AUC in the detection of self-anomalous traffic events by different methods (%)

[0112] model ST AH LA OC TC VP VO OO UK AnoPred 69.9 73.6 75.2 69.7 73.5 63.3 N / A N / A N / A FOL-STD 67.3 77.4 71.1 68.6 69.2 65.1 N / A N / A N / A FOL-Ensemble 73.3 81.2 74.0 73.4 75.1 70.1 N / A N / A N / A Movad 75.1 71.8 70.3 72.2 73.5 70.7 73.2 69.8 / 73.0 65.1 STFE(SOTA) 75.2 84.5 72.1 77.3 72.8 71.9 N / A N / A N / A Ours 86.3 84.1 82.6 80.3 84.6 78.6 77.8 85.1 / 89.2 72.9

[0113] Table 3 Comparison of AUC in the detection of non-self-anomalous traffic events by different methods (%)

[0114] model *ST *AH *LA *OC *TC *VP *VO *OO *UK AnoPred 70.9 62.6 60.1 65.6 65.4 64.9 64.2 57.8 N / A FOL-STD 75.1 69.8 68.1 74.1 72.0 69.7 63.8 69.2 N / A FOL-Ensemble 77.5 69.8 68.1 76.7 73.9 71.2 65.2 69.6 N / A Movad 62.8 62.2 62.8 61.6 62.9 61.9 63.2 63.1 / 67.8 57.8 STFE(SOTA) 80.6 65.6 69.9 76.5 74.2 N / A 75.6 70.5 N / A Ours 79.5 78.5 78.4 80.2 80.4 75.4 77.7 76.1 / 82.2 69.6

[0115] Simulation Experiment 3:

[0116] From below Figure 7(1) It can be seen that when the car is driving normally on the road, the anomaly score tends to 0. When the car deviates from the road, the anomaly score rises significantly. When a collision occurs, the anomaly score is close to 1. This shows that the method proposed in the present invention can detect anomalies and locate the start and end time of the anomaly. Figure 7 In this paper, the present invention selects visualization examples of traffic incident detection of self-abnormal category and non-abnormal self category, and compares and analyzes them with the visualization result graph of SOTA. Figure 7 In the abnormal region in (1), the anomaly score of the SOTA method rises slowly and remains at the peak position for a short time before rapidly decreasing, showing a certain degree of volatility. In contrast, the anomaly score of the method proposed in the present invention reaches 1 more quickly in the abnormal region and remains stable, with smaller overall fluctuations. Figure 7 In (2), the anomaly score of the SOTA method fluctuates significantly in the anomaly area, including multiple small peaks, reflecting the instability of the detection. The anomaly score obtained by the method proposed in the present invention remains relatively stable in the anomaly area. Although there are some fluctuations, it presents a relatively stable high value overall. Figure 7 In (3), the anomaly score obtained by the SOTA method shows a gradual upward trend in the abnormal area, but drops rapidly after approaching 1, accompanied by large fluctuations. The anomaly score obtained by our method quickly reaches 1 in the abnormal area and remains stable without significant fluctuations. Figure 7 In (4), the output curve of the SOTA method fluctuates greatly in the abnormal area and fails to stably reach a high value, showing the instability of the detection of this type of traffic incident. In contrast, the anomaly score obtained by the method proposed in the present invention remains stable in the abnormal area and is close to 1 with almost no significant fluctuation. In summary, the two-stage spatiotemporal information fusion method shows a smoother anomaly score curve in all four scenarios given, showing higher accuracy and stability in traffic incident detection and classification.

[0117] Simulation experiment 4:

[0118] In order to better verify the effectiveness of the performance of the dual-stage spatiotemporal information combined network on the DOTA dataset, the present invention conducted an ablation experiment as shown in Table 4. First, when the spatiotemporal information perception network is used alone, the AUC of the model is 74.17% and the accuracy is 73.04%. Secondly, when only the temporal feature learning network is used. The AUC of the model is 72.08% and the ACC is 71.33%. When both the spatiotemporal information perception network and the temporal feature learning network are included, the AUC of the model is 80.9% and the ACC is 76.05%. At this time, the performance based on the deep spatiotemporal memory and interactive network is the highest. This shows that the model proposed by the present invention can extract more effective spatiotemporal information features from the video frame under the joint action of the spatiotemporal information perception network and the temporal feature learning network, and capture the dependency of different video frames in the spatiotemporal domain, so as to better classify traffic events and locate the start time of traffic events.

[0119] Table 4 Ablation experiments on the DoTA dataset

[0120] Spatiotemporal Information Awareness Network Temporal feature learning network AUC (%) Accuracy(%) w w / o 74.17 73.04 w / o w 72.08 71.33 w w 80.9 76.05

Claims

1. A traffic incident detection method based on deep spatiotemporal memory and interactive network, characterized in that: Construct a detection model that combines the spatiotemporal information perception network and the temporal feature learning network, taking video frames as input to locate traffic accidents and predict the start time of occurrence, including the following steps: Step 1: Build a spatiotemporal information perception network, take video frames as input, and introduce the spatial domain fully connected layer MarphFC in the spatial dimension s Mining spatial semantic information and introducing the temporal fully connected layer MorphFC on the temporal channel t Capture the relationship between the input video features and time, and remap the features of the time dimension to the spatial dimension, so as to extract the spatial and temporal information features in the video frame; Step 2: Build a temporal feature learning network to enhance the learning ability of temporal information features in video frames and comprehensively model the temporal information in video frames; Step 3: Train the detection model using weighted cross entropy loss.

2. A traffic incident detection method based on deep spatiotemporal memory and interactive network according to claim 1, characterized in that: Step 1 is as follows: Introduce the spatial domain fully connected layer MarphFC in the MFCMLP module s , to hierarchically expand the receptive field of FC and make it run from a small area to a large area; the referenced spatial domain fully connected layer MarphFC s Process each frame of video independently in horizontal and vertical paths; Taking the horizontal path as an example, given an input video frame X∈R HW×C , which is projected into a token sequence, first splits X horizontally, setting the block length to L, thus obtaining X i ∈R L×C (1) where i∈{1,…,HW / L}, to reduce computational cost, each X is transformed along the channel dimension. i Divide into multiple groups, each of which has D channels, and get segmented blocks, each block is X i k ∈R LD (2) Where k∈{1,…,C / D}, each data block is flattened into a 1D vector and the FC weight matrix W∈R is applied LD×LD Each data block is transformed, that is, Y is After feature transformation, all blocks Y i k Reshape to the original dimension Y∈R H×W×C , the same is true for the vertical direction, except that the tag sequence is split vertically; the FC layer is applied to process each tag separately so that the groups can communicate along the channel dimension; finally, the horizontal, vertical and channel features are summed element by element to get the output result; as the network goes deeper, the block length L will increase hierarchically, so that the FC filter gradually discovers more core semantics from small to large spatial regions; Introducing the temporal fully connected layer MorphFC in MFCMLP t This module transforms the input feature map in the time dimension, flattens it, fully connects it, and then restores it to a high-dimensional space.

3. A traffic incident detection method based on deep spatiotemporal memory and interactive network according to claim 2, characterized in that: The temporal fully connected layer MorphFC is introduced in MFCMLP. t This module transforms the input feature map in the time dimension, flattens it, fully connects it, and then restores it to a high-dimensional space. Specifically: Given an input video segment labeled X∈R H×W×T×C , first divide X into several groups along the channel dimension to reduce the computational cost, with D channels in each group, and get X k ∈R H×W×T×D (4) Among them, k∈{1,…,C / D}, for each spatial position s, the features of all frames are concatenated into a block where s∈{1,…,HW}; Reference FC matrix W∈R TD×TD Transform the time features to obtain: All the blocks Resize back to the original tokens dimension and output Y∈R H×W×T×C .

4. The traffic incident detection method based on deep spatiotemporal memory and interactive network according to claim 1 is characterized in that: Step 2 is as follows: In the input time series, assume that the input time series is Where T represents the time step, after a 1×1 convolution transformation, we get: Wherein, W1 represents the weight parameter, which is the weight matrix of the above linear transformation, and b1 represents the bias in the above linear transformation; The parallel processing part of the dilated convolution mainly includes two parallel convolution processing blocks. Each convolution block contains two dilated convolution layers. For each dilated convolution layer, the output of the lth layer is expressed as: X (l) =DilatedConv(X (l-1) ,dilation=d l ) (7) Among them, d l refers to the expansion rate of the lth layer; Assume that the output of the first dilated convolution block is X1, and the output of the second dilated convolution block is X2. The specific processing process of each block is as follows: X1=DilatedConv1(W1Z+b1) (8) X2=DilatedConv2(W2Z+b2) (9) Among them, W2 is the weight parameter in the dilated convolution operation, and b2 represents the bias in the second dilated convolution layer; The information obtained after residual connection and subsequent processing is: X=X1+X2 (10) Then after a series of nonlinear activation and regularization processing, we get: X norm =WeightNorm(X) (11) X relu =Relu(X) (12) X dropout =Dropout(X) (13) The final output is: Y=W3Dropout(Relu(WeightNorm(X1+X2)))+b3 (14) Among them, W3 represents the weight matrix of the final output layer, and b3 represents the bias of the final output layer.

5. The traffic incident detection method based on deep spatiotemporal memory and interactive network according to claim 1 is characterized in that: Step 3 is as follows: For each video frame F[t], the final output anomaly score S[t]∈[0,1] is obtained, where 0 indicates no anomaly and 1 indicates an anomaly in the frame. A higher weight is given to the anomaly class that reflects the data distribution, and a weighted cross entropy loss is selected to train the model. For each weight ω i , the formula used is as follows: ω i =e / e i (15) Where e is the total number of examples in the dataset, e i is the number of examples of class i.

6. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method of claim 1.

7. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.

Citation Information

Patent Citations

  • Deep space-time hybrid cloud data center network flow real-time detection method

    CN115348074A

  • Video anomaly detection method and device based on space-time memory network

    CN117011753A

  • Traffic event detection method of cross-time-sequence fusion memory network

    CN118397511A

Cited By

  • Road accident and illegal parking detection early warning method and system based on intelligent traffic

    CN121011078A