Study on traffic video anomaly detection method based on enhanced space-time dependent Mama model
By enhancing the spatiotemporal dependency Mamba model, the problems of data scarcity and difficulty in handling spatiotemporal dependencies in traffic video anomaly detection are solved, achieving efficient and accurate anomaly detection, which is applicable to scenarios such as intelligent transportation and urban security.
Patent Information
- Application Number
- CN202510972781.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-31
AI Technical Summary
Traffic video anomaly detection suffers from problems such as the scarcity of frame-level accurately labeled data and difficulties in handling spatiotemporal dependencies, resulting in low detection accuracy.
We employ an enhanced spatiotemporal dependency-based Mamba model, which uses adaptive hierarchical feature reconstruction, multi-head local self-attention mechanism, and state space model, combined with bidirectional spatiotemporal scanning, to capture long-term spatiotemporal dependencies in video data and enhance the model's spatiotemporal modeling capabilities.
It improves the accuracy and efficiency of traffic video anomaly detection, effectively identifies abnormal behaviors in complex traffic scenarios, and enhances the model's discrimination ability and generalization performance in the target domain.
Smart Images

Figure HDA0005500489320000011 
Figure HDA0005500489320000021 
Figure HDA0005500489320000022
Abstract
Description
Technical Field
[0001] This invention relates to a traffic video anomaly detection method in the field of computer vision technology, and in particular to a traffic video anomaly detection method based on an enhanced spatiotemporal dependency state space model. Background Technology
[0002] In increasingly complex urban traffic environments, timely detection of traffic anomalies is crucial for ensuring smooth traffic flow and improving traffic management efficiency. However, the current field of traffic anomaly detection faces a severe challenge: the scarcity of frame-level precisely labeled training data and inadequate handling of the spatiotemporal dependencies of the data. In the field of road monitoring systems, automatic event detection methods have shown promise in the rapid and accurate detection of traffic anomalies, especially in dealing with various rare traffic anomaly events. Addressing these events generally presents the following problems:
[0003] 1) Compared to the broader WS-VAD domain, the availability of traffic anomaly video data is significantly limited, focusing solely on traffic scenarios and presenting challenges in obtaining representative features from training on general video tasks.
[0004] 2) Limitations of traditional multi-instance methods in video anomaly detection, especially their insufficient ability to handle variable-length videos.
[0005] 3) The model does not accurately capture the spatiotemporal relationships in the input data during training, resulting in poor model performance. Summary of the Invention
[0006] This invention relates to the field of traffic video anomaly detection technology, specifically a traffic video anomaly detection method based on enhanced spatiotemporal dependency Mamba. Currently, traffic video anomaly detection faces technical challenges such as data scarcity and difficulties in handling spatiotemporal dependencies. In particular, frame-level precisely labeled data in traffic videos is difficult to obtain, and existing models struggle to efficiently capture long spatiotemporal dependencies in video data, resulting in low detection accuracy. Therefore, this invention aims to provide a novel method based on weakly supervised learning and enhanced spatiotemporal modeling capabilities, effectively improving the accuracy and efficiency of traffic video anomaly detection.
[0007] like Figure 1 The technical solution of this invention is as follows: A traffic video anomaly detection method based on enhanced spatiotemporal dependency Mamba is provided, with specific steps including:
[0008] S1: The starting point of the entire process is the "Traffic Video Anomaly Detection" module. The goal is to analyze traffic videos through a series of processing modules to identify potential abnormal events.
[0009] S2: The system first receives video data input, which is the basic data source for detecting anomalies;
[0010] S3: Video frames are divided into non-overlapping spatiotemporal patches to facilitate subsequent processing and ensure that each spatiotemporal patch contains key spatial and temporal information;
[0011] S4: Video data enters multiple AHFRB modules (which may be feature extraction or data transformation modules). These modules extract useful features from the video data to prepare for subsequent processing.
[0012] S5: After passing through multiple AHFRB modules, the data enters the MHLSA module (which may be a multi-layer spatiotemporal feature modeling or self-attention module) to further perform spatiotemporal dependency modeling on the video data and extract spatiotemporal features to enhance the understanding of time-series video data.
[0013] S6: Enhanced Spatiotemporal Dependencies. The Mamba module processes video data through bidirectional spatiotemporal scanning (spatial priority and temporal priority), capturing spatial and temporal dependencies.
[0014] S7: Use a state-space model to recursively update the system state to capture spatiotemporal dependencies in videos, thereby understanding long-term dependencies in video sequences;
[0015] S8: To help the model understand the relative positions of spatiotemporal blocks, spatial location coding and temporal location coding were added, enabling the model to maintain the spatiotemporal order of the video sequence;
[0016] S9: The final video classification is performed using the classification token (CLS token). The data enters the anomaly detection and classification module, where the model identifies and classifies abnormal behaviors in the video.
[0017] In the model, the input is a video sequence, denoted as . The 3 represents the three color channels of the RGB image, T represents the number of frames in the video, which is the time dimension of the video sequence, and H and W are the height and width of each frame, respectively. Attached Figure Description
[0018] Figure 1 This is an overall flowchart of the traffic video anomaly detection method based on enhanced spatiotemporal dependency Mamba of the present invention, which shows the complete steps from video input to anomaly detection output.
[0019] Figure 2 This is a flowchart of traffic video data preprocessing and spatiotemporal block partitioning in the method of the present invention, illustrating the process of generating spatiotemporal blocks from video frames through 3D convolution operations.
[0020] Figure 3This is a schematic diagram of the Adaptive Hierarchical Feature Reconstruction Module (AHFRB), which is used to alleviate the domain offset problem and enhance feature representation capabilities.
[0021] Figure 4 This is a schematic diagram of the multi-head local self-attention mechanism (MHLSA), illustrating how contextual information from adjacent time segments is fused when constructing context-dependent features.
[0022] Figure 5 This is a schematic diagram of the Mamba architecture based on state space modeling, illustrating its core mechanisms of recursive state updates and spatiotemporal dependency modeling.
[0023] Figure 6 The overall architecture diagram of the bidirectional Mamba module for enhancing spatiotemporal dependencies illustrates the processing flow from spatiotemporal location encoding and bidirectional state propagation to anomaly classification output. Detailed Implementation
[0024] Input data shape:
[0025] The shape of the raw video data is This is a 4D tensor containing RGB image information for each frame.
[0026] The model then divides the video data into small spatiotemporal patches, each containing a certain number of frames and spatial regions. Assuming we divide each video frame into P×P patches, the input data becomes... in:
[0027] 1) L is the number of small blocks in each frame, usually This means that each video segment is broken down into multiple smaller blocks through spatial and temporal divisions.
[0028] 2) C is the feature dimension of each small block, which is the number of features after processing by convolutional networks or other network layers.
[0029] Data shape:
[0030] The input video data becomes At this point, each spatiotemporal block will serve as the input to the network.
[0031] Furthermore, to better handle spatiotemporal information, the model adds classification tokens (CLStoken) to each spatiotemporal block and incorporates location encodings (including spatial and temporal location encodings) into the input to help the model understand the positional relationships between video segments. These location encodings not only identify the locations of the spatiotemporal blocks but also help the model learn temporal and spatial dependencies.
[0032] The addition of position encoding:
[0033] The input becomes X = [X cls ,X P ]+p s +p t ,in:
[0034] X cls It is a classification label
[0035] X P It is spatiotemporal block data.
[0036] p s ,p t These are learnable location codes in the spatial and temporal dimensions, respectively.
[0037] The final data shape becomes (L+1,C), which is the input sequence containing classification labels and all spatiotemporal block features.
[0038] Output:
[0039] The model's output is a video classification result, typically a probability distribution of a classification label. At the end of the model, by stacking multiple B-Mamba modules, the final output is a linearly transformed class prediction.
[0040] Output = Linear(X) cls )
[0041] The output shape is (num_classes), where num_classesnum is the number of video categories.
[0042] Model structure and principle
[0043] The core structure of the model is designed to process video data through a series of modules, ultimately achieving anomaly detection in traffic videos. First, the model receives input traffic video data, which enters the system through designed modules. After input, the video data first passes through multiple AHFRB modules, which are responsible for feature extraction and preprocessing, transforming the raw video data into a form suitable for subsequent analysis. Next, the data processed by the AHFRB modules enters the MHLSA module. This module further captures the spatiotemporal dependencies in the video. Typically, such modules use methods like temporal modeling and feature extraction to help the model understand the temporal and spatial relationships between video frames. Further, the data passes through the Mamba module, which enhances spatiotemporal dependencies. This module is designed to handle long-term spatiotemporal dependencies in the video, effectively capturing long-term changes or behavioral patterns across frames, enhancing the model's ability to identify traffic anomalies.
[0044] After spatiotemporal dependency modeling, the data enters the anomaly detection and classification module. This module detects and classifies abnormal behaviors in the video based on the features learned by the model. Through the model's output, the system can determine whether an abnormal event exists. If an anomaly is detected, the system outputs "Yes," indicating that an abnormal event has been detected; if no anomaly is detected, it outputs "No." Finally, the system summarizes the detection results, outputs the final conclusion, and completes the entire traffic video anomaly detection process.
[0045] This model structure, through the collaborative work of multiple modules, especially the design of modules such as AHFRB, MHLSA, and Mamba, can effectively extract features from videos and capture long-term spatiotemporal dependencies, ultimately achieving efficient and accurate anomaly detection.
[0046] Adaptive Hierarchical Feature Reconstruction
[0047] There is a significant distributional difference between traffic videos and general datasets used by pre-trained models, such as human motion recognition datasets, known as domain shift. Directly using pre-trained features from these general datasets can lead to a decline in the model's performance in traffic anomaly detection.
[0048] In traffic anomaly detection tasks, a core challenge is often encountered: the "domain shift" problem. Specifically, current mainstream visual models typically rely on pre-training on large-scale general datasets (such as Kinetics, UCF101, and HMDB51), which mainly focus on everyday human action recognition or common scene events. However, the scenes, motion patterns, and event structures in traffic videos differ significantly from these general datasets. For example, traffic scenes involve more vehicles, pedestrians, and non-standard behaviors (such as driving against traffic, illegal parking, and accidents), and the camera is usually fixed with a stable background. This difference in distribution often leads to feature mismatch when directly transferring general pre-trained models, significantly reducing detection accuracy and generalization ability in target traffic scenes.
[0049] To alleviate the aforementioned problems and improve the discriminative ability of pre-trained models in the target domain, this study introduces an Adaptive Hierarchical Feature Reconstruction Module. The core idea of this module is to make the original pre-trained features more closely match the feature distribution of the target task through multi-layer feature extraction and reconstruction operations.
[0050] like Figure 3First, in the multi-layer feature extraction stage, we select multiple layers of different depths from the pre-trained model to extract features, constructing a hierarchical representation system. Shallower features retain richer local details and spatial structure information, such as edges, textures, and contours; while deeper features focus more on abstract semantic concepts, such as behavioral patterns or event categories. By fusing features from different layers, we can not only balance details and semantics but also provide the model with a more comprehensive scene understanding capability.
[0051] Next, the feature reconstruction stage begins. This module first normalizes the raw features extracted from each layer, using methods such as spatial pooling to compress spatial dimensions and linear transformation to unify channel dimensions. This process ensures that features from different layers have the same scale and format, facilitating subsequent fusion. After unifying dimensions, these features are reorganized and enhanced using learnable transformation functions (such as attention mechanisms or weighted fusion networks) to better align with the semantic structure and distribution characteristics of traffic videos. In this way, the original cross-domain features are mapped to a more discriminative feature space.
[0052] To maintain the robustness of the original features in the pre-trained model, we introduce a residual connection mechanism at the end of the reconstruction module. Specifically, the reconstructed new features are additively fused with the original deep features to form the residual feature output. This design is inspired by deep residual network structures such as ResNet, effectively preserving the general representation capabilities already present in the original model while incorporating feature information adapted to the target domain, thus achieving effective integration of new and old knowledge.
[0053] Overall, this method constructs a robust domain-adaptive feature optimization process through multi-layer extraction, reconstruction, fusion, and residual connections. It not only solves the problem of pre-trained feature failure caused by domain offset but also significantly improves the model's ability to perceive and identify complex abnormal behaviors in traffic scenarios, laying a technical foundation for building a more efficient and general traffic anomaly detection system.
[0054] Feature combinations extracted from different levels:
[0055]
[0056] Where S represents the number of layers extracted, and f represents the features of the S-th layer.
[0057] The reconstructed features are represented as follows:
[0058]
[0059] Finally, the reconstructed features and the original features are connected via residuals to generate the final feature representation:
[0060]
[0061] Multi-Head Local Self-Attention Mechanism
[0062] Traditional Multiple Instance Learning (MIL) methods for processing video data typically rely on a crucial assumption: that videos can be divided into segments, each independent and identically distributed. While this assumption may be valid for image processing or certain static tasks, it clearly fails to hold true for temporal tasks based on video, such as traffic anomaly detection. In reality, traffic videos exhibit significant temporal continuity and contextual dependencies. For instance, anomalies such as a vehicle suddenly changing lanes or a pedestrian crossing the road are rarely isolated events; they are often accompanied by a series of actions and situations preceding or following the event. These adjacent segments carry highly relevant contextual information. If the model ignores this temporal dependency, it will lead to the accumulation of errors in local judgments, affecting the overall anomaly detection performance.
[0063] To address this issue, we introduced a context-aware feature construction mechanism into the model, such as... Figure 4 As shown in the diagram. Specifically, for each target video segment, we no longer use the segment's own feature representation in isolation, but instead introduce feature information from several time steps before and after it, constructing a feature matrix containing temporal context through a concatenation operation. This approach captures local temporal continuity, ensuring that the representation of each segment not only depends on the image or motion pattern at the current instant but also incorporates evolutionary information from previous and subsequent time steps, thereby more accurately understanding its semantic role in the global behavioral sequence.
[0064] To further model the temporal dependencies between video segments, we introduce a multi-head self-attention mechanism. This mechanism dynamically calculates the relevance weights between different segments using multiple learnable attention heads, enabling interactive modeling of global information. The implicit representation of each segment not only comes from its own original information but is also influenced by the features of other related segments (especially those at adjacent time points). This mechanism breaks the limitations of fixed windows or finite state transitions in traditional temporal models, allowing the model to flexibly perceive long-distance dependencies and enhancing its ability to understand the logic before and after anomalous events.
[0065] In the output stage, we employ a residual connection structure to fuse the attention-enhanced context features with the original fragment features. The introduction of residual connections not only ensures the integrity of the original information and prevents information loss during multi-layer transformations, but also provides the model with stronger expressive power, enabling it to retain existing structural information while learning new features. The resulting fragment representation is more discriminative and beneficial for the accurate identification of abnormal behaviors.
[0066] The construction of contextual features is as follows:
[0067]
[0068] Where N is the size of the time window, representing the number of consecutive segments used.
[0069] The final output of the bullish self-attention:
[0070] MHLSA = Concat(head) l ,...,head j ,...,head h W o (3-5)
[0071] The self-attention result is residually concatenated with the original fragment features to obtain the final feature representation:
[0072]
[0073] State-space model (SSM)
[0074] In complex temporal modeling tasks such as traffic anomaly detection, relying solely on convolutional neural networks or traditional self-attention mechanisms is insufficient to fully capture long-range spatiotemporal dependencies. This is especially true in videos, where the causes and consequences of anomalous events often span multiple time steps, exhibiting significant delays and dynamic context switching. To address this, we introduce a mechanism to enhance spatiotemporal modeling capabilities—the Mamba architecture based on state-space modeling. Its core idea is as follows: Figure 5 As shown.
[0075] Mamba's key innovation lies in introducing a State Space Model (SSM) as a fundamental building block to model complex temporal dependencies in video data. Compared to traditional temporal modeling methods, the State Space Model possesses inherent memory capabilities and a state recursion mechanism, enabling it to continuously update its internal state during sequence processing, thereby achieving a dynamic perception of temporal continuity and historical information. This capability is particularly suitable for tasks in traffic scenarios where abnormal behaviors need to be "identified" within a context, such as vehicles gradually deviating from their lanes, sudden stops, and pedestrians crossing boundaries—all implicit evolutionary patterns.
[0076] Specifically, in the Mamba architecture, each time step of the system maintains a hidden state. This state is not updated in isolation, but evolves continuously through the interaction of a recursive state transition function and input features. In other words, the output at the current moment depends not only on the current input, but is also strongly influenced by the hidden state at the previous moment. This recursive structural design can continuously accumulate historical information, achieving compression and memorization of the global context.
[0077] Furthermore, Mamba, in its spatiotemporal modeling, not only focuses on the flow of information in the time dimension but also embeds spatial information into the state update process, thereby achieving true spatiotemporal fusion modeling. This mechanism not only captures sudden events locally but also models the evolutionary trajectory of abnormal events before and after their occurrence in the time dimension, giving the model stronger discriminative ability and deeper semantic perception.
[0078] Finally, to ensure the stability and efficiency of model training, Mamba employs an optimized parameterization method and an efficient parallel mechanism, resulting in good scalability and computational efficiency in the solution and update process of the state space. This is particularly crucial when dealing with long sequences and large-scale traffic video data.
[0079] Its mathematical formula is:
[0080] h′(t)=Ah(t)+Bx(t)
[0081] h(t) is the hidden state at time step t (representing the feature at the current moment).
[0082] x(t) is the input data at the current time (e.g., spatiotemporal block features).
[0083] A and B are the learned parameter matrices, representing the state transition and the effect of the input on the state, respectively.
[0084] Output:
[0085] y(t)=Ch(t)
[0086] Where C is a mapping matrix that maps the hidden state h(t) to the output y(t), which is the model's final prediction.
[0087] Discretization
[0088] In real-world tasks, time-series data such as traffic videos are typically sampled frame by frame, thus representing a discrete-time series. This presents a mismatch with the continuous-time system model established by traditional state-space models. To address this issue, we need to discretize the continuous system form in the state-space model to accommodate the data modeling requirements at discrete time steps.
[0089] Specifically, continuous systems are typically defined by a set of differential equations, and their state evolution follows continuous-time variables. However, in discrete-time series, we cannot obtain data at arbitrary points in time; we can only observe sampled values at fixed intervals (such as video frames). Therefore, in order to apply state-space models at discrete time steps, we need to convert the differential expressions in the continuous form into difference expressions, thereby achieving an efficient mapping from the continuous domain to the discrete domain.
[0090] Commonly used methods in this process include Euler discretization, bilinear transformation, or the more precise Runge-Kutta method. These discretization techniques can preserve the continuous dynamic characteristics of the system in a discrete time step in an approximate manner, enabling the model to better adapt to the actual data structure in sequence modeling while maintaining the original behavioral characteristics of the system.
[0091] Through this discretization process, the state-space model can be seamlessly integrated into video frame-level data analysis tasks, thereby enabling the modeling of long-range dependencies and dynamic evolution patterns in discrete time series, providing stronger modeling support for the detection and recognition of abnormal behavior in videos. This method not only preserves the theoretical interpretability and stability of continuous systems but also takes into account the adaptability and efficiency of practical discrete data processing.
[0092] Discretization is performed using the zero-order hold method:
[0093] A = exp(ΔA), B = (ΔA) -1 (exp(ΔA)-I)·ΔB
[0094] Here, ΔA and ΔB are the increments of the state transition matrix. Through discretization, the model can efficiently update the state at each time step.
[0095] Enhanced Spatiotemporal Dependency Mamba
[0096] like Figure 6 ,like Figure 6As shown, in the Mamba model architecture that enhances spatiotemporal dependence, the entire modeling process starts from the input video frame and sequentially goes through modules such as preprocessing, spatiotemporal segmentation, location encoding, and Mamba modeling, ultimately achieving accurate identification and capture of abnormal events in traffic videos. The overall process design of the model balances modeling capability, computational efficiency, and practical application adaptability, making it a highly efficient architecture for long-sequence, high-resolution video processing.
[0097] First, the input video data typically has high temporal and spatial resolution. Therefore, a preprocessing stage is set up at the input end. The main task of this stage is to reduce the dimensionality and structure of the raw high-dimensional video data to reduce the computational cost of subsequent modeling. In this stage, the model uses 3D convolution to extract local features simultaneously from both spatial and temporal dimensions. Specifically, 3D convolution can not only perceive the dynamic changes between video frames but also extract the static structure within spatial regions, thus providing a solid foundation for spatiotemporal modeling.
[0098] Through 3D convolution operations, the input video is divided into multiple non-overlapping spatio-temporal patches, which are compact representation sequences consisting of t×h×w patches, where t, h, and w represent scaling factors in the temporal, spatial height, and width directions, respectively. This partitioning method effectively reduces the spatio-temporal dimension of the input, significantly reducing computational resource consumption. Furthermore, it preserves important motion and structural information of the video within a local range, laying a solid foundation for subsequent feature modeling and temporal sequence construction.
[0099] After obtaining the spatiotemporal patch sequence, the model further introduces a temporal and spatial location encoding mechanism. This mechanism encodes the temporal and spatial location information of each patch within the original video and fuses it with its feature vector. This step is particularly crucial for video data, helping the model understand "what changes occurred in which region at which time," thereby enhancing the model's ability to discriminate spatial dynamic structures and temporal evolution paths. For example, in traffic videos, vehicle behavior is often highly correlated with its trajectory position and the time it enters the frame; the location encoding mechanism can enhance the model's sensitivity to and representational ability regarding these details.
[0100] Next, the model moves to its core component: the Bi-directional Enhanced Mamba Block. This block extends the basic Mamba state-space modeling architecture, incorporating bi-directional modeling capabilities. It can simultaneously propagate information from the past time step to the future and backtrack contextual dependencies from the future time step to the past. This bi-directional approach significantly enhances the model's ability to perceive temporal dependencies, making it particularly suitable for data such as traffic videos that exhibit long-term background evolution and behavioral continuity. Spatially, Mamba also possesses state propagation capabilities, enabling the modeling of information transmission and structural dependencies between regions, achieving synchronous modeling in both the spatiotemporal domains.
[0101] This design gives the Mamba model a natural advantage when facing complex traffic video scenarios, enabling it to accurately identify fine-grained dynamic abnormal behaviors such as vehicles making U-turns, driving against traffic, prolonged stops, and abnormal driving trajectories, thereby improving the model's ability to identify abnormal events and its response speed.
[0102] It is worth emphasizing that, thanks to Mamba's linear complexity state-space update mechanism, the model can maintain low computational and memory overhead even with long sequences and high-resolution inputs, avoiding the exponential resource consumption problem faced by traditional self-attention mechanisms when dealing with long temporal inputs. This efficiency not only gives the model better scalability and practical deployment capabilities, but also provides strong support for its application in real-world scenarios such as intelligent transportation, urban security, and video surveillance.
Claims
1. A traffic video anomaly detection method based on enhanced spatiotemporal dependency Mamba, characterized in that, The steps include: obtaining a traffic video dataset; 1) Construct a traffic video anomaly detection model based on enhanced spatiotemporal dependency Mamba and train it using a traffic video dataset; 2) Input the traffic video data to be detected into the traffic video anomaly detection model, and the model outputs the traffic anomaly detection results.
2. The traffic video anomaly detection method based on enhanced spatiotemporal dependency Mamba as described in claim 1, characterized in that, The detection model includes adaptive hierarchical feature reconstruction blocks and a multi-head local self-attention mechanism to capture spatiotemporal dependencies in videos.
3. The traffic video anomaly detection method based on enhanced spatiotemporal dependency Mamba as described in claim 1, characterized in that, The model includes a multi-head local self-attention module, designed to enhance the accuracy of abnormal event detection by capturing the temporal dependencies between video segments.
4. The traffic video anomaly detection method based on enhanced spatiotemporal dependency Mamba according to claim 1, characterized in that, The enhanced spatiotemporal dependency Mamba model includes spatiotemporal segmentation and a state-space model, which are used to efficiently process long temporal video data and model spatiotemporal dependencies.
5. The traffic video anomaly detection method based on enhanced spatiotemporal dependency Mamba according to claim 4, characterized in that, The spatiotemporal dependent Mamba model improves the ability to model the spatial and temporal synchronization of traffic videos through bidirectional processing.
6. The traffic video anomaly detection method based on enhanced spatiotemporal dependency Mamba according to claim 1, characterized in that, The method employs weakly supervised learning, training the model with video-level labels to reduce reliance on frame-level precise annotation data.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.