Pedestrian anomaly detection and prediction method
Through the improved model Bip-ITFA, which integrates time-frequency domain characteristics and attention mechanism, the robustness of pedestrian anomaly detection in complex backgrounds is solved, the detection accuracy and computing efficiency are improved, and it is suitable for fields such as intelligent video surveillance and public safety.
Patent Information
- Application Number
- CN202510523451.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-22
AI Technical Summary
The existing pedestrian anomaly detection technology is not robust enough in complex backgrounds and multiple scenarios, making it difficult to accurately identify abnormal behaviors, especially in actual scenarios with multiple scenes with multiple pedestrians.
A bidirectional prediction model for pedestrian anomaly detection using a fusion time frequency domain feature and attention mechanism is used to enhance feature extraction and information capture capabilities and improve detection accuracy through the improved model of residual ASB-GRU interactive convolution module and adaptive weight attention module.
It improves the accuracy of pedestrian anomaly detection and the robustness of the model in complex environments, adapts to multi-scene multi-peeder detection tasks, and reduces computing resource consumption.
Smart Images

Figure CN120356263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and specifically relates to a method for pedestrian anomaly detection and prediction. Background Art
[0002] As an important branch of video anomaly detection, Pedestrian Anomaly Detection (PAD) aims to identify abnormal behaviors or events of pedestrians from video sequences. This technology has broad application prospects in many fields such as intelligent video surveillance and public safety. However, in application scenarios involving pedestrian abnormal behaviors (such as autonomous driving, video security and surveillance, public safety supervision, etc.), due to the complex background features contained in the video scene, the dynamic change of light intensity, and the frequent occlusion phenomenon of target objects, it significantly increases the complexity and computational difficulty of the abnormal behavior event detection algorithm. The computational complexity is too high to store long-term relationships. Specifically, it is difficult to detect abnormal behaviors and identify long-term anomalies, ultimately leading to a significant decrease in the accuracy of pedestrian anomaly detection.
[0003] Existing improvement schemes can be divided into three categories:
[0004] 1. Methods based on spatio-temporal feature modeling: Such methods may improve the model's ability to extract temporal and spatial features, such as using 3D CNN, Transformer, etc. to process the temporal information of videos and solve the long-term dependence problem.
[0005] 2. Model compression techniques based on lightweight design: Aiming at the problem of high computational complexity, model lightweight methods such as knowledge distillation, model pruning, quantization, etc. are adopted to reduce the consumption of computational resources and facilitate actual deployment.
[0006] 3. Multimodal fusion methods: Try to combine information such as RGB, optical flow, depth, etc. to improve the robustness features of complex environments, but the generalization ability for drastic environmental changes is insufficient and it is easily affected by background noise.
[0007] It should be noted that the above-mentioned solutions still have insufficient feature robustness for person occlusion and environmental changes, so they cannot guarantee the identification of abnormal behaviors, especially in the actual scenarios of multiple scenes and multiple pedestrians, which have great limitations. To solve the above problems, the present invention proposes a two-way prediction method for pedestrian anomaly detection that fuses time-frequency domain features and attention mechanisms. Summary of the Invention
[0008] The purpose of the present invention is to provide a two-way prediction model for pedestrian anomaly detection that fuses time-frequency domain features and attention mechanisms to solve the problems raised in the above background art. It includes: a video anomaly detection dataset and a pedestrian anomaly detection Bip-ITFA prediction model
[0009] The video anomaly detection dataset selects four publicly available datasets, namely CUHK Avenue, ShanghaiTech Campus, HR-ShanghaiTech Campus, and HR-Avenue, as the dataset for this invention:
[0010] (1) The CUHK Avenue dataset is a well-known dataset for video anomaly detection, consisting of 16 groups of training videos and 21 groups of test videos, with a total of 30,652 frames of images (15,328 and 15,324 frames for the training and test sets respectively). This dataset completely records the normal activities of pedestrians and three types of typical abnormal behaviors from a fixed perspective: sudden high-speed movement (such as running at high speed), unconventional handling of objects (such as throwing actions), and unstructured movement trajectories (such as turning back and wandering). On the basis of retaining the complexity of the real scene (such as limb occlusion when the crowd is dense, and the difference in day and night lighting), the video data strengthens the discriminative expression of behavior features under dynamic environmental interference, providing multi-dimensional benchmark support for the verification of algorithm robustness.
[0011] (2) The ShanghaiTech dataset is constructed in 13 complex scenes with significant dynamic lighting changes and multi-view acquisition characteristics, belonging to the leading large-scale video anomaly detection benchmark library in the industry. Its training set integrates 330 segments of pure normal behavior videos, with a total of more than 270,000 high-resolution images; the test set covers 107 videos in 12 scenarios, including 130 manually annotated abnormal behavior instances. The abnormal behaviors mainly focus on the deviation of human activity patterns, and typical examples include: five high-risk scenarios such as uncontrolled running, violent snatching behavior, crossing and jumping, physical conflict, and riding in restricted areas, which precisely correspond to the key detection requirements in public security monitoring.
[0012] (3) The two datasets, HR-ShanghaiTech Campus and HR-Avenue, are subsets of ShanghaiTech Campus and CUHK Avenue, removing the videos irrelevant to human behavior in the original datasets and further optimizing the experimental materials.
[0013] The Bip-ITFA model includes a residual ASB-GRU interaction convolution module (AGI) and an adaptive weight attention module (ANA).
[0014] The residual ASB-GRU interaction convolution module (AGI) is composed of an adaptive spectrum module (ASB), a GRU module, and an interaction convolution module in a residual connection manner.
[0015] The adaptive weight attention module (ANA) is composed of a The self-attention approximation module is composed of learnable residual connections.
[0016] The pedestrian anomaly detection and prediction method provided by the present invention includes:
[0017] Preprocess the video anomaly detection dataset, construct a variational autoencoder bidirectional prediction RNN model (Bip-ITFA) for pedestrian anomaly detection and prediction, input the preprocessed video anomaly detection dataset into the network for training, input the anomaly detection data to be detected into the trained network for pedestrian anomaly detection, and finally output the predicted human skeleton diagram.
[0018] Among them, the preprocessed video anomaly detection dataset includes four datasets: CUHK Avenue, ShanghaiTech Campus, HR-ShanghaiTech Campus, and HR-Avenue.
[0019] Extract video data by preprocessing the pedestrian video dataset, cropping the size, unifying the duration and frame rate, and inputting it into AlphaPose to obtain skeleton information, etc., to obtain the final human skeleton dataset;
[0020] Furthermore, the preprocessing includes:
[0021] Decode the input video, convert the video file into frame images. Different formats of videos may need to be processed, and use OpenCV to convert the video data into frame pictures.
[0022] Furthermore, the cropping of the size, unifying the duration and frame rate includes:
[0023] First, locate the pedestrian bounding box through an object detection model (such as YOLO) and crop the core area, and adjust the image to a unified size (256×256) by direct scaling or padding black edges while maintaining the aspect ratio to eliminate background interference and adapt to the model input; secondly, standardize all videos to a fixed frame rate (30FPS) by frame extraction (reducing high frame rate) or frame interpolation (completing low frame rate) to ensure the consistency of the time dimension; finally, truncate ultra-long videos or circularly fill short videos according to a preset duration (such as 5 seconds) to align the durations of all samples for batch training. Key details include bounding box expansion to prevent over-tight cropping and inter-frame smoothing filtering to avoid jitter.
[0024] Furthermore, the inputting into AlphaPose to obtain skeleton information includes:
[0025] First, take the standardized video file (mp4) or the directory of frame images named in sequence as the input. Specify the pre-trained model (fast_res50_256x192) and the configuration file (such as the 256x192_res50 configuration in COCO format) through the command-line interface of AlphaPose, and perform frame-by-frame pedestrian detection and pose estimation; the algorithm sequentially locates the human body bounding box and predicts the coordinates and confidence of 17 joint points (such as shoulders, elbows, knees, etc.), and supports associating cross-frame IDs through a tracking algorithm (PoseFlow) in multi-pedestrian scenarios to maintain temporal consistency; finally, output the skeleton data in JSON or TXT format for each frame, including the normalized two-dimensional / three-dimensional joint positions, confidence scores, and pedestrian IDs, and optionally generate a visualized annotation video for verification.
[0026] Furthermore, the constructed pedestrian anomaly detection Bip-ITFA prediction model includes:
[0027] Constructed a residual ASB-GRU interactive convolution module (AGI) and an adaptive weight attention module (ANA). The improved model includes:
[0028] The input skeleton data is mined and classified for features through a GRU network. Subsequently, a variational autoencoder network is introduced to generate the output F2. At the same time, the input feature F1 is processed by the residual ASB-GRU interactive convolution module (AGI) to obtain the output F3. To enhance the diversity of features and improve the training and inference efficiency, F2 and F3 are fused by parallel multiplication. Finally, in the training stage, F3 is used to predict the trajectory through a bidirectional prediction RNN network to obtain the predicted trajectory F4, and F4 passes through the adaptive weight attention module (ANA) to obtain the output predicted skeleton map and the anomaly score of the anomaly time frame.
[0029] Furthermore, the parallel introduction of the residual ASB-GRU interactive convolution module AGI in the original variational autoencoder network includes:
[0030] Stack the human pose sequences. After stacking, obtain the time encoding through GRU. The joint points in these data are combined with the time information. For a single joint point, it can be regarded as a discrete time series S∈R C×L , S will be divided into several groups and then mapped to another dimension to add position embedding for enhancement. After enhancement, the dimension size is p′, C and L are the number of channels and the sequence length respectively, and L′ is the length of the transformed sequence, which is usually not equal to L.
[0031] Perform a fast Fourier transform on the sequence S to obtain the frequency domain representation F, and calculate the power spectrum P = |F| 2 , finally set a trainable threshold parameter θ, and the frequencies with power higher than the threshold are retained, and other frequencies are removed to obtain Ffiltered , the obtained F filtered and F need to be filtered through learnable global weights and local weights W G and W L to obtain F G and F L , and add them to obtain F integrated , perform an inverse fast Fourier transform on F integrated to obtain the sequence S′. The sequence S′ will be added to the original sequence S, and after fusing time information through a GRU, S″ is obtained. Finally, a transposed convolution is performed.
[0032] Furthermore, an adaptive weight attention module (ANA) module is introduced after the bidirectional prediction RNN network, including:
[0033] Design a self-attention approximation module, which reconstructs the self-attention mechanism based on the low-rank approximation theory. Compared with the traditional multi-head self-attention algorithm with a computational complexity of O(n 2 )(n is the sequence length), while the self-attention reduces the complexity to O(nm) (m << n) through matrix low-rank decomposition. While maintaining the key attention patterns, it significantly reduces the memory and computational resource consumption, ensures real-time performance, and improves the algorithm accuracy, which is very important for the pedestrian anomaly detection task.
[0034] Advantages of the present invention: The present invention makes improvements based on the original variational autoencoder bidirectional prediction RNN network, mainly including introducing a residual ASB-GRU transposed convolution module and an adaptive weight attention module (ANA) module. Through the above improvement methods, the present invention can enhance the feature extraction and information capture capabilities of the model and improve the accuracy of the model in pedestrian anomaly detection prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a structural diagram of the human skeleton provided in this specification;
[0036] Figure 2 is a structural diagram of the overall framework provided in this specification;
[0037] Figure 3 is a structural diagram of the residual ASB-GRU transposed convolution module (AGI) provided in this specification;
[0038] Figure 4 is a structural diagram of the adaptive weight attention module (ANA) module provided in this specification;
[0039] Figure 5 is the output human skeleton diagram provided in this specification. Specific implementation method
[0041] The following further describes the present invention patent in detail in conjunction with the accompanying drawings and specific embodiments, but it should be understood that the protection scope of the present invention is not limited by the specific implementation manners. According to the embodiments of the present invention, all other embodiments obtained by professionals in the field without performing innovative work shall fall within the protection scope of the present invention. Embodiment:
[0042] A pedestrian anomaly detection and prediction method specifically includes the following steps:
[0043] Step 1. Establishment of video anomaly detection dataset and enhancement of human skeleton data
[0044] Under the research framework of computer vision driven by deep learning, the training effect of the algorithm is significantly correlated with the data quality. To achieve reliable performance of the model in complex scenarios, the training process relies on a large amount of diverse image and video data resources, especially a sample set with rich scene complexity. Such high-quality data can effectively support the deep network to perform multi-level feature learning, thereby ensuring that the model has excellent generalization ability and robustness. It should be noted that in actual application scenarios, the original data often has the defect that the data volume is too large and the model is difficult to process, which promotes data preprocessing to become a key optimization strategy in the construction process of the pedestrian anomaly detection model.
[0045] Step 1.1. Establishment of video anomaly detection dataset
[0046] The pedestrian image dataset under the dressing-changing background requires that the video anomaly behaviors therein include elements of randomness, suddenness, and unpredictability. In the present invention, the publicly available video anomaly detection datasets CUHK Avenue, ShanghaiTech Campus, and HR-ShanghaiTech Campus, HR-Avenue are used as the datasets for model training.
[0047] The CUHK Avenue video anomaly detection dataset is a classic video analysis benchmark dataset released in 2013, mainly used for the detection and localization of abnormal behaviors in surveillance scenarios. It contains 37 fixed-view campus corridor and street surveillance videos (21 training videos and 16 test videos), with a total of more than 30,000 frames. The normal behaviors in the dataset are mainly pedestrians walking, and the abnormal behaviors include running, throwing objects, suddenly falling to the ground, etc. This dataset has become an important benchmark for algorithm robustness testing due to dynamic background interference (such as leaf shaking, light change) and the diversity of abnormal events.
[0048] The ShanghaiTech Campus dataset is a large-scale benchmark dataset in the field of video anomaly detection, constructed by the ShanghaiTech University team. This dataset contains 13 different campus scenes (covering both indoor and outdoor areas), with a total of 437 surveillance videos (330 normal videos in the training set and 107 videos with anomalies in the test set), and the total number of frames exceeds 270,000. Abnormal behaviors include complex events such as fighting, cycling, sudden running, dropping packages, and climbing over railings. Its characteristics lie in multi-scenes, multi-views, dynamic light changes, and complex background interference (such as crowded people and occlusion), making it suitable for the research of unsupervised or weakly supervised learning methods. It is one of the largest and most scene-rich benchmarks in the current field of video anomaly detection.
[0049] HR-ShanghaiTech Campus and HR-Avenue remove abnormal behaviors unrelated to humans and are high-resolution versions optimized for the resolution limitations of the original dataset in the field of video anomaly detection, providing a more practical high-quality data benchmark for refined anomaly detection and behavior analysis. The normal dataset still cannot be input into the model after preprocessing and requires data augmentation:
[0050] In the design of the data augmentation process, first, the video sequence is input into the AlphaPose pose estimation model frame by frame. The YOLOv3 object detector is used to locate all human bounding boxes in the scene, and the coordinates and confidence of 17 standard COCO key points (including the tip of the nose, shoulders, elbows, wrists, etc.) are predicted for each detected human instance. Subsequently, an initial skeleton data containing human ID, bounding box parameters (x_min, y_min, width, height), and the three-dimensional coordinates (x, y, confidence) of 17 joints is constructed. On this basis, by calculating the geometric center coordinates of four key points (left eye, right eye, left ear, right ear) in the head region, the 18th custom joint point is generated as the "head center point". This augmentation design aims to enhance the expression ability of head movement features. Finally, the multi-object skeleton data of each frame is output in a hierarchical nested JSON format, where each human instance contains the normalized coordinates of 18 joints, the relative time sequence frame number, and the associated kinematic feature vector. The human skeleton diagram used in this invention is as Figure 1 shown.
[0051] Step 2: Construct an improved Bip-ITFA model for pedestrian anomaly detection prediction
[0052] Step 2.1: Original prediction model
[0053] The original model is a neural network obtained by connecting a variational autoencoder network in parallel with a forward RNN network and a backward RNN network. It adopts the idea of Residual Learning, which helps it overcome the common problems of vanishing or exploding gradients in deep networks during training. The prior and recognition networks of the variational autoencoder are respectively composed of an encoder formed by a multi-layer perceptron, whose main function is data dimensionality reduction and also includes operations such as batch normalization and ReLU activation. The prediction part obtains the data of the autoencoder part through the encoder and inputs it into the forward-backward RNN network for prediction through skip connections (residual connections), which also alleviates the problem of vanishing gradients in deep networks and makes the network easier to train. Finally, the RNN network outputs the final skeleton graph of the skeleton and the final frame-level anomaly score.
[0054] Step 2.2, Improved Prediction Model Bip-ITFA
[0055] Compared with the prediction methods of traditional pedestrian anomaly detection, the method of predicting future frames not only has a large computational complexity but also has poor prediction performance and low accuracy. Therefore, the ability of the model to identify and judge different anomalies is the key to improving the recognition accuracy. The present invention improves the original prediction model and enhances its feature fusion by introducing a pedestrian anomaly detection prediction model (Bip-ITFA) with residual ASB-GRU interactive convolution and self-adaptive weight attention. The overall framework diagram of the model is as Figure 2 shown.
[0056] Step 2.2.1, Parallel Introduction of Residual ASB-GRU Interactive Convolution Module AGI into the Original Variational Autoencoder Network
[0057] In the pedestrian anomaly detection task, the dynamic changes in the environment (such as sudden changes in lighting, weather interference), occlusion of key human body parts (such as joints being invisible due to backpacks / umbrellas), scarcity of abnormal samples, and the computational complexity of high-dimensional spatio-temporal data together constitute the four core challenges of multi-modal feature representation and fusion. Traditional methods (time-domain modeling of LSTM or independent frequency-domain analysis) are often limited by the perceptual bias of a single modality: the time-domain model is not sensitive enough to sudden high-frequency anomalies (such as limb twitching), while the frequency-domain method is difficult to capture long-period behavior patterns (such as wandering trajectories); in addition, the missing joint data caused by occlusion will lead to discontinuous feature spaces, increasing the risk of the model overfitting to noise. For this reason, this paper proposes a residual ASB-GRU interactive convolution module AGI, whose design goal is to construct a differentiable and interpretable time-frequency domain joint representation framework, and achieve the coordinated optimization of occlusion robustness and computational efficiency through a dynamic graph topology reasoning and cross-domain feature complementary mechanism. Its architecture is as Figure 3 shown, and the formula is as follows.
[0058] F = Γ[S] ∈ CC×L′
[0059] F filtered = F⊙(P > θ)
[0060] F G = W G ⊙F
[0061] F L = W L ⊙F filtered
[0062] F integrated = F G + F L
[0063] S′ = Γ -1 [F integrated ∈ R C×p′
[0064] where Γ represents the fast Fourier transform, ⊙ represents element-wise multiplication, and Γ -1 represents the inverse fast Fourier transform. The sequence S′ will be added to the original sequence S, and after passing through a GRU to fuse the time information, S″ is obtained. Finally, through a transposed convolution, the calculation method is shown as follows:
[0065] A1 = φ(Conv1(S″)⊙(Conv2(S″)
[0066] A2 = φ(Conv2(S″)⊙(Conv1(S″)
[0067] O = Conv3(A1 + A2)
[0068] where φ represents the GELU activation and O represents the output.
[0069] The AGI module first applies the time-frequency domain fusion method to the pedestrian anomaly detection task. It regards the human body pose sequence as a discrete time series, innovatively performs a frequency domain transformation on the human body pose sequence, and at the same time uses GRU and transposed convolution for fusion, which enhances the feature processing ability of the model and also enhances the generalization of the model.
[0070] Step 2.2.2, introduce an adaptive weight attention module (ANA) after the bidirectional prediction RNN network
[0071] Traditional self-attention mechanisms face severe challenges in processing long sequences due to the quadratic growth of computational complexity with the sequence length. The parallel processing of multi-joint and multi-channel data causes the video memory requirements to increase exponentially, even triggering memory overflow on conventional GPUs. At the same time, the computational cost of matrix multiplication increases cubically with the sequence length, and the ineffective attention weights approaching zero between distant nodes (such as the joint movements separated by 10 seconds) further exacerbate the waste of computational resources, forcing the system to compromise between modeling long-range dependencies and maintaining real-time performance, severely restricting the practicality in high-response scenarios such as security and autonomous driving.
[0072] To this end, the present invention designs an ANA module, as Figure 4 shown. It achieves efficient computation through a multi-stage feature compression and dynamic sampling strategy: First, the predicted skeleton graph data is stacked in the temporal dimension and linearly mapped for dimensionality reduction to extract lightweight query matrix Q and key matrix K. Subsequently, a hybrid sampling strategy is adopted to select feature subsets with strong representational capabilities from Q and K respectively as positioning reference points, and a low-rank projection matrix is constructed through singular value decomposition to capture the core structure of the global attention distribution. On this basis, the original input sequence is divided into non-overlapping consecutive sub-segments of a fixed length, and the internal features of each sub-segment are aggregated through mean pooling operation to generate local feature vectors to suppress noise and retain semantic information. Finally, the similarity relationship between the reference points and the pooled features is modeled through a kernel function, and the local details and the global low-rank projection results are fused by combining the residual connection mechanism to form an approximate expression of the attention weight matrix that takes into account both computational efficiency and modeling accuracy. Its attention calculation formula is as follows:
[0073]
[0074] where represents the pseudo-inverse matrix of, and the module also needs to process the final obtained skeleton graph.
[0075] First, the human skeleton graph sequence data generated by the recurrent neural network (RNN) prediction network is stacked along the time dimension to construct a four-dimensional tensor with spatio-temporal joint features (dimension composition: three-dimensional data of spatial coordinate x + time step), and in this way, the skeleton prediction results at discrete time steps are integrated into a continuous spatio-temporal representation. Subsequently, a fully connected layer is used to perform feature compression and dimensionality reduction on the four-dimensional tensor, and project it into a three-dimensional time series space that meets the input requirements of the ANA module (feature dimension x time step x number of channels).
[0076] In the attention calculation stage, the ANA module dynamically recalibrates the features of the time series through a learnable adaptive weight mechanism. Specifically, this module calculates the correlation weights between different time steps and the importance distribution of channel features through parallel temporal attention branches and channel attention branches respectively, and finally generates an attention weight matrix through parameter fusion. The feature tensor weighted by attention is added to the original input features through a cross-layer identity mapping via a residual connection. This design not only preserves the integrity of the original temporal information but also effectively enhances the model's ability to capture key motion features.
[0077] Finally, the feature tensor enhanced by attention restores the spatial dimension through a deconvolution upsampling operation and is reconstructed into a human skeleton graph output with attention weights via a fully connected layer. The skeleton graph output is as Figure 5 shown. This processing flow can be formally expressed as:
[0078] H(x) = F(x) + αx
[0079] where H(x) represents the output of the residual, x represents the input, and F(x) represents the result after x passes through the ANA module. The α parameter is learnable. The entire process formula is as follows:
[0080]
[0081] FC↓(·) is a dimensionality reduction fully connected layer, ANA(·) is the adaptive attention calculation, represents the residual addition operation, and FC↑(·) is a dimensionality increase fully connected layer. This architecture design significantly improves the model's ability to model the temporal features of human motion while maintaining computational efficiency.
[0082] For parts or structures not specifically described in the present invention, existing technologies or existing products can be adopted and will not be elaborated here.
[0083] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made using the content of the specification of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A pedestrian anomaly detection and prediction method, characterized in that, Including: Convert the videos in the preprocessed video anomaly detection dataset into picture frames as the dataset of the model. Improve and construct a pedestrian anomaly detection prediction model based on the variational autoencoder bidirectional prediction RNN network. Input the preprocessed pictures into AlphaPose to obtain human skeleton diagrams and input them into the improved model (Bip-ITFA) for training. Input the human skeleton data of the abnormal behavior to be tested into the improved model for pedestrian anomaly detection, and output the anomaly curve graph of the video and the predicted human skeleton. Among them, the videos in the preprocessed video anomaly detection dataset include: (1) Select the publicly available video anomaly detection dataset on the network as the experimental dataset of the present invention; (2) Extract the video data by performing operations such as preprocessing, cropping the size, unifying the duration and frame rate on the pedestrian video data, and inputting it into AlphaPose to obtain skeleton information, so as to obtain the final human skeleton dataset; (3) Input the finally obtained dataset into the improved variational autoencoder bidirectional prediction RNN network for training. After the training is completed, input the human skeleton data of the abnormal behavior to be tested into the Bip-ITFA model for pedestrian anomaly detection.
2. The pedestrian anomaly detection and prediction method according to claim 1, wherein The selected dataset includes: (1) Select ShanghaiTech Campus taken by Shanghai Jiao Tong University as the dataset of this experiment. This dataset contains video sequences captured in 13 scenes with complex lighting conditions and diverse camera perspectives, belonging to a large-scale anomaly detection dataset. The training set consists of more than 270,000 frames including 330 videos, and the test set consists of 107 videos containing 130 abnormal events in 12 test scenes. The training videos only contain normal behaviors, while the test videos contain normal and abnormal behavior sequences. Most of the abnormal events in the test videos are related to human behaviors. Examples of abnormal behaviors include running, robbing, jumping, fighting, and cycling in the pedestrian area. (2) Select the CUHKAvenue dataset taken on the campus of the Chinese University of Hong Kong as the second dataset of this experiment. The CUHKAvenue dataset is a well-known dataset for video anomaly detection. This dataset contains 16 training video clips and 21 test video clips taken on the campus avenue of the Chinese University of Hong Kong, with a total of 30,652 frames (15,328 frames in the training set and 15,324 frames in the test set). The video clips in the CUHKAvenue dataset capture various scenes on the campus avenue, including normal behaviors and abnormal behaviors. Abnormal behaviors may include sudden running, dropping items, abnormal walking paths, etc. (3) Select the HR-ShanghaiTech Campus and HR-Avenue datasets as the third and fourth datasets of this experiment. The two datasets are subsets of ShanghaiTech Campus and CUHKAvenue, removing the videos irrelevant to human behaviors in the original datasets to further optimize the experimental materials.
3. The pedestrian anomaly detection and prediction method according to claim 1, characterized in that, The operations of preprocessing, cropping size, unifying duration, frame rate, and inputting into the AlphaPose algorithm to obtain skeleton information, etc. are used to extract data from video data, including: First, crop the video to unify the duration, then convert the video into 256×256 picture frames at a frame rate of 30 FPS, extract the skeletons from the picture frames in combination with the AlphaPose algorithm, and unify all the obtained skeleton information to form the human skeleton dataset corresponding to the human body dataset.
4. The pedestrian abnormal detection and prediction method according to claim 1, wherein The improved Bip-ITFA model includes: A variational autoencoder bidirectional prediction RNN network including an improved residual ASB-GRU interaction convolution module (AGI) and an adaptive weight attention module (ANA) is constructed. The improved network includes: The input skeleton data is mined and classified for features through the GRU network. Subsequently, a variational autoencoder network is introduced to generate the output F2. At the same time, the input feature F1 is processed by the residual ASB-GRU interaction convolution module (AGI) to obtain the output F3. To enhance the diversity of features and improve the training and inference efficiency, F2 and F3 are fused by parallel multiplication. Finally, in the training stage, F3 is used to predict the trajectory through the bidirectional prediction RNN network to obtain the predicted trajectory F4, and F4 passes through the adaptive weight attention module (ANA) to obtain the output predicted skeleton map and the abnormal score of the abnormal time frame.
5. The pedestrian anomaly detection and prediction method according to claim 1, characterized in that, The parallel introduction of the residual ASB-GRU interaction convolution module AGI into the original variational autoencoder network includes: Stack the human body pose sequences. After stacking, obtain the temporal encoding through GRU. Combine the joint points and temporal information in these data to form the activity sequence S. Perform a fast Fourier transform on the sequence S to obtain the frequency domain representation F. Calculate the power spectrum P of F. Finally, set a trainable threshold parameter θ. The frequencies with power higher than the threshold are retained, and the rest need to be filtered and added through the learnable global weight and local weight W G and W L Perform filtering and addition, and finally perform an inverse fast Fourier transform to obtain the sequence S'. The sequence S' will be added to the original sequence S. After fusing the temporal information through a GRU, obtain S'', and finally obtain the final output sequence O through a transposed convolution 6. The pedestrian anomaly detection and prediction method according to claim 1, characterized in that The introduction of an adaptive weight attention module (ANA) after the bidirectional prediction RNN network includes: This module uses method to perform low-rank approximation on the self-attention mechanism: First, stack the predicted skeleton graphs, and reduce the dimension of the stacked vectors through dimension conversion. In the data after dimension reduction Select representative subsets from the query matrix Q and the key matrix K respectively and As the positioning reference points, construct a low-rank projection matrix through singular value decomposition; secondly, divide the input sequence into non-overlapping regions of equal length, and generate feature vectors for each region through mean pooling operation; finally, construct an approximate expression of the attention weight matrix based on the reference points and the pooled feature vectors 7. The pedestrian anomaly detection and prediction method according to claim 1, wherein, The added residual structure includes: The residual structure takes the value obtained by adding the input data after convolution, normalization, and activation function processing to the equivalent mapping of the input data as the output and conveys it to the next layer structure.
Citation Information
Cited By
Lightweight few-sample man-machine interaction action recognition method, system and equipment
CN121096028A