Video Anomaly Detection Method and System Based on Multidimensional Second-Order Memory Guided Network
Through multi-dimensional second-order memory guidance network, the problem of insufficient detection accuracy in the prior art is solved, and more efficient abnormal detection effect is achieved.
Patent Information
- Application Number
- CN202211317988.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-10-26
AI Technical Summary
Existing video anomaly detection methods are difficult to effectively integrate the features of spatial and temporal dimensions, resulting in insufficient detection accuracy and the information of the memory network is too monotonous to accurately learn the data feature distribution.
A multidimensional second-order memory guidance network is used to extract the time-dimensional features between consecutive frames through implicit calculations, and four dimensional expansion convolution operations are performed. Combined with the multidimensional second-order memory guidance module to guide feature extraction at different stages, using intensity loss, feature attraction loss and rejection loss for training. In the test stage, the image quality is measured by PSNR to judge abnormalities.
It improves the accuracy and sensitivity of video anomaly detection, and can more effectively identify abnormal data, especially on complex data sets.
Smart Images

Figure CN115601679B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a video anomaly detection method and system based on a multi-dimensional second-order memory-guided network. Background Art
[0002] Video anomaly detection is an increasingly important research direction in the field of computer vision and is widely applied in fields such as industrial manufacturing and public safety. However, such tasks face huge challenges. Due to the wide application scope of anomaly detection, the definition of abnormal events is vague, the types of abnormal events are almost infinite and it is almost impossible to obtain all of them, and the positive and negative samples are extremely unbalanced. This makes it difficult for researchers to obtain an abnormal data set that can be used for training. Therefore, this task is converted into an unsupervised single-classification problem, that is, modeling among a large number of normal samples, enabling the model to learn the inherent patterns in such normal samples, and marking strange samples as abnormal during inference. Currently, the method of detecting anomalies based on reconstructing video frames has attracted much attention. The reconstruction-based method reconstructs features from the spatial dimension and obtains a reconstruction error for anomaly detection. The disadvantage of this method is very obvious, that is, it ignores the connection in the time dimension between consecutive video frames. Some methods based on the optical flow network can accurately display the motion features between video frames through complex calculations, but the motion features obtained by this method cannot be integrated into the feature reconstruction process. There are also some studies that have considered internalizing this complex external optical flow calculation process so that the motion features can be integrated into the feature reconstruction process, but their detection results often have poor accuracy. Another disadvantage of the reconstruction-based method is that it cannot accurately learn the feature distribution of the data. This is because the convolutional neural network has a powerful expression ability, so powerful that it can also well reconstruct abnormal times during the detection process. Some studies use memory networks to learn the data feature distribution and use this to constrain the powerful generalization ability of the convolutional neural network. However, the information they remember is too monotonous and has little improvement on the feature reconstruction effect.
[0003] According to the above analysis, it is necessary to provide a method for mixing and reconstructing spatial and temporal features, and using a memory model to guide features of different natures in stages. This method enables the reconstruction process to consider not only the information in the frame spatial dimension but also the information in the time dimension between consecutive frames, thereby improving the accuracy of anomaly detection. Summary of the Invention
[0004] In order to overcome the deficiencies of the above-mentioned prior art, the main object of the present invention is to provide a video anomaly detection method and system based on a multi-dimensional second-order memory-guided network to improve the accuracy of video anomaly detection.
[0005] To achieve the above invention object, the present invention adopts the following technical solutions:
[0006] Video anomaly detection method based on a multi-dimensional second-order memory-guided network, including an anomaly detection model training stage and an anomaly detection model testing stage:
[0007] In the anomaly detection model training stage, an implicit calculation method is used to extract features in the time dimension between consecutive frames, obtaining the initial liquidity feature f; four-dimensional expansion convolution operations are performed on the initial liquidity feature f to obtain the fused feature map F t ′; The corrected liquidity feature is extracted from the initial liquidity feature f and the feature map F t ′, and a multi-dimensional second-order memory-guided module is used to guide the process of liquidity feature extraction;
[0008] In the anomaly detection model testing stage, the entire video frame is input into the anomaly detection model, an anomaly score is given to the video frame, and the anomaly score is used to determine whether there is an anomaly in the video.
[0009] Furthermore, as a preferred technical solution of the present invention, the process of using the implicit calculation method to extract features in the time dimension between consecutive frames is specifically as follows:
[0010] Take the feature maps of 4 original consecutive frames and split them into four parts, and perform liquidity mapping between adjacent frames respectively, where the last frame and the first frame are subjected to liquidity mapping, as shown in formula (1):
[0011] C t =F t ·M t (1)
[0012] Where C t represents the quantified liquidity feature, F t is the original feature map of the t-th frame, and M t is the mapping area in the adjacent feature map corresponding to each pixel point in F t ;
[0013] C t contains all possible spatio-temporal displacements representing liquidity between consecutive frames, and kernel-soft-argmax is used to estimate the optimal displacement flow; kernel-soft-argmax uses a two-dimensional Gaussian kernel mask to suppress noise outliers, and the calculation formula for estimating the optimal displacement flow is as shown in formula (2):
[0014]
[0015] Where μ t (p) is the estimated optimal displacement flow, k t (·) is the two-dimensional Gaussian center, c t (·) is the correlation score, p and p′ are the abscissa and ordinate of each pixel point respectively, and ρ is an adjustable parameter;
[0016] Estimated optimal displacement flow μ t (p) Calculate F t The optimal displacement flow coordinates (x, y) of each pixel point on, and the maximum correlation value obtained in the implicit calculation process is v; store the optimal displacement flow coordinates (x, y) and the maximum correlation value v in the initial mobility feature f = (X, Y, V).
[0017] Further as a preferred technical solution of the present invention, perform four-dimensional expansion convolution operations on the initial mobility feature f = (X, Y, V), and each dimension expansion includes two convolution processes; after dimension expansion, obtain the feature map E t , where E t represents the mobility feature possessed by the initial mobility feature after dimension expansion; the E t is fused with the original feature map F t to obtain the feature map F t '; the feature map F t ' includes appearance features in the spatial dimension and mobility characteristics in the time dimension, and serves as the feature map input to the decoder.
[0018] Further as a preferred technical solution of the present invention, the information stored in the first-order memory module in the multi-dimensional second-order memory guidance module is the initial mobility feature f, and the second-order memory module stores the feature map F t ';
[0019] The first-order memory module includes the Lo-M module and the Ra-M module; input the optimal displacement flow coordinates (x, y) stored in the initial mobility feature f into the Lo-M module, and input the maximum correlation value v stored in the initial mobility feature f into the Ra-M module; during the training stage, Lo-M stores the typical distribution of the positions where mobility is generated in normal events, and Ra-M stores the typical distribution of the magnitude of the mobility change when normal events occur; during the testing stage, calculate the Euclidean distance between the initial mobility feature f and each memory item in the first-order memory module, and perform a concat operation with the most matching memory item to obtain the corrected mobility feature; where the Lo-M module calculates the Euclidean distance as shown in formula (3), and the Ra-M module calculates the Euclidean distance as shown in formula (4):
[0020]
[0021]
[0022] Among them, f t (x, y) is the optimal mobility coordinate of the t-th frame image, f t (v) is the maximum correlation value from t to t + 1, L' gra1Denote f t The Euclidean distance between (x, y) and the memory items in the Lo-M module, L′ gra2 Denote f t The Euclidean distance between f(v) and the memory items in the Ra-M module, g′ p1 Represents the memory item in the Lo-M module that has the closest Euclidean distance to f t (x, y), g′ p2 Represents the memory item in the Ra-M module that has the closest Euclidean distance to f t (v);
[0023] If there is no memory item in the Lo-M module close to f t (x, y), it indicates that the liquidity occurs at an unexpected position; if there is no memory item in the Ra-M module close to f t (v), it represents that the liquidity change is not caused by normal events;
[0024] The second-order memory module includes a combined compensation memory module Co-M; the combined compensation memory module Co-M constrains F t ′ and forces F t ′ to conform to the distribution of normal data; in the training stage, the Euclidean distance between the input data and the memory items of the combined compensation memory module Co-M is calculated as shown in formula (5):
[0025]
[0026] Among them, L″ gra Represents the Euclidean distance between F′ t and the memory items in the Co-M module, r t Represents the query item of F′ t , g″ p Represents the memory item in the combined compensation memory module Co-M that is closest to r t ;
[0027] Furthermore, as a preferred technical solution of the present invention, intensity loss, feature attraction loss, and repulsion loss are used as loss functions during the training process; the calculation of the intensity loss is as shown in formula (6):
[0028]
[0029] Among them Represents the predicted next frame, I t+1 Represents the real next frame;
[0030] The feature attraction loss forces the query item to match the most similar memory item in the multi-dimensional second-order memory guidance module, and the calculation formula of the feature attraction loss L gra is as shown in formula (7):
[0031] L gra = λ1L′ gra1 + λ2L′ gra2 + λ3L″ gra (7)
[0032] where λ1, λ2, and λ3 are adjustable weight parameters;
[0033] The rejection loss L rej forces the distance between respective memory items to be widened by maximizing the distance between the first match and the second match of the query item and the memory items in the multi-dimensional second-order memory guidance module, as shown in formula (8):
[0034]
[0035] where l′ g represents the inter-class distance of the first-order memory module, and l″ g represents the inter-class gap of the second-order memory module; are adjustable parameters; the larger l′ g , l″ are, the greater the difference in the data distributions represented by each memory item; the calculation formulas for l′ g , l″ g are shown in formulas (9) and (10) respectively:
[0036]
[0037]
[0038] where g′ s1 represents the second-nearest memory item in the Lo-M module in terms of the Euclidean distance from f t (x, y), and g′ s2 represents the second-nearest memory item in the Ra-M module in terms of the Euclidean distance from f t (v), and g″ p represents the nearest memory item in the joint compensation memory module Co-M to r t , and g″ s represents the second-nearest memory item in the joint compensation memory module Co-M to r t ;
[0039] Adding the intensity loss, feature attraction loss, and rejection loss together gives the objective function as shown in formula (11):
[0040] L pre = L int + L gra + L rej (11)
[0041] Among them, L pre represents the objective function of the multi-dimensional second-order memory guidance module.
[0042] Furthermore, as a preferred technical solution of the present invention, PSNR is used to measure the quality of the predicted image during the testing phase. The calculation of PSNR is shown in formula (12):
[0043]
[0044] Among them represents the predicted next frame, and I t+1 represents the real next frame, and N represents the number of pixel points. Considering that when the distribution of abnormal data is quite different from the memory items of the multi-dimensional second-order memory guidance module, it should be judged as abnormal. The abnormal evaluation score is shown in formula (13):
[0045]
[0046] Among them, S t represents the abnormal score, N(·) represents the normalization process, and L′ gra represents the feature loss function of the first-order memory module, and α and β are adjustable parameters.
[0047] Furthermore, the present invention also proposes a video anomaly detection system based on a multi-dimensional second-order memory guidance network. This system includes:
[0048] Anomaly detection model training unit, which extracts the features in the time dimension between consecutive frames by using an implicit calculation method to obtain the initial liquidity feature f; performs four-dimensional expansion convolution operations on the initial liquidity feature f to obtain the fused feature map F t ′; extracts the corrected liquidity feature from the initial liquidity feature f and the feature map F t ′, and uses the multi-dimensional second-order memory guidance module to guide the process of liquidity feature extraction;
[0049] Anomaly detection model testing unit, which inputs the entire video frame into the anomaly detection model, gives an anomaly score to the video frame, and uses the anomaly score to judge whether there is an anomaly in the video.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] 1. During the process of feature extraction, the present invention pays attention to the liquidity features in the time dimension between consecutive frames, and proposes a strategy of dimension expansion and feature fusion to fuse the liquidity features and appearance features across dimensions for subsequent feature reconstruction, improving the authenticity of the reconstructed image, and thus improving the detection accuracy;
[0052] 2. The present invention adopts a multi-dimensional second-order memory-guided module guidance strategy, learning data distributions representing normal data flow characteristics in different forms at different stages, thereby guiding and standardizing feature extraction during the training stage and more effectively identifying abnormal data during the testing stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the drawings and specific embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:
[0054] Figure 1 is the overall framework diagram of the method of the present invention;
[0055] Figure 2 is the video detection result diagram in the UCSD Ped2 dataset provided by the embodiment of the present invention;
[0056] Figure 3 is the video detection result diagram in the CHUK Avenue dataset provided by the embodiment of the present invention;
[0057] Figure 4 is the video detection result diagram in the ShanghaiTech dataset provided by the embodiment of the present invention;
[0058] Figure 5 is the ROC value of different components in the multi-dimensional second-order memory-guided module of the present invention on different datasets provided by the embodiment;
[0059] Figure 6 is the corresponding fps and AUC values of 7 models tested on the Ped2 dataset provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] The following will describe the present invention in detail with reference to the drawings for further explanation, so that those skilled in the art can understand the present invention more deeply and be able to implement it. However, the following is only used to explain the present invention by reference to examples and is not a limitation of the present invention.
[0061] As Figure 1 , the video anomaly detection method based on a multi-dimensional second-order memory-guided network includes a model training stage and a testing stage:
[0062] In the model training stage, an implicit calculation method is used to extract features in the time dimension between consecutive frames to obtain initial liquidity features f; convolution operations are used to obtain a fused feature map F t'; Use a multi-dimensional second-order memory guidance module to guide the process of extracting flow features and obtain corrected flow features; in the test phase, input the entire video frame into the trained model, assign an anomaly score to the video frame, and use the anomaly score to determine whether there is an anomaly in the video.
[0063] The meaning of implicit calculation is to map the flow between frames into quantifiable values and internalize this process in a single network. The specific process of using the implicit calculation method to extract features in the time dimension between consecutive frames is as follows:
[0064] Take the feature maps of 4 original consecutive frames and split them into four parts, respectively perform flow mapping between adjacent frames, and perform flow mapping between the last frame and the first frame, as shown in formula (1):
[0065] C t =F t ·M t (1)
[0066] Where C t represents the quantified flow feature, F t is the original feature map of the t-th frame, and M t is a mapping area of a certain size in the adjacent feature map corresponding to each pixel point in F t ;
[0067] C t contains all possible spatio-temporal displacements representing the flow between consecutive frames, but C t contains a large amount of invalid information and needs to use kernel-soft-argmax to estimate the optimal displacement flow; kernel-soft-argmax uses a two-dimensional Gaussian kernel mask to suppress noise outliers, and the calculation formula for estimating the optimal displacement flow is as shown in formula (2):
[0068]
[0069] Where μ t (p) is the estimated optimal displacement flow, k t (·) is the two-dimensional Gaussian center, c t (·) is the correlation score, p and p′ are the horizontal and vertical coordinates of each pixel point respectively, and ρ is an adjustable parameter;
[0070] The estimated optimal displacement flow μ t (p) is calculated to obtain F tThe optimal displacement flow coordinates (x, y) of each pixel on it, and the maximum correlation value obtained during the implicit calculation process is v; store the optimal displacement flow coordinates (x, y) and the maximum correlation value v in the initial liquidity feature f = (X, Y, V); the optimal displacement flow coordinates (x, y) represent the displacement from the origin of each pixel to the predicted point, which contains the liquidity with the concept of time series, and the maximum correlation value v represents the confidence of this prediction. The size of the confidence depends on whether the event occurs significantly in the original image, reflecting the severity of the liquidity change.
[0071] Perform four-dimensional expansion convolution operations on the initial liquidity feature f = (X, Y, V), and each dimensional expansion contains two convolution processes; after dimensional expansion, the feature map E is obtained t , where E t represents the liquidity feature obtained after dimensional expansion of the initial liquidity feature, which can better express the liquidity feature; the E t is fused with the original feature map F t to obtain the feature map F t '; the feature map F t ' is used as the feature map input to the decoder. F t ' contains the appearance feature in the spatial dimension and the liquidity feature in the time dimension, so that during the image reconstruction process, while retaining the static appearance feature, it makes up for the lack of dynamic changes in the time dimension of the image, making the reconstructed image more real. And in the test stage, the method of the present invention is more sensitive to abnormalities. Whether it is an appearance abnormality in space or an abnormal behavior of time change will cause the network to be unable to normally reconstruct the predicted video frame.
[0072] The information stored in the first-order memory module in the multi-dimensional second-order memory-guided module (MTM-net) is the initial liquidity feature f, and the second-order memory module stores the feature map F' t ;
[0073] The first-order memory module includes the Lo-M module and the Ra-M module; input the optimal displacement flow coordinates (x, y) stored in the initial liquidity feature f into the Lo-M module, and input the maximum correlation value v stored in the initial liquidity feature f into the Ra-M module; during the training stage, Lo-M stores the typical distribution of the positions where liquidity is generated in normal events, and Ra-M stores the typical distribution of the magnitude of liquidity change when normal events occur; in the test stage, calculate the Euclidean distance between the initial liquidity feature f and each memory item in the first-order memory module, and perform a concat operation with the most matching memory item to obtain the corrected liquidity feature; the Euclidean distance calculated by the Lo-M module is shown in formula (3), and the Euclidean distance calculated by the Ra-M module is shown in formula (4):
[0074]
[0075]
[0076] Among them, f t (x, y) is the best fluidity coordinate of the t-th frame picture, f t (v) is the maximum correlation value from t to t + 1, L′ gra1 represents f t (x, y) is the Euclidean distance between the memory item in the Lo-M module, L′ gra2 represents f t (v) is the Euclidean distance between the memory item in the Ra-M module, g′ p1 represents the memory item in the Lo-M module that is closest to f t (x, y) in terms of Euclidean distance, g′ p2 represents the memory item in the Ra-M module that is closest to f t (v) in terms of Euclidean distance;
[0077] If there is no memory item in the Lo-M module close to f t (x, y), it means that the fluidity occurs at an unexpected position; if there is no memory item in the Ra-M module close to f t (v), it means that the fluidity change is not caused by normal events;
[0078] F′ t includes spatial appearance features and temporal fluidity features. A first-order memory module cannot store features of both forms and requires a second-order memory module for constraint; the second-order memory module includes a combined compensation memory module Co-M; the combined compensation memory module Co-M constrains F′ t and forces F′ t to conform to the distribution of normal data; in the training stage, the Euclidean distance between the input data and the memory item of the combined compensation memory module Co-M is calculated as shown in formula (5):
[0079]
[0080] Among them, L″ gra represents the Euclidean distance between F′ t and the memory item in the Co-M module, r t represents the query item of F′ t , g″ p represents the memory item in the combined compensation memory module Co-M that is closest to r t ;
[0081] During the training process, intensity loss, feature attraction loss, and repulsion loss are used as loss functions; the calculation of intensity loss is as shown in formula (6):
[0082]
[0083] where represents the predicted next frame, and I t+1 represents the true next frame;
[0084] The feature attraction loss forces the query term to match the most similar memory term in the multi-dimensional second-order memory guidance module. The formula for the feature attraction loss L gra is shown in Equation (7) as follows:
[0085] L gra = λ1L′ gra1 + λ2L′ gra2 + λ3L″ gra (7)
[0086] where λ1, λ2, and λ3 are adjustable weight parameters;
[0087] The rejection loss L rej forces the distance between the respective memory terms to be widened by maximizing the distance between the best match and the second-best match of the query term and the memory terms in the multi-dimensional second-order memory guidance module, as shown in Equation (8):
[0088]
[0089] where l′ g represents the inter-class distance of the first-order memory module, and l″ g represents the inter-class distance gap of the second-order memory module. is an adjustable parameter; the larger l′ g and l″ g are, the greater the difference in the data distributions represented by each memory term; the calculation formulas for l′ g and l″ g are shown in Equations (9) and (10) respectively:
[0090]
[0091]
[0092] where g′ s1 represents the second-nearest memory term in terms of the Euclidean distance from f t (x, y) in the Lo-M module, g′ s2 represents the second-nearest memory term in terms of the Euclidean distance from f t (v) in the Ra-M module, g″ p represents the memory term closest to r t in the joint compensation memory module Co-M, and g″ sThe second nearest memory item among the co-compensation memory modules Co-M and r t The second nearest memory item;
[0093] Add the intensity loss, feature attraction loss, and rejection loss to obtain the objective function as shown in formula (11):
[0094] L pre = L int + L gra + L rej (11)
[0095] Where L pre Represents the objective function of the multi-dimensional second-order memory guidance module.
[0096] In the test phase, PSNR is used to measure the quality of the predicted image. The calculation of PSNR is shown in formula (12):
[0097]
[0098] Where Represents the predicted next frame, I t+1 Represents the real next frame, N represents the number of pixel points; considering that when the distribution of abnormal data is significantly different from the memory items of the multi-dimensional second-order memory guidance module, it should be judged as abnormal. The abnormal evaluation score is shown in formula (13):
[0099]
[0100] Where, S t Represents the abnormal score, N(·) represents the normalization process, L′ gra Represents the loss function of the first-order memory module, and α and β are adjustable parameters.
[0101] The effectiveness of the present invention is verified through specific experiments below. The experimental content includes:
[0102] (1) Introduction to the test dataset
[0103] This embodiment conducts evaluation experiments on three datasets commonly used for video anomaly detection, including the UCSDPedestrian 2 dataset (Ped2), the CUHK Avenue dataset (Avenue), and the ShanghaiTech dataset (ShanghaiTech). The Ped2 dataset contains 16 training videos and 12 test videos, among which there are 12 abnormal events. The abnormalities in Ped2 are all about abnormal walking events, such as the entry of cars and bicycles, and pedestrians skateboarding. The CUHK Avenue dataset contains 16 training videos and 21 test videos, including 47 abnormal events involving behaviors such as running and throwing objects. Due to the change in the camera shooting angle, the size of pedestrians in this dataset will change. Therefore, compared with Ped2, the abnormalities in Avenue are more difficult to detect. The ShanghaiTech dataset is the most challenging existing anomaly detection dataset, which consists of 330 training videos (274K frames) and 107 test videos (42K frames) taken in 13 scenarios, including 130 complex and diverse abnormal events.
[0104] (2) Parameter settings
[0105] The model of this embodiment is optimized by the Adam optimizer, with an initial learning rate of lr = 2e-4. The parameters λ1, λ2, and λ3 in the feature attraction loss L gra are 0.1, 0.1, and 0.1 respectively. The parameters rej and γ″ in the repulsion loss L are 0.1, 0.1, 1.2, and 1.0 respectively. The number of memory items in the first-order memory module is 10, and the number of channels is 2 and 1 respectively. The number of memory items in the second-order memory module is 10, and the number of channels is 512. We train for 40, 40, and 10 epochs on the Ped2, Avenue, and ShanghaiTech datasets respectively. The single end-to-end network of this embodiment is implemented using Pytorch, and all experiments are conducted on an NVIDIA RTX GPU.
[0106] (3) Model framework process
[0107] This embodiment uses an end-to-end single-branch network framework based on future frame prediction, pays special attention to the mobility changes in different stages between consecutive frames during the feature extraction process, and designs a multi-dimensional two-stage memory-guided module for the features in different spatial and temporal dimensions, as well as the different morphological features in different stages of the same-dimensional mobility features, to guide the adjustment of feature extraction. The overall framework diagram is as shown in Figure 1 shown. The consecutive frames I1 to I t are the inputs, and the predicted frames is the output. An implicit calculation and a multi-dimensional two-stage memory-guided module guiding process are included in the middle of the encoder (E) and the decoder (D). W, H, and C represent the dimensions of the encoder output tensor. R j is the initial result of the fluidity feature calculated in the implicit calculation process, and R' j is the initial fluidity feature after being guided by the one-stage memory module. g' p1 , g' p2 and g'' p are the best-matched memory items read by the Lo-M, Ra-M, and Co-M modules respectively.
[0108] (4) Model performance comparison
[0109] The present invention determines the performance of the model by comparing the area under the ROC curve, that is, the AUC(%) value. The higher the AUC value, the better the performance of the model. The AUC value of the method of the present invention on the CUHK Avenue dataset is 72.73%. In this embodiment, the model of the present invention is compared with many advanced models, as shown in Table 1:
[0110] Table 1 AUC results using different methods on UCSD Ped2, CUHK Avenue, and ShanghaiTech
[0111]
[0112]
[0113] The overall AUC of the method of the present invention on the ped2 dataset reaches the highest value of 97.6%, which is significantly higher than other algorithms and 12.6% higher than the baseline method autoencoder. On the Avenue dataset, the overall AUC obtained by the method of the present invention is 88.2%, which is 8.2% higher than the autoencoder. On the most complex ShanghaiTech dataset, the method of the present invention obtains 74.1%. This proves the effectiveness of using the present invention for anomaly detection.
[0114] The following is illustrated by specific videos. Figure 2 , Figure 3 , Figure 4 For visualizing several examples of the method of the present invention for anomaly detection on three datasets, where the broken lines in the figure represent the scores of a certain continuous frame in the test video sequence, and the lower the score, the more likely it is an anomaly. The shaded part represents the ground truth anomaly, the arrow points to the corresponding ground truth video frame, and the boxed part is the anomaly event. It can be seen from the figure that under normal circumstances, the anomaly score maintains a relatively stable level with slight fluctuations up and down. When an anomaly event occurs, the anomaly score will increase with the appearance of the anomaly (such as Figure 2riding a bicycle, Figure 3 running and throwing objects, Figure 4 (fighting and throwing objects) decreased sharply and maintained a relatively stable level, while the abnormal score increased in a timely manner after the abnormal event ended. This proves the sensitivity of the method of the present invention to abnormal events during abnormal detection.
[0115] During testing, the abnormal score will have some unexpected large fluctuations, especially for the Avenue and ShanghaiTech datasets, mainly due to some unavoidable noises in the datasets (such as camera jitter, rare normal events). The appearance of this kind of noise increases the complexity of the dataset, but in this complex and noisy environment for testing, the model of the present invention still shows good discrimination ability, reflecting the good performance of the model of the present invention.
[0116] (5) Ablation experiment
[0117] In this embodiment, the multi-dimensional two-stage memory-guided module model is split into two models with single-stage memory-guided modules, and the improvements brought by the memory-guided mobility features in different stages are demonstrated.
[0118] Table 2 Ablation experiment data for UCSD Ped2, CUHK Avenue and ShanghaiTech datasets
[0119]
[0120] As can be seen from Table 2, only using the first-stage memory-guided module to memory-guide the initially extracted mobility change features brought an improvement of 0.4% to Ped2, an improvement of 0.8% to Avenue, and an improvement of 1.3% to the ShanghaiTech dataset. Only using the second-stage memory-guided module to memory-guide the extended and fused mobility features brought an improvement of 0.9% to Ped2, an improvement of 1.2% to Avenue, and an improvement of 0.6% to the ShanghaiTech dataset. This shows that memory-guiding the mobility features of different dimensions in different stages will lead to the improvement of abnormal detection, proving the importance of the present invention in memory-guiding the features with different characteristics and dimensions.
[0121] The multi-dimensional second-order memory-guided module used in the present invention achieved the best results (97.6% & 88.2% & 74.1%) on both the Avenue dataset and the Ped2 dataset, which strongly proves that the multi-dimensional second-order memory guidance is superior to the single-stage memory guidance. As Figure 5 shown in the curve, it more intuitively shows the influence of different components in the multi-dimensional second-order memory-guided module.
[0122] (6) Detection speed
[0123] The model of this embodiment was trained on an NVIDIA TiTan RTX GPU, and the frames per second (fps) and AUC of the model during the test process were recorded as comprehensive indicators to measure the model efficiency. As Figure 6 shown, the number of frames of the 7 models represents the running speed of the models. It can be seen from the figure that the model of this embodiment has a high fps, and the AUC reaches the highest 97.6, indicating a good trade-off between speed and accuracy.
[0124] The video anomaly detection system based on the multi-dimensional second-order memory-guided network proposed in the present invention includes:
[0125] Anomaly detection model training unit, which uses an implicit calculation method to extract features in the time dimension between consecutive frames to obtain the initial liquidity feature f; performs four-dimensional expansion convolution operations on the initial liquidity feature f to obtain the fused feature map F t '; extract the corrected liquidity feature from the initial liquidity feature f and the feature map F t ', and use the multi-dimensional second-order memory-guided module to guide the process of extracting the liquidity feature;
[0126] Anomaly detection model testing unit, which inputs the entire video frame into the anomaly detection model, gives an anomaly score to the video frame, and uses the anomaly score to determine whether there is an anomaly in the video.
[0127] It should be noted that the description of the system in the embodiments of the present invention is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments, so it will not be repeated here.
[0128] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, so that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0129] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A video anomaly detection method based on a multi-dimensional second-order memory-guided network, characterized in that, It includes an abnormal detection model training stage and an abnormal detection model testing stage: In the abnormal detection model training stage, an implicit calculation method is adopted to extract features in the time dimension between consecutive frames, obtaining the initial liquidity feature f; four-dimensional expansion convolution operations are performed on the initial liquidity feature f to obtain the fused feature map F t '; The corrected liquidity feature is extracted from the initial liquidity feature f and the feature map F t '; and a multi-dimensional second-order memory guidance module is used to guide the process of liquidity feature extraction; In the abnormal detection model testing stage, the entire video frame is input into the abnormal detection model to assign an abnormal score to the video frame, and the abnormal score is used to determine whether there is an abnormality in the video. The information stored in the first-order memory module in the multi-dimensional second-order memory guidance module is the initial liquidity feature f, and the second-order memory module stores the feature map F t ′; The first-order memory module includes a Lo-M module and a Ra-M module; the optimal displacement flow coordinates (x, y) stored in the initial mobility feature f are input into the Lo-M module, and the maximum correlation value v stored in the initial mobility feature f is input into the Ra-M module. During the training stage, the Lo-M stores the typical distribution of the positions where mobility is generated in normal events, and the Ra-M stores the typical distribution of the magnitudes of the mobility changes that occur when normal events occur; during the testing stage, the Euclidean distance between the initial mobility feature f and each memory item in the first-order memory module is calculated, and the most matching memory item is concatenated with it to obtain a corrected mobility feature; the Euclidean distance calculation of the Lo-M module is shown in formula (3), and the Euclidean distance calculation of the Ra-M module is shown in formula (4): Among them, f t (x, y) is the optimal fluidity coordinate of the t-th frame image, f t (v) is the maximum correlation value from t to t + 1, L′ gra1 represents f t (x, y) is the Euclidean distance between the memory item in the Lo-M module, L′ gra2 represents f t (v) is the Euclidean distance between the memory item in the Ra-M module, g′ p1 represents the memory item in the Lo-M module that is closest to f t (x, y) in terms of Euclidean distance, g′ p2 represents the memory item in the Ra-M module that is closest to f t (v) in terms of Euclidean distance; If there is no memory item in the Lo-M module close to f t (x,y), it means that the liquidity occurs at an unexpected position; if there is no memory item in the Ra-M module close to f t (v), it represents that the change in liquidity is not caused by normal events. The second-order memory module includes a combined compensation memory module Co-M, and the combined compensation memory module Co-M constrains F t ′, and forcibly guides F′ t to conform to the distribution of normal data; in the training phase, calculate the Euclidean distance between the input data and the memory items of the combined compensation memory module Co-M, as shown in formula (5): Among them, L″ gra represents the Euclidean distance between F′ t and the memory items in the Co-M module, and r t represents the query item of F′ t , and g″ p represents the memory item in the joint compensation memory module Co-M that is closest to r t .
2. The video anomaly detection method based on the multi-dimensional second-order memory-guided network according to claim 1, wherein The process of extracting features in the time dimension between consecutive frames using the implicit calculation method specifically includes: The feature maps of 4 original consecutive frames are split into four parts, and mobility mapping is performed between adjacent frames respectively, where the last frame and the first frame are subjected to mobility mapping, as shown in formula (1): C t = F t · M t (1) Among them, C t represents the quantized liquidity feature, F t is the original feature map of the t-th frame, M t is the mapping region in the adjacent feature map corresponding to each pixel point in F t ; C t It contains all possible spatio-temporal displacements representing fluidity between consecutive frames, and uses kernel-soft-argmax to estimate the optimal displacement flow; kernel-soft-argmax uses a two-dimensional Gaussian kernel mask to suppress noise outliers, and the calculation formula for estimating the optimal displacement flow is shown in Formula (2): where μ t (p) is the estimated optimal displacement flow, k t (·) is the two-dimensional Gaussian center, c t (·) is the correlation score, p and p′ are the abscissa and ordinate of each pixel point respectively, and ρ is an adjustable parameter; Estimated optimal displacement flow μ t (p) Calculate to obtain F t The optimal displacement flow coordinates (x, y) of each pixel point on, and the maximum correlation value obtained during the implicit calculation process is v; store the optimal displacement flow coordinates (x, y) and the maximum correlation value v in the initial mobility feature f = (X, Y, V).
3. The video anomaly detection method based on the multi-dimensional second-order memory-guided network according to claim 2, wherein Perform four - dimensional expansion convolution operations on the initial liquidity feature f=(X, Y, V), where each dimensional expansion contains two convolution processes; after the dimensional expansion, the feature map E is obtained. t , the feature map E t represents the liquidity feature after dimensional expansion of the initial liquidity feature; the E t is fused with the original feature map F t to obtain the feature map F t '; the feature map F t ' contains the appearance feature in the spatial dimension and the liquidity characteristic in the temporal dimension, and serves as the feature map input to the decoder.
4. The video anomaly detection method based on the multi-dimensional second-order memory-guided network according to claim 3, characterized in that, During the training process, the intensity loss, feature attraction loss, and rejection loss are used as loss functions; the calculation of the intensity loss is shown in formula (6): Among them represents the predicted next frame, I t+1 represents the true next frame; The feature attraction loss forces the query item to match the most similar memory item in the multi-dimensional second-order memory guidance module. The feature attraction loss L gra is calculated as shown in Equation (7): L gra = λ1L' gra1 + λ2L' gra2 + λ3L'' gra (7) Where λ1, λ2, and λ3 are adjustable weight parameters; Rejection loss L rej By maximizing the distance between the first match and the second match of the query term and the memory term in the multi-dimensional second-order memory guidance module, the distance between the respective memory terms is forced to be widened, as shown in formula (8): where \(l'\) g represents the inter-class distance of the first-order memory module, and \(l''\) g represents the inter-class distance gap of the second-order memory module. , , \(\gamma'\), \(\gamma''\) are adjustable parameters; the larger \(l'\) g and \(l''\) g are, the greater the difference in the data distributions represented by each memory item; the calculation formulas for \(l'\) g and \(l''\) g are shown in Equations (9) and (10) respectively: Among them, g′ s1 represents the second-closest memory item in the Lo-M module to the Euclidean distance of f t (x,y), and g′ s2 represents the second-closest memory item in the Ra-M module to the Euclidean distance of f t (v). g″ p represents the closest memory item in the combined compensation memory module Co-M to r t , and g″ s represents the second-closest memory item in the combined compensation memory module Co-M to r t ; The intensity loss, feature attraction loss, and rejection loss are added together to obtain the objective function as shown in formula (11): L pre = L int + L gra + L rej (11) Among which L pre represents the objective function of the multi-dimensional second-order memory guiding module.
5. The video anomaly detection method based on a multi-dimensional second-order memory-guided network according to claim 4, wherein During the testing stage, PSNR is used to measure the quality of the predicted image, and the calculation of PSNR is shown in formula (12): Among them represents the predicted next frame, I t+1 represents the real next frame, N represents the number of pixel points; considering that when the distribution of abnormal data is quite different from the memory items of the multi-dimensional second-order memory guidance module, it should be judged as abnormal, and the abnormal evaluation score is shown in formula (13): Among them, S t represents the anomaly score, N(·) represents the normalization process, and L′ gra represents the feature loss function of the first-order memory module, and α and β are adjustable parameters.
6. Video anomaly detection system based on a multi-dimensional second-order memory-guided network, characterized in that, It includes: Anomaly detection model training unit, which uses an implicit calculation method to extract features in the time dimension between consecutive frames to obtain an initial liquidity feature f; performs four-dimensional expansion convolution operations on the initial liquidity feature f to obtain a fused feature map F t '; extracts the corrected liquidity feature from the initial liquidity feature f and the feature map F t '; and uses a multi-dimensional second-order memory guidance module to guide the process of liquidity feature extraction An abnormal detection model testing unit that inputs the entire video frame into the abnormal detection model to assign an abnormal score to the video frame, and uses the abnormal score to determine whether there is an abnormality in the video.
Citation Information
Patent Citations
Implicit motion compensation video object segmentation method and device
CN115147765A
Prefrontal modulation of context-specific memory encoding and retrieval in the hippocampus
US10664749B1