Video Anomaly Detection Method Based on Spatiotemporal Enhanced Associative Memory
By using the method of space-time enhancement of associated memory in video anomaly detection, recording and learning the prototypes of normal events and their prototype relationships, and using motion characteristics to enhance appearance characteristics, the problem of poor detection performance in the prior art is solved, and more efficient video anomaly detection is achieved.
Patent Information
- Application Number
- CN202310950812.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-07-31
AI Technical Summary
When using memory network to save prototypes of normal events, existing video anomaly detection methods fail to fully explore the rich relationships between prototype memories, resulting in poor detection performance.
The video anomaly detection method based on space-time enhancement associated memory is adopted to record and learn the prototypes of normal events and their prototype relationships, adjust features, and use motion features to enhance appearance features to achieve space-time semantic enhancement.
By recording and learning the prototypes of normal events and their prototype relationships, normal frames can be predicted more accurately, thereby improving the performance of video anomaly detection. The use of motion characteristics to enhance appearance characteristics, make full use of the intrinsic connections between space-time characteristics, and further improve the effectiveness of detection.
Smart Images

Figure CN116958878B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video anomaly detection, and particularly to a video anomaly detection method based on spatio-temporal enhanced associative memory. Background Art
[0002] With the wide application of surveillance cameras in public places, video anomaly detection has received increasing attention. However, video anomaly detection is a challenging task because the definition of anomalies is usually ambiguous, and anomaly events are very rare so that it is impossible to collect all anomaly events. Therefore, video anomaly detection is generally regarded as an unsupervised task, that is, only normal data is used to train the model during training, and frames that do not conform to the model are regarded as anomaly frames during testing.
[0003] In the prior art, unsupervised video anomaly detection usually uses reconstruction-based or prediction-based methods. Existing solutions propose to predict the next frame with consecutive multi-frame images and add optical flow constraints to extract motion features. There are also solutions that use a dual-generator-based framework to learn the interactions between normal scenes and capture global dependencies in the spatial and temporal domains by introducing motion position attention. However, although deep neural networks (DNNs) have advantages in representation ability, there is a risk that the network can accurately reconstruct or predict anomalies, resulting in a reduction in the overall effectiveness of anomaly detection.
[0004] To solve the above problems, some recent solutions propose to use memory networks to save the prototypes of normal events to reduce the representation ability of DNNs. In the existing solutions, a memory-augmented autoencoder (MemAE) is introduced, which uses appearance features to query the memory bank of the memory module to obtain the most relevant feature prototypes, and then aggregates the features and decodes them using a decoder. There are also solutions that propose to integrate a storage module with an update mechanism at the network bottleneck to record the prototype patterns of different normal events, thereby enhancing the model's prediction ability for normal video frames and suppressing the prediction ability for anomaly video frames. There are also solutions that suggest using denoising tasks and reconstruction tasks to record the appearance and optical flow normal patterns respectively, and exploring their correlations through adversarial learning.
[0005] However, the above memory networks are all based on the memory of the prototype content of normal events, without exploring the rich relationships between these prototype memories, resulting in poor performance in video anomaly detection. At the same time, existing solutions generally process appearance features and motion features separately, without directly modeling the spatio-temporal semantic representation of normal events and without making full use of the internal connection of spatio-temporal information, resulting in poor effectiveness in video anomaly detection. Therefore, how to improve the performance of video anomaly detection is a technical problem that needs to be solved urgently. Summary of the Invention
[0006] Aiming at the deficiencies of the above-mentioned existing technologies, the technical problem to be solved by the present invention is: how to provide a video anomaly detection method based on spatio-temporal enhanced associative memory, which can adjust features by recording and learning the prototypes of normal events and their prototype relationships, and can use motion features to enhance appearance features to achieve spatio-temporal semantic enhancement, thereby improving the performance of video anomaly detection.
[0007] To solve the above technical problem, the present invention adopts the following technical solutions:
[0008] A video anomaly detection method based on spatio-temporal enhanced associative memory, including:
[0009] S1: Obtain the video frame sequence to be detected;
[0010] S2: Input the video frame sequence into the trained anomaly detection model, and output the corresponding anomaly prediction value;
[0011] S3: Use the anomaly prediction value output by the anomaly detection model as the anomaly detection result of the video frame sequence to be detected;
[0012] In step S2, the anomaly detection model is trained through the following steps:
[0013] S201: Obtain the training data including the training video frame sequence and the anomaly ground truth, and generate the optical flow sequence of the training video frame sequence;
[0014] S202: Input the training video frame sequence and its optical flow sequence into the anomaly detection model;
[0015] S203: Extract the appearance features and motion features of the training video frame sequence and its optical flow sequence through the appearance encoder and the motion encoder;
[0016] S204: Use the motion features to fuse and enhance the appearance features through the spatio-temporal enhancement module to obtain the fused features;
[0017] S205: Perform associative retrieval based on the fused features through the associative memory module to obtain the relationships between the normal event prototypes;
[0018] S206: Adjust the fused features according to the relationships between the normal event prototypes to generate the final features;
[0019] S207: Decode the final features through the decoder to obtain the corresponding anomaly prediction value;
[0020] S208: Calculate the model loss according to the anomaly prediction value and the corresponding anomaly ground truth, and optimize the model parameters;
[0021] S209: Repeat steps S201 to S208 until the anomaly detection model converges.
[0022] Preferably, the appearance encoder, the motion encoder, and the decoder are all improved U-Net networks;
[0023] The improvements to the appearance encoder and the motion encoder are as follows: the pooling layer of the U-Net network is changed to a strided convolutional layer; the improvement to the decoder is that a deformable convolutional network layer is added before the transposed convolution of the U-Net network.
[0024] Preferably, the spatio-temporal enhancement module weights the appearance features based on the motion variance attention map of the motion features to enhance the appearance features through the motion features.
[0025] Preferably, the processing steps of the spatio-temporal enhancement module are as follows:
[0026] S2041: Use the appearance features and the motion features as the input of the spatio-temporal enhancement module;
[0027] S2042: Input the motion features into a 1×1 convolutional layer and a softmax layer to obtain the optical flow global context features;
[0028] S2043: Input the optical flow global context features into the variance attention module to generate the motion variance attention map;
[0029] S2044: Perform a Hadamard product operation on the appearance features and the motion variance attention map to generate the motion-enhanced appearance features;
[0030] S2045: After sequentially inputting the motion-enhanced appearance features into a 1×1 convolutional layer, a LayerNorm function layer, a ReLU function layer, and a 1×1 convolutional layer, add them to the appearance features to obtain the fused features.
[0031] Preferably, the variance attention module calculates and generates the motion variance attention map through the following formula:
[0032]
[0033] In the formula: att v represents the motion variance attention map; f m represents the optical flow global context features; d represents the size of the channels; D represents the number of spatial dimensions.
[0034] Preferably, the associative memory module records the relationships between normal event prototypes through a memory network; among them, it includes a content memory module and a relationship memory module for respectively recording the prototypes of normal events and the relationships between prototypes.
[0035] Preferably, the processing steps of the associative memory module are as follows:
[0036] S2051: Use the fused features as the input of the associative memory module;
[0037] S2052: Read the content memory module and the relationship memory module of the previous state;
[0038] The formula description is:
[0039]
[0040]
[0041] In the formula: X 1 represents the content memory module and the relationship memory module of the previous state; x represents the fused feature; f 1 , f 2 , f 3 represents a linear layer; represents the outer product;
[0042] S2053: Update the content memory module and the relationship memory module of the current state according to the read content memory module and relationship memory module of the previous state;
[0043] The formula description is:
[0044]
[0045]
[0046] Among them,
[0047] In the formula: M r , M i represent the updated content memory module and relationship memory module; M r-1 represents the content memory module and the relationship memory module of the previous state; f 4 represents a linear layer; SAM represents the self-attention associative memory operation;
[0048] S2054: Retrieve the updated relationship memory module to obtain the relationships between normal event prototypes.
[0049] Preferably, the final feature is generated through the following steps:
[0050] S2061: Input the relationships between normal event prototypes into the forward propagation network to obtain a content-addressable memory;
[0051] The formula description is:
[0052] M = f(M r );
[0053] In the formula: M represents the content-addressable memory; M rRepresents the relationship between normal event prototypes; f represents the forward propagation network;
[0054] S2062: Calculate the cosine similarity between the content-addressable memory and the input fusion feature, and perform a softmax operation on the cosine similarity to obtain the similarity weights between the fusion feature and each entry in the content-addressable memory;
[0055] The formula description is:
[0056]
[0057]
[0058] In the formula: w i Represents the similarity weight; m i , m j Represents the entries in the content-addressable memory M; x represents the fusion feature; d(x, m i ), d(x, m j ) represents the cosine similarity between the fusion feature and the entries m i and m j ;
[0059] S2063: Multiply the similarity weights and the entries of the content-addressable memory to obtain the corresponding prototype features;
[0060] The formula description is:
[0061]
[0062] In the formula: Represents the prototype feature; w represents the similarity weight matrix composed of the similarity weights w i ;
[0063] S2064: Concatenate the fusion feature and the prototype feature in dimension to obtain the final feature.
[0064] Preferably, calculate the model loss through the following formula:
[0065]
[0066] In the formula: L rec Represents the model loss; I t Represents the ground truth of the abnormality of the video frame; Represents the predicted value of the abnormality of the video frame.
[0067] Preferably, evaluate the performance of the anomaly detection model through the anomaly score;
[0068] Calculate the anomaly score of the anomaly detection model through the following formula:
[0069]
[0070] Among them,
[0071]
[0072] In the formula: S(I t ) represents the anomaly score; V represents the stack of maximum prediction errors at different scales; v i represents the maximum prediction error at scale i; N represents the total number of scales included in the error pyramid.
[0073] Compared with the prior art, the video anomaly detection method based on spatio-temporal enhanced associative memory in the present invention has the following beneficial effects:
[0074] In the present invention, an anomaly detection model after training is used to detect anomaly frames from the input video frame sequence. After the anomaly detection model generates the fused features, the relationships between normal event prototypes are retrieved through the fused features, and the fused features are adjusted through the relationships between normal event prototypes. However, the relationships between normal event prototypes include the prototype content of normal events and the mutual relationships between the prototype contents. That is to say, the present invention not only records and learns the prototype content of normal events, but also records and learns the relationships between prototype contents containing higher-order information and higher-level semantics. Furthermore, the rich relationships between prototype memories can be used to more accurately predict normal frames, thereby improving the performance of video anomaly detection.
[0075] Based on adjusting the fused features by combining the relationships between normal event prototypes, the anomaly detection model of the present invention uses two encoders to respectively extract the spatio-temporal features (i.e., appearance features and motion features) of the video frame sequence and its optical flow sequence. Furthermore, by imposing global motion context constraints on the appearance, the internal connection of spatio-temporal features can be fully utilized, that is, the motion features can be used to enhance the appearance features to achieve spatio-temporal semantic enhancement, so that the generated fused features can more accurately predict normal frames, thereby further improving the effectiveness of video anomaly detection. Finally, the effectiveness of the method proposed in the present invention is verified through a large number of experiments on three existing datasets. Description of the Drawings
[0076] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the drawings, where:
[0077] Figure 1 is the logic block diagram of the video anomaly detection method based on spatio-temporal enhanced associative memory;
[0078] Figure 2 is the network structure diagram of the anomaly detection model;
[0079] Figure 3 Network structure diagram of the associative memory module;
[0080] Figure 4 Comparison of the frame-level ROC curves on the UCSD ped2 and CHUK Avenue datasets;
[0081] Figure 5 Partial normality scores of the method of the present invention on the CHUK Avenue and ShanghaiTech datasets. Detailed implementation manners
[0082] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the figures herein can be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present invention provided herein is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0083] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the figures, or the orientation or positional relationship in which the product of the invention is customarily placed. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. In addition, terms such as "horizontal" and "vertical" do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined. In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "install", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0084] The following is a more detailed description through specific embodiments:
[0085] Embodiment:
[0086] This embodiment discloses a video anomaly detection method based on spatio-temporal enhanced associative memory.
[0087] As Figure 1 shown, the video anomaly detection method based on spatio-temporal enhanced associative memory includes:
[0088] S1: Obtain the video frame sequence to be detected;
[0089] S2: Input the video frame sequence into a trained anomaly detection model (also known as an associative memory model with spatio-temporal enhancement: Associative Memory with Spatio-Temporal Enhancement, AMSTE) to output the corresponding anomaly prediction value;
[0090] In this embodiment, a threshold is set: when the abnormal prediction value exceeds the threshold, it is determined that there is an abnormal frame in the input video frame sequence; otherwise, it is determined that there is no abnormal frame in the input video frame sequence.
[0091] S3: Use the abnormal prediction value output by the abnormal detection model as the abnormal detection result of the video frame sequence to be detected;
[0092] Combined with Figure 2 As shown, the abnormal detection model is trained through the following steps:
[0093] S201: Obtain training data including a training video frame sequence and abnormal ground truth, and generate an optical flow sequence of the training video frame sequence; in this embodiment, the video frame sequence is converted into a corresponding optical flow sequence by existing means, and the conversion method disclosed in the paper "Ilg E, Mayer N, Saikia T, et al. FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks" can be specifically referred to.
[0094] S202: Input the training video frame sequence and its optical flow sequence into the abnormal detection model;
[0095] S203: Extract the appearance features and motion features of the training video frame sequence and its optical flow sequence through an appearance encoder and a motion encoder;
[0096] S204: Use the motion features to fuse and enhance the appearance features through a Spatio-Temporal Enhancement Module (STEM) to obtain fused features;
[0097] S205: Perform associative retrieval based on the fused features through an associative memory module to obtain the relationships between normal event prototypes;
[0098] In this embodiment, the prototype content of normal events and the relationships between prototype contents are stored through an existing memory network. The prototype content refers to the embedding vector of the normal event prototype; the mutual relationship between prototype contents refers to the relationship between normal event prototypes, that is, calculating the relationships between these embedding vectors. Among them, the relationship between normal event prototypes refers to the relationship between the prototype contents of normal events.
[0099] S206: Adjust the fused features according to the relationships between normal event prototypes to generate final features;
[0100] S207: Decode the final features through a decoder to obtain corresponding abnormal prediction values;
[0101] S208: Calculate the model loss based on the anomaly prediction value and the corresponding true anomaly value, and optimize the model parameters;
[0102] S209: Repeat steps S201 to S208 until the anomaly detection model converges.
[0103] The anomaly detection model of the present invention detects anomaly frames from the input video frame sequence after training. After the anomaly detection model generates the fused features, it retrieves the relationships between the normal event prototypes through the fused features, and adjusts the fused features based on the relationships between the normal event prototypes. However, the relationships between the normal event prototypes include the prototype content of the normal events and the mutual relationships between the prototype contents. That is to say, the present invention not only records and learns the prototype content of the normal events, but also records and learns the relationships between the prototype contents containing higher-order information and higher-level semantics. Furthermore, it can use the rich relationships between the prototype memories to more accurately predict normal frames, thereby improving the performance of video anomaly detection.
[0104] Based on adjusting the fused features by combining the relationships between the normal event prototypes, the anomaly detection model of the present invention uses two encoders to respectively extract the spatio-temporal features (i.e., appearance features and motion features) of the video frame sequence and its optical flow sequence. Furthermore, by imposing global motion context constraints on the appearance, it can make full use of the internal connections of the spatio-temporal features, that is, it can use the motion features to enhance the appearance features to achieve spatio-temporal semantic enhancement, so that the generated fused features can more accurately predict normal frames, thereby further improving the effectiveness of video anomaly detection. Finally, the effectiveness of the method proposed in the present invention is verified through a large number of experiments on three existing data sets.
[0105] In the specific implementation process, since the video sequence contains spatial information and temporal information, which cannot be fully captured only by appearance information. Therefore, we use two encoders to extract spatial information (appearance features) and temporal information (motion features) respectively, and use a decoder to reconstruct the next frame. We use the improved U-Net (from O. Ronneberger, P. Fischer, and T. Brox, "U-net: Convolutional networks for biomedical image segmentation") network as the encoder and decoder, and the two encoder structures are the same. Different from the original U-Net, we change the pooling to strided convolution in the encoder. At the same time, we add a deformable convolutional network (DCN) (from X. Zhu, H. Hu, S. Lin, and J. Dai, "Deformable convnets v2: More deformable, better results") before the deconvolution in the decoder. This kind of convolution can make the shape of the convolutional kernel closer to the feature by adding direction vectors.
[0106] Specifically, the appearance encoder, motion encoder, and decoder are all improved U-Net networks; the improvement of the appearance encoder and motion encoder is that the pooling layer of the U-Net network is changed to a strided convolution layer; the improvement of the decoder is that a deformable convolutional network layer is added before the U-Net network performs deconvolution.
[0107] In the present invention, the appearance encoder, motion encoder, and decoder are all improved U-Net networks. Among them, in the appearance encoder and motion encoder, the pooling layer of the U-Net network is changed to a strided convolution layer, and in the decoder, a deformable convolutional network layer is added before the U-Net network performs deconvolution. By replacing the pooling layer with a strided convolution layer, the operation mode and characteristics of the network can be changed to a certain extent. The strided convolution layer uses a convolution operation to replace the pooling operation for downsampling, enabling the network to learn the process of feature extraction and downsampling simultaneously, thereby improving the feature representation ability. At the same time, the newly added deformable convolutional network can make the shape of the convolutional kernel closer to the feature by adding direction vectors, realizing the generation of offsets using features aggregated from high-level features and low-level features, which helps the features after convolution to imitate the different sizes of shapes appearing in the features, thereby improving the accuracy of video anomaly detection. In addition, the present invention uses skip connections connecting the appearance encoder and the decoder to provide multi-scale and multi-level information for video frame prediction.
[0108] In the specific implementation process, there is a strong correlation between the appearance and motion of normal events in the video, and anomalies are usually more related to moving objects. Therefore, we propose to use the global attention map of optical flow to weight the appearance features to enhance the appearance features of normal events. To obtain the global attention map of optical flow, we use 1×1 convolution to implement global context modeling. Since abnormal events usually involve fast motion, we design a variance attention module to highlight fast-moving objects in the video. Among them, the spatio-temporal enhancement module weights the appearance features based on the motion variance attention map of the motion features to enhance the appearance features through the motion features.
[0109] The processing steps of the spatio-temporal enhancement module are as follows:
[0110] S2041: Take the appearance features and motion features as the inputs of the spatio-temporal enhancement module;
[0111] S2042: Input the motion features into a 1×1 convolutional layer and a softmax layer to obtain the global context features of optical flow
[0112] S2043: Input the global context features of optical flow into the variance attention module (composed of variance and softmax) to generate a motion variance attention map
[0113] Among them, the variance attention module calculates and generates the motion variance attention map through the following formula:
[0114]
[0115] In the formula: att v represents the motion variance attention map; softmax represents the softmax function operation; f m represents the global context features of optical flow after global context modeling; b represents the batch size; d represents the channel dimension; D = h×w represents the number of spatial dimensions.
[0116] S2044: After transforming the appearance features (from to ), perform a Hadamard product operation (i.e., matrix multiplication) with the motion variance attention map to generate motion-enhanced appearance features (with a dimension of );
[0117] S2045: After transforming the motion-enhanced appearance features (from to ) After that, a 1×1 convolutional layer (the same as the 1×1 convolutional layer in GCNet), a LayerNorm function layer, a ReLU function layer, and a 1×1 convolutional layer are sequentially input, and then added to the appearance features to obtain the fused features.
[0118] It should be noted that the existing solution proposes motion attention to clearly represent the interconnections between crowd movements, while the STEM proposed in the present invention uses motion features to enhance appearance features. Specifically, the motion attention module in the existing solution uses the magnitude and angle features of optical flow as inputs, and obtains the attention values for each pixel in the frame through multiplication and softmax operations. In contrast, the STEM of the present invention uses appearance features and optical flow features as inputs, and adopts convolution and variance attention to impose global motion context constraints on the appearance.
[0119] The spatio-temporal enhancement module in the present invention makes full use of the internal connection of spatio-temporal features by imposing global motion context constraints on the appearance, that is, it can use motion features to enhance appearance features to achieve spatio-temporal semantic enhancement, so that the generated fused features can more accurately predict normal frames.
[0120] In the specific implementation process, the associative memory module consists of a content memory module and a relational memory module The content memory module and the relational memory module are respectively used to record and learn the prototype content of normal events and their mutual relationships. Here, n is the number of memory modules, and d is the dimension of the memory modules. The relational memory is calculated from the outer product of the content memory, which is one order higher than the content, and thus has more advanced and rich semantic features. The input of the associative memory module is where b represents the batch size and d is the number of channels. The detailed update and read processes of the associative memory module are as Figure 3 shown.
[0121] In this embodiment, the associative memory module records the relationships between normal event prototypes through a memory network. In the field of video anomaly detection, it is a common method to use a Memory Network to record normal event prototypes. The memory network can learn the features of normal events without supervision, so as to handle various different types of abnormal situations.
[0122] A memory network is a neural network model that has a memory unit for storing and retrieving information. In video anomaly detection, a memory network can be used to memorize the features of normal events. Specifically, by training on a large number of normal videos, the memory network can learn the feature representations of normal events. Then, when a new video is input into the system, the memory network can compare the video with the prototype of normal events learned previously. If the input video is similar to the prototype of normal events, it is considered normal; if the input video does not match the prototype of normal events, it is considered potentially abnormal. By using the memory network to record the prototype of normal events, the accuracy of video anomaly detection can be improved.
[0123] Among them, the prototype content refers to the embedded vector of the prototype of normal events; the mutual relationship between prototype contents refers to the relationship between the prototypes of normal events, that is, calculating the relationship between these embedded vectors. The retrieved prototype specifically refers to the mutual relationship between prototype contents. Because the order of the mutual relationship between prototype contents is higher than the order of prototype contents (the meaning of the order in first-order and second-order in advanced mathematics), the semantic meaning contained is more advanced, and at the same time it also contains the information of prototype contents. Therefore, the prototype relationship can be retrieved and output.
[0124] The processing steps of the associative memory module are as follows:
[0125] S2051: Use the fused features as the input of the associative memory module;
[0126] S2052: Read the content memory module and the relationship memory module of the previous state; different from using cosine similarity to retrieve appropriate memory module memory items in the prior art, the present invention distributes the input data to each memory item.
[0127] The formula is described as:
[0128]
[0129] X 1 = softmax(f 3 (x) T M r-1 f 2 (x));
[0130] In the formula: X 1 represents the content memory module and the relationship memory module of the previous state read; x represents the fused features; softmax represents the softmax function operation; f 1 , f 2 , f 3 represents the linear layer; represents the outer product;
[0131] S2053: Update the content memory module and the relationship memory module of the current state according to the content memory module and the relationship memory module of the previous state read; the relationship memory module and the content memory module, like other networks, have gradients and can be updated through training. Only by continuously updating the network through training can the prototype content and prototype content relationships stored in it better represent normal events.
[0132] Since M r records the relationships between the content prototypes of normal events in M i , we use SAM to process the reading results of the associative memory module and X 1 , and update the relationship memory module M r by adding the processed results to the reading results.
[0133] The formula description is as follows:
[0134]
[0135]
[0136] where
[0137] In the formula: M r , M i represent the updated content memory module and relationship memory module; M r-1 represents the content memory module and relationship memory module of the previous state; f 4 represents the linear layer; SAM (Self-attentive Associative Memory, SAM) represents the self-attentive associative memory operation.
[0138] In this embodiment, the self-attentive associative memory operation SAM adopted is an existing operation module, which comes from the SAM module in the paper (S. Zhang, M. Gong, Y. Xie, A. K. Qin, H. Li, Y. Gao, and Y.-S. Ong, “Influence-aware attention networks for anomaly detection in surveillance videos”).
[0139] Specifically, the specific calculation process of SAM is as follows:
[0140] M m = LN(W m M m ), m = q, k, v;
[0141]
[0142]
[0143] Where: W m represents the parameter weight; M represents the memory entry, and the dimension is specifically refers to ) m represents the variables used to represent q, k, v; q, k, v represent three different branches, and these three branches have the same calculation form, so m is used for formula combination; represents the defined outer product operation. k i , v i , represents M k , M v in one entry; M q , M k , M v represents the expression M m = LN(W m M), m = the output result of q, k, v; LN represents the linear normalization operator; ⊙ represents element-wise multiplication; represents the tanh activation function.
[0144] In SAM, each batch corresponds to one memory. The dimension of the relational memory module is actually b × n × d × d, and the dimension of the content memory module is actually b × d × d, where b is the batch size, d is the memory dimension, and n is the number of memory entries. In M r conversion, SAM uses a feed-forward neural network architecture, while AMM further expands it by combining content-addressable memory and cosine similarity, which enables AMM to aggregate multiple normal event prototypes with appropriate weights.
[0145] S2054: Retrieve the updated relational memory module to obtain the relationships between normal event prototypes. Continuously updating the relational memory module can make the relationships stored in the relational memory more representative of the relationships between normal event prototypes. The relational memory module is essentially a network with gradients, and the purpose of updating it is the same as updating other neural networks.
[0146] The associative memory module of the present invention retrieves the relationships between normal event prototypes by fusing features and adjusts the fused features through the relationships between normal event prototypes. However, the relationships between normal event prototypes include the prototype content of normal events and the mutual relationships between the prototype contents. That is to say, the present invention not only records and learns the prototype content of normal events but also records and learns the relationships between prototype contents containing higher-order information and more advanced semantics. Furthermore, it can use the rich relationships between prototype memories to more accurately predict normal frames, thereby improving the performance of video anomaly detection.
[0147] In the specific implementation process, since higher-level semantic information is recorded in M r , we convert it into a content addressable memory (CAM) and address the CAM in the same way as (D. Gong, L. Liu, V. Le, B. Saha, M. R. Mansour, S. Venkatesh, and A. v. d. Hengel, “Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection”).
[0148] Specifically, the final features are generated through the following steps:
[0149] S2061: Input the relationship between the retrieved normal event prototypes into the forward propagation network f to obtain the content addressable memory M( which can be regarded as n entries, and each entry is a vector);
[0150] The formula is described as:
[0151] M = f(M r );
[0152] In the formula: M represents the content addressable memory; M r represents the relationship between normal event prototypes; f represents the forward propagation network;
[0153] S2062: Calculate the cosine similarity between the content addressable memory and the input fusion feature , and perform a softmax operation on the cosine similarity to obtain the similarity weights between the fusion feature and each entry in the content addressable memory M (the dimension of this weight is );
[0154] In this embodiment, the fusion feature is dimensionally transformed into (which can be regarded as bhw entries, and each entry is a vector).
[0155] The formula is described as:
[0156]
[0157]
[0158] Where: w i represents the similarity weight; m i , m j represent entries in the content-addressable memory M; x represents the fused feature; d(x, m i ), d(x, m j ) represent the cosine similarity between the fused feature and the entries m i and m j ;
[0159] S2063: Multiply the similarity weight (with dimension ) and the entries of the content-addressable memory (with dimension ) to obtain the corresponding prototype feature;
[0160] In this embodiment, the result of the multiplication operation is After dimension transformation, the final output
[0161] The formula description is:
[0162]
[0163] Where: represents the prototype feature; w represents the similarity weight matrix composed of the similarity weights w i ;
[0164] S2064: Concatenate the fused feature and the prototype feature in dimension to obtain the final feature.
[0165] In this embodiment, the fused feature passes through the associative memory module to output the prototype feature Concatenate the two in dimension to obtain the final feature as the input to the decoder.
[0166] The present invention converts the relationship between the retrieved normal event prototypes into prototype features for dimension concatenation with the fused feature, enabling the adjustment of the fused feature through the relationship between the normal event prototypes, and thus being able to use the rich relationships between the prototype memories to more accurately predict normal frames, thereby improving the performance of video anomaly detection.
[0167] Specifically in the implementation process, the model of the present invention is trained with the reconstruction loss L rec as the objective function, that is, minimizing the l norm between the predicted frame t and its corresponding real frame I 2 . Calculate the model loss through the following formula:
[0168]
[0169] Where: L rec represents the model loss; I t represents the ground truth of the video frame anomaly; represents the predicted value of the video frame anomaly.
[0170] The present invention calculates the model loss through the above formula, and then optimizes the parameters of the anomaly detection model according to the model loss, so as to ensure the training effect and performance of the anomaly detection model.
[0171] In the test stage, the present invention uses an evaluation strategy to score the anomaly situation. The present invention calculates the Peak Signal-to-Noise Ratio (PSNR) between the predicted frame and its ground truth, that is, evaluates the performance of the anomaly detection model through anomaly scoring. The anomaly score of the anomaly detection model is calculated by the following formula:
[0172]
[0173] Wherein,
[0174]
[0175] Where: S(I t ) represents the anomaly score; V represents the stack of maximum prediction errors at different scales; v i represents the maximum prediction error of scale i; N represents the total number of scales included in the error pyramid. In this embodiment, the error maps of each resolution are stacked into a pyramid structure.
[0176] The present invention calculates the anomaly score of the anomaly detection model through the above formula to evaluate the model performance, so as to ensure the training effect and performance of the anomaly detection model.
[0177] To better illustrate the advantages of the technical solution of the present invention, the following experiment is disclosed in this embodiment.
[0178] 1. Experimental setup
[0179] Datasets: Experiments were conducted on three benchmark datasets. 1) UCSD Ped2 (from V. Mahadevan, W. Li, V. Bhalodia, and N. Vasconcelos, “Anomaly detection in crowded scenes”) contains 16 training videos and 12 test videos, including 12 abnormal events. 2) CUHK Avenue (from C. Lu, J. Shi, and J. Jia, “Abnormal event detection at 150fps in matlab”) consists of 16 training videos and 21 test videos, including 47 irregular activities. 3) ShanghaiTech (from W. Luo, W. Liu, and S. Gao, “A revisit of sparse coding based anomaly detection in stacked rnn framework”) contains 330 training videos and 107 test videos, covering 130 abnormal events.
[0180] Evaluation Metrics: The area under the frame-level ROC curve (Area Under the Curve, AUC) was used as the main metric. The higher the AUC score, the better the performance.
[0181] Implementation Details: The resolution of all frames was scaled to 256×256, and their pixel intensities were normalized to [-1,1]. The number of memory bars n and the dimension d were set to 30 and 512 respectively. We used the Adam optimizer (from D.P. Kingma and J. Ba, “Adam: A method for stochastic optimization”) with β 1 = 0.9 and β 2 = 0.9, and the initial learning rate was set to 5e-5 by the cosine annealing method (from I. Loshchilov and F. Hutter, “Sgdr: Stochastic gradient descent with warm restarts”). The number of training epochs for UCSD Ped2, CUHK Avenue, and ShanghaiTech were set to 60, 60, and 10 respectively. In the test phase, we stacked the error maps with resolutions of 32×32, 64×64, and 128×128 into a pyramid structure. All models were implemented using PyTorch on an NVIDIA RTX 3060.
[0182] 2. Comparison with State-of-the-Art Methods
[0183] In this experiment, the proposed method was compared with existing state-of-the-art methods on the UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets. As shown in Table 1, the model proposed in the present invention was the most prominent on Ped2 and Avenue, with frame AUCs of 98.38% and 88.74% respectively, demonstrating the effectiveness of the model.
[0184] Figure 4 The ROC curves shown also indicate that our method performs better than other methods on Ped2 and Avenue. Although the AUC of the proposed method on ShanghaiTech reached 74.22%, which is lower than 75.70% of DSM-Net, DSM-Net has greater computational requirements, meaning that DSM-Net requires higher hardware resources. Overall, the method of the present invention achieved good detection performance on the three benchmark datasets.
[0185] In Table 1: Future Frame Pred is from (W. Liu, W. Luo, D. Lian, and S. Gao, “Future frame prediction for anomaly detection – a new baseline”).
[0186] MemAE is from (D. Gong, L. Liu, V. Le, B. Saha, M. R. Mansour, S. Venkatesh, and A. v. d. Hengel, “Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection”)
[0187] AnomalyNet is from (J. T. Zhou, J. Du, H. Zhu, X. Peng, Y. Liu, and R. S. M. Goh, “Anomaly net: An anomaly detection network for video surveillance”).
[0188] Appearance-motion is from (T.-N. Nguyen and J. Meunier, “Anomaly detection in video sequence with appearance-motion correspondence”).
[0189] AnoPCN is from (M. Ye, X. Peng, W. Gan, W. Wu, and Y. Qiao, “Anopcn: Video anomaly detection via deep predictive coding network”).
[0190] MNAD-P is from (H. Park, J. Noh, and B. Ham, “Learning memory-guided normality for anomaly detection”).
[0191] AMMC-Net is from (R. Cai, H. Zhang, W. Liu, S. Gao, and Z. Hao, “Appearance-motion memory consistency network for video anomaly detection”).
[0192] Context-related is from (D. Li, X. Nie, X. Li, Y. Zhang, and Y. Yin, “Context-related video anomaly detection via generative adversarial network”).
[0193] STD is from (A. Guo, L. Guo, R. Zhang, Y. Wang, and S. Gao, “Self-trained prediction model and novel anomaly score mechanism for video anomaly detection”).
[0194] STCEN is from (Y. Hao, J. Li, N. Wang, X. Wang, and X. Gao, “Spatiotemporal consistency-enhanced network for video anomaly detection”)
[0195] SIGnet is from (Z. Fang, J. Liang, J. T. Zhou, Y. Xiao, and F. Yang, “Anomaly detection with bidirectional consistency in videos”).
[0196] DESDnet is from (Y. Zhong, X. Chen, J. Jiang, and F. Ren, “Reverse erasure guided spatio-temporal autoencoder with compact feature representation for video anomaly detection”).
[0197] DSM-Net is from (Z. Wang and Y. Chen, “Anomaly detection with dual-stream memory network”).
[0198] Table 1 Frame-level AUC comparison (%) of different methods on UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets
[0199]
[0200] Note: The best and the second-best results are represented in bold and underlined respectively.
[0201] 3. Qualitative analysis
[0202] The normality score curve of the method proposed by the present invention is as Figure 5 shown, and some representative test samples are provided. Obviously, when an abnormal event occurs or the event changes from abnormal to normal, the normality score curve will decline or rise. For example, in the UCSD ped2 dataset, when a normal pedestrian has a cycling event, the curve of the normal score will decline sharply. In addition, during the cycling stage, the normal score remains at a low level. When the event changes from cycling to a normal pedestrian, the abnormal score rises sharply, and the normal score remains at a high level during the normal event stage. These qualitative results further prove the effectiveness of the method proposed by the present invention.
[0203] 4. Ablation experiment
[0204] In this experiment, detailed ablation experiments were carried out on UCSD ped2 and CHUK Avenue, as shown in Table 2. We used the improved U-Net with a multi-scale anomaly assessment strategy as the baseline, as shown in the first row of Table 2. The AUC of this baseline was only 95.41% on ped2 and only 84.59% on Avenue. After adding AMM, STEM, and DCN, performance improvements of 1.74%, 1.47%, and 1.19% were obtained on UCSD ped2, and performance improvements of 2.64%, 1.93%, and 2.49% were obtained on Avenue. The performance of any combination of AMM, SCEM, and DCN was also higher than that of a single module. By adopting the proposed method, the AUC values on the UCSD ped2 and Avenue datasets were increased by 2.97% and 4.15% respectively, verifying the effectiveness of the proposed method.
[0205] Table 2 Ablation Study Based on UCSD ped2 and CHUK Avenue
[0206]
[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the technical solutions. Those of ordinary skill in the art should understand that any modifications or equivalent replacements made to the technical solutions of the present invention without departing from the purpose and scope of the present technical solution shall be covered by the scope of the claims of the present invention.
Claims
1. A video anomaly detection method based on spatio-temporal enhanced associative memory, characterized in that, it includes: S1: Obtain the video frame sequence to be detected; S2: Input the video frame sequence into the trained anomaly detection model, and output the corresponding anomaly prediction value; S3: Use the anomaly prediction value output by the anomaly detection model as the anomaly detection result of the video frame sequence to be detected; In step S2, the anomaly detection model is trained through the following steps: S201: Obtain the training data including the training video frame sequence and the anomaly ground truth, and generate the optical flow sequence of the training video frame sequence; S202: Input the training video frame sequence and its optical flow sequence into the anomaly detection model; S203: Extract the appearance features and motion features of the training video frame sequence and its optical flow sequence through the appearance encoder and the motion encoder; In step S203, the appearance encoder, the motion encoder and the decoder are all improved U-Net networks; The improvement of the appearance encoder and the motion encoder is: changing the pooling layer of the U-Net network into a stride convolutional layer; the improvement of the decoder is: adding a deformable convolutional network layer before the deconvolution of the U-Net network; S204: Use the spatio-temporal enhancement module to fuse and enhance the appearance features by using the motion features to obtain the fused features; In step S204, the spatio-temporal enhancement module weights the appearance features based on the motion variance attention map of the motion features to achieve enhancing the appearance features by the motion features; The processing steps of the spatio-temporal enhancement module are as follows: S2041: Use the appearance features and the motion features as the input of the spatio-temporal enhancement module; S2042: Input the motion features into a 1×1 convolutional layer and a softmax layer to obtain the optical flow global context features; S2043: Input the optical flow global context features into the variance attention module to generate the motion variance attention map; S2044: Perform a Hadamard product operation on the appearance features and the motion variance attention map to generate the motion-enhanced appearance features; S2045: After sequentially inputting the motion-enhanced appearance features into a 1×1 convolutional layer, a LayerNorm function layer, a ReLU function layer and a 1×1 convolutional layer, add them to the appearance features to obtain the fused features; The variance attention module calculates and generates the motion variance attention map through the following formula: Where: att v represents the motion variance attention map; f m represents the optical flow global context feature; d represents the size of the channel; D represents the number of spatial dimensions; S205: Perform associative retrieval based on the fused features through the associative memory module to obtain the relationship between normal event prototypes; S206: Adjust the fused features according to the relationship between normal event prototypes to generate the final features; S207: Decode the final features through the decoder to obtain the corresponding anomaly prediction value; S208: Calculate the model loss according to the anomaly prediction value and the corresponding anomaly ground truth, and optimize the model parameters; S209: Repeat steps S201 to S208 until the anomaly detection model converges.
2. The video anomaly detection method based on spatio-temporal enhanced associative memory according to claim 1, characterized in that, In step S205, the associative memory module records the relationships between normal event prototypes through a memory network; among them, there are a content memory module and a relationship memory module respectively used to record the prototypes of normal events and the relationships between prototypes.
3. The video anomaly detection method based on spatio-temporal enhanced associative memory according to claim 2, wherein, the processing steps of the associative memory module are as follows: S2051: Take the fused feature as the input of the associative memory module; S2052: Read the content memory module and the relationship memory module of the previous state; The formula description is: X 1 = softmax(f 3 (x) T M r-1 f 2 (x)); Wherein: X 1 represents the content memory module and the relationship memory module of the previous state; x represents the fused feature; f 1 , f 2 , f 3 represents a linear layer; represents an outer product; S2053: Update the content memory module and the relationship memory module of the current state according to the read content memory module and relationship memory module of the previous state; The formula description is: Among them, Where: M r and M i represent the updated content memory module and relationship memory module; M r-1 represents the content memory module and relationship memory module in the previous state; f 4 represents a linear layer; SAM represents the self-attention associative memory operation; S2054: Retrieve the updated relationship memory module to obtain the relationships between normal event prototypes.
4. The video anomaly detection method based on spatio-temporal enhanced associative memory according to claim 1, wherein, in step S206, the final feature is generated through the following steps: S2061: Input the relationships between normal event prototypes into the forward propagation network to obtain a content addressable memory; The formula description is: M = f(M r ); Where: M represents a content addressable memory; M r represents the relationship between normal event prototypes; f represents a forward propagation network; S2062: Calculate the cosine similarity between the content addressable memory and the input fused feature, and perform a softmax operation on the cosine similarity to obtain the similarity weights between the fused feature and each entry in the content addressable memory; The formula description is: where: w i represents the similarity weight; m i , m j represent entries in the content-addressable memory M; x represents the fused feature; d(x, m i ) and d(x, m j ) represent the cosine similarity between the fused feature and entries m i and m j ; S2063: Perform a multiplication operation on the similarity weights and the entries of the content addressable memory to obtain the corresponding prototype features; The formula description is: In the formula: represents the prototype feature; w represents the similarity weight matrix composed of the similarity weights w i ; S2064: Concatenate the fused feature and the prototype feature in the dimension to obtain the final feature.
5. The video anomaly detection method based on spatio-temporal enhanced associative memory according to claim 1, wherein, calculate the model loss through the following formula: Where: L rec represents the model loss; I t represents the ground truth of the abnormality of the video frame; represents the predicted value of the abnormality of the video frame.
6. The video anomaly detection method based on spatio-temporal enhanced associative memory according to claim 5, wherein, evaluate the performance of the anomaly detection model through an anomaly score; calculate the anomaly score of the anomaly detection model through the following formula: Among them, where: s(I t ) represents the anomaly score; V represents the stack of maximum prediction errors at different scales; v i represents the maximum prediction error at scale i; N represents the total number of scales included in the error pyramid.