Video anomaly detection method fusing double-flow network and memory enhancement module
By integrating the dual-stream network and the memory enhancement module, the video anomaly detection method solves the problem of ignoring motion information in the existing technology, realizes the effective fusion of appearance and motion features in video sequences, and improves the accuracy of detection.
Patent Information
- Application Number
- CN202510746408.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-12
AI Technical Summary
Existing video anomaly detection methods often ignore motion information when detecting events with continuous motion changes, resulting in inaccurate detection results.
The dual-stream network and memory enhancement module are integrated to extract the appearance features and motion features of the video sequence respectively through a dual encoder architecture, and a variety of normal features are learned in the memory enhancement module, combined with the attention decoder for prediction.
Improved the accuracy of video anomaly detection, enabling more accurate detection of events with continuous motion changes.
Smart Images

Figure CN120635774A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image or video recognition methods, and relates to a video anomaly detection method integrating a dual-stream network and a memory enhancement module. Background Art
[0002] Video Anomaly Detection (VAD) aims to detect unexpected behavior in surveillance videos. With the widespread use of surveillance cameras in public places, massive amounts of surveillance data have been generated. Manually processing this data consumes considerable time and effort, necessitating the development of intelligent video anomaly detection. Because the probability of abnormal events is low and the definition of abnormality varies across scenarios, most video anomaly detection methods employ unsupervised approaches. These methods train the model using normal data, allowing the model to learn only the characteristics of normal events. During the testing phase, events that do not conform to the model's description are defined as abnormal.
[0003] Early research in video anomaly detection relied on manually designed features to extract normal and abnormal events, classify abnormal events, and detect them. However, these traditional methods relied on the quality of the manually extracted features, resulting in poor detection performance. With the development of deep learning, most current methods use frame reconstruction or prediction to learn the normality of normal features and then identify anomalies based on the reconstruction or prediction error. However, these methods assume that abnormal frames are not well reconstructed or predicted. Deep generative models only learn an "identity function" rather than selectively learning normal features of the video. Therefore, even if an abnormal frame is input to the model, it simply regenerates the frame instead of refining it to match the learned normal features. This high reconstruction performance leads to a certain degree of reconstruction of abnormal events, a problem known as "overgeneralization." Videos are considered a special type of time series due to their spatial and temporal characteristics, where the spatial and temporal dimensions correspond to appearance and motion information, respectively. However, some work only considers appearance information and ignores motion information, which can lead to inaccurate detection results for events with continuous motion changes. Summary of the Invention
[0004] The purpose of the present invention is to provide a video anomaly detection method that integrates a dual-stream network and a memory enhancement module, which can model the appearance features and motion features in video sequences, improve the accuracy of video anomaly detection, and solve the problem that the existing technology only focuses on appearance information and ignores motion information, resulting in inaccurate detection results when detecting some events with continuous motion changes.
[0005] The technical solution adopted by the present invention is a video anomaly detection method integrating a dual-stream network and a memory enhancement module, which is implemented according to the following steps: Step 1: Obtain a video dataset, divide the video dataset into a training set and a test set, and preprocess the video data in the training set and the test set; Step 2: Build a network model that integrates the dual-stream network and the memory enhancement module; Step 3: Construct a loss function and iteratively train the network model that integrates the dual-stream network and the memory enhancement module to obtain a video frame prediction model. Step 4: After preprocessing, the video to be detected is input into the video frame prediction model to obtain a predicted frame, and video anomaly detection is performed by calculating the error between the predicted frame and the real frame.
[0006] The present invention is also characterized in that: In step 1, the training set contains only normal video samples, and the test set contains normal video samples and abnormal video samples; The specific preprocessing of video data is as follows: Each video in the training set and the test set is divided into f consecutive video frames, and we get { ,…, , }, arrive Respectively represent the first to the second segment of the corresponding video f video frames; then use FlowNet2 optical flow network to extract f The optical flow information corresponding to the video frames is obtained f- 1 corresponding optical flow { ,…, , };.
[0007] The network model that integrates the dual-stream network and the memory enhancement module includes an appearance encoder and a motion encoder set in parallel. The outputs of the appearance encoder and the motion encoder are jointly connected to a variance-attention-based appearance-motion fusion module, which is in turn connected to a memory enhancement module and an attention-based single decoder.
[0008] The appearance encoder and the motion encoder have the same structure, including four groups of combination modules. A maximum pooling layer is set between two connected combination modules. Each group of combination modules includes a convolution layer, a normalization layer, and an activation function connected in sequence. The input of the maximum pooling layer is the output of the activation function of the corresponding previous combination module, and the output of the maximum pooling layer is the input of the convolution layer of the corresponding next combination module. The inputs of the appearance encoder and the motion encoder are { ,…, , }and{ ,…, , }.
[0009] The input of the variance attention-based appearance-motion fusion module is the appearance features of the appearance encoder and the motion encoder. and motion characteristics , =1,2,3,4; The variance-attention-based appearance-motion fusion module first uses Convolution performs global context modeling on motion features to obtain motion features after global context modeling; Then use the features after global context modeling to generate variance attention :
[0010] in, represents the motion features after global context modeling, t=1,2,3,4, b is the batch size, represents the number of spatial dimensions, 、 are the height and width of the feature respectively, and softmax is the softmax function; Then the calculated variance attention is used to weight the appearance features, and then 1×1 convolution and addition operations are used to aggregate the appearance global context features of the features at each position to obtain the aggregated features. :
[0011] in, represents the appearance features extracted by the appearance encoder, 、 and They represent 1×1 convolution processing, normalization processing, and activation function respectively.
[0012] The memory enhancement module consists of two parts: reading and updating. First, the high-dimensional fusion features obtained by the appearance-motion fusion module based on variance attention are aggregated. Classified as queries , ; Among them, the query features are memorized and enhanced by the memory items in the storage module. The storage module has Items are used to record different features, Represents the memory item of the storage module, ; The reading process is: First calculate each query and all memory items The cosine similarity between them is used to measure the similarity between the two, and then the corresponding weight is calculated. :
[0013] in, represents the natural exponential function; Then the memory item Perform weighted summation to obtain aggregate features :
[0014] Finally, the aggregation item With query items The read features are obtained by splicing along the channel dimension :
[0015] Among them, Concat is a concatenation operation; Will indivual Features are recombined to obtain the output of the memory module , by reading all memory items in the storage module, different normal features are learned; The update process is: Update memory item Select the query with the highest cosine similarity Update, for Memory items, select Largest query ,Will Add the index set corresponding to the query ; Calculate each memory item and all queries The cosine similarity between them is calculated, and then the corresponding weight is calculated. :
[0016] Then use the index collection To update the memory item and renormalize the weights:
[0017] Finally, the queries in the index set are weighted and summed with the normalized weights to update the memory items in the storage module.
[0018] Based on the attention single decoder using deconvolution operation Upsample and integrate features of other scales 、 and Jump connection to the decoder, and finally get the predicted frame ; The decoder includes a deconvolution operation module, an ECA channel attention module, a deconvolution operation module, an ECA channel attention module, a deconvolution operation module, an ECA channel attention module, and an output module. The deconvolution operation module includes two groups of combination modules connected in sequence, a deconvolution layer, a normalization layer, and an activation function; the output module includes two groups of combination modules connected in sequence, a convolution layer, and an activation function; As the input of the first group of deconvolution operation modules, after being processed by the first group of deconvolution operation modules, the first group of deconvolution operation modules are connected to the first group of deconvolution operation modules by jump. Perform splicing operation, and then process it through the first ECA channel attention module to obtain the feature ; feature As the input of the second group of deconvolution operation modules, after being processed by the second group of deconvolution operation modules, the Perform splicing operation, and then process it through the second ECA channel attention module to obtain the feature ; feature As the input of the third group of deconvolution operation modules, after being processed by the third group of deconvolution operation modules, it is connected to the third group of deconvolution operation modules by jump. Perform splicing operation, and then process it through the second ECA channel attention module to obtain the feature ; As the input of the output module, the output is processed by the output module to obtain the predicted frame Specifically:
[0019]
[0020]
[0021]
[0022] in, is the total operation function of the deconvolution operation module, is the total operation function of the output module, ECA is the channel attention operation, and Concat is the splicing operation.
[0023] The process of training the network model that integrates the two-stream network and the memory enhancement module is as follows: Use the loss function L as the loss function to train the model, and input { ,…, , } is used as input to predict the next frame, which is the predicted frame ; The loss function L is specifically:
[0024] in, To predict losses, is the optical flow loss, is the feature compactness loss, is the feature separation loss, and is the weight parameter corresponding to feature compactness loss and feature separation loss
[0025]
[0026]
[0027]
[0028] Where, Used to measure the predicted frame Corresponding real frame The similarity between them, optical flow loss Used to measure the predicted frame Estimated playing field With real frame Estimated playing field the degree of similarity between them; represents the FlowNet2 optical flow network, For real frame The previous frame, is margin; Is a query The index of the most recent item, defined as:
[0029] in, and yes The nearest and second nearest terms of Represents a query The index of the next most recent item: .
[0030] During the model training process, the memory enhancement module updates the memory items in each training process; During the model testing process, the memory enhancement module updates the memory items according to the regularization score. To decide, when Above threshold , the corresponding video frame is considered to be an abnormal frame, and the memory item is not updated, otherwise the memory item is updated; Calculated according to the following formula:
[0031] in, , Spatial index, for:
[0032] in, Represents the predicted frames and real frame The pixel value at the (x,y) coordinate point.
[0033] In step 4, the pre-processing of the video to be detected is specifically as follows: the video to be detected is divided into f Continuous video frames; then use FlowNet2 network to extract f The video frames correspond to f -1 optical flow information, f consecutive video frames and f -1 optical flow information is input into the video frame prediction model to obtain the predicted frame; The specific method for video anomaly detection is to calculate the error between the predicted frame and the real frame: The prediction error map is obtained by calculating the difference between the predicted frame and the real frame, and then it is scaled to three different scales to form a pyramid architecture. The maximum prediction error at each scale is calculated by mean pooling. Calculate the peak signal-to-noise ratio between the predicted frame and the true frame :
[0034] Where, represent function, Representation scale i The maximum prediction error in yes The real frame of the moment, yes The predicted frame at time instant; Then, normalization is performed to obtain the range value and smooth the anomaly score through a Gaussian filter :
[0035] The probability of anomaly occurrence is determined based on the anomaly score. The higher it is, the greater the probability of an abnormality occurring.
[0036] The beneficial effects of the present invention are: The present invention first extracts the appearance features and motion features of video sequences respectively through a dual encoder architecture, and fuses the two types of features at the same scale. Secondly, the high-dimensional fused features are fed into a memory enhancement module with an update strategy to learn diverse normal features. Finally, a skip connection mechanism is used to feed the multi-scale fused features and memory-enhanced features into an attention-based decoder to predict future frames. This fully combines the appearance features with the motion features, improves the model's ability to learn normal features, and thus effectively detects abnormal events in videos. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is the overall network architecture of the video anomaly detection method integrating the dual-stream network and the memory enhancement module of the present invention; Figure 2 This is a workflow diagram of the memory enhancement module in the video anomaly detection method integrating the dual-stream network and the memory enhancement module of the present invention; Figure 3a This is a visualization diagram of the AUC results of the video anomaly detection method of the present invention that integrates the dual-stream network and the memory enhancement module on the UCSD ped2 dataset; Figure 3b This is a visualization diagram of the AUC results of the video anomaly detection method that integrates the dual-stream network and the memory enhancement module on the CUHK Avenue dataset.
[0038] Figure 3c This is a visualization diagram of the AUC results of the video anomaly detection method that integrates the dual-stream network and the memory enhancement module on the ShanghaiTech dataset.
[0039] Figure 4a This is the error visualization graph of the video anomaly detection method of the present invention that integrates the dual-stream network and the memory enhancement module on the UCSD ped2 dataset; Figure 4b This is an error visualization diagram of the video anomaly detection method integrating the dual-stream network and the memory enhancement module on the CUHK Avenue dataset; Figure 4c This is an error visualization diagram of the video anomaly detection method based on the fusion of the dual-stream network and the memory enhancement module on the ShanghaiTech dataset. DETAILED DESCRIPTION
[0040] The following describes it in detail with reference to specific implementation methods.
[0041] Example 1 The video anomaly detection method of the present invention, which integrates a dual-stream network and a memory enhancement module, is implemented according to the following steps: Step 1: Obtain a video dataset, divide the video dataset into a training set and a test set, and preprocess the video data in the training set and the test set; Step 2: Build a network model that integrates the dual-stream network and the memory enhancement module; Step 3: Construct a loss function and iteratively train the network model that integrates the dual-stream network and the memory enhancement module to obtain a video frame prediction model. Step 4: After preprocessing, the video to be detected is input into the video frame prediction model to obtain a predicted frame, and video anomaly detection is performed by calculating the error between the predicted frame and the real frame.
[0042] Example 2 On the basis of Example 1, in step 1, the training set only includes normal video samples, and the test set includes normal video samples and abnormal video samples; The specific preprocessing of video data is as follows: Each video in the training set and the test set is divided into f consecutive video frames, and we get { ,…, , }, as the input of the appearance encoder, arrive Respectively represent the first to the second segment of the corresponding video f video frames; then use FlowNet2 optical flow network to extract f The optical flow information corresponding to each video frame is obtained because one optical flow can be extracted from every two consecutive video frames. f- 1 corresponding optical flow { ,…, , }, as the input of the motion encoder.
[0043] Example 3 Based on Example 2, the network model that integrates the dual-stream network and the memory enhancement module includes an appearance encoder and a motion encoder arranged in parallel, and the outputs of the appearance encoder and the motion encoder are jointly connected to a variance-attention-based appearance-motion fusion module, which is in turn connected to a memory enhancement module and an attention-based single decoder.
[0044] The appearance encoder and motion encoder have the same structure, as shown in Figure 1As shown, it includes four groups of combination modules, and a maximum pooling layer MaxPool is set between two connected combination modules. Each group of combination modules includes a convolution layer Conv, a normalization layer BN, and an activation function Tanh connected in sequence. The input of the maximum pooling layer is the output of the activation function of the corresponding previous combination module, and the output of the maximum pooling layer is the input of the convolution layer of the corresponding next combination module. The inputs of the appearance encoder and the motion encoder are { ,…, , }and{ ,…, , }.
[0045] The input of the variance attention-based appearance-motion fusion module is the appearance features of the appearance encoder and the motion encoder. and motion characteristics , =1,2,3,4; The appearance-motion fusion module structure based on variance attention is as follows: Figure 1 As shown, first use Convolution performs global context modeling on motion features to obtain motion features after global context modeling. Since abnormal events often involve fast motion, variance attention is introduced to highlight fast-moving objects in the video, and variance attention is generated for the features after global context modeling. :
[0046] in, represents the motion features after global context modeling, t=1,2,3,4, b is the batch size, represents the number of spatial dimensions, 、 are the height and width of the feature respectively, and softmax is the softmax function; Then the calculated variance attention is used to weight the appearance features, and then 1×1 convolution and addition operations are used to aggregate the appearance global context features of the features at each position to obtain the aggregated features. :
[0047] in, represents the appearance features extracted by the appearance encoder, 、 and They represent 1×1 convolution processing, normalization processing, and activation function respectively.
[0048] The memory enhancement module structure is as follows Figure 2As shown, it includes two parts: reading and updating. First, the high-dimensional fusion feature obtained by aggregating the appearance-motion fusion module based on variance attention is Classified as queries , ; Among them, the query features are memorized and enhanced by the memory items in the storage module. The storage module has Items are used to record different features, Represents the memory item of the storage module, ; The reading process structure is as follows Figure 2 (a), specifically: First calculate each query and all memory items The cosine similarity between them is used to measure the similarity between the two, and then the corresponding weight is calculated. :
[0049] in, represents the natural exponential function; Then the memory item Perform weighted summation to obtain aggregate features :
[0050] Finally, the aggregation item With query items The read features are obtained by splicing along the channel dimension :
[0051] Among them, Concat is a concatenation operation; Will indivual Features are recombined to obtain the output of the memory module , by reading all memory items in the storage module, different normal features are learned; In order to enable the model to better learn different features, an update strategy is introduced in the memory enhancement module to continuously update the memory items in the storage module: The update process is as follows Figure 2 (b), specifically: Update memory item Select the query with the highest cosine similarity Update, for Memory items, select Largest query ,Will Add the index set corresponding to the query ,like Figure 2(c) Calculate each memory item and all queries The cosine similarity between them is calculated, and then the corresponding weight is calculated. :
[0052] Then use the index collection To update the memory item and renormalize the weights:
[0053] Finally, the queries in the index set are weighted and summed with the normalized weights to update the memory items in the storage module.
[0054] Based on the attention single decoder using deconvolution operation Upsample and integrate features of other scales 、 and Jump connection to the decoder, and finally get the predicted frame ; The decoder includes a deconvolution operation module, an ECA channel attention module, a deconvolution operation module, an ECA channel attention module, a deconvolution operation module, an ECA channel attention module, and an output module. The deconvolution operation module includes two groups of combination modules connected in sequence, a deconvolution layer ConvTranspose, a normalization layer BN, and an activation function Tanh. The output module includes two groups of combination modules connected in sequence, a convolution layer Conv, and an activation function Tanh. As the input of the first group of deconvolution operation modules, after being processed by the first group of deconvolution operation modules, the first group of deconvolution operation modules are connected to the first group of deconvolution operation modules by jump. Perform splicing operation, and then process it through the first ECA channel attention module to obtain the feature ; feature As the input of the second group of deconvolution operation modules, after being processed by the second group of deconvolution operation modules, the Perform splicing operation, and then process it through the second ECA channel attention module to obtain the feature ; feature As the input of the third group of deconvolution operation modules, after being processed by the third group of deconvolution operation modules, it is connected to the third group of deconvolution operation modules by jump. Perform splicing operation, and then process it through the second ECA channel attention module to obtain the feature ; As the input of the output module, the output is processed by the output module to obtain the predicted frame Specifically:
[0055]
[0056]
[0057]
[0058] in, is the total operation function of the deconvolution operation module, is the total operation function of the output module, ECA is the channel attention operation, and Concat is the splicing operation.
[0059] The multiple fused features in the decoder are not sampled from the same feature map, but are multiple independent features. The fused features at each scale fully integrate the appearance information and motion information.
[0060] Example 4 Based on Example 3, the process of training the network model integrating the dual-stream network and the memory enhancement module is as follows: using the loss function L as the loss function to train the model, ,…, , } is used as input to predict the next frame, which is the predicted frame ; The loss function L is specifically:
[0061] in, To predict losses, is the optical flow loss, is the feature compactness loss, It is a feature separation loss. The feature compactness loss forces the query to be close to the memory item with the highest similarity in the storage module, reducing the intra-class difference; the feature separation loss forces the memory items in the storage module to be sufficiently different.
[0062] and are the weight parameters corresponding to feature compactness loss and feature separation loss;
[0063]
[0064]
[0065]
[0066] Where, Used to measure the predicted frame Corresponding real frame The similarity between them, optical flow loss Used to measure the predicted frame Estimated playing field With real frame Estimated playing field the degree of similarity between them; represents the FlowNet2 optical flow network, For real frame The previous frame, is margin; Is a query The index of the most recent item, defined as:
[0067] in, and yes The nearest and second nearest terms of Represents a query The index of the next most recent item: .
[0068] Example 5 On the basis of Example 4, since the normal features of the training set and the test set are different, and the test set contains abnormal frames, this will cause the model to also record abnormal features well, so a regular score is introduced To solve this problem; During the model training process, the memory enhancement module updates the memory items in each training process; During the model testing process, the memory enhancement module updates the memory items according to the regularization score. To decide, when Above threshold , the corresponding video frame is considered to be an abnormal frame, and the memory item is not updated, otherwise the memory item is updated; Calculated according to the following formula:
[0069] in, , Spatial index, for:
[0070] in, Represents the predicted frames and real frame The pixel value at the (x,y) coordinate point.
[0071] Example 6 On the basis of Example 5, the pre-processing of the video to be detected in step 4 is specifically as follows: the video to be detected is divided into f Continuous video frames; then use FlowNet2 network to extract f The video frames correspond to f -1 optical flow information, f consecutive video frames and f -1 optical flow information is input into the video frame prediction model to obtain the predicted frame; The specific method for video anomaly detection is to calculate the error between the predicted frame and the real frame: The prediction error map is obtained by calculating the difference between the predicted frame and the real frame, and then it is scaled to three different scales to form a pyramid architecture. The maximum prediction error at each scale is calculated by mean pooling. Calculate the peak signal-to-noise ratio between the predicted frame and the true frame :
[0072] Where, represent function, Representation scale i The maximum prediction error in yes The real frame of the moment, yes The predicted frame at time instant; Then, normalization is performed to obtain the range value and smooth the anomaly score through a Gaussian filter :
[0073] The probability of anomaly occurrence is determined based on the anomaly score. The higher it is, the greater the probability of an abnormality occurring.
[0074] Example 7 Based on Example 6, simulation experiments were conducted on a single RTX 4090 GPU using the PyTorch deep learning framework. The results were evaluated using the Area Under the Curve (AUC) metric. The experiments were conducted on three unsupervised video anomaly detection datasets: UCSD Ped2, CUHK Avenue, and Shanghaitech. All datasets consisted of training and test sets. The training set contained only normal events, while the test set contained both normal and abnormal events.
[0075] This simulation experiment compared the detection method of the present invention with several currently advanced video anomaly detection methods: MNAD (Memory-guided Normality for Anomaly Detection), ASTNET (Attention-based Spatio-Temporal Network), DEDDnet (Decomposition-driven Network), PDM-Net (Prototype-guided Dynamics Matching Network), and MGAN-CL (Multi-branch Generative Adversarial Network with Context Learning). The results are shown in Table 1. The experimental results show that the present invention can significantly improve detection accuracy compared with other methods.
[0076] Table 1 Comparison of AUC values of the present invention with other methods
[0077] Figure 3a-3cThe present invention is evaluated on the UCSD Ped2, CUHK Avenue and ShanghaiTech datasets. Three methods, MemAE (Memory-augmented Autoencoder), Frame-Pred (Frame Prediction) and AMMC (Appearance-Motion Memory Consistency Network), are selected from the reconstruction-based, single-stream prediction-based and dual-stream prediction-based methods respectively. The ROC curve is drawn using the results of the reproduced results and the results of the present invention. The larger the AUC, the better the detection result. Figure 3a-3c As can be seen in Figure 2, the AUC of the method of the present invention is higher than that of the other methods on the three data sets. Therefore, the method of the present invention has higher detection accuracy.
[0078] Figure 4 further shows the error visualization results for the UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets. For each dataset, the first row shows the actual video frame, the second row shows the predicted frame generated by the model, and the third row shows the error between the predicted and actual frames. The first column of each dataset visualization shows the prediction results under normal conditions, where the prediction error is low. The red boxes mark the areas where abnormal events occur.
[0079] like Figure 4a As shown in the figure, the UCSD Ped2 dataset is shot in a single scene with a fixed camera position. There are fewer types of abnormal events, and the model can accurately detect abnormalities. Figure 4b As shown in the figure, the abnormal area of the CUHK Avenue dataset gradually increases from left to right. In the face of such multi-scale abnormal changes, the model can still accurately detect the abnormal area in each frame, so the model has the ability to detect multi-scale anomalies. Figure 4c As shown in the figure, the ShanghaiTech dataset has multiple scenes with different lighting conditions and many types of abnormal events. However, even in complex environments, the model can still detect abnormal events relatively accurately, indicating that the model has good robustness in handling complex scenes and multiple abnormal types of tasks.
Claims
1. A video anomaly detection method integrating a dual-stream network and a memory enhancement module, characterized in that: Follow these steps to implement: Step 1: Obtain a video dataset, divide the video dataset into a training set and a test set, and preprocess the video data in the training set and the test set; Step 2: Build a network model that integrates the dual-stream network and the memory enhancement module; Step 3: Construct a loss function and iteratively train the network model that integrates the dual-stream network and the memory enhancement module to obtain a video frame prediction model. Step 4: After preprocessing, the video to be detected is input into the video frame prediction model to obtain a predicted frame, and video anomaly detection is performed by calculating the error between the predicted frame and the real frame.
2. The video anomaly detection method integrating a dual-stream network and a memory enhancement module according to claim 1 is characterized in that: In step 1, the training set only contains normal video samples, and the test set contains normal video samples and abnormal video samples; The specific preprocessing of video data is as follows: Each video in the training set and the test set is divided into f consecutive video frames, and we get { ,…, , }, arrive Respectively represent the first to the second segment of the corresponding video f video frames; then use FlowNet2 optical flow network to extract f The optical flow information corresponding to the video frames is obtained f- 1 corresponding optical flow { ,…, , };.
3. The video anomaly detection method integrating a dual-stream network and a memory enhancement module according to claim 2 is characterized in that: The network model that integrates the dual-stream network and the memory enhancement module includes an appearance encoder and a motion encoder arranged in parallel, the outputs of the appearance encoder and the motion encoder are jointly connected to a variance-attention-based appearance-motion fusion module, and the variance-attention-based appearance-motion fusion module is sequentially connected to a memory enhancement module and an attention-based single decoder.
4. The video anomaly detection method integrating a dual-stream network and a memory enhancement module according to claim 3 is characterized in that: The appearance encoder and the motion encoder have the same structure, including four groups of combination modules, and a maximum pooling layer is provided between two connected combination modules. Each group of combination modules includes a convolution layer, a normalization layer, and an activation function connected in sequence. The input of the maximum pooling layer is the output of the activation function of the corresponding previous combination module, and the output of the maximum pooling layer is the input of the convolution layer of the corresponding next combination module. The inputs of the appearance encoder and the motion encoder are { ,…, , }and{ ,…, , }.
5. The video anomaly detection method integrating a dual-stream network and a memory enhancement module according to claim 4 is characterized in that: The input of the variance attention-based appearance-motion fusion module is the appearance features of the appearance encoder and the motion encoder. and motion characteristics , =1,2,3,4; The variance attention-based appearance-motion fusion module first uses Convolution performs global context modeling on motion features to obtain motion features after global context modeling; Then use the features after global context modeling to generate variance attention : in, represents the motion features after global context modeling, t=1,2,3,4, b is the batch size, represents the number of spatial dimensions, 、 are the height and width of the feature respectively, and softmax is the softmax function; Then the calculated variance attention is used to weight the appearance features, and then 1×1 convolution and addition operations are used to aggregate the appearance global context features of the features at each position to obtain the aggregated features. : in, represents the appearance features extracted by the appearance encoder, 、 and They represent 1×1 convolution processing, normalization processing, and activation function respectively.
6. The video anomaly detection method integrating a dual-stream network and a memory enhancement module according to claim 5 is characterized in that: The memory enhancement module consists of two parts: reading and updating. First, the high-dimensional fusion features obtained by the appearance-motion fusion module based on variance attention are aggregated. Classified as queries , ; Among them, the query features are memorized and enhanced by the memory items in the storage module. The storage module has Items are used to record different features, Represents the memory item of the storage module, ; The reading process is: First calculate each query and all memory items The cosine similarity between them is used to measure the similarity between the two, and then the corresponding weight is calculated. : in, represents the natural exponential function; Then the memory item Perform weighted summation to obtain aggregate features : Finally, the aggregation item With query items The read features are obtained by splicing along the channel dimension : Among them, Concat is a concatenation operation; Will indivual Features are recombined to obtain the output of the memory module , by reading all memory items in the storage module, different normal features are learned; The update process is: Update memory item Select the query with the highest cosine similarity Update, for Memory items, select Largest query ,Will Add the index set corresponding to the query ; Calculate each memory item and all queries The cosine similarity between them is calculated, and then the corresponding weight is calculated. : Then use the index collection To update the memory item and renormalize the weights: Finally, the queries in the index set are weighted and summed with the normalized weights to update the memory items in the storage module.
7. The video anomaly detection method integrating a dual-stream network and a memory enhancement module according to claim 6 is characterized in that: The attention-based single decoder uses deconvolution operations to Upsample and integrate features of other scales 、 and Jump connection to the decoder, and finally get the predicted frame ; The decoder includes a deconvolution operation module, an ECA channel attention module, a deconvolution operation module, an ECA channel attention module, a deconvolution operation module, an ECA channel attention module, and an output module in sequence. The deconvolution operation module includes two groups of combination modules connected in sequence, a deconvolution layer, a normalization layer, and an activation function; the output module includes two groups of combination modules and a convolution layer, and an activation function connected in sequence; As the input of the first group of deconvolution operation modules, after being processed by the first group of deconvolution operation modules, the first group of deconvolution operation modules are connected to the first group of deconvolution operation modules by jump. Perform splicing operation, and then process it through the first ECA channel attention module to obtain the feature ; feature As the input of the second group of deconvolution operation modules, after being processed by the second group of deconvolution operation modules, the Perform splicing operation, and then process it through the second ECA channel attention module to obtain the feature ; feature As the input of the third group of deconvolution operation modules, after being processed by the third group of deconvolution operation modules, it is connected to the third group of deconvolution operation modules by jump. Perform splicing operation, and then process it through the second ECA channel attention module to obtain the feature ; As the input of the output module, the output is processed by the output module to obtain the predicted frame Specifically: in, is the total operation function of the deconvolution operation module, is the total operation function of the output module, ECA is the channel attention operation, and Concat is the splicing operation.
8. The video anomaly detection method integrating a dual-stream network and a memory enhancement module according to claim 7 is characterized in that: The process of training the network model that integrates the dual-stream network and the memory enhancement module is as follows: Use the loss function L as the loss function to train the model, and input { ,…, , } is used as input to predict the next frame, which is the predicted frame ; The loss function L is specifically: in, To predict losses, is the optical flow loss, is the feature compactness loss, is the feature separation loss, and is the weight parameter corresponding to feature compactness loss and feature separation loss Where, Used to measure the predicted frame Corresponding real frame The similarity between them, optical flow loss Used to measure the predicted frame Estimated playing field With real frame Estimated playing field the degree of similarity between them; represents the FlowNet2 optical flow network, For real frame The previous frame, is margin; Is a query The index of the most recent item, defined as: in, and yes The nearest and second nearest terms of Represents a query The index of the next most recent item: 。 9. The video anomaly detection method integrating a dual-stream network and a memory enhancement module according to claim 8 is characterized in that: During the model training process, the memory enhancement module updates the memory items in each training process; During the model testing process, the memory enhancement module updates the memory items according to the regularization score. To decide, when Above threshold , the corresponding video frame is considered to be an abnormal frame, and the memory item is not updated, otherwise the memory item is updated; Calculated according to the following formula: in, , Spatial index, for: in, Represents the predicted frames and real frame The pixel value at the (x,y) coordinate point.
10. The video anomaly detection method integrating a dual-stream network and a memory enhancement module according to claim 9 is characterized in that: In step 4, the video to be detected is pre-processed as follows: Divide the video to be detected into f Continuous video frames; then use FlowNet2 network to extract f The video frames correspond to f -1 optical flow information, f consecutive video frames and f -1 optical flow information is input into the video frame prediction model to obtain the predicted frame; The specific method for video anomaly detection is to calculate the error between the predicted frame and the real frame: The prediction error map is obtained by calculating the difference between the predicted frame and the real frame, and then it is scaled to three different scales to form a pyramid architecture. The maximum prediction error at each scale is calculated by mean pooling. Calculate the peak signal-to-noise ratio between the predicted frame and the true frame : Where, represent function, Representation scale i The maximum prediction error in yes The real frame of the moment, yes The predicted frame at time instant; Then, normalization is performed to obtain the range value and smooth the anomaly score through a Gaussian filter : Determine the probability of anomaly occurrence based on the anomaly score. The higher it is, the greater the probability of an abnormality occurring.