Video anomaly detection method based on mixed attention and meta-prototype network
By adopting a hybrid attention and meta-prototype network method in video surveillance, extracting the spatial characteristics of video frames and learning normal dynamic prototypes, the problem of abnormal behavior detection in video surveillance is solved, and higher detection accuracy and rapid adaptation of new scenarios is achieved.
Patent Information
- Application Number
- CN202411846376.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-09
AI Technical Summary
Abnormal behavior detection faces many challenges in video surveillance, including the frequency of normal events being higher than abnormal events, uncertainty of abnormal events, messy environment and dynamic occlusion of people, etc., which leads to the difficulty of identifying and locating abnormal events in the prior art.
Using a video anomaly detection method based on mixed attention and meta-prototype network, spatial features are extracted frame by frame through the Inception encoder, and channel attention and spatial attention mechanism are introduced during the encoding process to reduce the weight of background feature areas and increase the attention of the target area. At the same time, the meta-prototype network is used to learn normal dynamic prototypes to realize the prediction and abnormal detection of video frames.
This method can effectively reduce interference in background feature areas, improve the accuracy of video abnormality detection and the adaptability of new scenes, and significantly improve the ability to identify abnormal parts in video frames.
Smart Images

Figure CN119964045A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video abnormal behavior detection, and in particular relates to a video anomaly detection method based on hybrid attention and meta-prototype network. Background Art
[0002] Abnormal activity detection (VAD) in video surveillance is a hot topic in the field of computer vision, such as traffic accident detection, criminal activity identification, and illegal behavior monitoring. However, it is challenging to detect abnormal activities in a large number of normal scenes. First, it is very difficult to collect and annotate all types of abnormal events because normal events occur much more frequently than abnormal events and are relatively rare. Second, the uncertainty of abnormal events is also a challenge. For example, a certain activity may be considered abnormal in some cases, but normal in other cases. Furthermore, in crowded public scenes, it is also difficult to identify abnormal events due to the cluttered surrounding environment and the interference of dynamic occlusion of the crowd. In addition, manual viewing and analysis of large amounts of surveillance videos is time-consuming and inefficient. Automatic anomaly detection systems are particularly important in analyzing and identifying abnormal events in surveillance videos. These systems can significantly improve surveillance efficiency, reduce manpower requirements, and ensure timely detection of potential security threats.
[0003] Since surveillance videos are more easily available under normal circumstances, unsupervised frameworks trained only with normal samples have attracted widespread attention. Typically, unsupervised anomaly detection models are trained with normal video data. During the inference process, samples that cannot be explained by the normal model are considered anomalies. Such judgments are usually based on reconstruction errors or prediction errors. A common unsupervised approach is to use clustering techniques such as Gaussian mixture models (GMM) or hierarchical clustering to group normal samples and then detect anomalies based on whether the test sample fits these clusters. Traditional clustering algorithms require the number of clusters to be set in advance or a threshold to determine whether a sample belongs to a cluster. Large reconstruction errors that exceed the preset threshold are considered anomalies. Another anomaly detection technique relies on sparse representation, which reconstructs test samples by using a dictionary built from normal samples and identifies those samples whose reconstruction errors exceed the preset threshold as anomalies.
[0004] However, in practical applications, when using memory items to reconstruct normal samples, the error often tends to increase, which increases the difficulty of distinguishing normal from abnormal situations. In addition, the memory items are designed for specific video frame scenarios, which brings several limitations. First, memory-based autoencoders are only applicable to scenarios equipped with specific memory modules. Once the scene changes, its anomaly detection performance will drop significantly. Therefore, a dedicated memory module needs to be customized for each different scenario. Second, in large-scale surveillance environments, perspective distortion and the motion distortion it causes are more significant, which may cause the motion patterns of distant people to be misjudged as abnormal, while normal patterns nearby can be successfully reconstructed by memory items. This position sensitivity problem limits the ability of memory-based autoencoders to accurately locate anomalies. Summary of the invention
[0005] The purpose of the present invention is to provide a video anomaly detection method based on hybrid attention and meta-prototype network, which encodes normal dynamics into prototypes in real time and reduces the interference of background feature areas, realizes the rapid adaptation capability of new scenes, and improves the accuracy of video anomaly detection.
[0006] The technical solution adopted by the present invention is a video anomaly detection method based on hybrid attention and meta-prototype network, comprising the following steps: Step 1: Given a video frame sequence, build an Inception-based encoder network to extract the spatial features of the frame frame by frame and retain the features of each layer; Step 2: The spatial features extracted from each layer are input into the channel attention; Step 3: Input the final features into the meta-prototype network to learn the prototype for encoding normal dynamics; Step 4: Implement the spatial attention of the jump connection between the encoder network and the decoder network, increase the weight of the feature area representing the target and fuse the upsampled feature map of each layer with the encoder feature map of the corresponding channel; Step 5: Predict future frames by extracting features and explicitly modeling normal dynamics in the video sequence; Step 6: Construct the loss function of the prediction model; Step 7: Iteratively train the model and obtain a model that can predict frames close to the real ones; Step 8: Perform video anomaly detection by calculating the error between the predicted frame and the real frame.
[0007] The present invention is also characterized in that: The specific process of step 2 is as follows: a channel attention mechanism is added during the encoding process to reduce the weight of the channel representing the background feature, which increases the attention to the target during the training process. The size of the feature map input to the SE channel attention is determined by Indicates that is the number of channels, and are the height and width of the matrix respectively; the global average pooling layer is used to compress the spatial dimension of the feature map, and the output matrix It is expressed as shown in formula (A): (A); Among them, x belongs to Represents the row index, y belongs to Represents the column index, z belongs to represents the channel index; the dimension of the matrix SE is ; To feed the feature map into the fully connected (FC) layer, it is first resized to , and then the ReLU activation function is applied; after that, the feature map passes through the FC layer again and the size is changed back , and finally obtain the weight matrix through the Sigmoid function ; In this process, channel mapping The weight is expressed as , as shown in formula (B): (B); The result of the channel attention processing is to assign a smaller weight value to the background channel in the input feature map, and the weight value is less than 1.
[0008] The specific process of step 3 is: Step 3.1: Build a dynamic prototype set using highly differentiable attention techniques. Step 3.2: Reconstruct the standard encoding by searching for these prototypes; Step 3.3: Fuse the input encoding with the obtained standard encoding to produce the output.
[0009] The specific process of step 3.1 is as follows: the video frame is preprocessed by the Inception encoder to obtain t feature encoding maps , which is considered as a c-dimensional vector, ; In the application of the attention mechanism, M attention mapping functions are used is the encoding vector Assign normal weights; at each pixel location, the normal weight is used to measure the standard deviation or normal distribution range of the encoding vector; in, Represents the mth feature map, and obtains a prototype with a collection of encoding vectors , the integration process normalized normal weights, as shown in formula (C): (C); Similarly, by applying multiple attention functions, M prototypes are generated, which together constitute a prototype pool. ; During prototype retrieval, map the preprocessed encoding to the input encoding Make a query to retrieve the relevant items in the prototype pool and reconstruct the encoding The process is shown in formula (D): (D); in, Represents the Nth encoding vector The similarity score with the mth prototype item; through the channel summation operation, the obtained normal map is combined with the original code H as the final output result of the model; then, the code output by the DPU passes through the remaining layers of the autoencoder to continue to perform subsequent frame prediction tasks.
[0010] The specific process of step 4 is: add a spatial attention mechanism of skip connection between the encoder and decoder to reduce the weight of the background feature area and increase the focus on the target area during training; and the encoded feature map with the same number of channels Input to the spatial attention module, represented as , Expressed as the number of channels of the feature matrix, Expressed as the height of the feature matrix, Represented as the width of the feature matrix; the feature map and Input to Convolution performs linear transformation, and then inputs it into the activation function ReLU for nonlinear transformation; the response matrix output by the activation function , as shown in formula (E): (E); in, Indicates deviation, and represents the slope of the linear transformation, Input to The two-dimensional matrix compresses the number of channels to 1, recorded as As shown in formula (F): (F); in, Represents the slope of the linear transformation, which is then passed through the Sigmoid function to obtain the mixed attention coefficient , weighted to the feature map As shown in formula (G): (G); Under the processing of the spatial attention mechanism, the areas representing the background in the encoded feature map are given weights less than 1 for use in the decoding process.
[0011] The loss function of the prediction model in step 6 is shown in formula (H); (H); in is the balance parameter, is the frame prediction loss, is the frame reconstruction loss.
[0012] The specific process of step 7 is: the loss of frame prediction is calculated by calculating the true value With network prediction value The distance between them is defined as follows: (I); The feature reconstruction loss is designed to give the learned normal prototypes the characteristics of density and diversity. It uses two losses and Optimize these two properties separately, as shown in formula (J): (J); in, is the weight parameter, For dense reconstruction of normal coding, it measures the average L2 distance between the input encoding vector and its most relevant prototype, as shown in formula (K): (K); in, Reconstructing the normal encoding for diversity further promotes the diversity between prototype items by pushing the learned prototypes away from each other, as shown in formula (L): (L); By controlling the required margin between prototypes using γ, we exploit the advantages of the above two aspects and encourage prototype items to encode compact and diverse normal dynamics for realistic frame prediction.
[0013] The specific process of step 8 is as follows: The core of the anomaly detection mechanism is to generate corresponding anomaly scores by accurately evaluating the quality of the predicted frames, and then make decisions based on these scores; in order to quantify the quality of the predicted frames, the peak signal-to-noise ratio PSNR standard is adopted, and its specific calculation formula is shown in (M); (M); in, Indicates the maximum value of the image color. is the original image without noise, K is The noise approximation of In the calculation results of PSNR, a higher value often means that the quality of the predicted frame is better, and its characteristics are closer to the distribution of normal frames; on the contrary, a lower PSNR value usually indicates that the quality of the predicted frame is poor, which may hide some abnormality or deviation from the normal state; In order to convert these PSNR-based predicted frame quality scores into more comparable and practical anomaly scores, the maximum and minimum normalization method is used to convert the original quality scores into the interval of [0,1]; After obtaining the anomaly scores, the final anomaly detection results are determined and output based on these scores by setting different thresholds. By flexibly adjusting the thresholds, the performance of anomaly detection can be optimized according to actual needs and application scenarios.
[0014] The beneficial effects of the present invention are: The video anomaly detection method based on hybrid attention and meta-prototype network provided by the present invention encodes the spatial features of video frames, uses the Inception module, and realizes multi-scale extraction of image features by connecting convolution layers and pooling layers of different sizes in parallel, thereby obtaining richer feature representation; hybrid attention combining channel attention and spatial attention of jump connection is used to focus on information in different regions, thereby increasing the focus on targets during training; prototypes for encoding normal dynamics are learned in the meta-prototype network, and further aggregated with the encoding graph of the automatic encoder, without providing an additional memory module, so as to achieve rapid adaptation to new scenes. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is the overall network framework diagram of the video anomaly detection method based on hybrid attention and meta-prototype network of the present invention; Figure 2 It is an architecture diagram of the Inception automatic encoder-decoder based on the hybrid attention mechanism of the present invention; Figure 3 is a schematic diagram of the Inception module of the present invention; Figure 4 is a schematic diagram of a channel attention module of the present invention; Figure 5 is a schematic diagram of a spatial attention module of the present invention; Figure 6 It is a schematic diagram of the results of the anomaly detection process of the UCSD Ped2, CUHK Avenue and ShanghaiTech data sets of the present invention. DETAILED DESCRIPTION
[0015] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0016] Example 1 The video anomaly detection method based on hybrid attention and meta-prototype network proposed in this embodiment is as follows: Figure 1 As shown, the following steps are included: Step 1: Given a video frame sequence, build an Inception-based encoder network to extract the spatial features of the frame frame by frame and retain the features of each layer; Step 2: The spatial features extracted from each layer are input into the channel attention; Step 3: Input the final features into the meta-prototype network to learn the prototype for encoding normal dynamics; Step 4: Implement the spatial attention of the jump connection between the encoder network and the decoder network, increase the weight of the feature area representing the target and fuse the upsampled feature map of each layer with the encoder feature map of the corresponding channel; Step 5: Predict future frames by extracting features and explicitly modeling normal dynamics in the video sequence; Step 6: Construct the loss function of the prediction model; Step 7: Iteratively train the model and obtain a model that can predict frames close to the real ones; Step 8: Perform video anomaly detection by calculating the error between the predicted frame and the real frame.
[0017] Example 2 The video anomaly detection method based on hybrid attention and meta-prototype network proposed in this embodiment is as follows: Figure 1 As shown, the following steps are included: Step 1: Given a video frame sequence, build an Inception-based encoder network to extract the spatial features of the frame frame by frame and retain the features of each layer; Step 2: The spatial features extracted from each layer are input into the channel attention; The specific process is as follows: a channel attention mechanism is added during the encoding process to reduce the weight of the channel representing the background feature, which increases the focus on the target during the training process. The size of the feature map input to the SE channel attention is Indicates that is the number of channels, and are the height and width of the matrix respectively; the global average pooling layer is used to compress the spatial dimension of the feature map, and the output matrix It is expressed as shown in formula (A): (A); Among them, x belongs to Represents the row index, y belongs to Represents the column index, z belongs to represents the channel index; the dimension of the matrix SE is ; To feed the feature map into the fully connected (FC) layer, it is first resized to , and then the ReLU activation function is applied; after that, the feature map passes through the FC layer again and the size is changed back , and finally obtain the weight matrix through the Sigmoid function ; In this process, channel mapping The weight is expressed as , as shown in formula (B): (B); The result of the channel attention processing is to assign a smaller weight value to the background channel in the input feature map, and the weight value is less than 1; Step 3: Input the final features into the meta-prototype network to learn the prototype for encoding normal dynamics; Step 4: Implement the spatial attention of the jump connection between the encoder network and the decoder network, increase the weight of the feature area representing the target and fuse the upsampled feature map of each layer with the encoder feature map of the corresponding channel; Step 5: Predict future frames by extracting features and explicitly modeling normal dynamics in the video sequence; Step 6: Construct the loss function of the prediction model; Step 7: Iteratively train the model and obtain a model that can predict frames close to the real ones; Step 8: Perform video anomaly detection by calculating the error between the predicted frame and the real frame.
[0018] Example 3 The video anomaly detection method based on hybrid attention and meta-prototype network proposed in this embodiment is as follows: Figure 1 As shown, the following steps are included: Step 1: Given a video frame sequence, build an Inception-based encoder network to extract the spatial features of the frame frame by frame and retain the features of each layer; Step 2: The spatial features extracted from each layer are input into the channel attention; The specific process is as follows: a channel attention mechanism is added during the encoding process to reduce the weight of the channel representing the background feature, which increases the focus on the target during the training process. The size of the feature map input to the SE channel attention is Indicates that is the number of channels, and are the height and width of the matrix respectively; the global average pooling layer is used to compress the spatial dimension of the feature map, and the output matrix It is expressed as shown in formula (A): (A); Among them, x belongs to Represents the row index, y belongs to Represents the column index, z belongs to represents the channel index; the dimension of the matrix SE is ; To feed the feature map into the fully connected (FC) layer, it is first resized to , and then the ReLU activation function is applied; after that, the feature map passes through the FC layer again and the size is changed back , and finally obtain the weight matrix through the Sigmoid function ; In this process, channel mapping The weight is expressed as , as shown in formula (B): (B); The result of the channel attention processing is to assign a smaller weight value to the background channel in the input feature map, and the weight value is less than 1; Step 3: Input the final features into the meta-prototype network to learn the prototype for encoding normal dynamics; The specific process is: Step 3.1: Build a dynamic prototype set using highly differentiable attention techniques. The specific process is as follows: After the video frame is preprocessed by the Inception encoder, t feature coding maps are obtained. , which is considered as a c-dimensional vector, ; In the application of the attention mechanism, M attention mapping functions are used is the encoding vector Assign normal weights; at each pixel location, the normal weight is used to measure the standard deviation or normal distribution range of the encoding vector; in, Represents the mth feature map, and obtains a prototype with a collection of encoding vectors , the integration process normalized normal weights, as shown in formula (C): (C); Similarly, by applying multiple attention functions, M prototypes are generated, which together constitute a prototype pool. ; During prototype retrieval, map the preprocessed encoding to the input encoding Make a query to retrieve the relevant items in the prototype pool and reconstruct the encoding The process is shown in formula (D): (D); in, Represents the Nth encoding vector The similarity score with the mth prototype item; through the channel summation operation, the normal map obtained is combined with the original code H as the final output result of the model. The core idea is to use rich normal information to enhance the encoding process of the encoder, thereby improving the prediction ability of the normal part of the video frame and weakening the prediction of the abnormal part at the same time; then, the code output by the DPU passes through the remaining layers of the autoencoder to continue to perform subsequent frame prediction tasks; Step 3.2: Reconstruct the standard encoding by searching for these prototypes; Step 3.3, fusing the input code with the obtained standard code to generate output; Step 4: Implement the spatial attention of the jump connection between the encoder network and the decoder network, increase the weight of the feature area representing the target and fuse the upsampled feature map of each layer with the encoder feature map of the corresponding channel; Step 5: Predict future frames by extracting features and explicitly modeling normal dynamics in the video sequence; Step 6: Construct the loss function of the prediction model; Step 7: Iteratively train the model and obtain a model that can predict frames close to the real ones; Step 8: Perform video anomaly detection by calculating the error between the predicted frame and the real frame.
[0019] Example 4 The video anomaly detection method based on hybrid attention and meta-prototype network proposed in this embodiment is as follows: Figure 1 As shown, the following steps are included: Step 1: Given a video frame sequence, build an Inception-based encoder network to extract the spatial features of the frame frame by frame and retain the features of each layer; Step 2: The spatial features extracted from each layer are input into the channel attention; The specific process is as follows: a channel attention mechanism is added during the encoding process to reduce the weight of the channel representing the background feature, which increases the focus on the target during the training process. The size of the feature map input to the SE channel attention is Indicates that is the number of channels, and are the height and width of the matrix respectively; the global average pooling layer is used to compress the spatial dimension of the feature map, and the output matrix It is expressed as shown in formula (A): (A); Among them, x belongs to Represents the row index, y belongs to Represents the column index, z belongs to represents the channel index; the dimension of the matrix SE is ; To feed the feature map into the fully connected (FC) layer, it is first resized to , and then the ReLU activation function is applied; after that, the feature map passes through the FC layer again and the size is changed back , and finally obtain the weight matrix through the Sigmoid function ; In this process, channel mapping The weight is expressed as , as shown in formula (B): (B); The result of the channel attention processing is to assign a smaller weight value to the background channel in the input feature map, and the weight value is less than 1; Step 3: Input the final features into the meta-prototype network to learn the prototype for encoding normal dynamics; The specific process is: Step 3.1: Build a dynamic prototype set using highly differentiable attention techniques. The specific process is as follows: After the video frame is preprocessed by the Inception encoder, t feature coding maps are obtained. , which is considered as a c-dimensional vector, ; In the application of the attention mechanism, M attention mapping functions are used is the encoding vector Assign normal weights; at each pixel location, the normal weight is used to measure the standard deviation or normal distribution range of the encoding vector; in, Represents the mth feature map, and obtains a prototype with a collection of encoding vectors , the integration process normalized normal weights, as shown in formula (C): (C); Similarly, by applying multiple attention functions, M prototypes are generated, which together constitute a prototype pool. ; During prototype retrieval, map the preprocessed encoding to the input encoding Make a query to retrieve the relevant items in the prototype pool and reconstruct the encoding The process is shown in formula (D): (D); in, Represents the Nth encoding vector The similarity score with the mth prototype item; through the channel summation operation, the normal map obtained is combined with the original code H as the final output result of the model. The core idea is to use rich normal information to enhance the encoding process of the encoder, thereby improving the prediction ability of the normal part of the video frame and weakening the prediction of the abnormal part at the same time; then, the code output by the DPU passes through the remaining layers of the autoencoder to continue to perform subsequent frame prediction tasks; Step 3.2: Reconstruct the standard encoding by searching for these prototypes; Step 3.3, fusing the input code with the obtained standard code to generate output; Step 4: Implement the spatial attention of the jump connection between the encoder network and the decoder network, increase the weight of the feature area representing the target and fuse the upsampled feature map of each layer with the encoder feature map of the corresponding channel; The specific process is as follows: add a skip-connected spatial attention mechanism between the encoder and decoder to reduce the weight of the background feature area and increase the focus on the target area during training; and the encoded feature map with the same number of channels Input to the spatial attention module, represented as , Expressed as the number of channels of the feature matrix, Expressed as the height of the feature matrix, Represented as the width of the feature matrix; the feature map and Input to Convolution performs linear transformation, and then inputs it into the activation function ReLU for nonlinear transformation; the response matrix output by the activation function , as shown in formula (E): (E); in, Indicates deviation, and represents the slope of the linear transformation, Input to The two-dimensional matrix compresses the number of channels to 1, recorded as As shown in formula (F): (F); in, Represents the slope of the linear transformation, which is then passed through the Sigmoid function to obtain the mixed attention coefficient , weighted to the feature map As shown in formula (G): (G); Under the processing of the spatial attention mechanism, the area representing the background in the encoded feature map will be given a weight less than 1 so that it can be used in the decoding process; Step 5: Predict future frames by extracting features and explicitly modeling normal dynamics in the video sequence; Step 6: Construct the loss function of the prediction model; Step 7: Iteratively train the model and obtain a model that can predict frames close to the real ones; Step 8: Perform video anomaly detection by calculating the error between the predicted frame and the real frame.
[0020] Example 5 The video anomaly detection method based on hybrid attention and meta-prototype network proposed in this embodiment is as follows: Figure 1 As shown, the following steps are included: Step 1: Given a video frame sequence, build an Inception-based encoder network to extract the spatial features of the frame frame by frame and retain the features of each layer; Step 2: The spatial features extracted from each layer are input into the channel attention; The specific process is as follows: a channel attention mechanism is added during the encoding process to reduce the weight of the channel representing the background feature, which increases the focus on the target during the training process. The size of the feature map input to the SE channel attention is Indicates that is the number of channels, and are the height and width of the matrix respectively; the global average pooling layer is used to compress the spatial dimension of the feature map, and the output matrix It is expressed as shown in formula (A): (A); Among them, x belongs to Represents the row index, y belongs to Represents the column index, z belongs to represents the channel index; the dimension of the matrix SE is ; To feed the feature map into the fully connected (FC) layer, it is first resized to , and then the ReLU activation function is applied; after that, the feature map passes through the FC layer again and the size is changed back , and finally obtain the weight matrix through the Sigmoid function ; In this process, channel mapping The weight is expressed as , as shown in formula (B): (B); The result of the channel attention processing is to assign a smaller weight value to the background channel in the input feature map, and the weight value is less than 1; Step 3: Input the final features into the meta-prototype network to learn the prototype for encoding normal dynamics; The specific process is: Step 3.1: Build a dynamic prototype set using highly differentiable attention techniques. The specific process is as follows: After the video frame is preprocessed by the Inception encoder, t feature coding maps are obtained. , which is considered as a c-dimensional vector, ; In the application of the attention mechanism, M attention mapping functions are used is the encoding vector Assign normal weights; at each pixel location, the normal weight is used to measure the standard deviation or normal distribution range of the encoding vector; in, Represents the mth feature map, and obtains a prototype with a collection of encoding vectors , the integration process normalized normal weights, as shown in formula (C): (C); Similarly, by applying multiple attention functions, M prototypes are generated, which together constitute a prototype pool. ; During prototype retrieval, map the preprocessed encoding to the input encoding Make a query to retrieve the relevant items in the prototype pool and reconstruct the encoding The process is shown in formula (D): (D); in, Represents the Nth encoding vector The similarity score with the mth prototype item; through the channel summation operation, the normal map obtained is combined with the original code H as the final output result of the model. The core idea is to use rich normal information to enhance the encoding process of the encoder, thereby improving the prediction ability of the normal part of the video frame and weakening the prediction of the abnormal part at the same time; then, the code output by the DPU passes through the remaining layers of the autoencoder to continue to perform subsequent frame prediction tasks; Step 3.2: Reconstruct the standard encoding by searching for these prototypes; Step 3.3, fusing the input code with the obtained standard code to generate output; Step 4: Implement the spatial attention of the jump connection between the encoder network and the decoder network, increase the weight of the feature area representing the target and fuse the upsampled feature map of each layer with the encoder feature map of the corresponding channel; The specific process is as follows: add a skip-connected spatial attention mechanism between the encoder and decoder to reduce the weight of the background feature area and increase the focus on the target area during training; and the encoded feature map with the same number of channels Input to the spatial attention module, represented as , Expressed as the number of channels of the feature matrix, Expressed as the height of the feature matrix, Represented as the width of the feature matrix; the feature map and Input to Convolution performs linear transformation, and then inputs it into the activation function ReLU for nonlinear transformation; the response matrix output by the activation function , as shown in formula (E): (E); in, Indicates deviation, and represents the slope of the linear transformation, Input to The two-dimensional matrix compresses the number of channels to 1, recorded as As shown in formula (F): (F); in, Represents the slope of the linear transformation, which is then passed through the Sigmoid function to obtain the mixed attention coefficient , weighted to the feature map As shown in formula (G): (G); Under the processing of the spatial attention mechanism, the area representing the background in the encoded feature map will be given a weight less than 1 so that it can be used in the decoding process; Step 5: Predict future frames by extracting features and explicitly modeling normal dynamics in the video sequence; Step 6: Construct the loss function of the prediction model; The loss function of the prediction model is shown in formula (H); (H); in is the balance parameter, is the frame prediction loss, is the frame reconstruction loss; Step 7: Iteratively train the model and obtain a model that can predict frames close to the real ones; The specific process is: the loss of frame prediction is calculated by With network prediction value The distance between them is defined as follows: (I); The feature reconstruction loss is designed to give the learned normal prototypes the characteristics of density and diversity. It uses two losses and Optimize these two properties separately, as shown in formula (J): (J); in, is the weight parameter, For dense reconstruction of normal coding, it measures the average L2 distance between the input encoding vector and its most relevant prototype, as shown in formula (K): (K); in, Reconstructing the normal encoding for diversity further promotes the diversity between prototype items by pushing the learned prototypes away from each other, as shown in formula (L): (L); Using γ to control the required margin between prototypes, we take advantage of the above two items and encourage prototype items to encode compact and diverse normal dynamics for realistic frame prediction; Step 8: Detect video anomalies by calculating the error between the predicted frame and the real frame; The specific process is as follows: The core of the anomaly detection mechanism is to generate corresponding anomaly scores by accurately evaluating the quality of the predicted frames, and then make decisions based on these scores; in order to quantify the quality of the predicted frames, the peak signal-to-noise ratio (PSNR) standard is used, and its specific calculation formula is shown in (M); (M); in, Indicates the maximum value of the image color. is the original image without noise, K is The noise approximation of In the calculation results of PSNR, a higher value often means that the quality of the predicted frame is better, and its characteristics are closer to the distribution of normal frames; on the contrary, a lower PSNR value usually indicates that the quality of the predicted frame is poor, which may hide some abnormality or deviation from the normal state; In order to convert these PSNR-based predicted frame quality scores into more comparable and practical anomaly scores, the maximum and minimum normalization method is used to convert the original quality scores into the interval of [0,1]; After obtaining the anomaly scores, the final anomaly detection results are determined and output based on these scores by setting different thresholds. By flexibly adjusting the thresholds, the performance of anomaly detection can be optimized according to actual needs and application scenarios.
[0021] Example 6 The video anomaly detection method based on hybrid attention and meta-prototype network proposed in this embodiment is used for abnormal behavior monitoring in various scenarios, such as campus monitoring, traffic monitoring, and security monitoring. Figure 1 As shown, the following steps are included: Step 1, spatial feature extraction stage: like Figure 2 As shown in Figure 1, the spatial features are extracted frame by frame based on the Inception encoder to obtain a feature map containing various texture details and appearance features, such as Figure 3 As shown, the Inception module contains convolution, convolution, Convolutional layers of different scales such as convolution are composed in parallel, which improves the encoder's ability to learn information of different scales.
[0022] Step 2: Add a channel attention mechanism during encoding: like Figure 4 As shown, the size of the feature map input to the SE channel attention is given by Represented as, where C is the number of channels, H and W are the height and width of the matrix respectively. The global average pooling layer is used to compress the spatial dimensions of the feature map, and the output matrix SE is represented as formula (A): (A) x belongs to h for row index, y belongs to w for column index, and z belongs to c for channel index. The dimensions of the matrix SE are To feed the feature map into the fully connected (FC) layer, it is first resized to , and then the ReLU activation function is applied. After that, the feature map passes through the FC layer again and the size is changed back to , and finally obtain the weight matrix through the Sigmoid function In this process, the weight of the channel map M is expressed as , expressed as shown in formula (B): (B) The result of channel attention processing is to assign a smaller weight value (less than 1) to the background channel in the input feature map.
[0023] By introducing the channel attention mechanism, the channels of the detected target feature area are enhanced and the feature channels representing the background area are reduced.
[0024] Step 3, learn to encode the normal state in the meta-prototype network: The forward propagation mechanism of DPU involves three core steps: first, using highly differentiable attention technology to build a dynamic prototype set, second, reconstructing the standard encoding by searching these prototypes, and finally fusing the input encoding with the obtained standard encoding to produce the output. These three steps can be summarized as: attention mechanism application, integration processing, and prototype retrieval.
[0025] The dynamic prototype set is constructed using highly differentiable attention technology. Specifically, the video frame is preprocessed by the Inception encoder to obtain t feature encoding maps. , consider it as c-dimensional vector, In the application of the attention mechanism, M attention mapping functions are used is the encoding vector Assign normal weights. At each pixel location, the normal weight is used to measure the standard deviation or normal distribution range of the encoding vector. Represents the mth feature map, and obtains a prototype with a collection of encoding vectors , the integration process normalized normal weights, as shown in formula (C): (C) Similarly, by applying multiple attention functions, we can generate M prototypes, which together constitute a prototype pool. During prototype retrieval, the preprocessed encoding is mapped to the input encoding Make a query to retrieve the relevant items in the prototype pool and reconstruct the encoding The process is shown in formula (D): (D) in Represents the Nth encoding vector The similarity score with the mth prototype item. Through the channel summation operation, we combine the obtained normal map with the original encoding H as the final output of the model. The core idea is to use rich normal information to enhance the encoding process of the encoder, thereby improving the prediction ability of the normal part of the video frame and weakening the prediction of the abnormal part. Subsequently, the encoding output by the DPU will pass through the remaining layers of the autoencoder to continue to perform subsequent frame prediction tasks.
[0026] Step 4: Implement spatial attention jump connection during decoding: A spatial attention mechanism with skip connections is added between the encoder and decoder to reduce the weight of background feature areas and increase the focus on target areas during training. Figure 5 As shown, the upsampled decoded feature map and the encoded feature map with the same number of channels Input to the spatial attention module, represented as , Expressed as the number of channels of the feature matrix, Expressed as the height of the feature matrix, Represents the width of the feature matrix. and Input to Convolution performs linear transformation, and then inputs it into the activation function ReLU for nonlinear transformation. The response matrix output by the activation function is As shown in formula (E): (E) in Indicates deviation, and represents the slope of the linear transformation, Input to The two-dimensional matrix compresses the number of channels to 1, recorded as , expressed as shown in formula (F): (F) Represents the slope of the linear transformation, which is then passed through the Sigmoid function to obtain the mixed attention coefficient , weighted to the feature map On the other hand, it is expressed as shown in formula (G): (G) Under the processing of the spatial attention mechanism, the areas representing the background in the encoded feature map are given weights less than 1 for use in the decoding process.
[0027] Step 5, model training phase: To train the model, the overall loss function Depend on and Composition, of which is the balance parameter, is the frame prediction loss, is the frame reconstruction loss. As shown in formula (H): (H) The loss of frame prediction is calculated by calculating the true value With network prediction value It is defined by the distance between them, as shown in formula (I): (I) The feature reconstruction loss is designed to give the learned normal prototypes the characteristics of density and diversity. It uses two losses and The two properties are optimized respectively, and the formula is shown in (J): (J) in is the weight parameter, For dense reconstruction normal coding, it measures the average L2 distance between the input encoding vector and its most relevant prototype as shown in formula (K): (K) Reconstructing the normal encoding for diversity further promotes the diversity between prototype items by pushing the learned prototypes away from each other, as shown in formula (L): (L) By controlling the required margin between prototypes using γ, we take advantage of the above two aspects and encourage prototype items to encode compact and diverse normal dynamics for normal frame prediction.
[0028] Step 6, anomaly detection phase: like Figure 6 As shown in Figure 1, the core of the anomaly detection mechanism is to generate corresponding anomaly scores by accurately evaluating the quality of the predicted frames, and then make decisions based on these scores. In order to quantify the quality of the predicted frames, we use the peak signal-to-noise ratio (PSNR), a widely recognized standard, and its specific calculation formula is shown in (M).
[0029] (M) Indicates the maximum value of the image color. is the original image without noise, K is The noise approximation.
[0030] In the calculation results of PSNR, a higher value often means that the quality of the predicted frame is better and its characteristics are closer to the distribution of normal frames; on the contrary, a lower PSNR value usually indicates that the quality of the predicted frame is poor and may hide some abnormality or deviation from the normal state.
[0031] In order to convert these PSNR-based predicted frame quality scores into more comparable and practical anomaly scores, we use the maximum and minimum normalization method to convert the original quality scores to the interval [0,1].
[0032] After obtaining the anomaly scores, we can use these scores to determine and output the final anomaly detection results by setting different thresholds. By flexibly adjusting the thresholds, we can optimize the performance of anomaly detection according to actual needs and application scenarios.
Claims
1. A video anomaly detection method based on hybrid attention and meta-prototype network, characterized in that: The following steps are involved: Step 1: Given a video frame sequence, build an Inception-based encoder network to extract the spatial features of the frame frame by frame and retain the features of each layer; Step 2: The spatial features extracted from each layer are input into the channel attention; Step 3: Input the final features into the meta-prototype network to learn the prototype for encoding normal dynamics; Step 4: Implement the spatial attention of the jump connection between the encoder network and the decoder network, increase the weight of the feature area representing the target and fuse the upsampled feature map of each layer with the encoder feature map of the corresponding channel; Step 5: Predict future frames by extracting features and explicitly modeling normal dynamics in the video sequence; Step 6: Construct the loss function of the prediction model; Step 7: Iteratively train the model and obtain a model that can predict frames close to the real ones; Step 8: Perform video anomaly detection by calculating the error between the predicted frame and the real frame.
2. The video anomaly detection method based on hybrid attention and meta-prototype network according to claim 1 is characterized in that: The specific process of step 2 is as follows: a channel attention mechanism is added during the encoding process to reduce the weight of the channel representing the background feature, which increases the attention to the target during the training process. The size of the feature map input to the SE channel attention is Indicates that is the number of channels, and are the height and width of the matrix respectively; Use the global average pooling layer to compress the spatial dimension of the feature map, and the output matrix It is expressed as shown in formula (A): (A); Among them, x belongs to Represents the row index, y belongs to Represents the column index, z belongs to represents the channel index; the dimension of the matrix SE is ; To feed the feature map into the fully connected (FC) layer, it is first resized to , and then the ReLU activation function is applied; after that, the feature map passes through the FC layer again and the size is changed back , and finally obtain the weight matrix through the Sigmoid function ; In this process, channel mapping The weight is expressed as , as shown in formula (B): (B); The result of the channel attention processing is to assign a smaller weight value to the background channel in the input feature map, and the weight value is less than 1.
3. The video anomaly detection method based on hybrid attention and meta-prototype network according to claim 1 is characterized in that: The specific process of step 3 is as follows: Step 3.1: Build a dynamic prototype set using highly differentiable attention techniques. Step 3.2: Reconstruct the standard encoding by searching for these prototypes; Step 3.3: Fuse the input encoding with the obtained standard encoding to produce the output.
4. The video anomaly detection method based on hybrid attention and meta-prototype network according to claim 3 is characterized in that: The specific process of step 3.1 is as follows: After the video frame is preprocessed by the Inception encoder, t feature coding images are obtained. , which is considered as a c-dimensional vector, ; In the application of the attention mechanism, M attention mapping functions are used is the encoding vector Assign normal weights; at each pixel location, the normal weight is used to measure the standard deviation or normal distribution range of the encoding vector; in, Represents the mth feature map, and obtains a prototype with a collection of encoding vectors , the integration process normalized normal weights, as shown in formula (C): (C); Similarly, by applying multiple attention functions, M prototypes are generated, which together constitute a prototype pool. ; During prototype retrieval, map the preprocessed encoding to the input encoding Make a query to retrieve the relevant items in the prototype pool and reconstruct the encoding The process is shown in formula (D): (D); in, Represents the Nth encoding vector The similarity score with the mth prototype item; through the channel summation operation, the obtained normal map is combined with the original code H as the final output result of the model; then, the code output by the DPU passes through the remaining layers of the autoencoder to continue to perform subsequent frame prediction tasks.
5. The video anomaly detection method based on hybrid attention and meta-prototype network according to claim 1 is characterized in that: The specific process of step 4 is: adding a spatial attention mechanism of skip connection between the encoder and decoder, reducing the weight of the background feature area, and increasing the focus on the target area during training; and the encoded feature map with the same number of channels Input to the spatial attention module, represented as , Expressed as the number of channels of the feature matrix, Expressed as the height of the feature matrix, Represented as the width of the feature matrix; the feature map and Input to Convolution performs linear transformation, and then inputs it into the activation function ReLU for nonlinear transformation; the response matrix output by the activation function , as shown in formula (E): (E); in, Indicates deviation, and represents the slope of the linear transformation, Input to The two-dimensional matrix compresses the number of channels to 1, recorded as As shown in formula (F): (F); in, Represents the slope of the linear transformation, which is then passed through the Sigmoid function to obtain the mixed attention coefficient , weighted to the feature map As shown in formula (G): (G); Under the processing of the spatial attention mechanism, the areas representing the background in the encoded feature map are given weights less than 1 for use in the decoding process.
6. The video anomaly detection method based on hybrid attention and meta-prototype network according to claim 1 is characterized in that: The loss function of the prediction model described in step 6 is shown in formula (H); (H); in is the balance parameter, is the frame prediction loss, is the frame reconstruction loss.
7. The video anomaly detection method based on hybrid attention and meta-prototype network according to claim 1 is characterized in that: The specific process of step 7 is: the loss of frame prediction is calculated by With network prediction value The distance between them is defined as follows: (I); The feature reconstruction loss is designed to give the learned normal prototypes the characteristics of density and diversity. It uses two losses and Optimize these two properties separately, as shown in formula (J): (J); in, is the weight parameter, For dense reconstruction normal coding, it measures the average L2 distance between the input encoding vector and its most relevant prototype, as shown in formula (K): (K); in, Reconstructing the normal encoding for diversity further promotes the diversity between prototype items by pushing the learned prototypes away from each other, as shown in formula (L): (L); By controlling the required margin between prototypes using γ, we exploit the advantages of the above two aspects and encourage prototype items to encode compact and diverse normal dynamics for realistic frame prediction.
8. The video anomaly detection method based on hybrid attention and meta-prototype network according to claim 1 is characterized in that: The specific process of step 8 is as follows: the core of the anomaly detection mechanism is to generate corresponding anomaly scores by accurately evaluating the quality of the predicted frame, and then make decisions based on these scores; in order to quantify the quality of the predicted frame, the peak signal-to-noise ratio PSNR standard is adopted, and its specific calculation formula is shown in (M); (M); in, Indicates the maximum value of the image color. is the original image without noise, K is The noise approximation of In the calculation results of PSNR, a higher value often means that the quality of the predicted frame is better, and its characteristics are closer to the distribution of normal frames; on the contrary, a lower PSNR value usually indicates that the quality of the predicted frame is poor, which may hide some abnormality or deviation from the normal state; In order to convert these PSNR-based predicted frame quality scores into more comparable and practical anomaly scores, the maximum and minimum normalization method is used to convert the original quality scores into the interval of [0,1]; After obtaining the anomaly scores, the final anomaly detection results are determined and output based on these scores by setting different thresholds. By flexibly adjusting the thresholds, the performance of anomaly detection can be optimized according to actual needs and application scenarios.