Video anomaly detection method based on state space model

By constructing a video anomaly detection method based on a state-space model, combining future frame prediction and optical flow reconstruction, and using a U-net network with jump connections and a non-negative visual state-space module, the problems of insufficient video detection accuracy and limited real-time reasoning capabilities are solved, and efficient video anomaly detection is achieved.

CN120635773APending Publication Date: 2025-09-12XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510746395.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing video anomaly detection methods lack detection accuracy when dealing with complex scene changes and semantic changes, and the real-time reasoning capabilities of deep learning models are limited, making it difficult to achieve efficient and fast anomaly detection.

Method used

A video anomaly detection method based on the state space model is adopted. By constructing a future frame prediction model and an optical flow reconstruction model, combining the frame prediction loss function and the optical flow reconstruction loss function, using a symmetric U-net network structure with jump connections and non-negative visual state space blocks, video feature extraction and reconstruction are performed, and vector quantization and non-negative visual state space modules are introduced to accelerate feature aggregation and model convergence.

Benefits of technology

While ensuring high detection accuracy, the inference speed is significantly improved, with the FPS reaching 90 frames per second, realizing efficient video anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635773A_ABST
    Figure CN120635773A_ABST
Patent Text Reader

Abstract

The invention discloses a video anomaly detection method based on a state space model, and the method specifically comprises the following steps: 1, obtaining a video data set, dividing the video data set into a training set and a test set, and carrying out the preprocessing of each video segment in the video data set; 2, constructing a future frame prediction model; 3, designing a frame prediction loss function, and training a future frame prediction model by using the data in the training set; 4, constructing an optical flow reconstruction model, designing an optical flow reconstruction loss function, and performing optical flow reconstruction; and step 5, calculating a peak signal-to-noise ratio generated by a frame prediction error of the future frame prediction model and an optical flow reconstruction error of the optical flow reconstruction model, obtaining a mixed score of each segment of video, and performing anomaly detection according to the obtained mixed score. According to the method, the reasoning speed is remarkably increased while high detection precision is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image or video recognition methods, and relates to a video anomaly detection method based on a state space model. Background Art

[0002] In the security surveillance field, manual detection of video anomalies is extremely tedious and time-consuming, and achieving efficient and rapid anomaly detection is even more challenging for non-professionals. Faced with the urgent need to detect massive amounts of video data, the intelligent upgrade of Video Anomaly Detection (VAD) technology has become a key research direction in the field of computer vision.

[0003] Traditional research primarily uses hand-crafted features, such as local binary patterns and gradient histograms, to achieve automated detection. However, these methods are limited by prior knowledge and suffer from poor performance. With the continuous advancement of deep learning and computer vision, video anomaly detection based on convolutional neural networks (CNNs) and transformers has achieved tremendous success. However, due to the sparse and discrete distribution of anomaly samples and the high cost of manual labeling, most VADs rely on single-class classification (OCC) to detect anomalies.

[0004] OCC is defined as learning a latent feature space from normal samples during the training phase, and classifying data that deviates from this feature distribution as anomalies during the inference phase. Therefore, previous research has focused on how to improve the discrimination between normal and abnormal samples. In deep learning-based VAD, reconstruction-based methods focus on modeling spatial features, while prediction-based methods focus on modeling temporal features. In addition, a few methods use visual gap filling to capture high-level semantics to perform VAD, resulting in different methods having advantages and disadvantages when dealing with different types of anomalies. Therefore, some work combines reconstruction and prediction to develop hybrid VAD to achieve higher performance.

[0005] CNN networks can extract multi-scale local spatial features and capture contextual features using different convolutional kernels. However, due to limitations in their local receptive field, they lack sufficient expressive power to capture global spatial features or long-term features. This makes it difficult to handle the complex scene changes and semantic changes in video surveillance, resulting in insufficient detection accuracy. The Vision Transformer (ViT) converts an image into a series of image patches. Compared to CNNs, ViT is more capable of extracting long-range dependencies. However, as the sequence length increases, the model requires more memory and computing resources, hindering the real-time reasoning capabilities of anomaly detection. Summary of the Invention

[0006] The purpose of the present invention is to provide a video anomaly detection method based on a state-space model, which significantly improves the inference speed while ensuring high detection accuracy.

[0007] The technical solution adopted by the present invention is a video anomaly detection method based on a state space model, which is specifically implemented according to the following steps:

[0008] Step 1: Obtain a video dataset, divide it into a training set and a test set, and preprocess each video in the video dataset;

[0009] Step 2: Build a future frame prediction model;

[0010] Step 3: Design a frame prediction loss function and use the data in the training set to train the future frame prediction model;

[0011] Step 4: Build an optical flow reconstruction model, design an optical flow reconstruction loss function, and perform optical flow reconstruction;

[0012] Step 5: Calculate the peak signal-to-noise ratio (PSNR) of the future frame prediction model frame prediction error and the optical flow reconstruction model optical flow reconstruction error to obtain a mixed score for each video segment, and perform anomaly detection based on the obtained mixed score.

[0013] The present invention is also characterized in that:

[0014] The training set contains only normal video samples, and the test set contains normal video samples and abnormal video samples.

[0015] In step 1, each video in the video dataset is preprocessed as follows:

[0016] Each video is divided into n consecutive image frames, and any frame is recorded as I t , where t=1, 2, 3, ..., n, each image frame I is normalized by t The pixel values ​​are normalized to [-1, 1], and each image frame I t The pixel values ​​are normalized to obtain the input image data x.

[0017] The future frame prediction model in step 2 is a symmetric U-net network structure with skip connections, including a Patch Embedding layer, an encoder based on a non-negative visual state space block, a vector quantization layer, a decoder based on a non-negative visual state space block, and a projection layer connected in sequence. The encoder based on a non-negative visual state space block and the decoder based on a non-negative visual state space block are skip-connected.

[0018] The working process of the future frame prediction model specifically includes the encoding stage and the decoding stage. The encoding stage is specifically as follows: the Patch Embedding layer takes the input data x∈RH×W×C Split into non-overlapping blocks of size 4×4 and map the image to 64 channels to generate an embedded image x'∈R H / 4×W / 4×64 , where H, W, and C represent the height, width, and number of channels of the input data respectively;

[0019] The encoder consists of E connected in sequence n There are two non-negative visual state space blocks, and a Patch Merging downsampling layer is set between the two adjacent non-negative visual state space blocks. The output x' of the Patch Embedding layer is input to the first non-negative visual state space block, and the encoder feature z of the last non-negative visual state space block is processed in sequence. e .

[0020] The decoding stage is specifically as follows: Encoder feature z e Input to the vector quantization layer, the vector quantization layer is used to calculate z e After compression, the feature z is obtained q , the feature z q Input to the decoder based on non-negative visual state space blocks;

[0021] The vector quantization layer defines the codebook vector e∈R K×d , K is the size of the codebook, d is the dimension of the codeword, then the vector quantization layer is e After compression, the feature z is obtained q Specifically:

[0022] z q =argmin j ||z e -e j ||2

[0023] Where j = {1, 2, ..., K}, e j represents the codebook vector of the jth item;

[0024] The decoder based on non-negative visual state space blocks includes D n The non-negative visual state space blocks are connected in sequence, and the adjacent two non-negative visual state space blocks are also connected with the Patch Expansion upsampling layer. q Input into the first non-negative visual state space block, then the output features of the first non-negative visual state space block are coupled with the features of the corresponding layer of the encoder obtained by the jump connection and input into the second non-negative visual state space block. In this way, it is processed step by step in the decoder, and the decoder outputs the feature y';

[0025] The number of non-negative visual state space blocks in the encoder and decoder is the same. The output features of the first non-negative visual state space block of the encoder are down-sampled by Patch Merging and transmitted to the second non-negative visual state space block of the encoder. At the same time, it is jump-connected to the Patch Expansion upsampling layer before the last non-negative visual state space block of the decoder, and so on, skip-connecting step by step.

[0026] Set the decoder's D n The output y' of the non-negative visual state space block is input to the projection layer, which restores the feature y' to image data y with the same pixel value as the input x, and finally obtains the predicted future frame I' t+1 .

[0027] The non-negative visual state space block includes a visual state space module and a non-negative enhancement module connected in sequence. The non-negative enhancement module takes the output of the visual state space module as input, normalizes the input features through LayerNorm and linear mapping operations, and then introduces the ReLU activation function and inputs the non-negative features into the convolution to alleviate gradient disappearance and reduce overfitting.

[0028] The frame prediction loss function L in step 3 FP By the prediction loss L p 、VQ loss L vq and gradient loss L gd composition:

[0029] L FP =L p +L vq +L gd

[0030] Among them, the prediction loss L p 、VQ loss L vq and gradient loss L gd They are:

[0031] L p =||I t+1 -I' t+1 ||2

[0032] L vq =||sg(z e )-z q ||2+β||z e -sg(z q )||2

[0033]

[0034] Among them, I t+1 represents the real future frame, I't+1 represents the future frame predicted by the future frame prediction model, sg(·) represents the stop gradient operator defined as the identity in the forward computation, β is set to 0.25, and m and n represent the spatial indices of the frame;

[0035] The frame prediction loss function L FP The goal is to minimize the number of frames, and then set the training parameters to train the future frame prediction model.

[0036] Step 4 is as follows:

[0037] The optical flow reconstruction model has the same structure as the future frame prediction model. The trained future frame prediction model is used to guide the training of the optical flow reconstruction model. Specifically:

[0038] Use the trained future frame prediction model to predict the next frame and get the next image frame I' t+1 Then, use Flownet2.0 optical flow network to obtain the next frame of image frame I' t+1 The corresponding optical flow O t , that is, the real optical flow, through the normalization operation to optical flow O t The pixel value becomes [-1, 1], and the optical flow reconstruction loss function L is used during the training of the optical flow reconstruction model. FR The minimum is the goal, and the real optical flow O after pixel value normalization is used t As the input of the optical flow reconstruction model to reconstruct the optical flow O' t Training the optical flow reconstruction model for the output of the optical flow reconstruction model;

[0039] L FR By reconstruction loss L r 、vq loss L vq , similarity loss L sim and motion difference loss L md composition:

[0040] L FR =L r +L vq +L sim +0.01L md

[0041] L r =||O t -O' t ||2

[0042] L sim =1-SSIM(O t , O' t )

[0043]

[0044] Among them, SSIM represents the structural similarity index measure, which is used to calculate the true optical flow O t and reconstruct the optical flow O' t The similarity between M t Yes O t and O t-1 The difference in movement between t-1 The optical flow of the previous frame represents the real optical flow; M' t Yes O' t and O t-1 The motion difference between them is set to 0.001.

[0045] Step 5 is as follows:

[0046] Calculate the peak signal-to-noise ratio generated by the frame prediction error of the future frame prediction model and the optical flow reconstruction error of the optical flow reconstruction model:

[0047] S p =PSNR(I t+1 , I' t+1 )

[0048] S r =PSNR( t , O' t )

[0049] Among them, S p is the peak signal-to-noise ratio of the frame prediction error generated by the future frame prediction model, S r The peak signal-to-noise ratio of the optical flow reconstruction error generated by the optical flow reconstruction model;

[0050] Then, S is normalized p and S r Normalize to the range of [0, 1] to obtain the corresponding frame score and optical flow score, which are used as the anomaly score;

[0051] According to the obtained anomaly score, the fragment-level fusion strategy is used to fuse and obtain the mixed score S (i) , and finally calculate the area under the curve according to the mixture fraction to obtain the final anomaly detection result:

[0052]

[0053] Where i∈{1, 2, ..., N}, N is the number of video clips, S (i) represents all scores of the i-th video clip, S p (i) and S r (i) are all the scores of the i-th video segment in the prediction and reconstruction processes, respectively.

[0054] The beneficial effects of the present invention are:

[0055] This paper introduces vector quantization to enhance the discrimination between different samples by compressing normal features, introduces a non-negative visual state space module, and accelerates the aggregation of different features and model convergence through pre-activation. The result is a hybrid video anomaly detection method that integrates frame prediction and optical flow reconstruction. This method can significantly improve the inference speed while ensuring high detection accuracy, with the FPS reaching 90. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a network overall block diagram of the video anomaly detection method based on the state space model of the present invention;

[0057] Figure 2 It is a structural diagram of a future frame prediction model in a video anomaly detection method based on a state space model of the present invention;

[0058] Figure 3 Graphs showing qualitative prediction results of frame prediction and optical flow reconstruction in three datasets in Example 7 of the present invention;

[0059] Figure 4 4 is an abnormality score curve diagram on three data sets in Example 7 of the present invention. DETAILED DESCRIPTION

[0060] The following describes it in detail with reference to specific implementation methods.

[0061] Example 1

[0062] The video anomaly detection method based on the state space model of the present invention has the following process: Figure 1 As shown, the specific steps are as follows:

[0063] Step 1: Obtain a video dataset, divide it into a training set and a test set, and preprocess each video in the video dataset;

[0064] Step 2: Build a future frame prediction model;

[0065] Step 3: Design a frame prediction loss function and use the data in the training set to train the future frame prediction model;

[0066] Step 4: Build an optical flow reconstruction model, design an optical flow reconstruction loss function, and perform optical flow reconstruction;

[0067] Step 5: Calculate the peak signal-to-noise ratio (PSNR) of the future frame prediction model frame prediction error and the optical flow reconstruction model optical flow reconstruction error to obtain a mixed score for each video segment, and perform anomaly detection based on the obtained mixed score.

[0068] Example 2

[0069] On the basis of Example 1, the training set only includes normal video samples, and the test set includes normal video samples and abnormal video samples.

[0070] In step 1, each video in the video dataset is preprocessed as follows:

[0071] Each video is divided into n consecutive image frames, and any frame is recorded as I t , where t=1, 2, 3, ..., n, each image frame I is normalized by t The pixel values ​​are normalized to [-1, 1], and each image frame I t The pixel values ​​are normalized to obtain the input image data x.

[0072] Example 3

[0073] Based on Example 2, the future frame prediction model in step 2 is a symmetric U-net network structure with skip connections, such as Figure 2 As shown in the figure, (a) is the network framework diagram of the future frame prediction model; (b) is a schematic diagram of the non-negative visual state space module; (c) is a schematic diagram of the visual state space (VSS) and the SS2D module structure therein; the future frame prediction model includes a Patch Embedding layer, an encoder based on a non-negative visual state space block, a vector quantization layer, a decoder based on a non-negative visual state space block, and a projection layer connected in sequence, and the encoder based on the non-negative visual state space block and the decoder based on the non-negative visual state space block are jump-connected.

[0074] The working process of the future frame prediction model specifically includes the encoding stage and the decoding stage, wherein the encoding stage is specifically as follows:

[0075] The Patch Embedding layer takes the input data x∈R H×W×C Split into non-overlapping blocks of size 4×4 and map the image to 64 channels to generate an embedded image x'∈R H / 4×W / 4×64 , where H, W, and C represent the height, width, and number of channels of the input data respectively;

[0076] The encoder consists of E connected in sequence n There are two non-negative visual state space blocks (NVSS blocks), and a Patch Merging downsampling layer is set between the two adjacent non-negative visual state space blocks. The output x' of the Patch Embedding layer is input to the first non-negative visual state space block, and the encoder feature z of the last non-negative visual state space block is processed in sequence. e .

[0077] The decoding stage is as follows:

[0078] In order to reduce the feature dimension and provide high-quality latent features while reducing the normal data error and increasing the abnormal data error, the present invention inserts vector quantization (VQ) at the bottleneck to compress the features, and retains the basic information to the maximum extent without overfitting. The specific structure is as follows Figure 2 (b)

[0079] Encoder feature z e Input to the vector quantization layer (VQ), the vector quantization layer z e After compression, the feature z is obtained q , the feature z q Input to the decoder based on non-negative visual state space blocks;

[0080] The vector quantization layer defines the codebook vector e∈R K×d , K is the size of the codebook, d is the dimension of the codeword, then the vector quantization layer is e After compression, the feature z is obtained q Specifically:

[0081] z q =argmin j ||z e -e j ||2

[0082] Where j = {1, 2, ..., K}, e j represents the codebook vector of the jth item;

[0083] The decoder based on non-negative visual state space blocks includes D n The non-negative visual state space blocks are connected in sequence, and the adjacent two non-negative visual state space blocks are also connected with the Patch Expansion upsampling layer. q Input into the first non-negative visual state space block, then the output features of the first non-negative visual state space block are coupled with the features of the corresponding layer of the encoder obtained by the jump connection and input into the second non-negative visual state space block. In this way, it is processed step by step in the decoder, and the decoder outputs the feature y';

[0084] The number of non-negative visual state space blocks in the encoder and decoder is the same. The output features of the first non-negative visual state space block of the encoder are down-sampled by Patch Merging and transmitted to the second non-negative visual state space block of the encoder. At the same time, it is jump-connected to the Patch Expansion upsampling layer before the last non-negative visual state space block of the decoder, and so on, skip-connecting step by step.

[0085] Set the decoder's D nThe output y' of the non-negative visual state space block is input to the projection layer, which restores the feature y' to image data y with the same pixel value as the input x, and finally obtains the predicted future frame I' t+1 .

[0086] The structure of the non-negative visual state space block is as follows Figure 2 As shown in (b) and (c) in the figure, it includes a visual state space module (VSS block) and a non-negative enhancement module (NE) connected in sequence. The VSS block uses deep convolution for high-level feature extraction, and the other part uses linear mapping and activation function to calculate the gating signal. The core SS2D (2D-Selective Scan) block establishes a global receptive field through complementary traversal paths, allowing each pixel to efficiently collect information from all other pixels across multiple directions; the non-negative enhancement module is used to optimize the model for feature aggregation in specific patterns and accelerate convergence. The non-negative enhancement module takes the output of the visual state space module as input, first normalizes the input features through LayerNorm and linear mapping operations, and then introduces the ReLU activation function to input the non-negative features into the convolution to alleviate gradient disappearance and reduce overfitting.

[0087] Example 4

[0088] Based on Example 3, the frame prediction loss function L in step 3 is FP By the prediction loss L p 、VQ loss L vq and gradient loss L gd composition:

[0089] L FP =L p +L vq +L gd

[0090] Among them, the prediction loss L p 、VQ loss L vq and gradient loss L gd They are:

[0091] L p =||I t+1 -I' t+1 ||2

[0092] L vq =||sg(z e )-z q ||2+β||z e -sg(z q )||2

[0093]

[0094] Among them, I t+1 represents the real future frame, I' t+1 represents the future frame predicted by the future frame prediction model, sg(·) represents the stop gradient operator defined as the identity in the forward computation, β is set to 0.25, and m and n represent the spatial indices of the frame;

[0095] The frame prediction loss function L FP The goal is to minimize the number of frames, and then set the training parameters to train the future frame prediction model.

[0096] Example 5

[0097] Based on Example 4, step 4 is specifically as follows: the optical flow reconstruction model has the same structure as the future frame prediction model, and the trained future frame prediction model is used to guide the training of the optical flow reconstruction model, specifically:

[0098] Use the trained future frame prediction model to predict the next frame and get the next image frame I' t+1 Then, use Flownet2.0 optical flow network to obtain the next frame of image frame I' t+1 The corresponding optical flow O t , that is, the real optical flow, through the normalization operation to optical flow O t The pixel value becomes [-1, 1], and the optical flow reconstruction loss function L is used during the training of the optical flow reconstruction model. FR The minimum is the goal, and the real optical flow O after pixel value normalization is used t As the input of the optical flow reconstruction model to reconstruct the optical flow O' t Training the optical flow reconstruction model for the output of the optical flow reconstruction model;

[0099] L FR By reconstruction loss L r 、vq loss L vq , similarity loss L sim and motion difference loss L md composition:

[0100] L FR =L r +L vq +L sim +0.01L md

[0101] L r =||O t -O' t ||2

[0102] L sim =1-SSIM(O t , O' t )

[0103]

[0104] Among them, SSIM represents the structural similarity index measure, which is used to calculate the true optical flow O t and reconstruct the optical flow O' t The similarity between M t Yes O t and O t-1 The difference in movement between t-1 The optical flow of the previous frame represents the real optical flow; M' t Yes O' t and O t-1 The motion difference between them is set to 0.001.

[0105] Example 5

[0106] Based on Example 5, step 5 is specifically as follows:

[0107] Calculate the peak signal-to-noise ratio generated by the frame prediction error of the future frame prediction model and the optical flow reconstruction error of the optical flow reconstruction model:

[0108] S p =PSNR(I t+1 , I' t+1 )

[0109] S r =PSNR( t , O' t )

[0110] Among them, S p is the peak signal-to-noise ratio of the frame prediction error generated by the future frame prediction model, S r The peak signal-to-noise ratio of the optical flow reconstruction error generated by the optical flow reconstruction model;

[0111] Then, S is normalized p and S r Normalize to the range of [0, 1] to obtain the corresponding frame score and optical flow score, which are used as the anomaly score;

[0112] According to the obtained anomaly score, the fragment-level fusion strategy is used to fuse and obtain the mixed score S (i) , and finally calculate the area under the curve according to the mixture fraction to obtain the final anomaly detection result:

[0113] Due to the inherent temporal continuity of videos, the present invention uses a Gaussian filter to smooth the anomaly score, selects a better score between the frame score and the optical flow score according to the AUC of each video segment, and combines the selected segment scores to generate the final hybrid score:

[0114] The formula for calculating the mixture fraction is as follows:

[0115]

[0116] Where i∈{1, 2, ..., N}, N is the number of video clips, S (i) represents all scores of the i-th video clip, S p (i) and S r (i) are all the scores of the i-th video segment in the prediction and reconstruction processes, respectively.

[0117] Example 7

[0118] On the basis of Example 6, in order to verify the effectiveness of the present invention, the method of the present invention was simulated. All simulation experiments were carried out on a single RTX 4090 GPU. Based on the PyTorch deep learning framework, the experimental results were evaluated using the area under the curve (AUC) indicator. Experiments were carried out on three unsupervised video anomaly detection datasets: UCSDPed2, CUHK Avenue, and Shanghaitech. All datasets contain training sets and test sets. The training sets only contain normal events, and the test sets contain normal and abnormal events. (This simulation experiment compares the detection method of the present invention with several currently advanced single-task and hybrid video anomaly detection methods (Frame-Pred (Frame Prediction, frame prediction method), MemAE (memory-augmented autoencoder, memory self-encoder method), MNAD (Memory-guidedNormality for Anomaly Detection, memory-guided detection method), MAAM-Net (Memory-augmentedappearance-motion network memory enhanced appearance-motion network method), and the results are shown in Table 1. The experimental results show that compared with other methods, the present invention can significantly improve detection accuracy.

[0119] Table 1 Comparison of AUC values ​​with other methods

[0120]

[0121]

[0122] Figure 3Two visualization results of the present invention are shown on three datasets: UCSD Ped2, CUHK Avenue, and Shanghaitech (SHT), including predicted frames and reconstructed optical flows for normal and abnormal events. In the predicted frames, the present method well predicts normal events and successfully detects abnormal events through the prediction error. In the optical flow, the weak amplitude of normal motion prevents the generation of normal optical flow, while the reconstruction of abnormal optical flow will introduce significant errors. Therefore, whether it is video frames or optical flow, the present method performs relatively well in detecting video anomalies.

[0123] Figure 4 We further visualize anomaly curves from two test video clips from the three datasets. An adaptive threshold is used to calculate anomaly scores, and each test video frame is classified by normalizing the anomaly score with the threshold. Scores close to 0 indicate normality, while scores close to 1 indicate anomalies. For example, in the Ped2#2 and #6 test videos, two consecutive abnormal events (cycling) are detected. This demonstrates that our method can effectively detect anomalies.

Claims

1. A video anomaly detection method based on a state space model, characterized in that: The specific implementation steps are as follows: Step 1, obtain a video dataset, divide it into a training set and a test set, and preprocess each video in the video dataset; Step 2, build a future frame prediction model; Step 3, design a frame prediction loss function, and use the data in the training set to train the future frame prediction model; Step 4, build an optical flow reconstruction model, design an optical flow reconstruction loss function, and perform optical flow reconstruction; Step 5, calculate the peak signal-to-noise ratio generated by the frame prediction error of the future frame prediction model and the optical flow reconstruction error of the optical flow reconstruction model, obtain the mixed score of each video, and perform anomaly detection based on the obtained mixed score.

2. The video anomaly detection method based on the state space model according to claim 1, characterized in that: The training set only includes normal video samples, and the test set includes normal video samples and abnormal video samples.

3. The video anomaly detection method based on the state space model according to claim 2, characterized in that: The preprocessing of each video in the video dataset in step 1 is specifically as follows: Each video is divided into n consecutive image frames, and any frame is recorded as I t , where t=1,2,3,...,n, each image frame I is normalized t The pixel values ​​are normalized to [-1,1], and each image frame I t The pixel values ​​are normalized to obtain the input image data x.

4. The video anomaly detection method based on the state space model according to claim 3, characterized in that: The future frame prediction model in step 2 is a symmetric U-net network structure with skip connections, including a PatchEmbedding layer, an encoder based on a non-negative visual state space block, a vector quantization layer, a decoder based on a non-negative visual state space block, and a projection layer connected in sequence, wherein the encoder based on a non-negative visual state space block and the decoder based on a non-negative visual state space block are skip-connected.

5. The video anomaly detection method based on the state space model according to claim 4 is characterized in that: The working process of the future frame prediction model specifically includes the encoding stage and the decoding stage, wherein the encoding stage is specifically as follows: the PatchEmbedding layer converts the input data x∈R H×W×C Split into non-overlapping blocks of size 4×4 and map the image to 64 channels to generate an embedded image x'∈R H / 4×W / 4×64 , where H, W, and C represent the height, width, and number of channels of the input data respectively; The encoder comprises E n There are two non-negative visual state space blocks, and a Patch Merging downsampling layer is set between the two adjacent non-negative visual state space blocks. The output x' of the Patch Embedding layer is input to the first non-negative visual state space block, and the encoder feature z of the last non-negative visual state space block is processed in sequence. e .

6. The video anomaly detection method based on the state space model according to claim 5, characterized in that: The decoding stage is specifically as follows: Encoder feature z e Input to the vector quantization layer, the vector quantization layer is used to calculate z e After compression, the feature z is obtained q , the feature z q Input to the decoder based on non-negative visual state space blocks; The vector quantization layer defines the codebook vector e∈R K×d , K is the size of the codebook, d is the dimension of the codeword, then the vector quantization layer is e After compression, the feature z is obtained q Specifically: q =argmin j ||z e -e j ||2, where j = {1, 2, ..., K}, e j represents the codebook vector of the jth item; The decoder based on non-negative visual state space block includes D n The non-negative visual state space blocks are connected in sequence, and the adjacent two non-negative visual state space blocks are also connected with the Patch Expansion upsampling layer. q Input into the first non-negative visual state space block, then the output features of the first non-negative visual state space block are coupled with the features of the corresponding layer of the encoder obtained by the jump connection and input into the second non-negative visual state space block. In this way, it is processed step by step in the decoder, and the decoder outputs the feature y'; The number of non-negative visual state space blocks in the encoder and decoder is the same. The output features of the first non-negative visual state space block of the encoder are down-sampled by Patch Merging and transmitted to the second non-negative visual state space block of the encoder. At the same time, it is jump-connected to the Patch Expansion upsampling layer before the last non-negative visual state space block of the decoder, and so on, skip-connecting step by step. Set the decoder's D n The output y' of the non-negative visual state space block is input to the projection layer, which restores the feature y' to image data y with the same pixel value as the input x, and finally obtains the predicted future frame I' t+1 .

7. The video anomaly detection method based on the state space model according to claim 5, characterized in that: The non-negative visual state space block includes a visual state space module and a non-negative enhancement module connected in sequence. The non-negative enhancement module takes the output of the visual state space module as input, first normalizes the input features through LayerNorm and linear mapping operations, then introduces the ReLU activation function, and inputs the non-negative features into the convolution to alleviate gradient disappearance and reduce overfitting.

8. The video anomaly detection method based on the state space model according to claim 7, characterized in that: The frame prediction loss function L in step 3 FP By the prediction loss L p 、VQ loss L vq and gradient loss L gd composition: L FP =L p +L vq +L gd Among them, the prediction loss L p 、VQ loss L vq and gradient loss L gd They are: L p =||I t+1 -I' t+1 ||2 L vq =||sg(z e )-With q ||2+β||z e -sg(z q )||2 Among them, I t+1 represents the real future frame, I' t+1 represents the future frame predicted by the future frame prediction model, sg(·) represents the stop gradient operator defined as the identity in the forward computation, β is set to 0.25, and m and n represent the spatial indices of the frame; The frame prediction loss function L FP The goal is to minimize the number of frames, and then set the training parameters to train the future frame prediction model.

9. The video anomaly detection method based on the state space model according to claim 8, characterized in that: The step 4 is specifically as follows: The optical flow reconstruction model has the same structure as the future frame prediction model. The trained future frame prediction model is used to guide the training of the optical flow reconstruction model. Specifically: Use the trained future frame prediction model to predict the next frame and get the next image frame I' t+1 Then, use Flownet2.0 optical flow network to obtain the next frame of image frame I' t+1 The corresponding optical flow O t , that is, the real optical flow, through the normalization operation to optical flow O t The pixel value becomes [-1,1], and the optical flow reconstruction loss function L is used in the training of the optical flow reconstruction model. FR The minimum is the goal, and the real optical flow O after pixel value normalization is used t As the input of the optical flow reconstruction model to reconstruct the optical flow O' t Training the optical flow reconstruction model for the output of the optical flow reconstruction model; L FR By reconstruction loss L r 、vq loss L vq , similarity loss L sim and motion difference loss L md composition: L FR =L r +L vq +L sim +0.01L md IT r =O t -A' t2 L sim =1-SSIM(O t ,THE' t ) Among them, SSIM represents the structural similarity index measure, which is used to calculate the true optical flow O t and reconstruct the optical flow O' t The similarity between M t Yes O t and O t-1 The difference in movement between t-1 The optical flow of the previous frame represents the real optical flow; M' t Yes O' t and O t-1 The motion difference between them is set to 0.

001.

10. The video anomaly detection method based on the state space model according to claim 9, characterized in that: The step 5 is specifically as follows: Calculate the peak signal-to-noise ratio generated by the frame prediction error of the future frame prediction model and the optical flow reconstruction error of the optical flow reconstruction model: S p =PSNR(I t+1 ,IN' t+1 ) S r =PSNR(O t ,THE' t ) Among them, S p is the peak signal-to-noise ratio of the frame prediction error generated by the future frame prediction model, S r The peak signal-to-noise ratio of the optical flow reconstruction error generated by the optical flow reconstruction model; Then, S is normalized p and S r Normalize to the range of [0,1] to obtain the corresponding frame score and optical flow score, which are used as the anomaly score; According to the obtained anomaly score, the fragment-level fusion strategy is used to fuse and obtain the mixed score S (i) , and finally calculate the area under the curve according to the mixture fraction to obtain the final anomaly detection result: Where i∈{1,2,...,N}, N is the number of video clips, S (i) represents all scores of the i-th video clip, S p (i) and S r (i) are all the scores of the i-th video segment in the prediction and reconstruction processes, respectively.