Video anomaly detection method and system based on global local feature double-flow network

By using a global-local feature dual-stream network structure to fuse features from video frames and optical flow sequences, the problem of low detection accuracy caused by local bias in existing methods is solved, and efficient detection of abnormal behavior in videos is achieved.

CN120953882APending Publication Date: 2025-11-14HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511077316.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing video anomaly detection methods suffer from local bias when dealing with complex image recognition tasks, neglecting the overall scene and resulting in insufficient ability to reconstruct abnormal behavior. They also struggle to effectively integrate local and global features in video data, leading to low detection accuracy.

Method used

A global-local feature dual-stream network structure is adopted. Features of video frames and optical flow sequences are extracted through appearance branches and motion branches respectively to generate a spatiotemporal prediction map. The network parameters are optimized by using the joint loss function of appearance and motion. By combining the features of global path and local path, the error between the network and the real image is generated to detect anomalies.

Benefits of technology

It improves the accuracy of detecting abnormal behaviors in videos, suppresses the phenomenon of abnormal events being incorrectly predicted as normal, and comprehensively enhances the model's representation capabilities in appearance and motion dimensions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953882A_ABST
    Figure CN120953882A_ABST
Patent Text Reader

Abstract

The invention discloses a video anomaly detection method and system based on a global local feature double-flow network. The method comprises the following steps: 1.1, loading an original video training and test data set; 1.2, dividing a video frame sequence and constructing a corresponding optical flow frame sequence; 2.1, constructing a video anomaly detection model based on global and local feature fusion; 2.2, the model takes a video frame sequence and an optical flow frame sequence as input of an appearance branch encoder and an input of a motion branch encoder respectively, and appearance feature codes and motion feature codes are extracted respectively; 2.3, restoring the image by the appearance and motion decoder, and generating a prediction frame image matched with the corresponding mode; 2.4, calculating an appearance and motion joint loss function, performing back propagation, and optimizing a codec of a branch; 3.1, dividing the test video into a video frame sequence and an optical flow frame sequence; 3.2, inputting into a trained model to generate a prediction frame image at the next moment; 3.3, calculating abnormal scores of the appearance stream and the motion stream of each test video frame; and 3.4, when the abnormal score exceeds a threshold value, judging the current frame as an abnormal frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of video anomaly detection technology, and relates to a detection technology that combines computer vision with deep learning. Specifically, it is a video anomaly detection method and system that integrates a dual-stream network of global and local features. Background Technology

[0002] Video anomaly detection (VAD) is an important technology in the field of computer vision. It aims to simulate human understanding of video content, automatically identifying events in surveillance videos that do not conform to normal behavioral expectations, such as fights, stampedes, and traffic accidents. It has significant application value and promising prospects in areas such as traffic monitoring, public safety, and industrial production monitoring.

[0003] Video anomaly detection technology currently faces two main challenges: First, anomalous events are highly unbounded and ambiguous, with different application scenarios often leading to changes in the definition and classification of anomaly types. Second, anomalous events occur infrequently in real life, and the anomalies that occur vary across different scenarios. Therefore, using supervised binary classification methods to collect all anomaly types is impractical.

[0004] With the rapid development of deep learning technology, methods based on Convolutional Neural Networks (CNNs) have been widely applied to VAD tasks. The mainstream approach treats VAD as a semi-supervised learning problem, aiming to learn pattern contours containing only normal data, and treating samples deviating from the normal pattern as anomalies during the testing phase. Typical semi-supervised methods can be divided into two categories: reconstruction and prediction. Reconstruction-based methods take the current frame as input and output the corresponding reconstruction result, using the reconstruction error to detect anomalies. Prediction-based methods use multiple frames to predict future frames and detect anomalies by comparing the prediction differences.

[0005] However, CNNs exhibit significant local bias when handling complex image recognition tasks. Their small receptive field convolutional operations tend to capture local texture patterns, limiting their ability to learn global context. In VAD tasks, this limitation often leads existing methods to focus more on local individual changes in a single video frame, neglecting the overall scene, resulting in good reconstruction of anomalous behavior. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention discloses a video anomaly detection method and system based on a global-local feature dual-stream network. During the training phase, this invention constructs a parallel network structure consisting of an appearance branch and a motion branch, receiving video frames and optical flow sequences as inputs respectively, to extract and fuse spatial and temporal information. During the testing phase, the two branches generate corresponding spatiotemporal prediction maps, and anomalies are detected based on the error between these maps and the actual images. This invention solves the problem of how to fuse local and global features in video data to delineate the boundaries of anomalous events, improving the accuracy of detecting anomalous behaviors in videos.

[0007] The technical solution adopted by this invention to solve its technical problem is as follows:

[0008] A video anomaly detection method based on a global-local feature dual-stream network includes data preprocessing, training, and testing phases. The specific steps are as follows:

[0009] S1, Data Preprocessing Stage:

[0010] Step 1.1: Load the original video training and testing datasets; the training videos contain only normal events, while the testing videos contain both normal and abnormal events;

[0011] Step 1.2: Divide the video frame sequence and construct the corresponding optical flow frame sequence;

[0012] S2, Training Phase:

[0013] Step 2.1: Construct a video anomaly detection model based on the fusion of global and local features. The video anomaly detection model consists of an appearance branch and a motion branch. According to the different modalities of the input data, the appearance branch and motion branch networks extract the corresponding features respectively. Both branches consist of an encoder and a decoder, which are responsible for feature encoding extraction and decoding reconstruction respectively.

[0014] Step 2.2: The video anomaly detection model uses the video frame sequence as the input to the encoder in the appearance branch and the optical flow frame sequence as the input to the encoder in the motion branch, and extracts the corresponding appearance features and motion feature codes respectively.

[0015] Step 2.3: The appearance decoder and motion decoder decode and restore the image structure layer by layer to generate a prediction frame image that matches the corresponding pattern;

[0016] Step 2.4: Calculate the joint loss function for appearance and motion, backpropagate to update network parameters, and optimize the encoder and decoder of the two branches;

[0017] S3, Testing Phase:

[0018] Step 3.1: Divide the test video into a video frame sequence and an optical flow frame sequence;

[0019] Step 3.2: Input the pre-trained video anomaly detection model to generate the predicted frame image for the next time step;

[0020] Step 3.3: Calculate the anomaly scores for the appearance stream and motion stream of each test video frame;

[0021] Step 3.4: Set a threshold. When the abnormal score exceeds the threshold, the current frame is determined to be an abnormal frame.

[0022] Preferably, step 1.1 involves loading the original video training and testing datasets. The training video contains only normal events, while the testing video contains both normal and abnormal events. The size of each frame in the video frame sequence I is standardized to H×W, where H and W represent the height and width of the image, respectively. Simultaneously, optical flow images are extracted from every two adjacent video frames to generate the corresponding optical flow sequence F.

[0023] Preferably, step 1.2 involves dividing the video frame sequence and constructing the corresponding optical flow frame sequence. Each time, a video segment {I} from the previous t consecutive frames is selected. r ,I r+1 ,…,I r+t-1} and its corresponding optical flow segment {F r ,F r+1 ,…,F r+t-2}, which serve as the input data for the appearance and motion branches in the model, respectively. r represents the starting frame index of the current training video segment.

[0024] As a preferred option, step 2.1 involves constructing a video anomaly detection model based on the fusion of global and local features. The overall architecture consists of an appearance branch and a motion branch. Based on different modalities of the input data, the appearance branch and motion branch networks extract corresponding features respectively. Each branch consists of an encoder and a decoder, responsible for feature encoding extraction and decoding reconstruction, respectively.

[0025] Preferably, in step 2.2: the model uses the video frame sequence as input to the encoder in the appearance branch and the optical flow frame sequence as input to the encoder in the motion branch, extracting the corresponding appearance and motion feature codes respectively. In each iteration, {I} is selected from the input video segments. r ,I r+1 ,…,I r+t-2} and the corresponding optical flow segment {F r ,F r+1 ,…,F r+t-3}, respectively input to the appearance encoder E A and motion encoder E M To extract appearance and motion features from the video, where I r and F rThese are the starting index frames of the input training segment. Simultaneously, I... r+t-1 and F r+t-2 These are used as appearance and optical flow labels for the real frames, respectively, for subsequent prediction error calculation.

[0026] Each branch's encoder consists of three corresponding stacked global and local modules. The global path extracts coarse global information from the input image sequence, such as shape and spatial distribution; the local path focuses on capturing detailed local information, such as texture and edge features. Finally, information is exchanged through a top-down modulator to effectively fuse global and local features. Each layer progressively models appearance and motion patterns in normal events, thereby accurately representing the spatiotemporal regularities of the video sequence.

[0027] After the encoder extracts features, E A With E M Based on the correlation between appearance patterns and motion patterns, the final appearance feature codes and motion feature codes are generated and input into the corresponding appearance decoder D. A With motion decoder D M Perform decoding.

[0028] Preferably, step 2.3 involves the appearance decoder and motion decoder decoding and reconstructing the image structure layer by layer to generate a predicted frame image that matches the corresponding pattern. Decoder d A and D M By performing layer-by-layer decoding, the feature map is gradually enlarged and the spatial structure is restored, generating a predicted image that matches the normal appearance and motion pattern.

[0029] Preferably, step 2.4 involves calculating the joint loss function for appearance and motion, backpropagating to update network parameters, and optimizing the encoders and decoders for both branches. To achieve efficient model training, a joint optimization objective function composed of multiple loss functions is designed to simultaneously optimize the encoders and decoders for both the appearance and motion branches. This objective function includes four sub-loss terms: appearance prediction loss (Loss) Appearance Gradient loss Grad Motion prediction loss Motion Optical flow loss Flow The appearance codec is optimized using appearance prediction and gradient loss in the appearance branch. The appearance prediction loss minimizes the pixel difference between the predicted and actual frames and is defined as follows:

[0030]

[0031] in, I represents the predicted frame generated by the model. r+t-1 Represents the actual frame at the corresponding time. This represents the square of the l2 norm, used to measure pixel-level reconstruction error.

[0032] Gradient loss constrains the gradient difference between the predicted frame and the actual frame in two spatial dimensions, and is defined as:

[0033]

[0034] in, and I r+t-1 (i,j) represent the gray values ​​at pixel position (i,j) in the predicted frame and the real frame, respectively, and ||·||1 represents the l1 norm.

[0035] In the motion branch, the predictive loss L is also used. M To constrain the L2 distance between the predicted optical flow and the actual motion field, it is defined as:

[0036]

[0037] in, F represents the predicted optical flow map. r+t-2 This represents the true optical flow diagram.

[0038] Optical flow is an excellent method for capturing temporal motion information from video clips. The optical flow loss is defined as:

[0039]

[0040] The complete loss function considers all the above loss functions and combines them as follows:

[0041] Loss=λ A Loss Appearance +λ G Loss Grad +λ M Loss Motion +λ F Loss Flow (5)

[0042] λ A , λ G , λ M and λ F It is a weight that balances the importance of different loss functions.

[0043] Repeat steps 2.2 to 2.4, iteratively train for R rounds, continuously optimize the total loss function Loss until the loss value tends to converge, and finally obtain the model trained by this method.

[0044] Preferably, step 3.1 involves dividing the test video into a video frame sequence and an optical flow frame sequence. Similar to the training phase, the test video is divided into t consecutive frames prior to the current time step {I...}. v ,I v+1 ,…,I v+t-1} and the corresponding optical flow frame fragment {F v ,F v+1 ,…,F v+t-2 The dimensions are standardized to an H×W input model. v represents the starting frame index of the current test video segment, and I... v and F v These are the starting index frames of the input test segment.

[0045] Preferably, step 3.2: Input the video segment {I} into the trained model to generate the predicted frame image for the next time step. v ,I v+1 ,…,I v+t-2} and optical flow segment {F v ,F v+1 ,…,F v+t-3 The data is fed into the already trained model to generate the prediction frame for the next time step. and

[0046] Preferably, step 3.3: calculate the anomaly scores of the appearance stream and motion stream for each test video frame.

[0047] Exception score s for appearance branch A (t), based on the difference between the predicted frame and the actual frame, first calculate the Peak Signal-to-Noise Ratio (PSNR) for each frame. PSNR is used to measure the difference between the predicted frame and the actual frame. The pixel difference between the actual frame I and the actual frame I is calculated using the following formula:

[0048]

[0049] in, To predict the maximum pixel value of a frame image, I(x,y) and These are the pixel values ​​of the actual frame and the predicted frame, respectively. After obtaining the Peak Signal-to-Noise Ratio (PSNR) for each frame, the average PSNR value of that segment is calculated.

[0050] Abnormal score s for the motor branch M Similarly, based on the difference between the predicted optical flow and the actual optical flow, the PSNR value of the optical flow is calculated, and the obtained PSNR value is uniformly normalized to the interval [0,1].

[0051] The total anomaly score S(t) is calculated by weighted fusion of the anomaly scores from the external flow and the moving flow:

[0052] S(t)=αs A (t)+βs M (t) (7)

[0053] α and β are weighted hyperparameters used to balance the contributions of outlier scores from appearance flow and motion flow.

[0054] Preferably, step 3.4 involves setting a threshold. When the anomaly score exceeds this threshold, the current frame is determined to be an anomaly frame. Based on the calculated total anomaly score S(t_), a threshold θ is set to determine whether the current frame is an anomaly. When the total anomaly score exceeds this threshold, the current frame is determined to be an anomaly frame.

[0055] As a preferred option, the global-local modules are described in detail below:

[0056] The Global-Local Block (GLB) consists of three parts: a global pathway (GP), a local pathway (LP), and a top-down modulator.

[0057] The global path employs a special Transformer encoder structure. First, the input 3D feature map is divided into N P×P image blocks by a feature-to-token conversion unit, which then converts it into N one-dimensional token sequences. Here, N = (H×W) / (P×P), where H and W represent the height and width of the 3D feature map. The feature-to-token conversion unit consists of a 2D convolutional layer with a kernel size of 1 and a stride of 1, a batch normalization layer, and a ReLU activation function, along with an average pooling layer with a kernel size of 2 and a stride of 2. After tokenization, the token sequences are normalized by a normalization layer, and positional encoding is added to preserve spatial location information. Subsequently, a global context representation of the N one-dimensional token sequences is generated through a multi-head self-attention mechanism, and then normalized again to stabilize the feature distribution. Finally, the token-to-feature conversion unit restores the 3D feature map, with the number of channels and size downsampled to half the original feature map. The Token to Feature Transformation Unit (TVT) consists of a 2D convolutional layer with a kernel size of 1 and a stride of 1, a batch normalization layer, and a ReLU activation function. Finally, a residual connection consisting of a 2D convolutional layer with a kernel size of 1 and a stride of 2 is added between the input and output of the TVT to Feature Transformation Unit to generate the output of the global path.

[0058] Local paths can be categorized into Appearance Local Pathway (ALP) and Motion Local Pathway (MLP) based on the type of input features. ALP consists of three cascaded 2D convolutional units. Each unit includes a 2D convolutional layer with a kernel size of 3 and a stride of 1, a batch normalization layer, and a ReLU activation function. Finally, a max pooling layer with a kernel size of 2 and a stride of 2 is used for double spatial downsampling, while simultaneously saving the pooling position index for subsequent unpooling operations. MLP consists of a 3D convolutional layer with a kernel size of 3 and a stride of (1,2,2), a batch normalization layer, and a ReLU activation function.

[0059] The top-down modulator achieves dynamic weighted modulation of local feature maps by applying learnable weight matrices along the channel dimension of the global feature map. During model initialization, a set of learnable weight matrices is randomly initialized using a uniform distribution. During training, gradients are calculated based on backpropagation and the loss function, and then dynamically updated by the optimizer to adaptively optimize the weight matrices and generate modulation coefficients. These coefficients are non-linearly normalized using the sigmoid activation function and then used to perform element-wise modulation on the local feature maps, effectively guiding and fusing local features with global context information. Subsequently, by combining a global path with appearance-local paths, motion-local paths, and the top-down modulator, appearance-global-local and motion-global-local modules in a two-stream network can be constructed.

[0060] Appearance branch network: consisting of an appearance encoder E A and appearance decoder D A It consists of two parts. Among them, E... A It is constructed by cascading three global and local appearance modules. The output tensor shapes of each module are 64×64×128, 32×32×256, and 16×16×512, respectively. A The algorithm is constructed by concatenating three 2x upsampling blocks. Each upsampling block first passes through an unpooling layer with a kernel size of 2 and a stride of 2, followed by three convolutional units consisting of a 2D convolutional layer with a kernel size of 3 and a stride of 1, a batch normalization layer, and a ReLU activation function. Each unpooling layer receives pooling position indices from the encoder, gradually recovering the spatial structure. The output tensor shapes are 32×32×256, 64×64×128, and 128×128×64, respectively. Finally, a 2D convolutional layer with a kernel size of 3 and a stride of 1 generates a feature map of size 256×256×3 to predict normal appearance patterns.

[0061] Motion branch network: consisting of a motion encoder EM and motion decoder D M It consists of two parts. Among them, E... M It is constructed by cascading three global and local motion modules. The output tensor shapes of each module are 64×64×128, 32×32×256, and 16×16×512, respectively. D M The system is constructed by cascading three 3D deconvolutional modules. Each module contains a 3D deconvolutional layer with a kernel size of 3 and a stride of (1,2,2), a batch normalization layer, and a ReLU activation function. The output tensor shapes are 32×32×256, 64×64×128, and 128×128×64, respectively. Finally, a 256×256×2 feature map is generated by a 3D deconvolutional layer with a kernel size of 3 and a stride of (1,2,2) to predict normal motion patterns.

[0062] To adapt to the dimensionality changes of the network's input and output tensors, the motion branch network employs a three-dimensional convolutional structure to model joint spatial-temporal features. Specifically, the original input tensor has dimensions [B, C1, H, W], where B represents the batch size, C1 represents the number of channels, and H and W represent the height and width of the image, respectively. Before inputting to the encoder and decoder, this four-dimensional tensor needs to be reshaped into a five-dimensional tensor [B, C2, T, H, W], where T represents the number of input frames, C2 represents the number of channels in a single frame, and C1 = T × C2 is satisfied to adapt to the input requirements of three-dimensional convolution. After obtaining the output tensor through forward propagation, its dimensions are restored to the original [B, C1, H, W].

[0063] This invention also discloses a video anomaly detection system based on a global-local feature dual-stream network for performing the above method, which includes the following modules:

[0064] Data preprocessing module: used to load the original video training and testing datasets; training videos contain only normal events, while testing videos contain both normal and abnormal events; segmenting video frame sequences and constructing corresponding optical flow frame sequences;

[0065] Training Module: Used to construct a video anomaly detection model based on the fusion of global and local features. The video anomaly detection model consists of an appearance branch and a motion branch. According to different modalities of the input data, the appearance branch and motion branch networks extract corresponding features respectively. Each branch consists of an encoder and a decoder, which are responsible for feature encoding extraction and decoding reconstruction respectively. The video anomaly detection model uses the video frame sequence as the input of the encoder in the appearance branch and the optical flow frame sequence as the input of the encoder in the motion branch to extract the corresponding appearance features and motion features respectively. The appearance decoder and motion decoder decode and reconstruct the image structure layer by layer to generate a prediction frame image that matches the corresponding mode. The joint loss function of appearance and motion is calculated, and the network parameters are updated by backpropagation to optimize the encoder and decoder of the two branches.

[0066] The testing module is used to divide the test video into a sequence of video frames and a sequence of optical flow frames; input the data into a pre-trained video anomaly detection model to generate the predicted frame image for the next moment; calculate the anomaly scores of the appearance flow and motion flow for each test video frame; and set a threshold, when the anomaly score exceeds the threshold, the current frame is determined to be an anomaly frame.

[0067] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0068] This invention proposes a video anomaly detection method and system based on a global-local module dual-stream network. The design incorporates a global path, a local path, and a top-down modulator, effectively combining fine local features and global contextual patterns in the video sequence. This suppresses the overgeneralization of anomalous events by the neural network, preventing anomalous events from being incorrectly predicted as normal. Furthermore, this invention constructs a dual-stream network that combines appearance and motion features, fusing image appearance information and optical flow motion information from the input video. This comprehensively enhances the model's ability to represent normal patterns in both appearance and motion dimensions, thereby improving model performance. Attached Figure Description

[0069] To make the technical solution of the present invention clearer, the accompanying drawings required in the embodiments are described below:

[0070] Figure 1 This is a flowchart of a preferred embodiment of the present invention for a video anomaly detection method based on a global local feature dual-stream network.

[0071] Figure 2 This is a schematic diagram of the network structure of a video anomaly detection method based on a global local feature dual-stream network according to a preferred embodiment of the present invention.

[0072] Figure 3 This is a schematic diagram of the global local module network structure of a video anomaly detection method based on a global local feature dual-stream network according to a preferred embodiment of the present invention.

[0073] Figure 4 This is a schematic diagram of the global-local module network structure used for appearance branch in a video anomaly detection method based on a global-local feature dual-stream network according to a preferred embodiment of the present invention.

[0074] Figure 5 This is a schematic diagram of the global-local module network structure used for motion branches in a video anomaly detection method based on a global-local feature dual-stream network according to a preferred embodiment of the present invention.

[0075] Figure 6 This is a schematic diagram of the structure of the appearance decoder and motion decoder in a video anomaly detection method based on a global local feature dual-stream network according to a preferred embodiment of the present invention.

[0076] Figure 7 This is a structural diagram of the Transformer encoder used for global path in a video anomaly detection method based on a global local feature dual-stream network according to a preferred embodiment of the present invention.

[0077] Figure 8 This is a block diagram of a video anomaly detection system based on a global-local feature dual-stream network, according to a preferred embodiment of the present invention. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Through the content disclosed in this specification, those skilled in the art can easily understand other advantages and effects of the present invention.

[0079] This embodiment presents a video anomaly detection method based on a global-local feature dual-stream network. During the training phase, a parallel network structure consisting of appearance and motion branches is constructed, receiving video frames and optical flow sequences as inputs respectively to extract and fuse spatial and temporal information. During the testing phase, the two branches generate corresponding spatiotemporal prediction maps, and anomalies are detected based on the error between these maps and the real images. This method solves the problem of how to fuse local and global features in video data to delineate the boundaries of anomalous events, thus improving the accuracy of detecting abnormal video behavior.

[0080] like Figure 1 As shown in the figure, this embodiment presents a video anomaly detection method based on a global local feature dual-stream network, which includes three stages: data preprocessing, training, and testing.

[0081] First, we will introduce the global and local modules involved in this embodiment in detail:

[0082] The Global-Local Block (GLB) architecture described above is as follows: Figure 3As shown, it consists of three parts: a global pathway (GP), a local pathway (LP), and a top-down modulator.

[0083] The global path employs a special Transformer encoder structure, such as... Figure 7 As shown, the input 3D feature map is first divided into N P×P image blocks by a feature-to-token conversion unit, and then converted into N one-dimensional token sequences, where N = (H×W) / (P×P), where H and W represent the height and width of the 3D feature map, and P = 2. The feature-to-token conversion unit consists of a 2D convolutional layer with a kernel size of 1 and a stride of 1, a batch normalization layer, and a ReLU activation function, and an average pooling layer with a kernel size of 2 and a stride of 2. After tokenization, the token sequences are first normalized by a normalization layer, and positional encoding is added to preserve spatial location information. Subsequently, a global context representation of the N one-dimensional token sequences is generated through a multi-head self-attention mechanism, and then normalized again to stabilize the feature distribution. Finally, the token-to-feature conversion unit restores the 3D feature map, with the number of channels and size downsampled to half of the original feature map. The Token to Feature Transformation Unit (TVT) consists of a 2D convolutional layer with a kernel size of 1 and a stride of 1, a batch normalization layer, and a ReLU activation function. Finally, a residual connection consisting of a 2D convolutional layer with a kernel size of 1 and a stride of 2 is added between the input and output of the TVT to Feature Transformation Unit to generate the output of the global path.

[0084] Local paths can be categorized into Appearance Local Pathway (ALP) and Motion Local Pathway (MLP) based on the type of input features. ALP consists of three cascaded 2D convolutional units. Each unit includes a 2D convolutional layer with a kernel size of 3 and a stride of 1, a batch normalization layer, and a ReLU activation function. Finally, a max pooling layer with a kernel size of 2 and a stride of 2 is used for double spatial downsampling, while simultaneously saving the pooling position index for subsequent unpooling operations. MLP consists of a 3D convolutional layer with a kernel size of 3 and a stride of (1,2,2), a batch normalization layer, and a ReLU activation function.

[0085] The top-down modulator achieves dynamic weighted modulation of local feature maps by applying learnable weight matrices along the channel dimension of the global feature map. During model initialization, a set of learnable weight matrices is randomly initialized using a uniform distribution. During training, gradients are calculated based on backpropagation and the loss function, and then dynamically updated by the optimizer to adaptively optimize the weight matrices and generate modulation coefficients. These coefficients are non-linearly normalized using the sigmoid activation function and then used to perform element-wise modulation on the local feature maps, effectively guiding and fusing local features with global context information. Subsequently, by combining a global path with appearance-local paths, motion-local paths, and the top-down modulator, appearance-global-local and motion-global-local modules in a two-stream network can be constructed, such as... Figure 4 and 5 As shown.

[0086] Appearance branch network: consisting of an appearance encoder E a and appearance decoder D A It consists of two parts. Among them, E... A It is constructed by cascading three global and local appearance modules. The output tensor shapes of each module are 64×64×128, 32×32×256, and 16×16×512, respectively. A The algorithm is constructed by concatenating three 2x upsampling blocks. Each upsampling block first passes through an unpooling layer with a kernel size of 2 and a stride of 2, followed by three convolutional units consisting of a 2D convolutional layer with a kernel size of 3 and a stride of 1, a batch normalization layer, and a ReLU activation function. Each unpooling layer receives pooling position indices from the encoder, gradually recovering the spatial structure. The output tensor shapes are 32×32×256, 64×64×128, and 128×128×64, respectively. Finally, a 2D convolutional layer with a kernel size of 3 and a stride of 1 generates a feature map of size 256×256×3 to predict normal appearance patterns.

[0087] Motion branch network: consisting of a motion encoder E M and motion decoder D M It consists of two parts. Among them, E... M It is constructed by cascading three global and local motion modules. The output tensor shapes of each module are 64×64×128, 32×32×256, and 16×16×512, respectively. D MThe system is constructed by cascading three 3D deconvolutional modules. Each module contains a 3D deconvolutional layer with a kernel size of 3 and a stride of (1,2,2), a batch normalization layer, and a ReLU activation function. The output tensor shapes are 32×32×256, 64×64×128, and 128×128×64, respectively. Finally, a 256×256×2 feature map is generated by a 3D deconvolutional layer with a kernel size of 3 and a stride of (1,2,2) to predict normal motion patterns. Appearance decoder D A and motion decoder D M Structure as Figure 6 As shown.

[0088] To adapt to the dimensionality changes of the network's input and output tensors, the motion branch network employs a three-dimensional convolutional structure to model the joint spatial-temporal features. Specifically, the original input tensor has dimensions [B, C1, H, W], where B represents the batch size, C1 represents the number of channels, and H and W represent the height and width of the image, respectively. Before inputting to the encoder and decoder, this four-dimensional tensor needs to be reshaped into a five-dimensional tensor [B, C2, T, H, W], where T represents the number of input frames, C2 represents the number of channels in a single frame, and C1 = T × C2 is satisfied to adapt to the input requirements of three-dimensional convolution. Here, T is the number of input optical flow frames (3). After obtaining the output tensor through forward propagation of the network, its dimensions are restored to the original [B, C1, H, W].

[0089] like Figure 1-6 As shown in the figure, this embodiment of a video anomaly detection method based on a global-local feature dual-stream network includes data preprocessing, a training phase, and a testing phase. The three phases are described in detail below:

[0090] S1, Data Preprocessing Stage:

[0091] Step 1.1: Load the original video training and testing datasets. The training video contains only normal events, while the test video contains both normal and abnormal events. The size of each frame in video frame sequence I is standardized to H×W, where H and W represent the height and width of the image, respectively. In this embodiment, both H and W are set to 256. Simultaneously, optical flow images are extracted from every two adjacent video frames to generate the corresponding optical flow sequence F.

[0092] Step 1.2: Divide the video frame sequence and construct the corresponding optical flow frame sequence. Each time, select a video segment {I} that is t consecutive frames prior to the current time. r ,I r+1 ,…,I r+t-1} and its corresponding optical flow segment {F r ,F r+1 ,…,F r+t-2}, which serve as the input data for the appearance and motion branches in the model, respectively. r represents the starting frame index of the current training video segment. In this embodiment, r = {1, 6, 11, ...}, and the step size is t = 5.

[0093] S2, Training Phase:

[0094] Step 2.1: Construct a video anomaly detection model based on the fusion of global and local features. The overall architecture consists of an appearance branch and a motion branch. Based on different modalities of the input data, the appearance branch and motion branch networks extract corresponding features respectively. Each branch consists of an encoder and a decoder, responsible for feature encoding extraction and decoding reconstruction, respectively.

[0095] Step 2.2: The model uses the video frame sequence as input to the encoder in the appearance branch and the optical flow frame sequence as input to the encoder in the motion branch, extracting the corresponding appearance and motion feature codes respectively. In each iteration, {I} is selected from the input video segments. r ,I r+1 ,…,I r+t-2} and the corresponding optical flow segment {F r ,F r+1 ,…,F r+t-3}, respectively input to the appearance encoder E A and motion encoder E M To extract appearance and motion features from the video, where I r and F r These are the starting index frames of the input training segment. Simultaneously, I... r+t-1 and F r+t-2 These are used as appearance and optical flow labels for the real frames, respectively, for subsequent prediction error calculation.

[0096] Each branch's encoder consists of three corresponding stacked global and local modules. The global path extracts coarse global information from the input image sequence, such as shape and spatial distribution; the local path focuses on capturing detailed local information, such as texture and edge features. Finally, information is exchanged through a top-down modulator to effectively fuse global and local features. Each layer progressively models appearance and motion patterns in normal events, thereby accurately representing the spatiotemporal regularities of the video sequence.

[0097] After the encoder extracts features, E A With E M Based on the correlation between appearance patterns and motion patterns, the final appearance feature codes and motion feature codes are generated and input into the corresponding appearance decoder D. A With motion decoder D M Perform decoding.

[0098] Step 2.3: The appearance decoder and motion decoder decode and reconstruct the image structure layer by layer, generating a predicted frame image that matches the corresponding pattern. Decoder D A and D M By performing layer-by-layer decoding, the feature map is gradually enlarged and the spatial structure is restored, generating a predicted image that matches the normal appearance and motion pattern.

[0099] Step 2.4: Calculate the joint loss function for appearance and motion, backpropagate to update network parameters, and optimize the encoders and decoders of both branches. To achieve efficient model training, this embodiment designs a joint optimization objective function composed of multiple loss functions to simultaneously optimize the encoders and decoders of the appearance and motion branches. This objective function includes four sub-loss terms, namely the appearance prediction loss. Appearance Gradient loss Grad Motion prediction loss Motion Optical flow loss Flow The appearance codec is optimized using appearance prediction and gradient loss in the appearance branch. The appearance prediction loss minimizes the pixel difference between the predicted and actual frames and is defined as follows:

[0100]

[0101] in, I represents the predicted frame generated by the model. r+t-1 Represents the actual frame at the corresponding time. It represents the l2 norm (squared Euclidean distance), used to measure pixel-level reconstruction error.

[0102] Gradient loss constrains the gradient difference between the predicted frame and the actual frame in two spatial dimensions, and is defined as:

[0103]

[0104] in, and I r+t-1 (i,j) represent the gray values ​​at pixel position (i,j) in the predicted frame and the real frame, respectively, and ||·||1 represents the l1 norm.

[0105] In the motion branch, the predictive loss L is also used. M To constrain the L2 distance between the predicted optical flow and the actual motion field, it is defined as:

[0106]

[0107] in, F represents the predicted optical flow map. r+t-2 This represents the true optical flow diagram.

[0108] Optical flow is an excellent method for capturing temporal motion information from video clips. The optical flow loss is defined as:

[0109]

[0110] The complete loss function considers all the above loss functions and combines them as follows:

[0111] Loss=λ A Loss Appearance +λ G Loss Grad +λ M Loss Motion +λ F Loss Flow (5)

[0112] λ A , λ G , λ M and λ F These are the weights that balance the importance of different loss functions, and are 1, 1, 2, and 1 respectively.

[0113] Repeat steps 2 to 4, iteratively train for R = 50 rounds, continuously optimize the total loss function Loss until the loss value tends to converge, and finally obtain the model trained by this method.

[0114] S3, Testing Phase:

[0115] Step 3.1: Divide the test video into a video frame sequence and an optical flow frame sequence. Similar to the training phase, divide the test video into t consecutive frames preceding the current time step {I...}. v ,I v+1 ,…,I v+t-1} and the corresponding optical flow frame fragment {F v ,F v+1 ,…,F v+t-2 The dimensions are standardized to an H×W input model. v represents the starting frame index of the current test video segment, and I... v and I v These are the starting index frames of the input test segment. In this embodiment, v = {1, 6, 11, ...}, and the step size is t = 5.

[0116] Step 3.2: Input the video segment {I} into the trained model to generate the predicted frame image for the next time step. v ,I v+1 ,…,I v+t-2} and optical flow segment {F v ,F v+1 ,…,F v+t-3 The data is fed into the already trained model to generate the prediction frame for the next time step. and

[0117] Step 3.3: Calculate the anomaly scores for the appearance stream and motion stream of each test video frame.

[0118] Exception score s for appearance branch A (t), based on the difference between the predicted frame and the actual frame, first calculate the Peak Signal-to-Noise Ratio (PSNR) for each frame. PSNR is used to measure the difference between the predicted frame and the actual frame. The pixel difference between the actual frame I and the actual frame I is calculated using the following formula:

[0119]

[0120] in, To predict the maximum pixel value of a frame image, I(x,y) and These are the pixel values ​​of the actual frame and the predicted frame, respectively. After obtaining the Peak Signal-to-Noise Ratio (PSNR) for each frame, the average PSNR value of that segment is calculated.

[0121] Abnormal score s for the motor branch M Similarly, based on the difference between the predicted optical flow and the actual optical flow, the PSNR value of the optical flow is calculated, and the obtained PSNR value is uniformly normalized to the interval [0,1].

[0122] The total anomaly score S(t) is calculated by weighted fusion of the anomaly scores from the external flow and the moving flow:

[0123] S(t)=αs A (t)+βs M (t) (7)

[0124] α and β are weighted hyperparameters used to balance the contributions of the outlier scores of the outlier flow and the motion flow. In this embodiment, α = 0.9 and β = 0.1.

[0125] Step 3.4: Set a threshold. When the anomaly score exceeds this threshold, the current frame is determined to be an anomaly frame. Based on the calculated total anomaly score (S9t), a threshold θ is set to determine whether the current frame is an anomaly. When the total anomaly score exceeds this threshold, the current frame is determined to be an anomaly frame. In this embodiment, θ = 0.5.

[0126] To verify the effectiveness of the method of this invention, it was compared with existing representative methods on several mainstream VAD datasets, including UCSD Ped2, CUHK Avenue, and ShanghaiTech. The detection performance was measured by the AUC metric, and the results are shown in Table 1. In Table 1, bold text indicates the optimal metric, and N / A indicates that no experimental results were provided for this method.

[0127] Table 1

[0128]

[0129] As can be seen from the experimental results in the table above, the method of the present invention achieves better detection results than the prior art on multiple datasets, thus verifying the significant improvement of anomaly detection performance brought about by the global-local feature fusion and dual-stream structure of the present invention.

[0130] like Figure 8 As shown, this embodiment discloses a video anomaly detection system based on a global-local feature dual-stream network for performing the above method, which includes the following modules:

[0131] Data preprocessing module: used to load the original video training and testing datasets; training videos contain only normal events, while testing videos contain both normal and abnormal events; segmenting video frame sequences and constructing corresponding optical flow frame sequences;

[0132] Training Module: Used to construct a video anomaly detection model based on the fusion of global and local features. The video anomaly detection model consists of an appearance branch and a motion branch. According to different modalities of the input data, the appearance branch and motion branch networks extract corresponding features respectively. Each branch consists of an encoder and a decoder, which are responsible for feature encoding extraction and decoding reconstruction respectively. The video anomaly detection model uses the video frame sequence as the input of the encoder in the appearance branch and the optical flow frame sequence as the input of the encoder in the motion branch to extract the corresponding appearance features and motion features respectively. The appearance decoder and motion decoder decode and reconstruct the image structure layer by layer to generate a prediction frame image that matches the corresponding mode. The joint loss function of appearance and motion is calculated, and the network parameters are updated by backpropagation to optimize the encoder and decoder of the two branches.

[0133] The testing module is used to divide the test video into a sequence of video frames and a sequence of optical flow frames; input the data into a pre-trained video anomaly detection model to generate the predicted frame image for the next moment; calculate the anomaly scores of the appearance flow and motion flow for each test video frame; and set a threshold, when the anomaly score exceeds the threshold, the current frame is determined to be an anomaly frame.

[0134] Other aspects of this embodiment can be found in the above method embodiments.

[0135] In summary, this invention discloses a video anomaly detection method and system based on a global-local feature dual-stream network. The invention processes the video to be trained into video frames and optical flow frames, which are then input into appearance and motion encoders respectively to extract appearance and motion features. The encoder is composed of multiple stacked global-local modules, capable of extracting detailed features and global appearance features of the image, and performing feature fusion and enhancement. During the training phase, the appearance and motion branches generate normal event prototype patterns corresponding to appearance and motion, respectively. During the testing phase, the model generates the next-time video frame and optical flow prediction frame, calculates the error between the real frame and the predicted frame to determine the anomaly score of the video frame, and determines whether an anomaly has occurred by setting an appropriate threshold. This invention significantly improves the accuracy and robustness of video anomaly detection by fusing global and local features from video appearance information and motion patterns.

[0136] The above embodiments are intended to illustrate the technical solutions and effects of the present invention, and are not intended to limit the technical solutions themselves. Any equivalent substitutions, modifications, or alterations made by those skilled in the art to the above embodiments without departing from the core ideas and technical principles of the present invention should be considered as covered within the scope of protection of the claims of the present invention.

Claims

1. A video anomaly detection method based on a global-local feature dual-stream network, characterized by: Includes the following steps: S1, Data Preprocessing Stage: Step 1.1: Load the original video training and testing datasets; the training videos contain only normal events, while the testing videos contain both normal and abnormal events; Step 1.2: Divide the video frame sequence and construct the corresponding optical flow frame sequence; S2, Training Phase: Step 2.1: Construct a video anomaly detection model based on the fusion of global and local features. The video anomaly detection model consists of an appearance branch and a motion branch. According to the different modalities of the input data, the appearance branch and motion branch networks extract the corresponding features respectively. Both branches consist of an encoder and a decoder, which are responsible for feature encoding extraction and decoding reconstruction respectively. Step 2.2: The video anomaly detection model uses the video frame sequence as the input to the encoder in the appearance branch and the optical flow frame sequence as the input to the encoder in the motion branch, and extracts the corresponding appearance features and motion feature codes respectively. Step 2.3: The appearance decoder and motion decoder decode and restore the image structure layer by layer, generating a predicted frame image that matches the corresponding pattern; Step 2.4: Calculate the joint loss function for appearance and motion, backpropagate to update network parameters, and optimize the encoder and decoder of the two branches; S3, Testing Phase: Step 3.1: Divide the test video into a video frame sequence and an optical flow frame sequence; Step 3.2: Input the pre-trained video anomaly detection model to generate the predicted frame image for the next time step; Step 3.3: Calculate the anomaly scores for the appearance stream and motion stream of each test video frame; Step 3.4: Set a threshold. When the abnormal score exceeds the threshold, the current frame is determined to be an abnormal frame.

2. The video anomaly detection method based on a global-local feature dual-stream network as described in claim 1, characterized in that, In step 1.1, the size of each frame in the video frame sequence I is standardized to H×W, where H and W represent the height and width of the image, respectively; at the same time, optical flow images are extracted from every two adjacent video frames to generate the corresponding optical flow sequence F.

3. The video anomaly detection method based on a global-local feature dual-stream network as described in claim 2, characterized in that, In step 1.2, each time a video segment {I} from the previous t frames is selected. r ,I r+1 ,…,I r+t-1 } and its corresponding optical flow segment {F r ,F r+1 ,…,F r+t-2 }, which serve as the input data for the appearance branch and motion branch in the video anomaly detection model, respectively; where r represents the starting frame index of the current training video segment.

4. The video anomaly detection method based on a global-local feature dual-stream network as described in claim 3, characterized in that, In step 2.2, during each iteration, {I} is selected from the input video clips. r ,I r+1 ,…,I r+t-2 } and the corresponding optical flow segment {F r ,F r+1 ,…,F r+t-3 }, respectively input to the appearance branch encoder E A and motion branch encoder E M To extract appearance and motion features from the video, where I r and F r These are the starting index frames of the input training segment; simultaneously, I... r+t-1 and F r+t-2 These serve as the appearance and optical flow labels for the actual frame, respectively. Both branches of the encoder consist of three corresponding global and local modules stacked together. The global path is responsible for extracting global information from the input image sequence, while the local path is used to capture local information. Finally, information is exchanged through a top-down modulator to effectively fuse global and local features. Each module progressively models the appearance and motion patterns in normal events. After the encoder extracts features, E A With E M Based on the correlation between appearance patterns and motion patterns, the final appearance feature codes and motion feature codes are generated and input into the corresponding appearance decoder D. A With motion decoder D M Perform decoding.

5. The video anomaly detection method based on a global-local feature dual-stream network as described in claim 4, characterized in that, In step 2.3, decoder D A and D M By performing layer-by-layer decoding, the feature map is gradually enlarged and the spatial structure is restored, generating a predicted image that matches the normal appearance and motion pattern.

6. The video anomaly detection method based on a global-local feature dual-stream network as described in claim 5, characterized in that, In step 2.4, a joint optimization objective function composed of multiple loss functions is designed to simultaneously optimize the encoder and decoder for the appearance branch and motion branch. This objective function includes four sub-loss terms, namely the appearance prediction loss. Appearance Gradient loss Grad Motion prediction loss Motion Optical flow loss Flow The appearance codec is optimized using appearance prediction and gradient loss in the appearance branch. The appearance prediction loss minimizes the pixel difference between the predicted and actual frames and is defined as follows: in, I represents the predicted frame generated by the model. r+t-1 Represents the actual frame at the corresponding time. Represents the square of the l2 norm; Gradient loss constrains the gradient difference between the predicted frame and the actual frame in two spatial dimensions, and is defined as: in, and I r+t-1 (i,j) represent the gray values ​​at pixel position (i,j) in the predicted frame and the real frame, respectively, and ||·||1 represents the l1 norm; In the motion branch, the predictive loss L is used. M To constrain the L2 distance between the predicted optical flow and the actual motion field, it is defined as: in, F represents the predicted optical flow map. r+t-2 Represents the true optical flow diagram; Optical flow loss is defined as: The complete loss function is: Loss=λ A Loss Appearance +λ G Loss Grad +λ M Loss Motion +λ F Loss Flow (5) λ A , λ G , λ M and λ F It is a weight that balances the importance of different loss functions; Repeat steps 2.2 to 2.4, iteratively train for R rounds, and continuously optimize the total loss function Loss until the loss value tends to converge, finally obtaining the trained video anomaly detection model.

7. The video anomaly detection method based on a global-local feature dual-stream network as described in claim 6, characterized in that, In step 3.1, the video segments {I} of the consecutive t frames preceding the current moment of the test video are selected. v ,I v+1 ,…,I v+t-1 } and the corresponding optical flow frame fragment {F v ,F v+1 ,…,F v+t-2 The dimensions are uniformly normalized to an H×W input model; v represents the starting frame index of the current test video segment, I v and F v These are the starting index frames of the input test segment.

8. The video anomaly detection method based on a global-local feature dual-stream network as described in claim 7, characterized in that, In step 3.2, the video segment {I v ,I v+1 ,…,I v+t-2 } and optical flow segment {F v ,F v+1 ,…,F v+t-3 The data is fed into a pre-trained video anomaly detection model to generate the predicted frame for the next time step. and 9. The video anomaly detection method based on a global-local feature dual-stream network as described in claim 8, characterized in that, In step 3.3, the anomaly score s for the appearance branch is... A (t), based on the difference between the predicted frame and the actual frame, the peak signal-to-noise ratio (PSNR) of each frame is first calculated using the following formula: in, To predict the maximum pixel value of a frame image, I(x,y) and The pixel values ​​of the actual frame and the predicted frame are respectively; after obtaining the peak signal-to-noise ratio of each frame, the average PSNR value of the segment is calculated; Abnormal score s for the motor branch M (t), based on the difference between the predicted optical flow and the actual optical flow, calculate the PSNR value of the optical flow, and normalize the obtained PSNR value to the interval [0,1]; The total anomaly score S(t) is calculated by weighted fusion of the anomaly scores from the external flow and the moving flow: S(t)=αs A (t)+βs M (t) (7) α and β are weight hyperparameters; In step 3.4, a threshold θ is set based on the obtained total anomaly score S(t) to determine whether the current frame is an anomaly; when the total anomaly score exceeds the threshold, the current frame is determined to be an anomaly frame.

10. A video anomaly detection system based on a global-local feature dual-stream network, used to perform the method as described in any one of claims 1-9, characterized in that, Includes the following modules: Data preprocessing module: used to load the original video training and testing datasets; training videos contain only normal events, while testing videos contain both normal and abnormal events; Divide the video frame sequence and construct the corresponding optical flow frame sequence; Training module: Used to build a video anomaly detection model based on the fusion of global and local features. The video anomaly detection model consists of an appearance branch and a motion branch. According to different modalities of the input data, the appearance branch and motion branch networks extract corresponding features respectively. Each branch consists of an encoder and a decoder, which are responsible for feature encoding and extraction and decoding and reconstruction respectively. The video anomaly detection model uses the video frame sequence as the input of the encoder in the appearance branch and the optical flow frame sequence as the input of the encoder in the motion branch to extract the corresponding appearance features and motion feature encoding respectively. The appearance decoder and motion decoder decode and restore the image structure layer by layer, generating a predicted frame image that matches the corresponding pattern; calculate the joint loss function of appearance and motion, backpropagate to update the network parameters, and optimize the encoder and decoder of the two branches; The testing module is used to divide the test video into a sequence of video frames and a sequence of optical flow frames; input the sequence into a pre-trained video anomaly detection model to generate the predicted frame image for the next time step; and calculate the anomaly scores of the appearance flow and motion flow for each test video frame. Set a threshold; when the abnormal score exceeds the threshold, the current frame is determined to be an abnormal frame.