Monitoring video anomaly detection method based on interactive learning dual network

Through the unsupervised video anomaly detection method based on interactive learning dual networks, the problem of relying on labeled data in the existing technology is solved, efficient and unsupervised abnormality detection is achieved, and the intelligence level of the monitoring system is improved.

CN120126046APending Publication Date: 2025-06-10GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510175102.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing monitoring video abnormality detection methods rely on a large amount of labeled data and are difficult to widely use in actual monitoring scenarios. The traditional manual observation is low efficiency, strong subjectivity, and poor real-time performance.

Method used

Unsupervised video anomaly detection method based on interactive learning dual networks is adopted. Through interactive collaborative training between generator and discriminator, pseudo-labels are generated for training discriminative networks. The pseudo-labels generated by the discriminative network are used to improve the generation network and realize self-optimization of the model.

Benefits of technology

Without any labeled data, efficient abnormal detection is achieved, detection efficiency and accuracy are improved, dependence on external labels is reduced, and the intelligence level of monitoring system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126046A_ABST
    Figure CN120126046A_ABST
Patent Text Reader

Abstract

The invention discloses a monitoring video anomaly detection method based on an interactive learning dual network, and the method comprises the following steps: video collection: collecting different types of monitoring videos of a plurality of scenes through employing a monitoring camera, and taking the collected videos as a training set; and feature extraction: dividing each video into a video clip f according to every 16 continuous video frames, and extracting vector representation of the video clip by using a pre-trained neural network model on the public data set. Through interactive cooperative training of two networks of a generator and a discriminator, the method can be used for extracting the vector representation of the video clip under the condition of not needing any annotation data; the method realizes efficient anomaly detection, is excellent in performance on a plurality of data sets, has a wide application prospect, can provide powerful technical support for an intelligent monitoring system, is applied to the fields of traffic management, public place safety, industrial production line monitoring and the like, helps monitoring personnel to timely discover and handle abnormal events, and improves the safety of the monitoring personnel. And the intelligent level of the monitoring system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of abnormal detection of surveillance videos, and specifically relates to a method for abnormal detection of surveillance videos based on an interactive learning dual network. Background Art

[0002] The task of abnormal detection of surveillance videos is an important research direction in the field of computer vision, aiming to automatically identify abnormal situations that are significantly different from normal behaviors or events in surveillance videos. With the acceleration of the urbanization process and the increasing demand for public safety, video surveillance systems have been widely applied in fields such as traffic management, public place safety, and industrial production line monitoring. However, traditional surveillance systems mainly rely on manual observation and manual analysis, and this method has many defects. First, the efficiency of manual observation is low. Especially in large-scale surveillance scenarios, surveillance personnel need to observe multiple cameras simultaneously, which is prone to visual fatigue, resulting in missed detections and false detections. Second, the subjectivity of manual observation is strong. Different surveillance personnel may have inconsistent judgment criteria for abnormal events, resulting in instability of detection results. Finally, the real-time performance of manual observation is poor. Surveillance personnel cannot observe videos continuously all day long, and it is difficult to detect and handle abnormal events in a timely manner. Therefore, how to achieve automated and intelligent video abnormal detection has become an urgent problem to be solved.

[0003] In recent years, with the rapid development of deep learning technology, video abnormal detection methods based on deep learning have made remarkable progress. Deep learning models can automatically learn spatio-temporal features in videos, thus better capturing the patterns of abnormal events. Compared with traditional rule-based or handcrafted feature-based methods, deep learning-based methods have the following advantages: 1. Deep learning models can automatically extract high-level features in videos, avoiding the cumbersome process of handcrafting features; 2. Deep learning models have strong generalization ability and can handle complex abnormal events, especially when the types and manifestations of abnormal events are diverse; 3. Deep learning models can continuously improve the detection accuracy through training with large-scale data and are applicable to actual surveillance scenarios.

[0004] Although deep learning methods have made remarkable progress in video anomaly detection, they still face some challenges. Among them, deep learning methods usually require a large amount of labeled data for training, while labeled data in video anomaly detection is difficult to obtain, especially in large-scale surveillance scenarios where the labeling cost is extremely high. To address the problem of difficult-to-obtain labeled data, researchers have proposed weakly supervised learning and semi-supervised learning methods, which train by using video-level labels or unlabeled normal video data, thus greatly reducing the labeling workload. Although these improved methods have alleviated the labeling cost problem in video anomaly detection to a certain extent, there are still some limitations. For example, weakly supervised learning methods, although reducing the labeling workload, still require video-level labels and are difficult to accurately locate anomaly events. Semi-supervised learning methods still need to manually check all events in the video to construct a training set of fully normal video data.

[0005] After retrieval, the patent with publication number CN114842371A discloses a semi-supervised video anomaly detection method. This invention only collects normal videos as training data and trains a locally sensitive hashing model through self-supervised learning. The trained locally sensitive hashing model can be used for video anomaly detection. However, how to achieve efficient and accurate video anomaly detection without relying on a large amount of labeled data is still an urgent problem to be solved.

[0006] In summary, with the explosive growth of surveillance video data, traditional methods based on manual observation and manual analysis can no longer meet the actual needs. Although deep learning-based methods have improved the detection efficiency to a certain extent, they still rely on a large amount of labeled data and are difficult to be widely applied in actual surveillance scenarios. Therefore, there is an urgent need for a more economical and effective completely unsupervised video anomaly detection method that uses videos without any labels and supervision information to train the model and deploy it to detect anomaly events in surveillance, so as to achieve a more automated, low-cost, and high-efficiency surveillance system.

[0007] To solve this problem, the present invention proposes an unsupervised surveillance video anomaly detection method based on an interactive learning dual network. Summary of the Invention

[0008] The object of the present invention is to provide a method for abnormal detection of surveillance videos based on an interactive learning dual network. Through the interactive collaborative training of two networks, namely a generator and a discriminator, it can achieve efficient abnormal detection without any labeled data, perform well on multiple datasets, have broad application prospects, provide strong technical support for intelligent surveillance systems, and be applied to fields such as traffic management, public place safety, and industrial production line monitoring. It helps surveillance personnel to timely discover and handle abnormal events, improves the intelligence level of the surveillance system, and solves the problems mentioned in the background technology.

[0009] To solve the above problems, the present invention provides a technical solution:

[0010] A method for abnormal detection of surveillance videos based on an interactive learning dual network, comprising the following steps:

[0011] S1. Video acquisition: Use surveillance cameras to collect different types of surveillance videos of multiple scenes, and use the collected videos as a training set;

[0012] S2. Feature extraction: Divide each video into a video segment f according to every 16 consecutive video frames, and use a neural network model pre-trained on a public dataset to extract the vector representation of the video segment, obtaining abstract high-dimensional video segment features;

[0013] S3. Build a dual network model: A dual network is composed of a generator network G and a discriminator network D;

[0014] S4. Pre-training and starting of the generator network G: Use some data after data cleaning to pre-train and start the generator network G;

[0015] S5. Interactive learning of the dual network: The generator network G and the discriminator network D are alternately trained to achieve interactive learning, and this is repeated iteratively until the training is completed;

[0016] S6. Model deployment and inference: The trained model is deployed to the corresponding computer service platform. For the video stream data captured by the actual surveillance camera, it is input into the deep learning network model for detection to determine whether there are abnormal events, and the detection results are obtained.

[0017] Preferably, the surveillance videos collected in S1 include indoor and outdoor scenes under day and night conditions, the total number of collected videos is 2000, the average frame rate of the videos is 30 FPS, and the video size is 240x320.

[0018] Preferably, the network model in S2 is the I3D model, and the dimension of the extracted feature vector is 2048 dimensions.

[0019] Preferably, the step of pre-training the generation network G in S4 is as follows:

[0020] S41. Training data cleaning: Use the L 2 distance between continuous feature vectors to clean the training data set for pre-training of G. When it is the case, use the video segment feature vector for pre-training;

[0021] S42. Generation network pre-training start: The cleaned training data is used for the reconstruction pre-training of the generation network G. The mean square error loss function of the loss function is adopted for 25 pre-training iterations.

[0022] Preferably, the interactive learning of the dual network in S5 includes: the generation network G guiding the discriminant network D for training, the discriminant network D assisting the generation network D for optimization, and the iterative alternating training of the dual network.

[0023] Preferably, the step of the generation network G guiding the discriminant network D for training is as follows;

[0024] A1. The generation network generates pseudo-labels:

[0025] A1-1. After the pre-training starts, the generation network G uses the video segment f i,j feature as input data and generates the reconstructed data of the feature as output. The reconstruction loss function of the segment is:

[0026]

[0027] A1-2. Regard the video segments with reconstruction error exceeding the error threshold as abnormal segments and mark the pseudo-label y = 1, otherwise regard them as normal video segments and mark the pseudo-label y = 0:

[0028]

[0029] A1-3. Infer all the training data video segments to obtain the pseudo-labels of the segments for subsequent guiding the discriminant network training;

[0030] A2. Discriminant network training based on pseudo-labels:

[0031] A2-1. The generation network G generates pseudo-labels for all video segments and conducts supervised training on the discriminant network. The loss function used for training is the binary classification loss function BCE-loss, and the mathematical definition of the loss function is:

[0032] Preferably, the step of the discriminant network assisting the generation network for optimization is as follows:

[0033] B1. The discriminative network generates pseudo-labels:

[0034] B1-1. Use all video clip features f i,j as the input of the discriminative network for inference and generate the anomaly probability of these features

[0035] B1-2. Consider the video clips with anomaly probability exceeding the error threshold as abnormal clips and mark the pseudo-label y = 1, otherwise consider them as normal video clips and mark the pseudo-label y = 0:

[0036]

[0037] B1-3. Perform inference on all training data video clips to obtain the pseudo-labels of the clips for subsequent guiding the training of the discriminative network;

[0038] B2. Optimization of the generative network based on pseudo-labels:

[0039] B2-1. Use the feature vectors of all predicted normal video clips with pseudo-label y = 0 to optimize the reconstruction network of the autoencoder.

[0040] Preferably, the steps of iterative alternating training of the dual network are:

[0041] C1. Forward propagation of the network model:

[0042] C1-1. Input the 2028-dimensional clip features into the network, and the reconstructed clip features f are output by the reconstructed word encoder model i g ,j to complete one forward propagation;

[0043] C1-2. Input the 2028-dimensional clip features into the network, and the fully connected discriminator generates the anomaly probability of the predicted clip to complete one forward propagation;

[0044] C2. Backward propagation of the network model:

[0045] C2-1. Calculate the loss function: Calculate the deviation using the pseudo-label y generated by the generative network G and the anomaly probability of the clip predicted by the forward propagation of the model for calculating the loss function;

[0046] C2-2. Use the pseudo-label y generated by the discriminative network D, select the feature of the video clip predicted as normal and the reconstructed clip feature output by the forward propagation of the model, and calculate the deviation for calculating the loss function;

[0047] C3. Iterative update of network weights: Calculate the partial derivatives of each variable in the loss function L obtained by the generator network and the discriminator network, and calculate the corresponding gradient value g;

[0048] C4. Iterative alternating training: Use the collaborative training interactive learning method for the forward and backward propagation processes of the above dual network, and iteratively train on the training dataset for 25 epochs to obtain a trained dual network model.

[0049] Preferably, the specific steps of S6 are as follows:

[0050] S61. Deploy the trained discriminator network D to the computer;

[0051] S62. Divide and extract features from the video stream transmitted by the monitoring camera through S2, and input the extracted features into the deployed model;

[0052] S63. The model infers the segment features to obtain the anomaly probability of each segment and determine whether the segment is abnormal.

[0053] The beneficial effects of the present invention are:

[0054] 1. Advantages of unsupervised learning: Provide an unsupervised monitoring video anomaly detection method based on an interactive learning dual network. The video data used to train the model does not require any manual inspection and annotation at all. The trained model can realize the automatic detection of abnormal events in the monitoring video stream data. Compared with the traditional method relying on manual visual inspection or feature detection based on constructing manual features, it can achieve more intelligent and automatic high-speed detection, greatly improving the detection efficiency and detection accuracy;

[0055] 2. Advantages of network structure: The generator network adopts an autoencoder structure and trains the network by minimizing the reconstruction loss, enabling it to better reconstruct the features of normal events, while the reconstruction of the features of abnormal events is poor. This design enables the model to effectively distinguish normal and abnormal events. The discriminator network adopts a fully connected classifier and is trained through the binary cross-entropy loss function, which can effectively estimate the probability that the input video segment feature vector is abnormal. This design enables the model to output high-precision anomaly detection results;

[0056] 3. Advantages of Dual Network Interactive Learning: Through the alternating training of the generator network (G) and the discriminator network (D), the pseudo-labels generated by the generator network are used to train the discriminator network, and the pseudo-labels generated by the discriminator network are used to improve the generator network. This collaborative training mechanism enables the two networks to promote each other and gradually improve their performance. The generator network generates pseudo-labels through the reconstruction error, and the discriminator network generates pseudo-labels through probability values. This mechanism of automatically generating pseudo-labels reduces the dependence on external annotation and can dynamically adjust the training process of the model, improving the robustness of the model. Description of the Drawings

[0057] For ease of explanation, the present invention will be described in detail by the following specific embodiments and drawings.

[0058] Figure 1 is the overall flowchart of the present invention;

[0059] Figure 2 is the schematic diagram of the dual network model structure of the present invention;

[0060] Figure 3 is the schematic diagram of the dual network interactive learning of the present invention. Detailed Embodiments

[0061] As Figures 1-3 shown, the following technical solutions are adopted in this specific embodiment:

[0062] Example:

[0063] A method for abnormal detection of surveillance videos based on an interactive learning dual network includes the following steps:

[0064] S1. Video acquisition: Use a surveillance camera to collect different types of surveillance videos of multiple scenes, and use the collected videos as a training set. The videos include normal videos and abnormal videos, but there is no need for manual inspection and marking of the video categories.

[0065] S2. Feature extraction: Divide each video into a video segment f according to every 16 consecutive video frames to represent a video event, that is, the smallest unit of a video event is a video segment. Then, use a neural network model pre-trained on a public dataset to extract the vector representation of the video segment and obtain the abstract high-dimensional video segment features.

[0066] S3. Build a dual network model: A dual network is composed of a generator network G and a discriminator network D, which is used for the training and inference deployment of video abnormal detection.

[0067] S4. Pre-training start of the generator network G: Use some data after data cleaning to pre-train and start the generator network G to enable it to have the initial ability of video abnormal detection.

[0068] S5. Interactive learning of the dual network: The generation network G and the discriminant network D are alternately trained to achieve interactive learning. After startup, the generation network G generates pseudo-labels to guide the training of the discriminant network D, and then the discriminant network D generates pseudo-labels to assist in optimizing the generation network D. This process is repeated iteratively until the training is completed.

[0069] S6. Model deployment and inference: The trained model is deployed to the corresponding computer service platform. For the video stream data captured by the actual monitoring camera, it is input into the deep learning network model for detection to determine whether there is an abnormal event and obtain the detection result.

[0070] Among them, the monitored videos collected in S1 include indoor and outdoor scenes under day and night conditions. The total number of collected videos is 2000, the average frame rate of the videos is 30 FPS, and the video size is 240x320.

[0071] Among them, the network model in S2 is the I3D model, and the dimension of the extracted feature vector is 2048. At the same time, every 16 frames are divided into a video segment, and the remaining parts that do not meet 16 frames are directly discarded. Then, a pre-trained public model is used as the feature extractor to extract the video segment features.

[0072] Among them, since the generation network adopts an autoencoder structure, it can capture the main representation of the training data. However, the abnormal features in the training data will interfere with the reconstruction representation of the normal features, thus unable to provide an effective startup for the generation network G. Therefore, it is necessary to clean the training data, and then use some of the cleaned data for the pre-training of the startup of the generation network G.

[0073] The steps for the pre-training of the generation network G are as follows:

[0074] S41. Cleaning of training data: From the perspective of segment features, compared with normal video segments, the feature vectors of abnormal segments often have higher feature amplitude fluctuation values. Because abnormal events often have intense manifestations, which show high amplitudes in adjacent segments in the feature vectors. Therefore, the L 2 distance between consecutive feature vectors is used to clean the training data set for the pre-training of G. Only when is satisfied, the video segment feature vectors are used for pre-training, where the superscripts t and t + 1 indicate the time order of the features in the video, and D th is the threshold, which is set to 0.7 in the present invention.

[0075] S42. Generation network pre-training start: The cleaned training data is used for the reconstruction pre-training of the generation network G. The mean square error loss function of the loss function is adopted, and 25 pre-training iterations are carried out. After pre-training, the generation network will have the ability to reconstruct normal feature segments well but have a poor reconstruction of abnormal feature segments, showing a large reconstruction error.

[0076] Among them, the interactive learning of the dual network in S5 includes: the generation network G guiding the training of the discriminant network D, the discriminant network D assisting the optimization of the generation network D, and the iterative alternating training of the dual network.

[0077] Among them, the steps for the generation network G to guide the training of the discriminant network D are as follows;

[0078] A1. The generation network generates pseudo-labels:

[0079] A1-1. After the pre-training start, the generation network G has the ability to reconstruct the features of normal segments. Therefore, G takes the video segment f i,j features as input data and generates the reconstructed data of the features as output. The reconstruction loss function of the segment is:

[0080]

[0081] A1-2. The video segments with a reconstruction error exceeding the error threshold are regarded as abnormal segments and marked with the pseudo-label y = 1, otherwise they are regarded as normal video segments and marked with the pseudo-label y = 0:

[0082]

[0083] A1-3. Infer all the training data video segments to obtain the pseudo-labels of the segments, which are used to guide the subsequent training of the discriminant network;

[0084] A2. Discriminant network training based on pseudo-labels:

[0085] A2-1. The generation network G generates pseudo-labels for all video segments and conducts supervised training on the discriminant network. The loss function used for training is the binary classification loss function BCE-loss, and the mathematical definition of the loss function is:

[0086] Among them, is the output of the discriminant network D, representing the probability of predicting that the video segment is abnormal.

[0087] Among them, the steps for the discriminant network to assist the optimization of the generation network are as follows:

[0088] B1. The discriminant network generates pseudo-labels:

[0089] B1-1. After being trained, the discriminative network D has the ability to predict the features of abnormal segments. Therefore, all video segments f i,j features are used as the input of the discriminative network for inference, and the abnormal probabilities of these features are generated

[0090] B1-2. Video segments with abnormal probabilities exceeding the error threshold are regarded as abnormal segments and marked with the pseudo-label y = 1, otherwise they are regarded as normal video segments and marked with the pseudo-label y = 0:

[0091]

[0092] B1-3. All training data video segments are inferred to obtain the pseudo-labels of the segments, which are used for subsequent training of the discriminative network;

[0093] B2. Optimization of the generative network based on pseudo-labels:

[0094] B2-1. The feature vectors of all predicted normal video segments with the pseudo-label y = 0 are used to optimize the reconstruction network of the autoencoder, enabling it to more accurately predict normal segments while generating a large reconstruction error for abnormal segment predictions.

[0095] Among them, the steps of the dual network iterative alternating training are as follows:

[0096] C1. Forward propagation of the network model:

[0097] C1-1. For the generative network G, the 2028-dimensional segment features are input into the network, and the reconstructed segment features are output by the reconstructed word encoder model to complete one forward propagation;

[0098] C1-2. For the discriminative network D, the 2028-dimensional segment features are input into the network, and the fully connected discriminator generates the abnormal probability of the predicted segment to complete one forward propagation;

[0099] C2. Backward propagation of the network model:

[0100] C2-1. Calculate the loss function: For the discriminative network D, the deviation is calculated using the pseudo-label y generated by the generative network G and the abnormal probability of the segment predicted by the forward propagation of the model to calculate the loss function, which is used to update and optimize the discriminative network model;

[0101] C2-2. For the generative network G, using the pseudo-label y generated by the discriminative network D, select the features of the video segments predicted as normal and the reconstructed segment features output by the forward propagation of the model, calculate the deviation to calculate the loss function, which is used to update and optimize the discriminative network model;

[0102] C3. Iterative update of network weights: Calculate the partial derivatives of each variable in the loss function L obtained by the generator network and the discriminator network, and calculate the corresponding gradient value g. In this embodiment, the SGD (Stochastic Gradient Descent) optimization method is used to optimize and update the weight parameters in the network model. The update formula is: w t+1 = w t - ηg t ;

[0103] where η is the learning rate, which takes the value of 0.002 in the present invention, and g t is the average gradient value of a batch of samples. In this embodiment, the number of samples in one batch is set to 4096;

[0104] C4. Iterative alternating training: For the training of the generator network and the discriminator network in the dual model, it mainly includes two parts: the forward propagation of the network model and the backward propagation of the network model. The dual network sequentially repeats steps A1 and A2 as an overall iteration in sequence. The initial pseudo-labels are generated by the generator network G started by pre-training, while the pseudo-labels generated by G in subsequent iterations are generated by the generator network optimized by interactive learning;

[0105] Each round of the generator network optimized by A2 will generate pseudo-labels for the data for interactive learning in training. Finally, the forward propagation and backward propagation processes of the above dual network are used in the way of collaborative training and interactive learning, and the trained dual network model is obtained by iteratively training 25 epochs on the training dataset.

[0106] The specific steps of S6 are as follows:

[0107] S61. Deploy the trained discriminator network D to the computer;

[0108] S62. Divide the video stream transmitted by the monitoring camera into segments and extract features through S2, and input the extracted features into the deployed model;

[0109] S63. The model infers the anomaly probability of each segment from the segment features and determines whether the segment is abnormal. When the anomaly probability of the segment is greater than 0.5, it is detected as an abnormal event, and when it is lower than 0.5, it is detected as a normal event.

[0110] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A surveillance video anomaly detection method based on interactive learning dual network, characterized in that: The steps include: S1. Video acquisition: Use surveillance cameras to collect different types of surveillance videos of multiple scenes and use the collected videos as training sets; S2, feature extraction: Divide each video into a video segment f according to every 16 consecutive video frames, and use the neural network model pre-trained on the public dataset to extract the vector representation of the video segment to obtain the abstract high-dimensional video segment features; S3. Build a dual network model: A dual network is formed by a generative network G and a discriminative network D; S4, pre-training and starting of the generated network G: using part of the data after data cleaning to pre-train and start the generated network G; S5. Interactive learning of dual networks: The generator network G and the discriminator network D are trained alternately to achieve interactive learning, and the training is repeated until the training is completed; S6. Model deployment and reasoning: The trained model is deployed to the corresponding computer service platform. The video stream data captured by the actual surveillance camera is input into the deep learning network model for detection to determine whether there are abnormal events and obtain the detection results.

2. According to claim 1, a monitoring video anomaly detection method based on interactive learning dual network is characterized in that: The surveillance videos collected in S1 include indoor scenes and outdoor scenes under daytime and nighttime conditions. The total number of collected videos is 2000, the average video frame rate is 30FPS, and the video size is 240x320.

3. According to claim 1, a monitoring video anomaly detection method based on interactive learning dual network is characterized in that: The network model in S2 is an I3D model, and the dimension of the extracted feature vector is 2048 dimensions.

4. According to claim 1, a monitoring video anomaly detection method based on interactive learning dual network is characterized in that: The steps of pre-training the generated network G in S4 are: S41, training data cleaning: Use the L2 distance between consecutive feature vectors to clean the training data set for pre-training of G. When using the video segment feature vector Conduct pre-training; S42, generating network pre-training start: the cleaned training data is used to generate reconstruction pre-training of the network G, using the mean square error loss function of the loss function, and performing 25 pre-training iterations.

5. The monitoring video anomaly detection method based on interactive learning dual network according to claim 1 is characterized in that: The interactive learning of the dual network in S5 includes: the generation network G guides the training of the discriminant network D, the discriminant network D assists the optimization of the generation network D, and the dual network iterative alternating training.

6. The monitoring video anomaly detection method based on interactive learning dual network according to claim 5 is characterized in that: The steps of generating the network G to guide the training of the discriminant network D are as follows: A1. Generate pseudo labels through the generative network: A1-1. After pre-training starts, the generative network G will generate the video clip f i,j The feature is used as input data and the reconstruction data of the feature is generated as output. The reconstruction loss function of the fragment is: A1-2. The reconstruction error exceeds the error threshold The video clip is regarded as an abnormal clip and marked with a pseudo label y = 1, otherwise it is regarded as a normal video clip and marked with a pseudo label y = 0: A1-3, infer all the training data video clips to obtain pseudo labels for the clips, which are used to guide the discriminant network training in the future; A2. Pseudo-label based discriminant network training: A2-1. Generate pseudo labels for all video clips through the generative network G. Perform supervised training on the discriminative network. The loss function used in the training is the binary classification loss function BCE-loss. The mathematical definition of the loss function is 7. The monitoring video anomaly detection method based on interactive learning dual network according to claim 5 is characterized in that: The steps of the discriminant network assisting in generating network optimization are: B1. Discriminant network generates pseudo labels: B1-1, all video clips f i,j The features are used as inputs to the discriminant network for inference, and the abnormal probability y of these features is generated i ; B1-2. The abnormal probability exceeds the error threshold The video clip is regarded as an abnormal clip and marked with a pseudo label y=1, otherwise it is regarded as a normal video clip and marked with a pseudo label y=0: B1-3, infer all the training data video clips to obtain pseudo labels for the clips, which are used to guide the discriminant network training in the future; B2. Generative network optimization based on pseudo labels: B2-1. All feature vectors of the predicted normal video clips with the pseudo label y=0 are used to optimize the reconstruction network of the autoencoder.

8. The monitoring video anomaly detection method based on interactive learning dual network according to claim 5 is characterized in that: The steps of iterative alternating training of the dual network are: C1. Forward propagation of the network model: C1-1. Input the 2028-dimensional segment feature into the network, and reconstruct the segment feature f output by the character encoder model i g ,j Complete one forward propagation; C1-2, input the 2028-dimensional segment features into the network, and the fully connected discriminator generates the abnormal probability y of the predicted segment, completing a forward propagation; C2. Backward propagation of the network model: C2-1. Calculate the loss function: Use the pseudo-label y generated by the generative network G and the probability of abnormality y of the segment predicted by the forward propagation of the model to calculate the deviation and calculate the loss function; C2-2, using the pseudo-label y generated by the discriminant network D, select the video segment features predicted to be normal and the reconstructed segment features output by the model forward propagation, calculate the deviation and calculate the loss function; C3, iterative update of network weights: calculate the partial derivatives of each variable in the loss function L calculated by the generator network and the discriminator network, and calculate the corresponding gradient value g; C4, iterative alternating training: The forward propagation and back propagation processes of the above dual network are trained in a collaborative training interactive learning manner, and 25 epochs are iteratively trained on the training data set to obtain a trained dual network model.

9. The monitoring video anomaly detection method based on interactive learning dual network according to claim 1 is characterized in that: The specific steps of S6 are as follows: S61, deploying the trained discriminant network D to a computer; S62, dividing the video stream transmitted by the surveillance camera into segments and extracting features through S2, and inputting the extracted features into the deployed model; S63. The model infers the fragment features to obtain the abnormal probability y of each fragment, and determines whether the fragment is abnormal.

Citation Information

Patent Citations

  • Unsupervised video anomaly detection method

    CN114842371A