Anomaly Detection Method and System for Surveillance Videos Based on Adversarial Learning
The adversarial learning-based method for video anomaly detection improves detection accuracy and efficiency by using a convolutional neural network with a memory module for feature comparison and reconstruction, addressing the limitations of existing methods and enabling effective real-time abnormal event identification.
Patent Information
- Application Number
- CN202211381511.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-04
AI Technical Summary
The detection accuracy and efficiency of existing surveillance video abnormal detection methods are not high. Traditional methods have low effects in special scenarios, making it difficult to effectively identify abnormal events.
The monitoring video anomaly detection method based on adversarial learning is adopted, and the prediction network and memory module network are constructed, and the classification loss model is constructed, and the abnormality of video frames is judged by the discriminator model.
It improves the accuracy and efficiency of monitoring video abnormality detection, can better extract the characteristic information of video sample frames, adapt to various scenarios, and realize real-time abnormality detection.
Smart Images

Figure CN115909144B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video anomaly detection, and particularly to a method and system for monitoring video anomaly detection based on adversarial learning. Background Art
[0002] Thanks to the digitization and informatization of modern society, as well as the improvement of people's public safety awareness, the monitoring network covers most of people's living and working environments. Monitoring devices are widely used in various corners of the city, especially in various places with large flows of people, such as shopping malls, hospitals, schools, streets, communities, airports, stations, etc. These cameras generate a vast amount of video data. Detecting anomalies in human behaviors in these video data can effectively monitor and collect evidence for abnormal situations such as illegal intrusion, robbery, theft, stampede, traffic accidents, etc. The vigorous development and wide application of video surveillance technology have played a huge role in maintaining economic prosperity.
[0003] Traditional video surveillance systems are passive systems, and their main functions are to record, store, and playback currently occurring events. However, in the absence of human supervision, traditional video surveillance systems cannot identify and promptly alarm some abnormal events such as fights, robberies, fires, etc. Relying solely on human eyes to observe surveillance videos will consume a large amount of human and material resources, and as the working hours increase, people's energy will decline to varying degrees, making it easy to misdetect or miss abnormal events. Therefore, introducing intelligent monitoring video anomaly detection technology into the monitoring system is an inevitable trend for future development.
[0004] With the great success achieved by deep learning algorithms in the field of computer vision in recent years, algorithms based on deep neural networks have gradually been applied to video anomaly detection tasks. Two types of methods have been derived, namely anomaly detection methods based on current frame reconstruction and future frame prediction. The method based on reconstructing the current frame distinguishes abnormal frames from normal frames based on the idea that the reconstruction error of abnormal frames is large. The method based on future frame prediction, on the other hand, makes a decision on the normality of future frames based on the idea that anomalies are difficult to predict. Although these two different anomaly detection methods have achieved some results, the idea of using proxy tasks to achieve anomaly detection has an inherent defect. That is, whether it is the reconstruction method or the prediction method, its essence is to output an image that is as similar as possible to the real frame. When the network is trained very well, due to the influence of its ideological defect, the discrimination between normal frames and abnormal frames of these two types of methods is not necessarily high, and the effect is relatively low in some special scenarios. Summary of the Invention
[0005] The main objective of the present invention is to provide a method and system for abnormal detection of surveillance videos based on adversarial learning, aiming to solve the technical problems of low detection accuracy and efficiency of existing methods for abnormal detection of surveillance videos.
[0006] To achieve the above objective, the present invention provides a method for abnormal detection of surveillance videos based on adversarial learning, and the method includes the following steps:
[0007] S1: Obtain video sample frames arranged in chronological order. Taking each video sample frame as a starting point, select k video sample frames in chronological order to construct a video sample frame group as the input of the prediction network.
[0008] S2: Based on a convolutional neural network, taking the video sample frame as the input and the corresponding feature map as the output, construct a prediction network.
[0009] S3: Taking the feature map as the input of the memory module network and the normal sample feature map of the same scale as the feature map as the output of the memory module network, perform end-to-end adversarial training without supervision.
[0010] S4: Based on the prediction network and the memory module network, construct a to-be-trained model for detecting abnormal video frames. At the same time, based on the participation training of each video sample frame, from the preliminary feature extraction network to the application of the deep feature extraction and classification network, construct a classification loss model by introducing reconstruction, adversarial, and memory losses.
[0011] S5: Based on the video sample frame groups constructed from the video sample frames and the labels corresponding to each video sample frame group respectively, taking the video sample frame as the input and the labels corresponding to each video sample frame group respectively as the output, combined with the classification loss model, train the to-be-trained model for detecting abnormal video frames to obtain a model for detecting abnormal video frames.
[0012] S6: For each video sample frame in each video sample frame group, through the discriminator model, determine the abnormal score of whether each video sample frame in the video sample frame group is normal or abnormal according to the reconstruction loss obtained by the model. Determine the video sample frame with an abnormal score greater than the preset value as an abnormal video frame, otherwise as a normal video frame.
[0013] Optionally, in step S2, the prediction network is a U-Net encoder.
[0014] Optionally, in step S3, the memory module network uses normal event samples during training and adds abnormal samples during testing.
[0015] Optionally, in step S3, the memory module network includes two operations: reading and updating. After obtaining the features of a new normal sample, a reading operation is performed on the memory module network to select the normal sample features that are most similar to itself; the memory module network is updated according to the features of the new normal sample.
[0016] Optionally, step S3 includes:
[0017] For the output of the deep feature extraction and classification network, a feature q of size H×W×C is obtained t . Where H is the height of the feature, W is the width of the feature, and C is the number of channels;
[0018] According to the matching algorithm of the memory module network, the feature p with the highest matching probability is obtained t , and the size is also H ×W×C;
[0019] The queried feature p t is concatenated with the extracted feature q t on the channels to obtain a new feature of size H×W ×2C to update the memory module network.
[0020] Optionally, step S4 specifically includes:
[0021] Send consecutive t frames of normal training samples X = {x1, x2,..., x t} into the prediction network;
[0022] The encoder of the prediction network extracts the features q of the t-frame video frames t , and the prediction network will, according to the similarity between q t and the normal sample features stored in the memory module, read the corresponding p t and concatenate it with q t to obtain the feature (q t , p t ) and update the memory module network;
[0023] Send the feature (q t , p t ) to the decoder of the prediction network, and finally obtain the predicted (t + 1)-th frame of the video frame
[0024] The prediction loss, memory loss, and adversarial loss are weighted to obtain the overall loss function Loss. Optionally, the expression of the overall loss function Loss is specifically:
[0025] Loss = L pred + λ m L mem + λ α L adv
[0026] Among them, λ m and λ α are coefficients used to balance the proportions of the memory loss and the adversarial loss in the entire loss function. L pred is the prediction loss, L mem is the memory loss, and L adv is the adversarial loss.
[0027] In addition, to achieve the above object, the present invention further provides a surveillance video anomaly detection system based on adversarial learning. The system includes:
[0028] A sample frame acquisition module that obtains video sample frames arranged in chronological order. Starting from each video sample frame, k video sample frames are selected in chronological order to construct a video sample frame group as the input of the prediction network;
[0029] A prediction network construction module that constructs a prediction network based on a convolutional neural network, with the video sample frames as the input and the corresponding feature maps as the output;
[0030] An adversarial training module that uses the feature maps as the input of the memory module network and the normal sample feature maps of the same scale as the feature maps as the output of the memory module network, and performs end-to-end adversarial training in an unsupervised manner;
[0031] A loss model construction module that constructs a video anomaly frame detection model to be trained based on the prediction network and the memory module network. At the same time, based on the participation training of each video sample frame, through the application of the preliminary feature extraction network to the deep feature extraction classification network, a classification loss model is constructed by introducing reconstruction, adversarial, and memory losses;
[0032] An anomaly detection model construction module that, based on the video sample frame groups constructed from the video sample frames and the labels corresponding to each video sample frame group respectively, uses the video sample frames as the input and the labels corresponding to each video sample frame group respectively as the output, and combines the classification loss model to train the video anomaly frame detection model to be trained to obtain a video anomaly frame detection model;
[0033] An anomaly scoring module that, for each video sample frame in each video sample frame group, determines the anomaly score of whether each video sample frame in the video sample frame group is normal or abnormal through the discriminator model according to the reconstruction loss obtained from the model reconstruction. A video sample frame with an anomaly score greater than a preset value is determined to be an abnormal video frame, otherwise it is a normal video frame.
[0034] An abnormal detection method and system for surveillance videos based on adversarial learning proposed in an embodiment of the present invention. This method includes two parts. The first part sends real-time video sample frames into a feature extraction network, compares the features of the samples with the features in the memory module, updates and reads the features in the memory module. The second part performs channel-based splicing on the features read from the memory module and the features obtained by the feature reading network, sends them into a decoder to obtain the reconstructed picture, and obtains an abnormal score for judging whether the video sample frames in this video sequence are normal or not through the reconstruction error. The present invention constructs a video abnormal frame detection model based on a preliminary feature extraction network, a deep feature extraction and classification network, and a fully convolutional neural network, and then applies the video abnormal frame detection model to complete the detection of whether the video sample frames are normal or abnormal, better extracts the feature information in the video sample frames, improves the adaptability of the video abnormal frame detection model, and solves the technical problems of low detection accuracy and efficiency of existing surveillance video abnormal detection methods at present. Description of the Drawings
[0035] Figure 1 It is a schematic flowchart of the abnormal detection method for surveillance videos based on adversarial learning in the present invention;
[0036] Figure 2 It is a schematic diagram of the video abnormal frame detection model in the present invention;
[0037] Figure 3 It is a schematic structural diagram of the prediction network in the present invention;
[0038] Figure 4 It is a schematic structural diagram of the memory module in the present invention.
[0039] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0040] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] An embodiment of the present invention provides an abnormal detection method for surveillance videos based on adversarial learning. Refer to Figure 1 , Figure 1 It is a schematic flowchart of the abnormal detection method for surveillance videos based on adversarial learning in the present invention.
[0042] In this embodiment, the abnormal detection method for surveillance videos based on adversarial learning includes the following steps:
[0043] Step 1. Obtain sample frames: Obtain video sample frames arranged in chronological order. Starting from each video sample frame, select 4 video sample frames in chronological order to construct a video sample frame group as the input of the prediction network.
[0044] Step 2. Construct a prediction network: Based on a convolutional neural network, construct a prediction network with video sample frames as input and the corresponding feature maps of the video sample frames as output;
[0045] Memory module: A structure located in the feature space of the prediction network, used to record the features of normal video samples. Its input is the feature map of consecutive video frames, and its output is the normal sample feature map of the same scale as the feature map. The entire network undergoes end-to-end adversarial training without supervision.
[0046] The said Step 3 includes the following steps:
[0047] Step 3-1: For the output of the deep feature extraction and classification network, obtain the feature q of size H×W×C t . Where H is the height of the feature, W is the width of the feature, and C is the number of channels.
[0048] Step 3-2: According to the matching algorithm of the memory module, obtain the feature with the highest matching probability, also of size H×W×C.
[0049] Step 3-3: Concatenate the queried feature p t with the extracted feature q t on the channels to obtain a new feature of size H×W×2C. After all query items q t query their corresponding p t , since q t is also a normal sample feature, the memory module learns new normal sample features, so at this time the memory module will update its own memory unit.
[0050] Loss model: Based on the prediction network and the memory module network, construct a video abnormal frame detection model to be trained. At the same time, based on the participation of each video sample frame in training, from the application of the preliminary feature extraction network to the deep feature extraction and classification network, construct a classification loss model by introducing reconstruction, adversarial, and memory losses;
[0051] The said Step 4 includes the following steps:
[0052] Step 4-1: Send consecutive t frames of normal training samples X = {x1, x2, …, x t} into the prediction network;
[0053] Step 4-2: The encoder of the prediction network extracts the feature q of t video frames t . At this time, the network will read the corresponding p t from the memory module according to the similarity between q t and the normal sample features stored in the memory module, and q tConcatenate to obtain the feature (q t , p t ) and update the memory module;
[0054] Step 4-3: Send the feature (q t , p t ) to the decoder of the prediction network, and finally obtain the predicted video frame of the (t + 1)-th frame
[0055] Step 4-4: Obtain the overall loss function Loss of this model by weighting the prediction loss, memory loss, and adversarial loss. The formula is as follows:
[0056] Loss = L pred + λ m L mem + λ α L adv
[0057] Among them, in the formula, λ m , λ α are coefficients used to balance the proportions of the memory loss and the adversarial loss in the entire loss function. L pred is the prediction loss, L mem is the memory loss, and L adv is the adversarial loss.
[0058] Step 5. Anomaly detection model: Based on the video sample frame groups constructed from the video sample frames, and the labels corresponding to each video sample frame group respectively, using the video sample frames as the input and the labels corresponding to each video sample frame group respectively as the output, combined with the classification loss model, train the model to be trained for video anomaly frame detection to obtain the video anomaly frame detection model;
[0059] Step 6. Anomaly score based on prediction error: After the input sample passes through the prediction network, some information will be lost, and the prediction error is used to quantify the lost information. The prediction network only uses normal event samples during training, learns the feature patterns of normal samples, and tries to predict normal event samples as accurately as possible. Therefore, during testing, the prediction network will generate a small prediction error for normal event samples, while the abnormal sample patterns are not learned by the network and will generate a large prediction error during the prediction process. Based on this idea, in the anomaly detection algorithm using the prediction network, the prediction error of the input sample is often used as the anomaly score, and the sample with a prediction error higher than the pre-set error threshold is judged as an abnormal sample, and vice versa as a normal sample.
[0060] The sizes of the predicted image and the original image obtained through the prediction network are the same, so the prediction error is represented by the mean square error between the pixels of the original sample and the predicted sample. For a video frame with a size of m × n, the calculation process of its prediction error is as follows:
[0061]
[0062] Among them, x represents the original video frame, represents its corresponding predicted video frame, and i, j respectively represent the spatial indices of the pixels on the video frame, where i = 1, 2, …, m and j = 1, 2, …, n. For each video sample frame in each video sample frame group, the discriminator model constructs an anomaly score for determining whether each video sample frame in the video sample frame group is normal or abnormal based on the reconstruction loss obtained by model reconstruction. A video sample frame with an anomaly score greater than a preset value is determined to be an abnormal video frame, otherwise it is a normal video frame.
[0063] The prediction network is a U-Net encoder. In step 3, only normal event samples are used during the training process of the model, and abnormal samples are added only during the testing process. The continuous t-frame normal training samples X = {x1, x2, …, x t} are sent into the prediction network, and the encoder of the prediction network extracts the features q t of these t video frames. The network for extracting the features of t video frames will, according to the similarity between q t and the features of normal samples stored in the memory module, read the corresponding p t and splice it with q t to obtain the feature (q t , p t ) and update the memory module. The feature (q t , p t ) is sent to the decoder of the prediction network, and finally the predicted (t + 1)-th video frame is obtained
[0064] The memory module includes two operations: reading and updating. When the model obtains the features of a new normal sample, it will perform a reading operation on the memory module to select the normal sample features that are most similar to itself; then, the memory module will be updated according to the features of the new normal sample.
[0065] This embodiment provides an abnormal detection method for surveillance videos based on adversarial learning. This method draws inspiration from the perspective of how the human brain recognizes, understands, and identifies abnormalities, and proposes a new abnormal detection method based on the idea of "what is seen is normal, what is not seen is abnormal". This method abandons the basic ideas of the two types of methods in the past, overcomes the defects of the original ideas, and realizes a new method for video abnormal detection by learning and understanding video content. The framework of this method is divided into two parts. The first part sends real-time video sample frames into the feature extraction network, compares the features of the samples with the features in the memory module, and updates and reads the features in the memory module. The second part performs channel-based splicing on the features read from the memory module and the features obtained from the feature reading network, sends them into the decoder to obtain the reconstructed picture, and obtains an abnormal score for judging whether the video sample frames in this video sequence are normal through the reconstruction error.
[0066] To further prove the superiority of the abnormal detection performance of the generative adversarial network model based on the memory module proposed in this application. Now, the performance metrics comparison of this method with other abnormal detection methods on three datasets, namely UCSD Ped2, Avenue, and ShanghaiTech, is provided. It can be seen from the experimental data that compared with the existing unsupervised abnormal detection algorithms, on the UCSD Ped2 dataset, the AUC metric of the generative adversarial network based on the memory module proposed in this application reaches 97.8%, and the model of this application is the algorithm with the best abnormal detection performance on the UCSD Ped2 dataset. On the Avenue dataset, the AUC metric of the algorithm in this application reaches 87.3%, second only to the algorithm proposed by Hyunjong et al. However, the abnormal detection performance of the algorithm in this application on the other two datasets is better than that of Hyunjong et al.'s algorithm. On the ShanghaiTech dataset, the AUC metric of the algorithm in this application reaches 73.5%, second only to the ALOCC algorithm. However, the algorithm in this application is better than the ALOCC algorithm on the UCSD Ped2 and Avenue datasets, and the ALOCC algorithm slices the input video frames, resulting in a relatively slow training and testing process for the model.
[0067] Experimental speed comparison: In an actual video surveillance system, real-time anomaly detection of surveillance videos is often required, so a high detection speed is required for anomaly detection algorithms. In this application, the detection speed of the model in the anomaly detection task is tested and compared with the detection speeds of other anomaly detection algorithms on the UCSD Ped2 dataset. The detection speed data of the comparison algorithms are from their corresponding original papers. The resolution of video frames in the UCSD Ped2 dataset is 240×360, and during the actual test process, it is processed into a size of 256×256. The anomaly detection speed of the model is positively correlated with the size of the input video frame. The larger the size of the video frame, the slower the processing speed of the algorithm.
[0068] Furthermore, regarding the speeds of various algorithms, the processing time of the anomaly detection model of this application for a single video frame is 0.028 seconds, which can meet the real-time requirements of the video surveillance system. Compared with anomaly detection models using adversarial learning such as ALOCC, the detection speed of the model in this application is faster, mainly for two reasons: one is that the model in this paper does not cut the input video frame, and the network processes the entire frame image faster; the other reason is that the model in this application abandons the discriminator during the test stage, making the structure of the model simpler.
[0069] To more clearly explain this application, a specific example of an anomaly detection method for surveillance videos based on adversarial learning is proposed.
[0070] Refer to Figure 2 , a method for anomaly detection of surveillance videos based on adversarial learning provided by the present invention, obtains a video anomaly frame detection model according to the following steps 1 - step 13, and then applies the video anomaly frame detection model to complete the normal or abnormal detection of video sample frames;
[0071] Step 1. Obtain video sample frames arranged in chronological order. Starting from each video sample frame, select 5 video sample frames in chronological order to construct a video sample frame group (the purpose is to predict the 5th frame based on the previous 4 frames)
[0072] Step 2. Based on a convolutional neural network, use the video sample frame as the input and the corresponding reconstructed graph of the video sample frame as the output to construct a prediction network;
[0073] Step 3. The prediction network consists of three upsampling layers and three downsampling layers. Each downsampling layer uses a maximum pooling layer with a window size of 2×2 to reduce the size of the feature map. Each downsampling layer performs two convolutional operations on the features and uses ReLU as the activation function, as Figure 3 shown.
[0074] Step 4. The memory module contains two operations: reading and updating. As Figure 4As shown, after the model obtains the features of a new normal sample, it will perform a read operation on the memory module to select the normal sample features that are most similar to itself; then, the memory module will be updated according to the features of the new normal sample. The specific processes of the read operation and update operation of the memory module are as follows:
[0075] Step 5. Read operation: Divide the features q of the input sample (continuous t-frame video frames) with size H×W×C t into H×W query items at the channel dimension The size of is 1×1×C. For each read the corresponding information from the memory module containing N memory units The matching probability is a two-dimensional correlation map of size M×K, which is calculated by applying the softmax function to the cosine similarity between it and its corresponding memory unit p n as follows:
[0076] Step 6. The calculation process is as follows:
[0077]
[0078] Step 7. Through the matching probability the concatenated features corresponding to can be calculated as follows:
[0079]
[0080] Step 8. When for each the corresponding is queried, all will form a feature of the same size as Then is concatenated with q t at the channel dimension to obtain a feature map of size H×W×2C This feature map is used for the subsequent learning of the network.
[0081] Step 9. Update operation: After all query items query their corresponding Since are also normal sample features, the memory module has learned the features of the new normal sample. Therefore, at this time, the memory module will update its own memory units. The specific update rule is that for each select the matching probability The largest memory unit is updated.
[0082] Step 10. The update method is as follows:
[0083]
[0084] Among them, f(·) represents the L2 norm, represents the index set of the query item with the largest cosine similarity to each memory item, n is the parameter after n is normalized, n, The calculation formulas of are as follows:
[0085]
[0086]
[0087] Step 11. To ensure that the query item is as similar as possible to the memory unit item in the memory module, and at the same time to ensure the diversity of the memory unit, the memory module will generate a feature tightness loss L compact and a feature separation loss L separate , which are shown as follows:
[0088]
[0089] Among them, p f is the memory item with the smallest cosine distance from the query item , p S is the memory item with the second smallest cosine distance from , and α is a margin value used to prevent from being too similar to the memory item and destroying the diversity of the memory unit.
[0090] Step 12. In this paper, a classic convolutional neural network is used as the basic architecture of the discriminator. The discriminator consists of 5 convolutional layers and 1 fully connected layer. The activation function used in the fully connected layer is Sigmoid. To make the training of the model more stable, a batch normalization (BN) layer is added after each convolutional layer.
[0091] Step 13. During the training process, the predicted image generated by the prediction network and the real image corresponding to the predicted image are fed into the discriminator for authenticity judgment. The first three convolutional layers use convolutional kernels of size 5×5, the last two convolutional layers use convolutional kernels of size 3×3, the stride is 2 for all, the activation function used is ReLU, and the number of channels of the features output by the 5 convolutional layers are 64, 128, 256, 512, and 512 respectively. Feature extraction is performed by these 5 convolutional layers and then fed into the fully connected layer for discriminative estimation. The fully connected layer uses the Sigmoid activation function. Therefore, the final output of the discriminator is a scalar value between [0,1], and this value represents the discriminative result of the discriminator on the authenticity of the input image.
[0092] In a preferred embodiment, an abnormal detection system for surveillance videos based on adversarial learning is also proposed, specifically including:
[0093] A sample frame acquisition module, which obtains video sample frames arranged in chronological order. Starting from each video sample frame, k video sample frames are selected in chronological order to construct a video sample frame group as the input of the prediction network;
[0094] A prediction network construction module, based on a convolutional neural network, takes video sample frames as input and outputs feature maps corresponding to the video sample frames to construct a prediction network;
[0095] An adversarial training module, which takes the feature map as the input of the memory module network and outputs a normal sample feature map of the same scale as the feature map as the output of the memory module network, and performs end-to-end adversarial training in an unsupervised manner;
[0096] A loss model construction module, based on the prediction network and the memory module network, constructs a video abnormal frame detection model to be trained, and at the same time, based on the participation training of each video sample frame, from the preliminary feature extraction network to the application of the deep feature extraction classification network, constructs a classification loss model by introducing reconstruction, adversarial, and memory losses;
[0097] An abnormal detection model construction module, based on the video sample frame groups constructed from video sample frames and the labels corresponding to each video sample frame group respectively, takes video sample frames as input and the labels corresponding to each video sample frame group as output, and combines with the classification loss model to train the video abnormal frame detection model to be trained to obtain a video abnormal frame detection model;
[0098] An abnormality scoring module, for each video sample frame in each video sample frame group, determines the abnormality score of whether each video sample frame in the video sample frame group is normal or abnormal through the discriminator model according to the reconstruction loss obtained from model reconstruction, and determines the video sample frame with an abnormality score greater than the preset value as an abnormal video frame, otherwise as a normal video frame.
[0099] For other embodiments or specific implementations of the abnormal detection system for surveillance videos based on adversarial learning in the text of the present invention, reference may be made to the above method embodiments, and details will not be elaborated herein.
[0100] In addition, an embodiment of the present invention further provides a storage medium, on which a program for the abnormal detection method of surveillance videos based on adversarial learning is stored. When the program for the abnormal detection method of surveillance videos based on adversarial learning is executed by a processor, the steps of the abnormal detection method of surveillance videos based on adversarial learning as described above are implemented. Therefore, details will not be elaborated herein. Additionally, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in this application, reference may be made to the description of the method embodiment of this application. By way of example, the program instructions can be deployed to be executed on one computing device, or on multiple computing devices located at one location, or alternatively, on multiple computing devices distributed at multiple locations and interconnected through a communication network.
[0101] Those of ordinary skill in the art can understand that all or part of the processes in implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0102] Through the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally speaking, functions accomplished by computer programs can be easily implemented by corresponding hardware, and there are also various specific hardware structures for implementing the same function, such as analog circuits, digital circuits or dedicated circuits. However, in most cases for the present invention, implementation by software programs is a better embodiment. Based on such understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
Claims
1. An abnormal detection method for surveillance videos based on adversarial learning, characterized in that The method includes the following steps: S1: Obtain video sample frames arranged in chronological order. Starting from each video sample frame, select k video sample frames in chronological order to construct a video sample frame group as the input to the prediction network; S2: Based on a convolutional neural network, construct a prediction network with the video sample frames as the input and the feature maps corresponding to the video sample frames as the output; S3: Use the feature maps as the input to the memory module network and the normal sample feature maps of the same scale as the feature maps as the output of the memory module network for end-to-end adversarial training without supervision; Step S3 includes: For the output of the deep feature extraction classification network, a feature q of size H×W×C is obtained t , where H is the height of the feature, W is the width of the feature, and C is the number of channels; Obtain the feature p with the highest matching probability according to the matching algorithm of the memory module network t , and the size is also H×W×C; The retrieved feature p t is concatenated with the extracted feature q t on the channel dimension to obtain a new feature of size H×W×2C, so as to update the memory module network; S4: Based on the prediction network and the memory module network, construct a video anomaly frame detection model to be trained. At the same time, based on the participation training of each video sample frame, from the preliminary feature extraction network to the application of the deep feature extraction and classification network, construct a classification loss model by introducing reconstruction, adversarial, and memory losses; Step S4 specifically includes: Feed the consecutive 𝑡 frames of normal training samples 𝑋={ , , … , } into the prediction network; The encoder of the prediction network extracts the features of 𝑡 video frames , and the prediction network will, according to the similarity with the features of normal samples stored in the memory module, read the corresponding and concatenate them to obtain the features ( , ) and update the memory module network; Send the features ( , ) to the decoder of the prediction network, and finally obtain the predicted video frame of the (t + 1)-th frame ; Obtain the overall loss function Loss by weighting the prediction loss, memory loss, and adversarial loss; S5: Based on the video sample frame groups constructed from the video sample frames and the labels corresponding to each video sample frame group, use the video sample frames as the input and the labels corresponding to each video sample frame group as the output, and combine with the classification loss model to train the video anomaly frame detection model to be trained to obtain a video anomaly frame detection model; S6: For each video sample frame in each video sample frame group, use the discriminator model to determine the anomaly score of whether each video sample frame in the video sample frame group is normal or abnormal according to the reconstruction loss obtained by model reconstruction. Determine the video sample frames with anomaly scores greater than the preset value as abnormal video frames, otherwise as normal video frames.
2. The anomaly detection method for surveillance videos based on adversarial learning according to claim 1, wherein, In step S2, the prediction network is a U-Net encoder.
3. The method for abnormal detection of surveillance videos based on adversarial learning according to claim 1, wherein, In step S3, the memory module network uses normal event samples during training and adds abnormal samples during testing.
4. The anomaly detection method for surveillance videos based on adversarial learning according to claim 1, characterized in that In step S3, the memory module network includes two operations: reading and updating. After obtaining the features of a new normal sample, a reading operation will be performed on the memory module network to select the normal sample features most similar to itself; The memory module network will be updated according to the features of the new normal sample.
5. The abnormal detection method for surveillance videos based on adversarial learning according to claim 1, characterized in that, The expression of the overall loss function Loss is specifically: Loss = + + Among them, , are coefficients used to balance the proportions of the memory loss and the adversarial loss in the entire loss function, is the prediction loss, is the memory loss, is the adversarial loss.
6. An abnormal detection system for surveillance videos based on adversarial learning, characterized in that, A system for implementing the method for monitoring video anomaly detection based on adversarial learning according to claim 1, the system includes: A sample frame acquisition module: Obtain video sample frames arranged in chronological order. Starting from each video sample frame, select k video sample frames in chronological order to construct a video sample frame group as the input to the prediction network; A prediction network construction module: Based on a convolutional neural network, construct a prediction network with the video sample frames as the input and the feature maps corresponding to the video sample frames as the output; An adversarial training module: Use the feature maps as the input to the memory module network and the normal sample feature maps of the same scale as the feature maps as the output of the memory module network for end-to-end adversarial training without supervision; Loss model construction module: Based on the prediction network and the memory module network, construct a video anomaly frame detection model to be trained. At the same time, based on the participation training of each video sample frame, from the initial feature extraction network to the application of the deep feature extraction and classification network, by introducing reconstruction, adversarial, and memory losses, construct a classification loss model; Anomaly detection model construction module: Based on the video sample frame groups constructed from video sample frames and the labels corresponding to each video sample frame group respectively, using the video sample frames as input and the labels corresponding to each video sample frame group as output, combined with the classification loss model, train the video anomaly frame detection model to be trained to obtain a video anomaly frame detection model; Anomaly scoring module: For each video sample frame in each video sample frame group, use the discriminator model to determine the anomaly score of whether each video sample frame in the video sample frame group is normal or abnormal according to the reconstruction loss obtained by model reconstruction. Determine the video sample frames with anomaly scores greater than the preset value as abnormal video frames, otherwise as normal video frames.