A semi-supervised abnormal behavior detection and prediction method

By building a bidirectional frame prediction network, combining scene encoder and conditional variational autoencoder, the problem of abnormal behavior detection based on scene dependence is solved, and abnormal behavior detection and prediction under semi-supervised conditions is realized, and the accuracy of detection and prediction is improved.

CN116597508BActive Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310456314.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-07-18
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

The prior art cannot effectively solve the problem of scene-dependent abnormal behavior detection, and lacks the ability to predict abnormal behavior, especially in semi-supervised conditions, it is difficult to distinguish between normal behavior and abnormal behavior.

Method used

A two-way frame prediction network is built, including a forward frame prediction network and a reverse frame prediction network, combined with a conditional variational autoencoder, and the scene dependence is processed through the scene encoder, and an abnormal behavior detection and prediction are achieved using forward and reverse frame prediction errors and KL divergence optimization models.

Benefits of technology

It performs better than existing methods on the scene-dependent anomaly behavior detection dataset and has the ability to predict future abnormal behaviors, which improves the accuracy of abnormal behavior detection and prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116597508B_ABST
    Figure CN116597508B_ABST
Patent Text Reader

Abstract

The present invention discloses a semi-supervised abnormal behavior detection and prediction method. First, a set number of video frames in a training or test video to be processed are obtained, and then fixed-size regions related to them are cropped from all the frames to form a person segment to be processed and a background image. The person segment to be processed and the background image are input into a forward frame prediction network, and the predicted person frames output by the forward frame prediction network and the encoded background image extracted are input into a backward frame prediction network, and the backward frame prediction network outputs a predicted image of the previous part of the frames of the person segment to be processed. An error is calculated based on the predicted person frames output by the forward frame prediction network and the observed real person frames to achieve abnormal behavior detection. An error is calculated based on the predicted person frames output by the backward frame prediction network and the observed real person frames to achieve abnormal behavior prediction. The present invention has a stronger abnormal behavior prediction ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pattern recognition, and particularly relates to a semi-supervised abnormal behavior detection and prediction method. Background Art

[0002] Normal behavior usually refers to the regular behavior that conforms to the regulations in daily life, and abnormal behavior refers to the behavior that is not expected to occur and even has harmfulness. The purpose of abnormal behavior detection is to estimate the probability of abnormal behavior at each observed moment in a video, so as to detect abnormal behavior according to the probability level. Abnormal behavior prediction refers to estimating the probability of abnormal behavior occurring in a future period of time in advance without observing future behavior, so as to give an early warning of possible abnormal behavior. The characteristics of semi-supervised abnormal behavior detection and prediction are that during the training of the model, only video data containing normal behavior is available, while during the testing and application of the model, the video contains both normal behavior and abnormal behavior, and the model needs to detect and predict in advance whether abnormal behavior occurs, without the need to distinguish what specific behavior it is. Video abnormal behavior detection and prediction technology has broad application prospects in the fields of intelligent monitoring systems, video analysis, etc.

[0003] The mainstream semi-supervised abnormal behavior detection methods can be divided into distance-based methods, reconstruction-based methods, and prediction-based methods. The distance-based methods construct a model to learn the representation of normal data, and determine the abnormal probability according to the deviation between the test data and the normal data. The commonly used data representation forms include the nearest neighbor samples, the decision boundary of one-class support vector machines, and the clustering centers. Among them, the method based on the nearest neighbor samples directly uses simple nearest neighbor search to calculate the distance between the test sample and its neighboring normal samples as the abnormal score, resulting in a huge amount of calculation; the method using one-class support vector machines calculates the decision boundary from the normal data, but when given new normal training data, there is a disadvantage of a large overhead for updating the model; the clustering-based method uses the clustering centers of normal data as the representation of normal patterns, which greatly reduces the number of sample representations, but the number of clusters is usually small and needs to be determined manually, and once determined, it remains fixed, so it is difficult to comprehensively represent normal patterns, resulting in poor performance.

[0004] The reconstruction-based methods aim to train a model that can reconstruct normal patterns, and determine the abnormal probability according to the reconstruction error. It can be divided into sparse coding reconstruction and video reconstruction methods. With the technological progress of deep neural networks, video reconstruction methods usually perform better. Such methods train a model that can accurately reconstruct normal frames but cannot accurately reconstruct abnormal frames, and obtain the abnormal probability according to the error between the reconstructed frames and the observed real frames. However, due to the powerful representation ability of neural networks, overfitting is likely to occur when the input and output are the same data, resulting in poor abnormal behavior detection effect.

[0005] The prediction-based method assumes that abnormal frames are difficult to predict. It trains a frame prediction model that takes a video clip as input and outputs the next frame after the clip. The anomaly probability is obtained based on the error between the predicted and the true next frame. The frame prediction method is less prone to overfitting than the video reconstruction method and usually has better anomaly behavior detection performance under the same model settings. Therefore, the prediction-based method has become the mainstream in recent years.

[0006] Although the above methods, especially the prediction-based method, have achieved good performance on general anomaly behavior detection datasets, they still cannot solve the problem of anomaly behavior detection with scene dependence. The scene dependence of anomaly behavior means that a certain behavior is normal in one scene but abnormal in another scene, that is, whether the behavior is abnormal is determined by the scene. For example, playing football on the playground is a normal behavior, while playing football on the road is an abnormal behavior. Existing methods do not consider the association between behavior and scene, so they cannot accurately distinguish whether a behavior is abnormal based on the scene.

[0007] In addition, the above methods only have the ability to detect abnormal behaviors and cannot be directly used for abnormal behavior prediction. Although there are behavior prediction methods in the field of action recognition, they are based on sufficient and labeled normal behavior types, and abnormal behaviors are usually difficult to obtain and label, resulting in the inability to apply the behavior prediction methods in the field of action recognition to abnormal behavior prediction. Therefore, there is currently no model that can handle semi-supervised abnormal behavior prediction. Summary of the Invention

[0008] To overcome the deficiencies of the prior art, the present invention provides a semi-supervised abnormal behavior detection and prediction method. First, a set number of video frames in the training or test video to be processed are obtained, and then fixed-size regions related to them are cropped from all the frames to form a person segment to be processed and a background image; the person segment to be processed and the background image are input into a forward frame prediction network, and the predicted person frame output by the forward frame prediction network and the encoded background image extracted are input into a reverse frame prediction network, which outputs a predicted image of the previous part of the person segment to be processed; the error is calculated based on the predicted person frame output by the forward frame prediction network and the observed true person frame to achieve abnormal behavior detection; the error is calculated based on the predicted person frame output by the reverse frame prediction network and the observed true person frame to achieve abnormal behavior prediction. The present invention has stronger abnormal behavior prediction ability.

[0009] The technical solution adopted by the present invention to solve its technical problems includes the following steps:

[0010] Step 1: Obtain a set number of a video frames in the video to be processed, and obtain a person segment to be processed centered on a person and a background image unrelated to the person through preprocessing; a is an odd number;

[0011] Step 1-1: Obtain the to-be-processed person segments centered on a person;

[0012] Use a target detector to detect the rectangular bounding boxes of all persons, take the geometric center of each rectangular bounding box as the center position of the corresponding person, and based on the center position of each person in the middle frame of a video frames, set a new bounding box with a fixed size, and crop the images centered on the person and with the size of the fixed-size bounding box in all video frames to form the to-be-processed person segments centered on a person;

[0013] Step 1-2: Obtain the background images irrelevant to people;

[0014] In the middle frame of a video frames, cover the bounding box areas of all persons detected by the target detector with black to form the background image I b ;

[0015] Step 2: Construct a bidirectional frame prediction network;

[0016] The bidirectional frame prediction network includes a forward frame prediction network and a backward frame prediction network, both having the same structure, and each network includes an image encoder, a conditional variational autoencoder, and an image decoder;

[0017] Step 3: Optimize the parameters of the bidirectional frame prediction network using the obtained to-be-processed training person segments and background images;

[0018] Feed the to-be-processed person segments and background images into the bidirectional frame prediction network, calculate the error between the predicted frame output and the corresponding ground truth as the first loss, and this loss consists of two parts: the forward frame prediction loss and the backward frame prediction loss; at the same time, calculate the divergence between the prior distribution and the posterior distribution in the conditional variational autoencoder as the second loss, and train the bidirectional frame prediction network by minimizing these two loss values through backpropagation; perform the same operation on all to-be-processed person segments; the specific steps include:

[0019] Step 3-1: Calculate the frame prediction error loss of the forward frame prediction network;

[0020] Assume that the current time is the t-th frame, the number of historical frames is n, and to predict abnormal behaviors within the next α frames, a to-be-processed training person segment contains a total of n + α + 1 frames, denoted as {f t-n , …, f t-1 , f t , f t+1 , … f t+α}, and the forward frame prediction network takes the observable f t-n , …, f t-1 frames as input and outputs the predicted Frame; For training the forward frame prediction network, calculate each predicted frame and the corresponding ground truth frame f t+i The sum of the mean squared error loss and the L1 loss between them:

[0021]

[0022] where

[0023] Predicted frame

[0024] f: Ground truth frame

[0025] λ L1 : The weight of the L1 loss

[0026] Step 3-2: Calculate the frame prediction error loss of the backward frame prediction network;

[0027] For the backward frame prediction network predicting the i-th frame, i ∈ [1, α], during the training phase, the predicted output frame of the forward network is combined with the observable frames f t , …, f t+i+1-n to form the first input, and the predicted ground truth frames f t+i, …, f t+1 of the forward network are combined with the observable frames f t , …, f t+i+1-n to form the second input, and these two inputs are fed into the backward frame prediction network; the predicted frames corresponding to these two inputs are respectively denoted as and Their corresponding predicted ground truth frames are both f t+i-n , calculate between and f t+i-n and between and f t+i-n The average mean squared error and L1 loss to train the backward frame prediction network:

[0028]

[0029] Step 3-3: Calculate the KL divergence between the prior distribution and the posterior distribution in the conditional variational autoencoder;

[0030] Obtain all scenes to be processed and assign a class label to each scene; then construct an image classification network as the scene encoder, use all the obtained scene images as inputs and the scene class labels as supervision information to train the scene encoder; after training, fix the parameters of the scene encoder; send the background image I b into the scene encoder to obtain the scene encoding;

[0031] The conditional variational autoencoder concatenates the image encoding output by the image encoder and the obtained scene encoding, and feeds them into the encoder of the conditional variational autoencoder to generate the parameters of the posterior distribution; then samples the latent variable from the posterior distribution through the reparameterization technique, concatenates it with the scene encoding again, and feeds it into the decoder of the conditional variational autoencoder to generate a feature map conditioned on the scene; assuming that the prior distribution is a standard Gaussian distribution, calculate the KL divergence between it and the posterior distribution:

[0032]

[0033] where,

[0034] the mean of the posterior Gaussian distribution

[0035] the variance of the posterior Gaussian distribution;

[0036] The total loss for training the entire bidirectional frame prediction network is the sum of the forward frame prediction, the backward frame prediction, and the KL divergence, and the weight of the KL divergence is λ KL ; calculate the total loss for all training human segments to be processed, and jointly train the entire model by minimizing the total loss to obtain an optimized bidirectional frame prediction network. The total loss is expressed as:

[0037] L total =L f +L b +λ KL L KL (4)

[0038] Step 4: Apply the optimized bidirectional frame prediction model to the obtained test video to be processed for anomaly behavior detection and prediction, and calculate the anomaly probability at the current moment and in a future period of time;

[0039] Assume that the current observation moment is t, and the observed frames are represented as {f t-n , …, f t-1 , f t}. For the anomaly behavior detection task, an anomaly probability needs to be calculated for the current frame f t ; for the anomaly behavior prediction task, an anomaly probability needs to be calculated for a future period of unobservable frames {f t+1 , …, f t+α}; specifically, it includes the following steps:

[0040] Step 4-1: Calculate the anomaly probability for anomaly behavior detection at the current moment;

[0041] For each test human segment {f t-n , …, f t-1} and the background image are input into the trained forward frame prediction network to obtain the predicted frame Then the predicted image is compared with the observed real image f t The error is calculated using Equation (1), i.e.:

[0042]

[0043] At time t, in order to infer the anomaly probability of the entire frame from all pending test person segments, the same processing is performed on all pending test person segments in the test video to calculate the prediction error, and the maximum error value is taken as the anomaly probability at time t;

[0044] Step 4-2: Calculate the anomaly probability of anomaly behavior prediction for a future segment of frames;

[0045] For each pending test person segment in the test video, in order to predict the anomaly probability of the (t + i) (i ∈ [1, α])-th frame, first the image predicted in Step 4-1 and some of the observed images {f t , …, f t+i+1-n} are concatenated to form which is then fed into the backward frame prediction network to output the backward predicted image Then is compared with the corresponding real frame f t+i-n The error is calculated according to Equation (1); in order to calculate the anomaly probability of this test person segment during the entire period [t + 1, t + α], the maximum error of the backward frame prediction during this period is calculated, representing the anomaly probability of the entire period:

[0046]

[0047] Similarly, in order to obtain the anomaly probability of the entire future period [t + 1, t + α] of this test video from all pending test person segments, the same processing is performed on all pending test person segments to calculate the backward prediction error, and the maximum error is taken as the anomaly probability at [t + 1, t + α].

[0048] Furthermore, the forward frame prediction network and the backward frame prediction network respectively input T in image frames into their respective image encoders and output T out image frames from the image decoders; the function of the image encoder is to encode the input image and feed it into the conditional variational autoencoder; the conditional variational autoencoder reconstructs the image encoding conditioned on the obtained background image encoding, and its function is to control the reconstruction effect of the input image conditioned on the scene; finally, the reconstructed image encoding is fed into the image decoder to generate the predicted image frames.

[0049] The beneficial effects of the present invention are as follows:

[0050] The present invention achieves the highest level of performance or exceeds most of the existing methods on the general scene-independent abnormal behavior detection datasets ShanghaiTech, Avenue, and Corridor. On the scene-dependent abnormal behavior detection datasets NWPU Campus and ShanghaiTech-sd, it leads other existing methods by 3.7% - 5.8%. On the NWPU Campus dataset that can be used for the abnormal behavior prediction task, it leads the method with only forward frame prediction by about 0.9% on average under different prediction duration settings. Description of the Drawings

[0051] Figure 1 It is a flowchart of the technical solution of the present invention.

[0052] Figure 2 It is a structural diagram of the bidirectional frame prediction network of the present invention.

[0053] Figure 3 It is a schematic diagram of using the bidirectional frame prediction model of the present invention for abnormal behavior detection and prediction.

[0054] Figure 4 It is a detailed structural schematic diagram of the bidirectional frame prediction network of the present invention. Detailed Embodiment

[0055] The present invention will be further described below in conjunction with the drawings and embodiments.

[0056] The existing abnormal behavior detection technologies cannot handle scene-dependent abnormal behaviors, and there is currently no abnormal behavior prediction technology. The purpose of the present invention is to provide a new semi-supervised abnormal behavior detection and prediction method, which has the ability to detect scene-dependent abnormal behaviors and can predict whether an abnormal behavior will occur in the future for a period of time, that is, it has the ability to predict abnormal behaviors at the same time.

[0057] To overcome the deficiencies of the prior art, the present invention provides a semi-supervised abnormal behavior detection and prediction method. First, a set number of video frames in the training or test video to be processed are obtained. Then, for each person that appears, a fixed-size region related to them is cropped from all frames to form a person segment to be processed. At the same time, all the people in the middle frame of the original video are covered to form a background image. The first part of the frames of the person segment to be processed is input into a forward frame prediction deep neural network. Meanwhile, the background image passes through an image classification network to output a class encoding, and this class encoding is also input into the forward frame prediction network as a conditional encoding. The forward frame prediction network outputs a predicted image of the second part of the frames of the person segment to be processed. Then, the predicted person frames output by the forward frame prediction network and the encoded background image are input into a reverse frame prediction network, which outputs a predicted image of the first part of the frames of the person segment to be processed. The error is calculated based on the predicted person frames output by the forward frame prediction network and the observed real person frames to achieve abnormal behavior detection. The error is calculated based on the predicted person frames output by the reverse frame prediction network and the observed real person frames to achieve abnormal behavior prediction. In the forward and reverse frame prediction networks, the scene image encoding controls the association between the output image and the scene as a condition, thereby dealing with scene-dependent abnormal behaviors.

[0058] As Figure 1 shown, a semi-supervised abnormal behavior detection and prediction method includes the following steps:

[0059] Step 1: Obtain a set number of video frames in the training and test videos to be processed, and preprocess to obtain a person segment to be processed centered on people and a background image unrelated to people.

[0060] Step 1-1: Obtain a person segment to be processed centered on people. Use an object detector to detect the bounding boxes of all people, and based on the center position of each person in the middle frame of the set number of videos, set a bounding box of a fixed size, and crop the images in the regions of these bounding boxes in all frames to form a person segment to be processed centered on people.

[0061] Step 1-2: Obtain a background image unrelated to people. In the middle frame of the set number of video frames, cover the bounding box regions of all people detected by the object detector with black to form the background image I b . The detailed structure of the bidirectional frame prediction network used is as Figure 4 shown.

[0062] Step 2: Construct a bidirectional frame prediction network. The bidirectional frame prediction includes a forward frame prediction network and a reverse frame prediction network, both having the same structure. Each network includes an image encoder, a conditional variational autoencoder, and an image decoder. A bidirectional frame prediction network structure is asFigure 2 As shown, where C, T, H, and W represent the number of channels of each image in the input video segment, the number of frames of the video segment, the image height, and the image width respectively, and γ ∈ {0, 1} is a hyperparameter.

[0063] The input of each network is a set number of frames, and the output is another set number of frames. Its role is to predict the output frames as accurately as possible based on the input frames. The role of the image encoder is to encode the input images and send them into the conditional variational autoencoder; at the same time, the conditional variational autoencoder reconstructs the image encoding conditioned on the obtained background image encoding, and its role is to control the reconstruction effect of the input images conditioned on the scene; finally, the reconstructed image encoding is sent into the image decoder to generate the predicted images.

[0064] Step 3: Use the obtained training person segments to be processed and the background images to optimize the parameters of the bidirectional frame prediction network. For a person segment to be processed and a background image, send them into the constructed bidirectional frame prediction network, calculate the error between the output predicted frames and the corresponding ground truths as the first loss, which consists of two parts: the forward frame prediction loss and the backward frame prediction loss; at the same time, calculate the divergence between the prior distribution and the posterior distribution in the conditional variational autoencoder as the second loss, and train the bidirectional frame prediction network by minimizing these two loss values through backpropagation. The same operation is performed on all training person segments to be processed. The specific steps include:

[0065] Step 3-1: Calculate the frame prediction error loss of the forward network. Assume that the current time is the t-th frame, the number of historical frames is n, and to predict abnormal behaviors within the next α frames, a training person segment to be processed contains a total of n + α + 1 frames, denoted as {f t-n , …, f t-1 , f t , f t+1 , … f t+α}. The forward network takes the observable f t-n , …, f t-1 frames as input and outputs the predicted frames. To train the forward network, calculate the mean squared error loss and the L1 loss between each predicted frame and the corresponding ground truth frame f t+i :

[0066]

[0067] where,

[0068] predicted frame

[0069] f: ground truth frame

[0070] λ L1 : weight of the L1 loss

[0071] Step 3-2: Calculate the frame prediction error loss of the reverse network. For the reverse network predicting the i-th (i ∈ [1, α]) frame, during the training phase, the frames predicted forward are respectively used and the real frame f t+i , …, f t+1 are combined with the observable frames f t , …, f t+i+1-n to form two kinds of inputs and feed them into the reverse network. In this way, the reverse network can make full use of the observed information to perform short-term anomaly prediction more accurately. The predicted frames corresponding to the outputs of these two inputs are respectively denoted as and Their corresponding predicted real frames are both f t+i-n . Calculate the average mean square error and L1 loss between and f t+i-n and between and f t+i-n to train the reverse network:

[0072]

[0073] where

[0074] L f : The loss function represented by formula (1)

[0075] Step 3-3: Calculate the Kullback-Leibler (KL) divergence between the prior distribution and the posterior distribution in the conditional variational autoencoder. First, train a scene image classification network as a scene encoder according to the category information of the known scene. After training, fix the parameters of the scene encoder, and send the currently obtained scene image I b into the scene encoder to obtain the scene encoding. The conditional variational autoencoder concatenates the feature map output by the image encoder and the obtained scene encoding, and feeds them into the encoder of the conditional variational autoencoder to generate the parameters of the posterior distribution. Then, sample the latent variable from the posterior distribution through the reparameterization technique, concatenate it with the scene encoding again, and send it into the decoder of the conditional variational autoencoder to generate the feature map conditional on the scene. Assume that the prior distribution is a standard Gaussian distribution, and calculate the Kullback-Leibler (KL) divergence between it and the posterior distribution:

[0076]

[0077] where

[0078] The mean of the posterior Gaussian distribution

[0079] The variance of the posterior Gaussian distribution.

[0080] There are conditional variational autoencoders in both the forward and reverse networks, and the loss is calculated in the same way, as shown in Equation (3).

[0081] The total loss for training the entire bidirectional frame prediction network is the sum of the forward frame prediction, reverse frame prediction, and KL divergence, with the weight of the KL divergence being λ KL . The total loss is calculated for all training human segments to be processed, and the entire model is jointly trained by minimizing the total loss to obtain an optimized bidirectional frame prediction network. The total loss is expressed as:

[0082] L total = L f + L b + λ KL L KL (4)

[0083] Step 4: Apply the trained bidirectional frame prediction model to the acquired test video to be processed for abnormal behavior detection and prediction, and calculate the abnormal probability at the current moment and for a period of time in the future. After the bidirectional frame prediction model is trained, it can be used for abnormal behavior detection and prediction. Assume that the current observation moment is t, and the observed frames are represented as {f t-n , …, f t-1 , f t}. For the abnormal behavior detection task, an abnormal probability needs to be calculated for the current frame f t ; for the abnormal behavior prediction task, an abnormal probability needs to be calculated for a period of unobservable future frames {f t+1 , …, f t+α}. A schematic diagram of using the bidirectional frame prediction model for abnormal behavior detection and prediction is shown in Figure 3 as follows, and this process specifically includes the following steps:

[0084] Step 4-1: Calculate the abnormal probability for abnormal behavior detection at the current moment. For each test human segment {f t-n , …, f t-1} and the background image in the test video, input them into the trained forward frame prediction network to obtain the predicted frame Then, use Equation (1) to calculate the error between the predicted image and the observed real image f t , that is:

[0085]

[0086] At time t, in order to infer the abnormal probability of the entire frame from all test human segments to be processed, the same process is performed on all test human segments in the test video to calculate the prediction error, and the maximum error value is taken as the abnormal probability at time t.

[0087] Step 4-2: Calculate the anomaly probability of abnormal behavior prediction for a future segment of frames. For each test person segment to be processed in the test video, in order to predict the anomaly probability of the (t+i)-th (i∈[1, α]) frame, first, the image predicted in Step 4-1 and some observed images {f t , …, f t+i+1-n} are concatenated to form which is then fed into the reverse frame prediction network to output the reverse predicted image Then, the error is calculated according to Equation (1) between and the corresponding true frame f t+i-n . To calculate the anomaly probability of this test person segment over the entire period [t+1, t+α], calculate the maximum error of the reverse frame prediction within this period, which represents the anomaly probability of the entire period:

[0088]

[0089] Similarly, to obtain the anomaly probability of the entire future period [t+1, t+α] of this test video from all test person segments to be processed, the reverse prediction errors are calculated by performing the same processing on all test person segments to be processed, and the maximum of these errors is taken as the anomaly probability at time [t+1, t+α]. Specific embodiments:

[0091] The embodiments of the present invention are implemented on the publicly available datasets ShanghaiTech, Avenue, Corridor, NWPU Campus and the dataset ShanghaiTech-sd reorganized for studying scene-dependent abnormal behavior detection.

[0092] The comparison of the embodiments of the present invention with other methods on general abnormal behavior detection datasets (ShanghaiTech, Avenue, and Corridor, none of which contain scene-dependent abnormal behavior) is shown in Table 1. The evaluation metric is the area under the receiver operating characteristic curve (AUC). It can be seen that the present invention achieves the highest or leading performance.

[0093] Table 1 Comparison on general abnormal behavior detection datasets (AUC, %)

[0094]

[0095] The comparison of the embodiments of the present invention with other methods on abnormal behavior detection datasets with scene-dependent abnormal behavior (NWPU Campus and ShanghaiTech-sd) is shown in Table 2. It can be seen that the present invention is significantly superior to other methods.

[0096] Comparison on the Scenario-Dependent Abnormal Behavior Detection Dataset (AUC, %)

[0097]

[0098] The performance of the embodiments of the present invention in the abnormal behavior prediction task on the abnormal behavior prediction dataset (NWPU Campus) is shown in Table 3. Among them, "only forward frame prediction" means that only a general frame prediction model is used to predict future frames, and the error between the predicted future frames and the last frame observed currently is calculated as the abnormal probability for abnormal prediction. It can be seen from the comparison that the present invention has stronger abnormal behavior prediction ability.

[0099] Table 3 Abnormal Behavior Prediction Results with Different Prediction Durations (AUC, %)

[0100]

Claims

1. A semi-supervised abnormal behavior detection and prediction method, characterized in that, It includes the following steps: Step 1: Obtain a set number \(a\) of video frames from the video to be processed. After preprocessing, obtain the person-centered person segments to be processed and the background images irrelevant to people; \(a\) is an odd number. Step 1-1: Obtain the person-centered person segments to be processed. Use a target detector to detect the rectangular bounding boxes of all people, take the geometric center of each rectangular bounding box as the central position of the corresponding person, and set a new bounding box of a fixed size based on the central position of each person in the middle frame of the \(a\) video frames. Crop the images of the fixed-size bounding boxes centered on people in all video frames to form the person-centered person segments to be processed. Step 1-2: Obtain the background images irrelevant to people. In the middle frame of a video frames, the bounding box regions of all persons detected by the target detector are covered with black to form a background image I b ; Step 2: Construct a bidirectional frame prediction network. The bidirectional frame prediction network includes a forward frame prediction network and a backward frame prediction network, both having the same structure. Each network contains an image encoder, a conditional variational autoencoder, and an image decoder. Step 3: Use the obtained person segments to be processed for training and the background images to optimize the parameters of the bidirectional frame prediction network. Send the person segments to be processed and the background images into the bidirectional frame prediction network, calculate the error between the predicted frame output and the corresponding ground truth as the first loss. This loss consists of two parts: the forward frame prediction loss and the backward frame prediction loss. At the same time, calculate the divergence between the prior distribution and the posterior distribution in the conditional variational autoencoder as the second loss. Minimize these two loss values through backpropagation to train the bidirectional frame prediction network. Perform the same operation on all person segments to be processed. The specific steps include: Step 3-1: Calculate the frame prediction error loss of the forward frame prediction network. Assume that the current moment is the t-th frame, the number of historical frames is n, and abnormal behaviors within the next α frames are to be predicted. A training person segment to be processed contains a total of n+α+1 frames, denoted as {f t-n ,…,f t-1 ,f t ,f t+1 ,…f t+α}. The forward frame prediction network takes the observable f t-n ,…,f t-1 frames as inputs and outputs the predicted frames; To train the forward frame prediction network, calculate the sum of the mean square error loss and the L1 loss between each predicted frame and the corresponding true frame f t+i : Where, Predicted frame \(f\): Ground truth frame λ L1 : Weight of the L1 loss Step 3-2: Calculate the frame prediction error loss of the backward frame prediction network. For the backward frame prediction network predicting the i-th, i ∈ [1, α] frames, during the training phase, the predicted output frames of the forward network are combined with the observable frames f t , …, f t+i+1-n to form the first type of input. The predicted ground-truth frames f t+i , …, f t+1 of the forward network are combined with the observable frames f t , …, f t+i+1-n to form the second type of input. These two types of input are fed into the backward frame prediction network. The predicted frames corresponding to these two types of input are denoted as and respectively, and their corresponding predicted ground-truth frames are both f t+i-n . Calculate the average mean square error and L1 loss between and f t+i-n and between and f t+i-n to train the backward frame prediction network: Step 3-3: Calculate the KL divergence between the prior distribution and the posterior distribution in the conditional variational autoencoder. Obtain all the scenes to be processed and assign a class label to each scene. Then construct an image classification network as the scene encoder. Use all the obtained scene images as inputs and the scene class labels as supervision information to train the scene encoder. After training, fix the parameters of the scene encoder. Send the background image \(I_b\) into the scene encoder to obtain the scene encoding. The conditional variational autoencoder concatenates the image encoding output by the image encoder and the obtained scene encoding, sends them into the encoder of the conditional variational autoencoder to generate the parameters of the posterior distribution. Then sample the latent variable from the posterior distribution through the reparameterization technique, concatenate it with the scene encoding again, and send it into the decoder of the conditional variational autoencoder to generate the feature map conditioned on the scene. Assume the prior distribution is a standard Gaussian distribution, and calculate the KL divergence between it and the posterior distribution: Where, Mean of the posterior Gaussian distribution Variance of the posterior Gaussian distribution; The total loss for training the entire bidirectional frame prediction network is the sum of the forward frame prediction, backward frame prediction, and KL divergence, with the weight of the KL divergence being λ KL ; Calculate the total loss for all training human segments to be processed, and jointly train the entire model by minimizing the total loss to obtain the optimized bidirectional frame prediction network. The total loss is expressed as: L total = L f + L b + λ KL L KL (4) Step 4: Apply the optimized bidirectional frame prediction model to the video to be processed for testing to perform abnormal behavior detection and prediction, and calculate the abnormal probability at the current moment and in a future period of time. Assume that the current observation time is t, and the observed frames are represented as {f t-n , …, f t-1 , f t}. For the abnormal behavior detection task, an abnormal probability needs to be calculated for the current frame f t ; for the abnormal behavior prediction task, an abnormal probability needs to be calculated for a future period of unobservable frames {f t+1 , …, f t+α}; specifically, it includes the following steps: Step 4-1: Calculate the abnormal probability of abnormal behavior detection at the current moment. For each test person segment {f t-n , …, f t-1} to be processed in the test video and the background image, input them into the trained forward frame prediction network to obtain the predicted frame Then, for the predicted image and the observed real image f t Calculate the error using Equation (1), that is: At time t, in order to infer the anomaly probability of the entire frame from all the to-be-processed test person segments, the same processing is performed on all the to-be-processed test person segments in the test video to calculate the prediction error, and the maximum error value is taken as the anomaly probability at time t; Step 4-2: Calculate the anomaly probability of the anomaly behavior prediction for a future segment of frames; For each test person segment to be processed in the test video, in order to predict the anomaly probability of the (t+i)-th (i∈[1,α]) frame, first, the image predicted in step 4-1 and some observed images {f t ,…,f t+i+1-n} are concatenated to form which is then fed into the reverse frame prediction network to output the reverse predicted image Then, is compared with the corresponding true frame f t+i-n to calculate the error according to Equation (1); in order to calculate the anomaly probability of this test person segment during the entire period [t+1,t+α], calculate the maximum error of the reverse frame prediction within this period, which represents the anomaly probability of the entire period: Similarly, in order to obtain the anomaly probability of the entire future period [t + 1, t + α] of the test video from all the to-be-processed test person segments, the same processing is performed on all the to-be-processed test person segments to calculate the backward prediction error, and the maximum error is taken as the anomaly probability at time [t + 1, t + α].

2. The semi-supervised anomaly behavior detection and prediction method according to claim 1, characterized in that The forward frame prediction network and the backward frame prediction network respectively input T image frames into their respective image encoders and output T image frames from the image decoders. The role of the image encoder is to encode the input image and send it into the conditional variational autoencoder. The conditional variational autoencoder reconstructs the image encoding conditioned on the obtained background image encoding, and its role is to control the reconstruction effect of the input image conditioned on the scene. Finally, the reconstructed image encoding is sent into the image decoder to generate the predicted image frames. in of image frames into their respective image encoders and output T out of image frames from the image decoder; the role of the image encoder is to encode the input image and send it into the conditional variational autoencoder; the conditional variational autoencoder reconstructs the image encoding conditioned on the obtained background image encoding, and its role is to control the reconstruction effect of the input image conditioned on the scene; finally, the reconstructed image encoding is sent into the image decoder to generate the predicted image frames.

Citation Information

Patent Citations

  • Timing data anomaly detection and correction

    US10609440B1

  • Image super-resolution and non-uniform blur removal method based on fusion network

    WO2020015167A1