Video anomaly detection method, device, equipment, storage medium and product

By integrating the appearance and motion characteristics of the video and simplifying the processing flow of the memory module, the existing video abnormality detection technology is solved, and efficient video abnormality detection is achieved.

CN119380237BActive Publication Date: 2025-05-27XIAN JIAOTONG LIVERPOOL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411407730.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-05-27
Estimated Expiration
2044-10-10

AI Technical Summary

Technical Problem

The existing video anomaly detection technology ignores the video characteristics. Model training requires high computing power and large parameters, resulting in low detection accuracy and efficiency.

Method used

By fusing the video appearance features with the video motion features and fusing the predicted frames with the predicted frame difference, the basic convolution kernel and automatic codec are used to simplify the processing flow of the memory module and realize the prediction of future frames for video abnormality detection.

Benefits of technology

It improves the accuracy and efficiency of video abnormality detection, reduces the requirements for computing power, can detect under high processing speed, and retains the complex characteristics of the original video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380237B_ABST
    Figure CN119380237B_ABST
Patent Text Reader

Abstract

The present invention discloses a video anomaly detection method, apparatus, device, storage medium and product. The method includes: obtaining a video frame sequence to be verified and a target verification video frame, and determining a difference sequence of frames to be verified according to the video frame sequence to be verified; respectively inputting the video frame sequence to be verified and the difference sequence of frames to be verified into a pre-trained video anomaly detection model to obtain a target predicted video frame output by the model; performing video anomaly detection on the target verification video frame according to the target verification video frame and the target predicted video frame to obtain an anomaly detection result. The technical solution of the embodiment of the present invention improves the accuracy and efficiency of video anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and in particular, to a video anomaly detection method, device, equipment, storage medium and product. Background Art

[0002] Video anomaly detection aims to identify abnormal situations in videos that may deviate from normal behaviors, such as traffic accidents, natural disasters, and abnormal behaviors. In the actual application field, video anomaly detection usually requires the detection technology to have both high accuracy and high processing speed. Therefore, an efficient video analysis model is required as the basis for anomaly detection.

[0003] The existing video anomaly detection mainly conducts detection from two aspects. One is the anomaly detection technology based on classification, and the other is the video anomaly detection technology based on prediction. However, the existing video anomaly detection technology ignores video characteristics, the model training process has high requirements for computing power, and the number of parameters is huge, resulting in low anomaly detection accuracy and low anomaly detection efficiency. Summary of the Invention

[0004] The present invention provides a video anomaly detection method, device, equipment, storage medium and product to improve the accuracy and detection efficiency of video anomaly detection.

[0005] According to one aspect of the present invention, a video anomaly detection method is provided. The method includes:

[0006] Obtain a sequence of video frames to be verified and a target verification video frame, and determine a sequence of frame differences to be verified according to the sequence of video frames to be verified;

[0007] Input the sequence of video frames to be verified and the sequence of frame differences to be verified into a pre-trained video anomaly detection model respectively to obtain a target predicted video frame output by the model;

[0008] Perform video anomaly detection on the target verification video frame according to the target verification video frame and the target predicted video frame to obtain an anomaly detection result.

[0009] According to another aspect of the present invention, a video anomaly detection device is provided. The device includes:

[0010] A target frame determination module, configured to obtain a sequence of video frames to be verified and a target verification video frame, and determine a sequence of frame differences to be verified according to the sequence of video frames to be verified;

[0011] A frame prediction module, configured to input the sequence of video frames to be verified and the sequence of frame differences to be verified into a pre-trained video anomaly detection model respectively to obtain a target predicted video frame output by the model;

[0012] A video anomaly detection module, configured to perform video anomaly detection on the target verification video frame according to the target verification video frame and the target prediction video frame, so as to obtain an anomaly detection result.

[0013] According to another aspect of the present invention, there is provided an electronic device, including:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the video anomaly detection method according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for implementing the video anomaly detection method according to any embodiment of the present invention when executed by a processor.

[0018] According to another aspect of the present invention, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the steps in the above video anomaly detection method are implemented.

[0019] The technical solution of the embodiment of the present invention adopts a two-time fusion scheme to fuse the spatio-temporal features of the past frame sequence, that is, the video appearance feature and the video motion feature, and fuse the difference between the prediction frames, so as to predict the second future frame to obtain the video anomaly detection result. By comparing the difference between the prediction frame and the actual frame, it is judged whether an abnormal event occurs in the video. Only basic convolutional kernels and autoencoders are used in feature extraction and prediction, and the processing flow of the memory module is simplified, with extremely low requirements for computing power and meeting the demand for high processing speed, thereby improving the video anomaly detection efficiency. In addition, this method retains all complex features of the original video to the greatest extent, obtains the relationship between the video static (appearance) feature and the dynamic (motion) feature by using the memory module, and then improves the video anomaly detection accuracy through the prediction of the second future frame.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0022] Figure 1A is a flowchart of a video anomaly detection method provided in Embodiment 1 of the present invention;

[0023] Figure 1B is a schematic structural diagram of a video anomaly detection model provided in Embodiment 1 of the present invention;

[0024] Figure 2A is an overall architecture diagram of a video anomaly detection method provided in Embodiment 2 of the present invention;

[0025] Figure 2B is an overall architecture diagram of an appearance encoding and decoding module provided in Embodiment 2 of the present invention;

[0026] Figure 2C is an overall architecture diagram of a motion encoding and decoding module provided in Embodiment 2 of the present invention;

[0027] Figure 2D is an overall architecture diagram of a memory module provided in Embodiment 2 of the present invention;

[0028] Figure 3 is a schematic structural diagram of a video anomaly detection device provided in Embodiment 3 of the present invention;

[0029] Figure 4 is a schematic structural diagram of an electronic device for implementing the video anomaly detection method of the embodiments of the present invention. Detailed implementation manners

[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0031] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] Embodiment 1

[0033] Figure 1A FIG. is a flowchart of a video anomaly detection method provided for Embodiment 1 of the present invention. This embodiment is applicable to the situation of performing anomaly detection on a video corresponding to a specific scenario, such as a traffic accident or a natural disaster. This method can be executed by a video anomaly detection device, which can be implemented in the form of hardware and / or software, and the video anomaly detection device can be configured in an electronic device. As Figure 1A shown, the method includes:

[0034] S110. Obtain a sequence of video frames to be verified and a target verification video frame, and determine a sequence of frame differences to be verified according to the sequence of video frames to be verified.

[0035] S120. Input the sequence of video frames to be verified and the sequence of frame differences to be verified into a pre-trained video anomaly detection model respectively, and obtain a target predicted video frame output by the model.

[0036] S130. Perform video anomaly detection on the target verification video frame according to the target verification video frame and the target predicted video frame, and obtain an anomaly detection result.

[0037] Among them, the target verification video frame can be a video frame that needs to be detected for anomalies, and the sequence of video frames to be verified can be video frames that have a sequence order relationship with the target verification video frame. For example, if the target verification video frame is the (t + 2)-th frame, then the sequence of video frames to be verified can be the 1st to t-th frames. By performing anomaly prediction on the sequence of video frames to be verified, it is verified whether the target verification video frame is abnormal, and further whether the video in the video section to which the target verification video frame belongs is abnormal.

[0038] Among them, the sequence of frame differences to be verified can be the differences obtained by subtracting consecutive video frames. Exemplarily, if the sequence of video frames to be verified is F 1 , F 2 , …, Ft+1 Then, subtract the pixels at the same pixel coordinate positions of two adjacent frames to obtain t consecutive frame differences X 1 , X 2 , …, X t . For example, subtract the F t+1 th frame from the F t th frame to obtain the X t th frame difference.

[0039] It should be noted that the frame sizes of each video frame in the video frame sequence to be verified are the same as those of each video frame in the frame difference sequence to be verified, for example, both are 256×256. Specifically, according to factors such as computing power, the frame size can be further compressed or enlarged, but they should all be kept consistent.

[0040] Among them, the video anomaly detection model can be an existing prediction model. To further improve the prediction accuracy, the video anomaly detection model can also be a network model pre-constructed and trained by relevant technical personnel according to actual needs.

[0041] As Figure 1B shown in the schematic diagram of the architecture of a video anomaly detection model. Among them, the video anomaly detection model 10 includes a motion encoding and decoding module 11, an appearance encoding and decoding module 12, and a memory module 13; among them, the motion encoding and decoding module 11 includes a motion encoder 110 and a motion decoder 111; the appearance encoding module 12 includes an appearance encoder 120 and an appearance decoder 121. The output ends of the motion encoder 110 and the appearance encoder 120 are respectively connected to the input end of the memory module 13; the output end of the memory module 13 is respectively connected to the input ends of the motion decoder 111 and the appearance decoder 121. The motion encoder 110 is used to receive the frame difference sequence to be verified, and the appearance encoder 120 is used to receive the video frame sequence to be verified.

[0042] The video anomaly detection model constructed in this embodiment is a prediction model based on a two-stream structure. Two encoders are used to extract and analyze the features of the frame video sequence and the frame difference sequence respectively, and a joint memory module is designed to fuse and match the features of the two sequences. Finally, the two prediction results are fused to obtain the target prediction video frame, so as to perform video anomaly detection.

[0043] Specifically, input the video frame sequence to be verified and the frame difference sequence to be verified into the pre-trained video anomaly detection model respectively to obtain the target prediction video frame output by the model.

[0044] In a specific embodiment, inputting the video frame sequence to be verified and the frame difference sequence to be verified into the pre-trained video anomaly detection model respectively to obtain the target prediction video frame output by the model includes:

[0045] Step a11: Input the video frame sequence to be verified into the appearance encoding and decoding module in the video anomaly detection model. The appearance encoder in the appearance encoding and decoding module performs image feature processing to obtain video appearance features.

[0046] Among them, the appearance encoding and decoding module is used to extract the appearance characteristics of static scenes and objects in video frames and predict the next frame by inputting several consecutive frame sequences of sampling. The appearance encoding and decoding module can adopt the result of the U-Net network model. Both the appearance encoder and the appearance decoder in the appearance encoding and decoding module are constructed based on three residual convolution modules.

[0047] Specifically, input the video frame sequence to be verified into the appearance encoder in the appearance encoding and decoding module for image feature processing to obtain the video appearance features output by the appearance encoder. Exemplarily, assume that the appearance encoder is denoted as E a It is represented that the video frame sequence to be verified is F 1 , F 2 , …, F t , then the expression of the video appearance feature Z a is as follows:

[0048]

[0049] Among them, represents the parameters of the appearance encoder.

[0050] Step a12: Input the frame difference sequence to be verified into the motion encoding and decoding module of the video anomaly detection model. The motion encoder of the motion encoding and decoding module performs image feature processing to obtain video motion features.

[0051] Among them, the motion encoding and decoding module is used to convert the input frame difference sequence for image feature processing and convert it into video motion features.

[0052] Exemplarily, assume that the motion encoder is denoted as E m It is represented that the frame difference sequence to be verified is X 1 , X 2 , …, X t , then the expression of the video motion feature Z m is as follows:

[0053]

[0054] Among them, represents the parameters of the motion encoder.

[0055] Step a2: Input the video appearance features and video motion features into the memory module of the video anomaly detection model for feature fusion processing to obtain video fusion features, and perform feature decomposition on the video fusion features to obtain target appearance features and target motion features.

[0056] Among them, the memory module is used to efficiently learn various normal patterns of appearance and motion. During the model training stage, features highly similar to the memory items are considered as features corresponding to normal events, so they are given higher weights, making the abnormal samples have a greater error from the finally predicted normal samples. Through the joint memory module, the appearance features and motion features can share information, enabling the model to more effectively detect anomalies in the video.

[0057] In a specific implementation, inputting the video appearance features and video motion features into the memory module of the video anomaly detection model for feature fusion processing to obtain video fusion features, and performing feature decomposition on the video fusion features to obtain target appearance features and target motion features includes:

[0058] Step a21: Concatenate the video appearance features and video motion features to obtain a fused query feature.

[0059] Specifically, it can be to fuse the video appearance feature Z a and the video motion feature Z m by connecting the head and tail to obtain the fused query feature Z k .

[0060] Step a22: Determine the feature similarity between the fused query feature and at least one memory item stored in the current time period.

[0061] Calculate the feature similarity degree ω k of the query feature Z i and each memory item m k,i according to the cosine similarity and softmax operation. It should be noted that the memory item can specifically be the fused query feature retained in the historical prediction time period and is continuously iteratively updated based on each prediction time period.

[0062] Step a23: Generate video fusion features based on the feature similarity, each memory item, and the fused query feature.

[0063] Specifically, adjust the fused query feature according to the feature similarity, that is, move Z k closer to m i to generate a new video fusion feature to make it closer to the stored memory items, that is, normal features. Among them, is expressed as follows:

[0064]

[0065] Among them, N is the total number of memory items.

[0066] Step a24: Perform feature decomposition on the video fusion feature to obtain the target appearance feature and the target motion feature.

[0067] Decompose the video fusion feature to obtain the target appearance feature and the target motion feature

[0068] Step a31: Input the target appearance feature into the appearance decoder of the appearance codec module to perform video frame prediction, and obtain the predicted video frame.

[0069] Specifically, the appearance decoder D a generates a predicted video frame according to the target appearance feature

[0070]

[0071] Among them, represents the parameters of the appearance decoder.

[0072] Step a32: Input the target motion feature into the motion decoder of the motion codec module to perform difference frame prediction, and obtain the predicted difference frame.

[0073] Specifically, the motion decoder D m generates a predicted difference frame according to the target motion feature

[0074]

[0075] Among them, represents the parameters of the motion decoder.

[0076] Step a4: Obtain the target predicted video frame according to the predicted video frame and the predicted difference frame.

[0077] For the above predicted video frame and the predicted difference frame directly perform frame fusion to obtain the target predicted video frame

[0078] Perform video anomaly detection on the target verification video frame according to the target verification video frame and the target predicted video frame, and obtain the anomaly detection result.

[0079] ​​Among them, the target verification video frame is the corresponding real frame of the target prediction video frame, and the abnormality degree of the video frame is quantified by calculating the difference between the target prediction video frame and its corresponding real frame (target verification video frame). In this step, the Peak Signal-to-Noise Ratio (PSNR) and its derivative forms can be used as the core indicators for evaluating abnormal events.

[0080] In a specific embodiment, according to the target verification video frame and the target prediction video frame, video abnormality detection is performed on the target verification video frame to obtain an abnormality detection result, including:

[0081] Step b1: Determine the peak signal-to-noise ratio according to the target verification video frame and the target prediction video frame.

[0082] Specifically, the target verification video frame and the target prediction video frame and the peak signal-to-noise ratio t+2 between them and the target verification video frame F are determined as follows:

[0083]

[0084] where N represents the number of pixels per frame.

[0085] Step b2: Determine the target memory item according to the feature similarity between at least one memory item stored in the memory module of the current video frame in the current time period and the fusion query feature.

[0086] Among them, the target memory item m δ is the memory item with the highest feature similarity to the fusion query feature Z among the memory items stored in the memory module of the current video frame in the current time period. k

[0087] Step b3: Determine the L2 distance according to the target memory item and the fusion query feature.

[0088] Specifically, the L2 distance D(z t , m) between the fusion query feature and the target memory item is determined as follows:

[0089]

[0090] where represents the fusion query feature of the t-th frame; m δ is the target memory item; and K is the number of frames.

[0091] It should be noted that if the distance between the fused query feature corresponding to the input video frame and the most similar memory item is too large, it can be considered that an abnormal event has occurred in the video at the t-th frame. This evaluation criterion can effectively avoid the problem that when there is an abnormal event in the video frames of the input video sequence, the query feature corresponding to the abnormal event is misinterpreted as a normal feature by the joint memory module, resulting in a decrease in the PSNR value corresponding to the subsequent abnormal frames.

[0092] Step b4: Determine the target outlier according to the peak signal-to-noise ratio and the L2 distance.

[0093] Among them, the target outlier can be the weighted sum of the peak signal-to-noise ratio (PSNR) and the L2 distance, which is used to evaluate the probability of an abnormal event in the video frame. Among them, β is a preset weight parameter. A high outlier indicates that there is a greater possibility of an abnormal event target in the corresponding frame. The method for determining the target outlier S(t + 2) of the predicted video frame at the (t + 2)-th frame is as follows:

[0094] S(t + 2) = β(1 - P(t + 2)) + (1 - β)D(t);

[0095] Step b5: Perform video anomaly detection on the target verification video frame according to the target outlier to obtain the anomaly detection result.

[0096] The larger the target outlier, the greater the possibility of abnormality of the video frame; the smaller the target outlier, the smaller the possibility of abnormality of the video frame.

[0097] Exemplarily, an outlier threshold can be preset in advance, which can be specifically preset according to actual needs. If the target outlier is greater than the preset outlier threshold, it is considered that there is an abnormality in the video sequence to which the target verification video frame belongs; if the target outlier is not greater than the preset outlier threshold, it is considered that there is no abnormality in the video sequence to which the target verification video frame belongs.

[0098] The technical solution of the embodiment of the present invention realizes the prediction of the second future frame to obtain the video anomaly detection result through a two-time fusion scheme, which fuses the spatio-temporal features of the past frame sequence, that is, the video appearance feature and the video motion feature, and fuses the difference between the predicted frame and the predicted frame. By comparing the difference between the predicted frame and the actual frame, it is judged whether an abnormal event has occurred in the video. Only basic convolutional kernels and autoencoders are used in feature extraction and prediction, and the processing flow of the memory module is simplified, with extremely low requirements for computing power, meeting the demand for high processing speed, thereby improving the video anomaly detection efficiency. In addition, this method retains all the complex features of the original video to the greatest extent, obtains the relationship between the static (appearance) feature and the dynamic (motion) feature of the video by using the memory module, and then improves the video anomaly detection accuracy through the prediction of the second future frame.

[0099] The present invention also provides a model training method for a video anomaly detection model, and the specific steps are as follows:

[0100] Step c1: Obtain a normal video frame sequence sample, and generate a normal difference frame sequence sample according to the normal video frame sequence sample.

[0101] Exemplarily, a video sequence can be extracted from a video that has never had an anomaly as a normal video frame sequence sample, and a difference calculation is performed on consecutive frames in the normal video frame sequence to obtain a normal difference frame sequence sample.

[0102] Step c21: Input the normal video frame sequence sample into the appearance encoding and decoding module in the pre-constructed initial anomaly detection model, and perform image feature processing by the appearance encoder in the appearance encoding and decoding module to obtain predicted appearance features.

[0103] Step c22: Input the normal difference frame sequence sample into the motion encoding and decoding module of the initial anomaly detection model, and perform image feature processing by the motion encoder in the motion encoding and decoding module to obtain predicted motion features.

[0104] Step c3: Input the predicted appearance features and the predicted motion features into the memory module in the initial anomaly detection model for feature fusion processing to obtain predicted fusion features, and perform feature decomposition on the predicted fusion features to obtain predicted appearance features and predicted motion features.

[0105] Step c41: Input the predicted appearance features into the appearance decoder of the appearance encoding and decoding module for video frame prediction to obtain a reference video frame.

[0106] Step c42: Input the predicted motion features into the motion decoder of the motion encoding and decoding module for difference frame prediction to obtain a reference difference frame.

[0107] Step c5: Determine a first loss value according to the predicted video frame and the corresponding normal video frame, and determine a second loss value according to the reference difference frame and the corresponding normal difference frame.

[0108] Specifically, if the predicted video frame is and its corresponding normal video frame is F t+1 , then the determination method of the first loss value L a of the appearance encoding and decoding module is as follows:

[0109]

[0110] If the reference difference frame is and its corresponding normal difference frame is X t+1 , then the determination method of the second loss value L m of the motion encoding and decoding module is as follows:

[0111]

[0112] Step c6: According to the first loss value and the second loss value, train the initial anomaly detection model until the preset model training end condition is met, and obtain the video anomaly detection model.

[0113] Among them, the model training end condition can be that the number of iterations reaches the threshold, or the loss value tends to be stable or reaches the preset loss value threshold.

[0114] Specifically, if both the first loss value and the second loss value are greater than the preset loss value threshold, it can be considered that the model training reaches the optimal state, end the model training, and obtain the video anomaly detection model.

[0115] Embodiment 2

[0116] Figure 2A This is the overall architecture diagram of a video anomaly detection method provided by Embodiment 2 of the present invention. Based on the above embodiment, this embodiment provides a preferred example.

[0117] The video anomaly detection model in this embodiment is an anomaly detection model based on a two-stream structure. The basic principle is to assume that most of the content of the video should be non-abnormal. Therefore, the future video frames predicted based on normal video segments should also tend to be normal. Finally, by comparing the difference between the predicted frame and the actual frame, it is judged whether an abnormal event occurs in the video.

[0118] The appearance encoding and decoding module is used to extract the appearance features of static scenes and objects in the video sequence frames, and predict the next frame by inputting several consecutive sampled frames. As Figure 2B shown in the overall architecture diagram of an appearance encoding and decoding module. This structure adopts the U-Net structure, where the two-dimensional (2D) convolutional neural network is used as the backbone network of the appearance encoding and decoding module. Both the appearance encoder and the appearance decoder are constructed based on three residual convolutional modules. For the appearance encoder, only batch normalization is used after the module. For the appearance decoder, the Pixel shuffle layer is used to replace the convolutional layer after each module to improve the resolution. In addition, all modules of the appearance encoder and the appearance decoder use LeakyReLU as the activation layer. Between the same-resolution feature layers of the appearance encoder and the appearance decoder, skip connections are used to suppress the gradient explosion problem.

[0119] Specifically, adjust F 1 , F 2 , …, F t into a stacked video frame sequence of 256×256, and the specific frame size can be further compressed or enlarged based on specific computing power considerations. The video frame sequence is F 1 , F2 ,…,F t Input to appearance encoder E a , after being processed by the appearance encoder, the video appearance feature Z is obtained a :

[0120]

[0121] in, Represents the parameters of the appearance encoder.

[0122] The motion encoder extracts features from the video difference sequence and outputs the video motion feature Z m , video appearance feature Z a and video motion feature Z m Generate fusion features through memory module Will Decompose to obtain target appearance features The detailed processing methods of the motion codec module and the memory module are described in the subsequent sections and will not be described here. a According to the target appearance characteristics Generate prediction frame The specific process is as follows:

[0123]

[0124] in, Represents the parameters of the appearance decoder.

[0125] During the training phase, the goal is to minimize the predicted frame With its true value F t+1 The L2 distance between them makes them closer, so the loss function L of the appearance encoding and decoding module a The definition is as follows:

[0126]

[0127] Through the loss value L a The appearance encoding and decoding module is continuously trained until the loss value stabilizes or reaches the set threshold.

[0128] like Figure 2C The overall architecture diagram of a motion codec module is shown in FIG. In the motion autoencoder, the technology firstly generates a motion codec in the continuous video frame F 1 ,F 2 ,…,F t ,F t+1 Subtract the pixels of the t+1th frame from the tth frame to obtain t consecutive frame differences X 1 ,X 2 ,…,X tAdjust the continuous frame differences to stacked differences of size 256×256. The frame size can be further compressed or enlarged based on specific computing power considerations, but it should be consistent with the frame size in the appearance encoding and decoding module. Also, the number of frames remains the same.

[0129] Specifically, input the continuous frame differences X 1 , X 2 , …, X t into the motion encoder E m and obtain the video motion feature Z after being processed by the motion encoder m :

[0130]

[0131] where represents the parameters of the motion encoder.

[0132] The video appearance feature Z a and the video motion feature Z m generate a fused feature through the memory module Decompose to obtain the target appearance feature Decompose to obtain the target motion feature The motion decoder D m generates a predicted frame according to the target appearance feature The specific process is as follows: Specifically:

[0133]

[0134] where represents the parameters of the motion decoder.

[0135] In the training stage, to make the predicted frame differences closer to the real differences X t+1 , control to minimize the L2 distance between them. Therefore, the loss function L of the motion encoding and decoding module m is defined as follows:

[0136]

[0137] Continuously train the model of the motion encoding and decoding module through the loss value L m until the loss value tends to be stable or reaches the set threshold.

[0138] For example Figure 2DThe overall architecture diagram of a memory module is shown. It is mainly used to efficiently learn various normal patterns of video appearance and video motion. During the training phase, features highly similar to the memory items can be considered as the features corresponding to normal events. Therefore, they are assigned higher weights, causing a larger error between abnormal samples and the normal samples finally predicted. Through the memory module, appearance and motion features can share information, enabling the model to more effectively detect abnormal situations in the video. In this embodiment, the structure of the memory module is simplified, and only one layer of the memory module is used to obtain the final fused feature.

[0139] The specific working principle of the memory module is as follows:

[0140] Query feature of the memory module: Fuse the video appearance feature Z a and the video motion feature Z m by connecting them head to tail to obtain the query feature Z k .

[0141] Memory items: The memory module contains multiple memory items m n , which are responsible for storing different normal event features.

[0142] Reading operation of the memory module: Calculate the similarity degree ω k between the query feature Z i and each memory item m k,i according to the cosine similarity and softmax operation, and adjust the query feature according to the similarity, that is, move Z k closer to m i to generate a new feature to make it closer to the stored memory items, that is, normal features. That is:

[0143]

[0144] where N is the total number of memory items.

[0145] Updating operation of the memory module: Calculate the similarity degree v i between each memory item m k and the query feature Z k,i according to the cosine similarity and softmax operation, and adjust the memory item according to the similarity. The operation process is opposite to the reading.

[0146] Fuse the and in the above steps directly to obtain a new predicted frame and use this as the final prediction result. Calculate the difference between it and the real frame F t+2 to quantify the abnormality degree of the video frame. In this step, the peak signal-to-noise ratio (PSNR) and its derivative forms are used as the core metrics for evaluating abnormal events. The specific operation is as follows:

[0147] Calculate the PSNR: where N is the number of pixels per frame. A lower PSNR value means that the frame is more likely to be abnormal.

[0148]

[0149] Where N represents the number of pixels per frame.

[0150] Calculate Z k and m δ The L2 distance between them: The memory items in the memory module store various normal patterns. During the test phase, when the input video frame is normal, its query feature should be highly similar to the memory items, while for abnormal samples, the opposite is true. Therefore, in addition to the PSNR value, the L2 distance between the query feature Z k and the nearest similar item m δ can also be calculated, that is:

[0151]

[0152] Where is the query feature of the t-th frame, and m δ is the memory item most similar to . If the distance between the query feature corresponding to the input video frame and the most similar memory item is too large, it can be considered that an abnormal event has occurred in the video at the t-th frame. This evaluation criterion can effectively avoid the problem that when there is an abnormal event in the video frames of the input video sequence, the query feature corresponding to the abnormal event is misinterpreted as a normal feature by the joint memory module, resulting in a decrease in the PSNR value corresponding to the subsequent abnormal frames.

[0153] The final abnormal score is the weighted sum of the PSNR and the L2 distance, which is used to evaluate the probability of an abnormal event in the video frame. Here, β is a weight parameter. A high abnormal score indicates that an abnormal event is more likely to exist in the corresponding frame. The target abnormal value S(t + 2) of the predicted video frame at the (t + 2)-th frame is determined as follows:

[0154] S(t + 2) = β(1 - P(t + 2))+(1 - β)D(t);

[0155] In this embodiment, through a two - fusion scheme, the spatio - temporal features (i.e., appearance features and motion features) of the past frame sequence are fused, and the difference between the predicted frames is fused, so as to predict the second future frame and obtain the video anomaly detection result. First, the video is pre - processed to obtain the sampled frame sequence and the difference sequence of the corresponding frames. Two encoders are used to extract and analyze the features of these two sequences respectively, and a joint memory module is designed to fuse and match the features of the two sequences. Then, the matched features are re - decomposed and applied to the two sequences respectively to obtain a predicted frame and a predicted difference. Finally, the predicted frame and the predicted difference are fused to obtain the final predicted frame. By comparing the difference between the predicted frame and the actual frame, it is judged whether an abnormal event occurs in the video. This method only uses basic convolutional kernels and auto - encoders for feature extraction and prediction, and greatly simplifies the processing flow of the memory module, with extremely low requirements for computing power and meeting the demand for high processing speed. In addition, this method retains all the complex features of the original video to the greatest extent, obtains the relationship between the static (appearance) features and dynamic (motion) features of the video using the hybrid memory module, and improves the anomaly detection accuracy through the prediction of the second future frame.

[0156] Compared with the traditional two - stream prediction model, in this embodiment, frame difference is used instead of optical flow as the input on the motion feature stream, greatly reducing the computing power requirements of the model. At the same time, the joint memory module is simplified, reducing the feature fusion and processing costs and accelerating the processing speed. Through two fusions, on the basis of feature fusion, the fusion of the predicted frame and the predicted difference is added, achieving efficient and accurate prediction.

[0157] Embodiment III

[0158] Figure 3 It is a schematic structural diagram of a video anomaly detection device provided in Embodiment III of the present invention. A video anomaly detection device provided in an embodiment of the present invention is applicable to the situation of anomaly detection of videos corresponding to specific scenarios, such as traffic accidents or natural disasters. This video anomaly detection device can be implemented in the form of hardware and / or software, such as Figure 3 As shown, this video anomaly detection device can be configured in an electronic device and specifically includes: a target frame determination module 301, a frame prediction module 302, and a video anomaly detection module 303. Among them,

[0159] The target frame determination module 301 is used to obtain the to - be - verified video frame sequence and the target verification video frame, and determine the to - be - verified frame difference sequence according to the to - be - verified video frame sequence;

[0160] The frame prediction module 302 is used to input the to - be - verified video frame sequence and the to - be - verified frame difference sequence into a pre - trained video anomaly detection model respectively to obtain the target predicted video frame output by the model;

[0161] The video anomaly detection module 303 is configured to perform video anomaly detection on the target verification video frame according to the target verification video frame and the target prediction video frame, so as to obtain an anomaly detection result.

[0162] The technical solution of the embodiment of the present invention adopts a two - stage fusion scheme to fuse the spatio - temporal features of the past frame sequence, that is, the video appearance feature and the video motion feature, and fuse the difference between the prediction frames, so as to predict the second future frame to obtain the video anomaly detection result. By comparing the difference between the predicted frame and the actual frame, it is determined whether an abnormal event has occurred in the video. Only basic convolutional kernels and auto - encoders are used in feature extraction and prediction, and the processing flow of the memory module is simplified, with extremely low requirements for computing power, meeting the demand for high processing speed, thereby improving the video anomaly detection efficiency. In addition, this method maximally retains all complex features of the original video, obtains the relationship between the static (appearance) feature and the dynamic (motion) feature of the video by using the memory module, and then improves the video anomaly detection accuracy through the prediction of the second future frame.

[0163] Optionally, the video anomaly detection model includes a motion encoding - decoding module, an appearance encoding - decoding module, and a memory module; the motion encoding - decoding module includes a motion encoder and a motion decoder; the appearance encoding module includes an appearance encoder and an appearance decoder.

[0164] Optionally, the frame prediction module 302 includes:

[0165] An appearance feature prediction unit, configured to input the sequence of video frames to be verified into the appearance encoding - decoding module in the video anomaly detection model, and perform image feature processing on the appearance encoder in the appearance encoding - decoding module to obtain the video appearance feature; and

[0166] A motion feature prediction unit, configured to input the sequence of differences between the frames to be verified into the motion encoding - decoding module in the video anomaly detection model, and perform image feature processing on the motion encoder in the motion encoding - decoding module to obtain the video motion feature;

[0167] A feature fusion unit, configured to input the video appearance feature and the video motion feature into the memory module of the video anomaly detection model for feature fusion processing, obtain the video fusion feature, and perform feature decomposition on the video fusion feature to obtain the target appearance feature and the target motion feature;

[0168] A video frame prediction unit, configured to input the target appearance feature into the appearance decoder of the appearance encoding - decoding module for video frame prediction to obtain a predicted video frame; and

[0169] A frame difference prediction unit for inputting the target motion feature into a motion decoder of a motion coding and decoding module to perform difference frame prediction and obtain a predicted difference frame;

[0170] A target frame determination unit for obtaining a target predicted video frame according to the predicted video frame and the predicted difference frame.

[0171] Optionally, the feature fusion unit is specifically configured to:

[0172] Perform feature splicing on the video appearance feature and the video motion feature to obtain a fused query feature;

[0173] Determine the feature similarity between the fused query feature and at least one memory item stored in the current time period;

[0174] Generate a video fusion feature according to the feature similarity, each memory item, and the fused query feature;

[0175] Perform feature decomposition on the video fusion feature to obtain a target appearance feature and a target motion feature.

[0176] Optionally, the video anomaly detection module 303 is specifically configured to:

[0177] Determine the peak signal-to-noise ratio according to the target verification video frame and the target predicted video frame;

[0178] Determine a target memory item according to the feature similarity between at least one memory item stored in the memory module of the current video frame in the current time period and the fused query feature;

[0179] Determine the L2 distance according to the target memory item and the fused query feature;

[0180] Determine a target outlier according to the peak signal-to-noise ratio and the L2 distance;

[0181] Perform video anomaly detection on the target verification video frame according to the target outlier to obtain an anomaly detection result.

[0182] Optionally, the device further includes a model training module, and the model training module is specifically configured to:

[0183] Obtain a normal video frame sequence sample, and generate a normal difference frame sequence sample according to the normal video frame sequence sample;

[0184] Input the normal video frame sequence sample into an appearance coding and decoding module in a pre-constructed initial anomaly detection model, and perform image feature processing by an appearance encoder in the appearance coding and decoding module to obtain a predicted appearance feature; and,

[0185] Input the normal difference frame sequence sample into the motion encoding and decoding module of the initial anomaly detection model. The motion encoder of the motion encoding and decoding module processes the image features to obtain predicted motion features;

[0186] Input the predicted appearance features and the predicted motion features into the memory module in the initial anomaly detection model for feature fusion processing to obtain predicted fusion features, and perform feature decomposition on the predicted fusion features to obtain predicted appearance features and predicted motion features;

[0187] Input the predicted appearance features into the appearance decoder of the appearance encoding and decoding module for video frame prediction to obtain a reference video frame; and,

[0188] Input the predicted motion features into the motion decoder of the motion encoding and decoding module for difference frame prediction to obtain a reference difference frame;

[0189] Determine a first loss value according to the predicted video frame and the corresponding normal video frame, and determine a second loss value according to the reference difference frame and the corresponding normal difference frame;

[0190] Train the initial anomaly detection model according to the first loss value and the second loss value until a preset model training end condition is satisfied to obtain a video anomaly detection model.

[0191] The video anomaly detection device provided by the embodiments of the present invention can execute the video anomaly detection method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0192] Embodiment 4

[0193] Figure 4 FIG. shows a schematic structural diagram of an electronic device 40 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0194] As Figure 4As shown, the electronic device 40 includes at least one processor 41 and a memory communicatively connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. The memory stores a computer program executable by the at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0195] Multiple components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disc, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0196] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the video anomaly detection method.

[0197] In some embodiments, the video anomaly detection method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the video anomaly detection method described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the video anomaly detection method by any other appropriate means (e.g., by means of firmware).

[0198] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0199] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0200] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0201] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0202] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0203] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0204] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0205] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A video anomaly detection method, characterized in that: include: Acquire a sequence of video frames to be verified and a target verification video frame, and determine a sequence of frame difference values ​​to be verified according to the sequence of video frames to be verified; Inputting the video frame sequence to be verified into the appearance codec module in the video anomaly detection model, and performing image feature processing by the appearance encoder in the appearance codec module to obtain video appearance features; and The frame difference sequence to be verified is input into the motion codec module of the video anomaly detection model, and the motion encoder of the motion codec module performs image feature processing to obtain video motion features; Inputting the video appearance features and the video motion features into the memory module of the video anomaly detection model for feature fusion processing to obtain video fusion features, and performing feature decomposition on the video fusion features to obtain target appearance features and target motion features; Inputting the target appearance feature into an appearance decoder of an appearance codec module to perform video frame prediction to obtain a predicted video frame; and Inputting the target motion feature into a motion decoder of a motion coding module to perform difference frame prediction to obtain a predicted difference frame; Obtaining a target predicted video frame according to the predicted video frame and the predicted difference frame; According to the target verification video frame and the target prediction video frame, video anomaly detection is performed on the target verification video frame to obtain an anomaly detection result.

2. The method according to claim 1, characterized in that The step of inputting the video appearance features and the video motion features into the memory module of the video anomaly detection model for feature fusion processing to obtain video fusion features, and performing feature decomposition on the video fusion features to obtain target appearance features and target motion features includes: Performing feature splicing on the video appearance feature and the video motion feature to obtain a fused query feature; Determining a feature similarity between the fused query feature and at least one memory item stored in the current time period; Generate a video fusion feature according to the feature similarity, each of the memory items and the fusion query feature; The video fusion features are decomposed to obtain target appearance features and target motion features.

3. The method according to claim 2, characterized in that The performing video anomaly detection on the target verification video frame according to the target verification video frame and the target prediction video frame to obtain an anomaly detection result includes: Determining a peak signal-to-noise ratio according to the target verification video frame and the target prediction video frame; Determine a target memory item according to a feature similarity between at least one memory item stored in a memory module of a current video frame in a current time period and a fusion query feature; Determining an L2 distance based on the target memory item and the fused query feature; Determining a target outlier according to the peak signal-to-noise ratio and the L2 distance; According to the target abnormal value, video abnormality detection is performed on the target verification video frame to obtain an abnormality detection result.

4. The method according to claim 1, characterized in that: The model training method of the video anomaly detection model is as follows: Acquire normal video frame sequence samples, and generate normal difference frame sequence samples according to the normal video frame sequence samples; Inputting the normal video frame sequence samples into the appearance codec module in the pre-built initial anomaly detection model, the appearance encoder in the appearance codec module performs image feature processing to obtain predicted appearance features; and, Inputting the normal difference frame sequence samples into the motion coding module of the initial anomaly detection model, and performing image feature processing by the motion encoder of the motion coding module to obtain predicted motion features; Inputting the predicted appearance feature and the predicted motion feature into the memory module in the initial anomaly detection model for feature fusion processing to obtain a predicted fusion feature, and performing feature decomposition on the predicted fusion feature to obtain a predicted appearance feature and a predicted motion feature; Inputting the predicted appearance features into an appearance decoder of an appearance codec module to perform video frame prediction to obtain a reference video frame; and Inputting the predicted motion features into a motion decoder of a motion coding module to perform difference frame prediction to obtain a reference difference frame; Determining a first loss value according to the predicted video frame and the corresponding normal video frame, and determining a second loss value according to the reference difference frame and the corresponding normal difference frame; According to the first loss value and the second loss value, the initial anomaly detection model is trained until a preset model training end condition is met to obtain a video anomaly detection model.

5. A video anomaly detection device, characterized in that: include: A target frame determination module is used to obtain a sequence of video frames to be verified and a target verification video frame, and determine a sequence of frame difference values ​​to be verified based on the sequence of video frames to be verified; The video anomaly detection model includes a motion codec module, an appearance codec module and a memory module; the motion codec module includes a motion encoder and a motion decoder; the appearance codec module includes an appearance encoder and an appearance decoder; A frame prediction module, used to input the video frame sequence to be verified and the frame difference sequence to be verified into a pre-trained video anomaly detection model, respectively, to obtain a target predicted video frame output by the model; A video anomaly detection module, used to perform video anomaly detection on the target verification video frame according to the target verification video frame and the target prediction video frame, to obtain an anomaly detection result; The frame prediction module, include: an appearance feature prediction unit, used for inputting the video frame sequence to be verified into the appearance codec module in the video anomaly detection model, and the appearance encoder in the appearance codec module performs image feature processing to obtain video appearance features; and A motion feature prediction unit, used for inputting the frame difference sequence to be verified into a motion codec module of a video anomaly detection model, and performing image feature processing by a motion encoder of the motion codec module to obtain video motion features; A feature fusion unit, used for inputting the video appearance feature and the video motion feature into the memory module of the video anomaly detection model for feature fusion processing to obtain a video fusion feature, and performing feature decomposition on the video fusion feature to obtain a target appearance feature and a target motion feature; A video frame prediction unit, configured to input the target appearance feature into an appearance decoder of an appearance codec module to perform video frame prediction to obtain a predicted video frame; as well as, A frame difference prediction unit, used for inputting the target motion feature into a motion decoder of a motion coding module to perform difference frame prediction to obtain a predicted difference frame; The target frame determination unit is used to obtain a target predicted video frame according to the predicted video frame and the predicted difference frame.

6. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the video anomaly detection method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the video anomaly detection method according to any one of claims 1 to 4 when executed.

8. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the video anomaly detection method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Abnormal behavior detection method based on appearance and action feature dual prediction

    CN113762007A