Unsupervised video anomaly detection method, device and equipment based on multi-level perception
Through the feature knowledge distillation technology based on multi-level perception, a normal model of multi-level feature video event is constructed, which solves the problem of difficulty in perceiving the difference in abnormal events in unsupervised video abnormality detection, and realizes high-precision video abnormality detection.
Patent Information
- Application Number
- CN202510526985.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
AI Technical Summary
The existing unsupervised video anomaly detection method is difficult to effectively perceive the differences between abnormal events and normal events at multiple levels, resulting in low detection accuracy.
Using a multi-level perception method, a high-quality multi-level feature video event normal model is constructed through feature knowledge distillation technology, and a pre-trained deep neural network is used as a teacher network to train a lightweight student network to produce multi-level features similar to the teacher network, thereby realizing video anomaly detection.
It realizes high-precision detection of abnormal events in video, can more accurately portray video events, reduce interference from background redundant information, make full use of multi-level features for modeling, and fit the process of human perception of abnormal events.
Smart Images

Figure CN120047879A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image or video recognition, and particularly relates to video anomaly detection technology, and more particularly to an unsupervised video anomaly detection method, device and equipment based on multi-level perception. Background Art
[0002] With the rapid development of social economy and the increasing demand for public safety, a large number of surveillance cameras have been deeply installed in various places, providing important support for security, safety, municipal management, information collection, etc. One of the core tasks of installing cameras for video surveillance is to monitor the occurrence of abnormal events and alarm for abnormal situations. Traditional video surveillance mainly relies on manual monitoring and recognition, and it is becoming increasingly difficult to continue in the face of the surging surveillance video data. Therefore, the development of video anomaly detection (VAD) technology that can automatically analyze video content and efficiently detect abnormal events has great social and economic benefits.
[0003] Due to the low frequency of abnormal events, it is difficult to collect sufficient abnormal events for analysis and modeling. Currently, the mainstream video anomaly detection methods follow a semi-supervised modeling process, which only collects a large number of redundant normal videos to construct a training data set, and trains a normal model that can describe normal events based on the normal video training set. Thus, during the abnormal inference process, events that do not conform to the description of the normal model are determined as abnormal. However, it is time-consuming, laborious and even unrealistic in the face of the surging surveillance videos. A more practical approach is to achieve unsupervised video anomaly detection, which aims to directly model and detect based on completely unlabeled surveillance video data. Unsupervised video anomaly detection does not rely on normal videos. Although certain progress has been made, the current unsupervised video anomaly detection methods still face a significant problem: it is difficult to perceive the differences between abnormal events and normal events at multiple levels. The current unsupervised methods generally input the original pixels or their corresponding features into a deep neural network, and perform modeling and detection based on the output and input of the network. However, this approach has obvious defects: on the one hand, the original pixels are sensitive to factors such as noise and environmental changes, and do not consider high-level semantics; at the same time, the separately extracted features may lose the key information that can distinguish abnormal events, and may ignore the lower-level differences due to overly high-level abstract representations; on the other hand, only relying on the neural network and input information in modeling and inference ignores the multi-level features generated by the hidden layer of the network, and cannot fully capture the differences of abnormal events at multiple levels.
[0004] Based on the multi-level perception-based feature modeling technology, the hierarchical representation learning ability of the deep neural network can be utilized to capture feature information at different levels. For example, the network layers closer to the input layer tend to capture lower-level edge information, while the higher network layers learn high-level semantics that are irrelevant to detailed changes. Therefore, it is necessary to provide an unsupervised video anomaly detection method based on multi-level perception, and construct an unsupervised anomaly detection model by modeling multi-level features to achieve high-precision detection of abnormal events in videos. Summary of the Invention
[0005] In order to effectively solve the above problems existing in the prior art, the present invention provides an unsupervised video anomaly detection method, device and equipment based on multi-level perception, which can construct a high-quality video event normal state model based on multi-level features based on unlabeled video data through feature-based knowledge distillation technology, and then achieve high-precision video anomaly detection based on multi-level perception.
[0006] The present invention provides an unsupervised video anomaly detection method based on multi-level perception, including: Step 110: Obtain unlabeled video data; the video events in the video data include normal events as the main part of the video data and abnormal events as the impurity part of the video data; Step 120: Use a pre-trained object detector and combine motion cues to locate foreground objects in the video data; Step 130: Extract image blocks of multiple consecutive frames before and after the located foreground object, stack them after scaling to a unified size, and construct a spatio-temporal cube to describe the video event; Step 140: Use feature-based knowledge distillation technology to distill the knowledge of the main part in the video data and construct a video event normal state model; the feature-based knowledge distillation technology includes using a pre-trained deep neural network as a teacher network and training a lightweight neural network described by the spatio-temporal cube as a student network to make the student network generate multi-level features similar to those of the teacher network; Step 150: Construct a spatio-temporal cube of the video data to be detected, and input the spatio-temporal cube into the teacher network and the student network trained during the construction of the video event normal state model; use the difference between the multi-level features output by the teacher network and the student network to perform anomaly scoring on each spatio-temporal cube to achieve anomaly detection of the video data to be detected.
[0007] On the other hand, the present invention provides an unsupervised video anomaly detection device based on multi-level perception. The device uses the foregoing method to implement unsupervised video anomaly detection based on multi-level perception. The device includes a video data acquisition module, a foreground target localization module, a video event description module, a feature-based knowledge distillation module, and an anomaly detection module, wherein: The video data acquisition module is used to acquire unlabeled video data; the video data is unlabeled and mixed with unlabeled video events, including normal events as the main part of the video data and abnormal events as the impurity part of the video data; The foreground target localization module is used to localize the foreground target in the video data by using a pre-trained object detector and combining motion cues; The video event description module is used to extract image patches of multiple consecutive frames before and after the localized foreground target, scale the image patches of the multiple consecutive frames to a unified size and stack them to construct a spatio-temporal cube for describing a video event; The feature-based knowledge distillation module is used to distill the knowledge of the main part of the video data by using feature-based knowledge distillation technology and construct a normal model of video events; the feature-based knowledge distillation technology includes using a pre-trained deep neural network as a teacher network and training a lightweight neural network as a student network to make the student network generate multi-level features similar to those of the teacher network; The anomaly detection module is used to construct a spatio-temporal cube of the video data to be detected, input the spatio-temporal cube into the teacher network and the student network trained when constructing the normal model of video events; and use the difference between the multi-level features output by the teacher network and the student network to perform anomaly scoring on each spatio-temporal cube to achieve anomaly detection of the video data to be detected.
[0008] The present invention also protects a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the foregoing unsupervised video anomaly detection method based on multi-level perception are implemented.
[0009] Compared with the prior art, an unsupervised video anomaly detection method, device, and equipment based on multi-level perception provided by the present invention have the following beneficial technical effects: (1) First, the present invention first marks the foreground target in the video data, and then constructs a spatio-temporal cube to more accurately depict the video event and reduce the interference of background redundant information.
[0010] (2) Secondly, the constructed spatio-temporal cube is input into the teacher network and the student network respectively. By minimizing the distance between the multi-level features extracted by the two networks, a high-quality video event normal model is constructed. To quickly reduce the training loss, due to the low frequency, novelty, and openness of abnormal events, the parameter optimization process of the student network generally gives priority to generating features similar to those of the teacher network for a large number of redundant normal events. At the same time, due to the limited capacity and ability of the student network, it is unable to generate multi-level features similar to those of the teacher network for abnormal events. Therefore, during the distillation process, the knowledge or information corresponding to the normal events that make up the majority of the unlabeled video data will be distilled into the small-scale student network, while the "impurity part" (information corresponding to abnormal events) will be retained, realizing the construction of a high-quality normal model based on the mixed unlabeled video data.
[0011] (3) The powerful hierarchical representation learning ability of deep neural networks makes it possible to capture feature perception information at different levels in modeling and detection. Therefore, during the process of constructing a video event normal model using feature-based knowledge distillation technology, multi-level features from low to high can be fully utilized for modeling. Moreover, this modeling method based on multi-level features enables anomaly detection to also combine multi-level features, thereby realizing video anomaly detection based on multi-level perception, which is more in line with the process of human perception of abnormal events. Description of the Drawings
[0012] Figure 1 It is a schematic flowchart of the unsupervised video anomaly detection method based on multi-level perception in the first embodiment of the present invention; Figure 2 It is a schematic flowchart of constructing a high-quality video event normal model using feature-based knowledge distillation technology in the first embodiment of the present invention, where, is the feature extracted from the th intermediate layer of the teacher network , is the feature extracted from the th network layer of the student network , , is the number of network layer groups selected from the teacher network and the student network, and are the conversion layers for the th network layer in the teacher network and the student network respectively, represents the distance between the multi-level features of the teacher network and the student network; Figure 3 It is a schematic structural diagram of the unsupervised video anomaly detection device based on multi-level perception in the third embodiment of the present invention; Figure 4Schematic diagram of the internal structure of a computer device in an embodiment of the present invention. Detailed implementation manners
[0013] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0014] The core objective of video anomaly detection is to analyze the motion patterns, spatio-temporal relationships, or context semantics, etc. of foreground objects in video data by using data extraction and analysis means, and to determine whether each video event of the foreground object is a normal event or an abnormal event. The foreground object refers to the target object worthy of attention in the video frame except for the video background, such as vehicles, pedestrians or other observed objects; the video event describes the motion or development process of the target object in a short period of time, such as the driving of a vehicle, the walking of a pedestrian, or the movement of other observed objects, etc.; the normal event refers to the video event that conforms to the convention and scene expectation, such as the normal driving of a vehicle, the normal walking of a pedestrian, etc.; the abnormal event refers to the video event that violates the convention and deviates from the expectation, such as cycling on the sidewalk, abnormal movement of the crowd, climbing over the fence, falling, etc.
[0015] Inspired by a deep learning technology called knowledge distillation, the present invention uses a feature-based knowledge distillation technology to construct a video event normal model and designs an unsupervised video anomaly detection method based on multi-level perception differences, aiming to distill the knowledge about normal events (the main part) from the video data mixed with normal events and abnormal events without manual screening (unlabeled), and to detect abnormal events by establishing a normal model and making full use of the feature (multi-level features) representations generated by each network layer during the distillation process.
[0016] The present invention will be further described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0017] In the first embodiment, as Figure 1 shown, the present invention provides an unsupervised video anomaly detection method based on multi-level perception, including: Step 110: Obtain unlabeled video data; the video events in the video data include normal events as the main part of the video data and abnormal events as the impurity part of the video data; Step 120: Use a pre-trained object detector and combine with motion cues to locate the foreground objects in the video data; Step 130: Extract the image blocks of multiple consecutive frames before and after the located foreground objects, scale the image blocks of the multiple consecutive frames before and after to the same size and stack them to construct a spatio-temporal cube for describing video events; Step 140: Use feature-based knowledge distillation technology to distill the knowledge of the main part in the video data and construct a normal model of video events; the feature-based knowledge distillation technology includes using a pre-trained deep neural network as the teacher network, and training a lightweight neural network described by a spatio-temporal cube with video events to make the student network generate multi-level features similar to those of the teacher network; Step 150: Construct a spatio-temporal cube of the video data to be detected, and input the spatio-temporal cube into the teacher network and the student network trained during the construction of the video event normal model; use the difference between the multi-level features output by the teacher network and the student network to perform anomaly scoring on each spatio-temporal cube to achieve anomaly detection of the video data to be detected.
[0018] Specifically, in step 120, for the input unlabeled video data, the pre-trained object detector combines motion cues to detect and locate the foreground objects in the video data frame by frame, and marks the foreground objects (such as humans, vehicles, etc.) with detection boxes, which helps to reduce the interference of the video background.
[0019] The pre-trained object detector can be selected from Fast R-CNN, YOLO, Mask R-CNN, etc.
[0020] Furthermore, in step 130, for each foreground object located by the detection box, obtain image patches of the video foreground including a total of frames before and after the current frame; after scaling the consecutive image patches to a unified size (for example, the pixel is set to ), construct a spatio-temporal cube (spatial temporal cube, STC) to describe the video event, and use to represent the video event, which serves as the basic unit for subsequent modeling and data processing.
[0021] Generally speaking, Knowledge Distillation refers to a method that can transfer knowledge or information from one network (the teacher network) to another network (the student network), following the teacher-student framework. Under this framework, the Teacher Network is a large pre-trained network, that is, a pre-trained neural network containing initial knowledge information; while the Student Network is a network with a smaller scale than the teacher network for receiving knowledge information, and its goal is to learn rich knowledge by mimicking the behavior of the teacher network and achieve performance similar to that of the teacher network. The core idea of this method is to guide the learning of the student network through the output of the teacher network (such as probability distribution), so that the student network can significantly reduce the computational cost while maintaining high performance.
[0022] Furthermore, Feature based Distillation uses the feature information extracted from each network layer in the deep neural network for distillation. It uses the feature information extracted by the teacher network in the intermediate layer as a supervision signal to guide the training of the student network, enabling the student network to learn the feature representation generated by the teacher network, thereby achieving the lightweight and high efficiency of the model (network) while maintaining or approaching the performance of the teacher network. The application of the above technologies enables the present invention to make full use of the feature representations at all levels of video data, which endows the present invention with the ability to detect abnormal events facing multi-level perception differences.
[0023] In order to extract multi-level features to achieve unsupervised video anomaly detection for multi-level perception, it is necessary to give full play to the powerful hierarchical representation learning ability of the deep neural network, which requires extracting the corresponding features of each network layer and combining them with the normal modeling and anomaly detection of video events.
[0024] The following considerations are mainly taken into account for such modeling: (1) During the training of the student network, due to the characteristics of abnormal events being low-frequency, novel, and open, the process of optimizing the parameters of the student network will give priority to a large number of redundant normal events and extract features similar to those of the teacher network for normal events. At the same time, due to limited capacity and ability, the student network often fails to extract multi-level features similar to those of the teacher network for abnormal events. Therefore, during the feature distillation process, the knowledge contained in the normal events in the unlabeled video set will be distilled and transferred to the student network, while the abnormal events will be retained, thereby constructing a normal model in the case of mixed normal and abnormal events.
[0025] (2) In feature - based knowledge distillation, multiple network - layer features are involved, and these features can capture perceptual information at different levels. When constructing a video event normal model based on feature - based knowledge distillation, the multi - level features from low to high of a deep neural network can be utilized for modeling. This lays a foundation for detecting abnormal events based on the differences in multi - level perception by leveraging the differences in multi - level features during anomaly inference.
[0026] (3) Select a pre - trained network (such as a model for tasks like action recognition and classification) as the teacher network to provide multi - level features for the student network to imitate and learn. The pre - trained deep neural network can provide rich high - level semantic features, enabling the construction of the video event normal model and anomaly inference to incorporate high - level semantic information, which provides the possibility for detecting abnormal events based on high - level semantics. For example, by using a pre - trained classification network, abnormal events with semantic deviation can be detected by capturing the differences in high - level features of different category objects.
[0027] Based on the above considerations, further in step 140, select a pre - trained network as the teacher network. For example, the pre - trained network can be selected as YOLOv4; use feature - based knowledge distillation technology to train a lightweight student network to generate multi - level features similar to those of the teacher network. This includes inputting the constructed spatio - temporal cube into the teacher network and the student network respectively, and minimizing the distance between the multi - level features generated by the two networks, thereby distilling the knowledge of the main part (i.e., normal events) in the unlabeled video data and transferring it to the student network, while the knowledge of the impurity part (i.e., abnormal events) is retained, thus constructing a high - quality video event normal model.
[0028] YOLOv4 (You Only Look Once v4) is a classic model in the field of object detection, proposed in 2020, suitable for training and inference on ordinary GPUs without the need for expensive hardware. YOLOv4 follows the detection head of YOLOv3 that contains three initial prediction layers, but has been optimized in many aspects based on YOLOv3, significantly improving the detection accuracy while maintaining the real - time detection speed.
[0029] Specifically, as Figure 2 shown, in this embodiment, the process of constructing a video event normal model using feature - based knowledge distillation technology includes: Step 141, select a deep neural network pre - trained on a public dataset (including classification, object detection, action recognition, etc.) as the teacher network. Based on multiple network layers of it, high - quality features at multiple levels from low to high of video events described by spatio - temporal cubes can be extracted. from low to high multi - level high - quality features.
[0030] Step 142: Use the feature information extracted by the teacher network in the intermediate layer as a supervision signal to guide the training of the student network, enabling the student network to learn the feature representation of the teacher network and generate multi-level features similar to those of the teacher network (multi-level feature alignment). Specifically, it includes: Input the video event into the teacher network and the student network, and extract multi-level features from the outputs of multiple network layers of these two networks respectively; Assume that the corresponding network layers of the teacher network and the student network are selected, and represent the teacher network and the student network respectively. Design a loss function using the distance between the multi-level features output by the teacher network and the student network to characterize the similarity between the multi-level features; If the feature set extracted by the teacher network is denoted as , and the feature set extracted by the student network is denoted as , then for a video event , the loss function is given by the following formula: ; where, and are the transformation layers for the th network layer in the teacher network and the student network respectively, which can generally be used to unify the size or shape of the features, , ; represents the distance between the multi-level features; During the training process, train the student network by minimizing the training loss characterized by the loss function, so that the parameters in the student network are gradually optimized and updated; The distance between the multi-level features output by the teacher network and the student network can adopt various measurement methods such as the distance based on mean square error, cosine similarity distance, etc.; Specifically, if the distance based on mean square error (MeanSquare Error, MSE) is selected, for a video event ; where, represents the distance between the multi-level features based on mean square error; Meanwhile, due to the rarity of abnormal events and the large redundancy of normal events, as well as the relatively small capacity and capabilities of the student network, in order to quickly reduce the training loss (smaller loss function), the student network will preferentially generate multi-level features similar to those of the teacher network for a large number of normal events, while it is difficult to generate equally high-quality features for distinctive abnormal events. Therefore, the knowledge of the main part contained in the teacher network is gradually distilled and transferred to the student network, thus constructing a high-quality video event normality model based on the student network to characterize the features of normal events.
[0031] Furthermore, in step 150, after obtaining the spatio-temporal cube of the video data to be detected according to step 120 and inputting it into the teacher network and the above-mentioned high-quality video event normality model, the video events in the video data to be detected can be detected for abnormal events. Since multiple levels of features are involved in the training of this video event normality model, the present invention makes full use of this feature for detection. The extracted multi-level features, from low-level color and texture features to high-level complex semantic features (including category, gender, etc.), can all be used to distinguish abnormal events. Since the student network tends to generate features similar to those of the teacher network for normal events, the anomaly score of the spatio-temporal cube and the determination of the degree of anomaly can be carried out by comparing the distance based on the mean square error, cosine similarity distance, etc. between the corresponding multi-level features output by the student network and the teacher network.
[0032] The anomaly score of the spatio-temporal cube is given by the following formula: ; where, , ; represents the distance between multi-level features, is the number of feature layers selected for calculating the anomaly score, and high-quality multi-level features are extracted from the teacher network and the student network. Specifically, taking the distance based on the mean square error MSE as an example, the calculation method of the anomaly score of video event is as follows: .
[0033] The feature representations of the video events in the video data to be detected at each level from low to high can all participate in the anomaly scoring process, and when the features of abnormal events at certain levels (such as color, texture or object category) are significantly different from those of normal events, this difference can be mined and reflected. If the anomaly score is greater than a pre-set threshold, the video event in the video data to be detected is determined to be an abnormal event, thus realizing video anomaly detection based on multi-level perception.
[0034] To obtain the anomaly score of a video frame data in the video data to be detected, take the maximum value of the anomaly scores of all spatio-temporal cubes corresponding to the image patches of the video foreground on the video frame as the anomaly score of the video frame: ; wherein, is the anomaly score of the spatio-temporal cube corresponding to the th image patch of the video foreground on the video frame, is the number of all foreground image patches included in the video frame. If the anomaly score of a video frame is greater than a preset threshold, the video frame is determined to be an abnormal frame.
[0035] In the second embodiment of the present invention, the pre-trained network YOLOv4 is used as the teacher network, and the student network can select various architectures that are isomorphic or heterogeneous to the teacher network. For example, the number of network layers in the original YOLOv4 network can be reduced to construct a lighter "simplified version" of the YOLOv4 network. The advantage of such a design is that the student network can easily generate feature representations that are the same as or similar to those of the teacher network, facilitating subsequent feature distance calculation and feature alignment. Of course, due to the existence of a feature conversion function in the method design, a convolutional neural network or other networks can also be selected, and the feature conversion function is used to unify the feature sizes of the teacher and student networks. In short, the student network can be any mainstream architecture, generally only need to meet the requirements of being lightweight and providing the same number of features as the teacher network.
[0036] In addition, the present invention also provides an unsupervised video anomaly detection device based on multi-level perception in the third embodiment, as Figure 3 shown. The device includes a video data acquisition module, a foreground target localization module, a video event description module, a feature-based knowledge distillation module, and an anomaly detection module, wherein: The video data acquisition module is used to acquire unlabeled video data; the video data is unlabeled and mixed with unlabeled video events, including normal events as the main part of the video data and abnormal events as the impurity part of the video data; The foreground target localization module is used to localize the foreground targets in the video data by using a pre-trained target detector and combining motion cues; The video event description module is used to extract the image patches of multiple consecutive frames before and after the localized foreground targets, scale the image patches of the multiple consecutive frames before and after to the same size and stack them to construct a spatio-temporal cube for describing a video event; Feature-based knowledge distillation module, which is used to distill the knowledge of the main part in video data by using feature-based knowledge distillation technology to construct a normal model of video events; the feature-based knowledge distillation technology includes using a pre-trained deep neural network as a teacher network and training a lightweight neural network as a student network to make the student network generate multi-level features similar to those of the teacher network. Anomaly detection module, which is used to construct a spatio-temporal cube of the video data to be detected, input the spatio-temporal cube into the teacher network and the student network trained during the construction of the video event normal model; and use the difference between the multi-level features output by the teacher network and the student network to perform anomaly scoring on each spatio-temporal cube to achieve anomaly detection of the video data to be detected.
[0037] The above-mentioned unsupervised video anomaly detection method and device based on multi-level perception provided by the present invention can, based on unlabeled video data, through feature-based knowledge distillation technology, and by constructing a high-quality video event normal model based on multi-level features, achieve video anomaly event detection based on multi-level perception.
[0038] Specifically, the present invention first marks the foreground objects in the video data, and then constructs a spatio-temporal cube to more accurately depict video events and reduce the interference of background redundant information; then inputs the obtained spatio-temporal cube into the teacher network and the student network respectively, and constructs a high-quality video event normal model by minimizing the distance between the multi-level features extracted from the two networks; and based on this video event normal model, perform anomaly scoring on the video events, and then make a judgment on whether the video events are abnormal.
[0039] Due to the low frequency, novelty and openness of abnormal events, during the distillation process, the knowledge or information corresponding to the normal events that account for the main part in the unlabeled video data will be distilled into the small-scale student network, while the knowledge or information of the impurity part (abnormal events) will be retained, realizing the construction of a high-quality normal model based on a mixed data set.
[0040] In addition, the powerful hierarchical representation learning ability of the deep neural network makes it possible to capture perceptual information at different levels. Therefore, during the process of constructing a video event normal model by using feature-based knowledge distillation technology, the multi-level features extracted from low to high by the network can be fully utilized for modeling; and this modeling method based on multi-level features enables anomaly detection to fuse multi-level perceptual differences and is more in line with the process of human perception of abnormal events.
[0041] In one embodiment, the present invention also provides a computer device, which can be a server or a terminal, and its internal structure diagram can be as Figure 4As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes an unsupervised video anomaly detection method based on multi-level perception. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0042] Those skilled in the art can understand that the description of the device technical features in the above embodiments and the structural schematic diagram as Figure 4 shown do not constitute a limitation on all devices to which the solution of the present invention is applied. The specific device may include more or fewer components, or combine certain components, or have different component arrangements.
[0043] In another embodiment, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it realizes the steps of the aforementioned unsupervised video anomaly detection method based on multi-level perception.
[0044] Those of ordinary skill in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0045] Matters not covered by the present invention are well-known techniques.
[0046] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope described in this specification.
[0047] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. An unsupervised video anomaly detection method based on multi-level perception, characterized in that: include: Step 110: Obtain unlabeled video data; Video events in the video data, including normal events as the main part of the video data, and abnormal events as the impurity part of the video data; Step 120: locating foreground objects in the video data using a pre-trained object detector and combining motion cues; Step 130: extracting image blocks of multiple consecutive frames before and after the located foreground object, scaling the image blocks of multiple consecutive frames before and after to a uniform size and stacking them to construct a space-time cube for describing the video event; Step 140: using a feature-based knowledge distillation technique to distill the main body of knowledge in the video data and construct a video event normality model; the feature-based knowledge distillation technique includes using a pre-trained deep neural network as a teacher network and training a lightweight neural network as a student network using video events described by a space-time cube, so that the student network generates multi-level features similar to those of the teacher network; Step 150: construct a space-time cube of the video data to be detected, input the space-time cube into the teacher network and the student network trained when constructing the video event normality model; use the difference in multi-level features output from the teacher network and the student network to perform anomaly scoring on each space-time cube to achieve anomaly detection of the video data to be detected.
2. The unsupervised video anomaly detection method based on multi-level perception according to claim 1 is characterized in that: In step 120, the process of locating the foreground target in the video data includes using a pre-trained target detector and combining it with motion cues to detect and locate the foreground target in the video data frame by frame, and marking the foreground target with a detection frame to reduce interference from the video background.
3. The unsupervised video anomaly detection method based on multi-level perception according to claim 2 is characterized in that: In step 140, the knowledge of the main part of the video data is distilled out by using feature-based knowledge distillation technology to construct a video event normality model, including: Step 141, select a deep neural network pre-trained on a public dataset as a teacher network, and extract video events described by a space-time cube based on multiple network layers of the teacher network Multi-level features from low to high; Step 142, taking the feature information extracted by the teacher network in the middle layer as the supervision signal, constructing the loss function; by minimizing the loss function, guiding the training of the student network, performing multi-level feature alignment, so that the student network learns the feature representation of the teacher network, and extracts multi-level features similar to the teacher network; at the same time, the main part of the knowledge contained in the teacher network is gradually distilled out and passed to the student network, and a video event normality model is constructed based on the student network.
4. The unsupervised video anomaly detection method based on multi-level perception according to claim 3 is characterized in that: Step 142 includes: Video Events Input them into the teacher network and the student network respectively, and extract multi-level features from the outputs of multiple network layers of the teacher network and the student network respectively; Select the middle layer of the teacher network network layers, and the student network corresponding to the teacher network The loss function is designed by using the distance between the multi-level features output by the teacher network and the student network to characterize the similarity of the multi-level features. , the loss function is given by: ; in, and denote the teacher network and the student network respectively, Teacher Network The extracted feature set, It is a student network The extracted feature set; and The first The conversion layer of the network layer, , ; Represents the distance between multi-level features; The student network is trained by minimizing the loss function, so that the parameters in the student network are gradually optimized and updated, thereby completing the construction of a video event normality model based on the student network.
5. The unsupervised video anomaly detection method based on multi-level perception according to claim 4 is characterized in that: The process of step 150 includes: According to step 120, a space-time cube of the video data to be detected is obtained; Input the space-time cube of the video data to be detected into the teacher network and the student network trained when constructing the normal model of the video event; select The feature layers are used to calculate the anomaly score and extract high-quality multi-level features from the teacher network and the student network; Using the distance between the multi-level features output from the teacher network and the student network, anomaly scoring is performed on each space-time cube, and the anomaly score of the space-time cube is obtained as follows: ; in, , ; Represents the distance between multi-level features; The anomaly score of the space-time cube is used to determine whether the video event corresponding to the space-time cube is an abnormal event: if the anomaly score of the space-time cube is greater than a preset threshold, the video event in the video data to be detected is determined to be an abnormal event.
6. The unsupervised video anomaly detection method based on multi-level perception according to claim 5 is characterized in that: The distance between the multi-level features Use mean square error based distance or cosine similarity distance.
7. The unsupervised video anomaly detection method based on multi-level perception according to claim 6 is characterized in that: In S150, the process of implementing the abnormality detection of the video data to be detected also includes obtaining the image blocks corresponding to all the video foregrounds on a video frame. The anomaly score of the space-time cube is The maximum value of the anomaly scores is taken as the anomaly score of a video frame: ; in, is the first The anomaly score of the space-time cube corresponding to the image patch of the video foreground; If the anomaly score of a video frame is greater than a preset threshold, the video frame is determined to be an abnormal frame.
8. An unsupervised video anomaly detection device based on multi-level perception, characterized in that: The device is used to implement the steps of the unsupervised video anomaly detection method based on multi-level perception as described in any one of claims 1 to 7, and the device includes a video data acquisition module, a foreground target positioning module, a video event description module, a feature-based knowledge distillation module and an anomaly detection module, wherein: The video data acquisition module is used to acquire unlabeled video data; video events in the video data include normal events as the main part of the video data and abnormal events as the impurity part of the video data; A foreground object localization module is used to locate foreground objects in video data using pre-trained object detectors combined with motion cues; A video event description module is used to extract image blocks of multiple consecutive frames before and after the located foreground target, scale the image blocks of multiple consecutive frames before and after to a uniform size, and then stack them to construct a space-time cube for describing the video event; A feature-based knowledge distillation module is used to distill the main part of the knowledge in the video data using the feature-based knowledge distillation technology to construct a video event normality model; the feature-based knowledge distillation technology includes using a pre-trained deep neural network as a teacher network and training a lightweight neural network as a student network using video events described by a space-time cube, so that the student network generates multi-level features similar to those of the teacher network; The anomaly detection module is used to construct a spatiotemporal cube of the video data to be detected, input the spatiotemporal cube into the teacher network and the student network trained when constructing the video event normality model; use the difference in multi-level features output from the teacher network and the student network to perform anomaly scoring on each spatiotemporal cube to achieve anomaly detection of the video data to be detected.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the unsupervised video anomaly detection method based on multi-level perception are implemented as described in any one of claims 1-7.
Citation Information
Patent Citations
Knowledge distillation-based unsupervised industrial image anomaly detection method and system
CN114240892A
Unsupervised end-to-end video abnormal event data identification method and device
CN114255447A
Unsupervised video anomaly detection method and device based on mask auto-encoder
CN114724060A
Video anomaly detection method based on combination of residual attention block and self-selection learning
CN116152722A
Video anomaly detection method and device based on semantic completion type blank-filling test auto-encoder
CN118366078A
Cited By
Scene-dependent video anomaly detection method and device
CN121600327A
A scene-dependent video anomaly detection method and device
CN121600327B