Video anomaly detection with analysis capability and model training method thereof

CN118692007BActive Publication Date: 2026-09-29HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410776255.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-17
Publication Date
2026-09-29
Estimated Expiration
2044-06-17

AI Technical Summary

Technical Problem

[0004]针对现有技术的以上缺陷或改进需求,本申请提供了一种具有分析能力的视频异常检测及其模型训练方法,其目的在于解决现有视频异常检测方法对视频异常判定缺乏解释分析的技术问题

Benefits of technology

[0030](1)本申请视频异常检测模型中通过采用多模态解码网络学习异常视频片段中的帧特征、指令文本特征和分析文本特征之间的特征关系,由此实现对异常视频片段的检测和分析,通过本申请视频异常检测模型不仅能高效检测出视频中的异常事件,同时也可以对检测出的异常事件进行解释分析。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118692007B_ABST
    Figure CN118692007B_ABST
Patent Text Reader

Abstract

The application discloses a video anomaly detection with analysis capability and a model training method thereof, and belongs to the technical field of video detection. The video anomaly detection model is composed of a video anomaly detection network and a multi-modal decoding network. The video anomaly detection network is trained for recognizing and training abnormal videos through full-supervision training data converted from weak-supervision training data, and then learns the feature relationship among abnormal image frame features, corresponding instruction text features and corresponding analysis text features through the multi-modal decoding network. Therefore, the video anomaly detection model can not only efficiently detect abnormal events in a video, but also realize explanation and analysis on the detected abnormal events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of video detection technology, and more specifically, relates to a video anomaly detection method with analytical capabilities and a model training method thereof. Background Technology

[0002] Video anomaly detection aims to identify anomalous events in videos and has received widespread research attention in recent years due to its significant applications in public safety and video content understanding. Current video anomaly detection methods based on training data can be broadly categorized into three types based on the type of training data annotation: unsupervised, weakly supervised, and fully supervised. Unsupervised methods are trained only on single-class normal videos or unlabeled normal / anomaly videos, while weakly supervised methods are trained on normal / anomaly videos with video-level labels. Since video data with frame-level labels is scarce, fully supervised detection methods are relatively few.

[0003] Because machine learning-based video anomaly detection methods extract and differentiate video features through machine learning algorithms, they lack an understanding of the human meaning of "anomaly." Therefore, when anomaly videos are detected, they cannot provide administrators with an explanation of the detected anomalies—that is, "what is abnormal" and "why is it considered abnormal." This lack of transparency limits human understanding of the detection system. Summary of the Invention

[0004] In view of the above-mentioned defects or improvement needs of the existing technology, this application provides a video anomaly detection and model training method with analytical capabilities, which aims to solve the technical problem that the existing video anomaly detection methods lack interpretation and analysis for video anomaly judgment.

[0005] To achieve the above objectives, firstly, this application provides a method for training a video anomaly detection model with analytical capabilities, comprising:

[0006] The initial video anomaly detection network is trained by taking the image frames and frame labels of the video segments in the training set as input and the anomaly score of the image frames as output. The video anomaly detection network is then used to evaluate the anomaly score of each image frame in the video segment and identify the abnormal image frames with anomaly scores higher than a threshold.

[0007] The image features of the abnormal image frame and the text features of the instruction text corresponding to the video segment where the abnormal image frame is located are used as inputs, and the analysis text corresponding to the video segment where the abnormal image frame is located is used as outputs to train the initial multimodal decoding network and obtain the multimodal decoding network.

[0008] The video anomaly detection model is composed of the video detection network and the multimodal decoding network. The video anomaly detection model combines the input instruction text to detect the video to be detected and outputs abnormal video segments and corresponding analysis text.

[0009] Preferably, the instruction text is the prompt text corresponding to the analysis text.

[0010] Preferably, the training set is obtained through the following method:

[0011] Select videos with video-level tags;

[0012] Randomly select one image frame from the abnormal events in the abnormal video as a labeled frame; use a pre-trained video anomaly detection network to evaluate the anomaly score of all image frames in the abnormal video, and divide normal or abnormal video segments with frame-level labels based on the anomaly score and the position of the labeled frame.

[0013] A pre-trained video description generation network is used to generate descriptive text based on the normal or abnormal video clips;

[0014] A large language model is used to generate analysis text based on whether the fragment is abnormal and the corresponding descriptive text.

[0015] The training set consists of a set of training data, which includes preset instruction text, analysis text, and corresponding normal or abnormal video clips.

[0016] Preferably, normal or abnormal video segments with frame-level labels are divided based on the anomaly score and the position of the labeled frame, specifically as follows:

[0017] Consecutive adjacent frames with anomaly scores greater than or equal to a threshold are grouped into anomaly segments, including labeled frames, wherein the frame labels of the consecutive adjacent frames are all anomaly; consecutive adjacent frames with anomaly scores less than a threshold are grouped into normal segments, including labeled frames, wherein the frame labels of the consecutive adjacent frames are all normal.

[0018] Preferably, the video description generation network is a multimodal visual language large model, Video-LLaVA.

[0019] Preferably, the large language model is specifically the large language model Llama3.

[0020] Preferably, the image features of the abnormal image frames are extracted by a video feature extraction network, specifically a video encoder using the multimodal alignment framework LanguageBind.

[0021] Preferably, the text features of the instruction text are extracted using a text feature extraction network, specifically the text encoder of the large language model Vicuna.

[0022] Preferably, the video anomaly detection network is specifically a video anomaly detection model UR-DMU.

[0023] Preferably, the multimodal decoding network is a large language model, Vicuna.

[0024] Secondly, this application provides a video anomaly detection method with analytical capabilities, including:

[0025] The system receives a video to be detected and a command text, and uses a pre-trained video anomaly detection model in conjunction with the command text to detect the video to be detected; it outputs anomaly video segments and corresponding analysis text; wherein the video anomaly detection model is trained according to any one of the methods described in the first aspect.

[0026] Thirdly, this application provides an electronic device, comprising: a memory for storing a program; and a processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method described in the first or second aspect.

[0027] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to perform the methods described in the first or second aspect.

[0028] Fifthly, this application provides a computer program product that, when run on a processor, causes the processor to perform the methods described in the first or second aspect.

[0029] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art:

[0030] (1) In the video anomaly detection model of this application, the feature relationship between frame features, instruction text features and analysis text features in the abnormal video segment is learned by using a multimodal decoding network, thereby realizing the detection and analysis of abnormal video segments. The video anomaly detection model of this application can not only efficiently detect abnormal events in the video, but also interpret and analyze the detected abnormal events.

[0031] (2) In the training set acquisition stage of the video anomaly detection model, this application mines frame labels near random time-series annotation points of video anomaly events, transforming videos with video-level labels into video segments with frame-level labels, thereby transforming weakly supervised training data into fully supervised training data, and the anomaly detection model trained thereafter has higher detection accuracy. Attached Figure Description

[0032] Figure 1This is a flowchart of the video anomaly detection model training method provided in the embodiments of this application;

[0033] Figure 2 This is a schematic diagram illustrating video anomaly detection using a video anomaly detection model provided in an embodiment of this application;

[0034] Figure 3 This is a schematic diagram of the training set acquisition process provided in an embodiment of this application;

[0035] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0037] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0038] In the description of the embodiments of this application, unless otherwise stated, "multiple sets" means two or more sets, for example, multiple sets of training data means two or more sets of training data, etc.

[0039] like Figure 1 As shown, an embodiment of this application provides a method for training a video anomaly detection model, which includes the following steps:

[0040] First, a training set is obtained, which includes multiple sets of training data. Each set of training data includes video clips with frame-level labels, instruction text corresponding to the video clips, and analysis text.

[0041] The initial video anomaly detection network is trained by taking the image frames and frame labels of the video segments in the training set as input and the anomaly score of the image frames as output. The video anomaly detection network is then used to evaluate the anomaly score of each image frame in the video segment and identify the abnormal image frames with anomaly scores higher than a threshold.

[0042] The image features of the abnormal image frame and the text features of the instruction text corresponding to the video segment where the abnormal image frame is located are used as inputs, and the analysis text corresponding to the video segment where the abnormal image frame is located is used as outputs to train the initial multimodal decoding network and obtain the multimodal decoding network.

[0043] The video anomaly detection model is composed of the video detection network and the multimodal decoding network. The video anomaly detection model combines the instruction text to detect the video to be detected and outputs abnormal video segments and corresponding analysis text.

[0044] like Figure 2 The diagram shown is a schematic diagram of video anomaly detection using a video anomaly detection model provided in an embodiment of this application. This model achieves anomaly detection and anomaly analysis through the following steps:

[0045] (1) Use a pre-trained video feature extraction network to extract the spatiotemporal characteristics of the video frame sequence from the video to be detected:

[0046]

[0047] Specifically, the video feature extraction network is the video encoder of the multimodal alignment framework LanguageBind, V i φ represents the input i-th frame image in the video to be detected. v This represents a video feature extraction network, representing the features of each frame of an image. Includes categorical features f i cls and block features f i patch patch = 1, 2, ..., N p N p d represents the number of blocks in the image, d represents the feature dimension, and N represents the number of video frames.

[0048] (2) Use a pre-trained text feature extraction network to extract text features from the input instruction text:

[0049] f text =φ t (T)

[0050] Specifically, the text feature extraction network is the text encoder of the large language model Vicuna, where T represents text and φ t f represents a text feature extraction network. text This indicates the text features of the output.

[0051] (3) A pre-trained video anomaly detection network is used to predict the anomaly score of each frame in the video to be detected, and the abnormal image frames with anomaly scores higher than the threshold are identified. The features corresponding to the abnormal image frames are retained, specifically:

[0052] The category features f of each frame image i cls Input a video anomaly detection network, and the network outputs the corresponding anomaly score s. i ;

[0053] Image frames are sampled based on a given threshold θ to identify anomaly scores higher than the threshold (s). i For abnormal image frames (>θ), retain the features corresponding to the abnormal image frames.

[0054] Specifically, the video anomaly detection network is the UR-DMU video anomaly detection model.

[0055] (4) The multimodal decoding network decodes the text features of the instruction text and the features of the abnormal image frames, and outputs the abnormal analysis text (interpretable text), which specifically includes the following steps:

[0056] Features corresponding to image frames The text features of the instruction text are concatenated and input into a multimodal text decoding network for autoregressive analysis to generate interpretable text.

[0057]

[0058] Specifically, the multimodal decoding network is the large language model Vicuna, φ proj Used to visualize feature F s Mapped to the text feature space, Decoder() is a multimodal decoding network, T 0:i For the text in the autoregressive process, including the input text and the already generated text, T i+1 This is the output analysis text.

[0059] like Figure 3 The diagram shown is a schematic representation of a training set acquisition process provided in an embodiment of this application, which specifically includes the following steps:

[0060] (1) Collect videos, including normal videos and abnormal videos containing abnormal events, specifically including the following steps:

[0061] First, select videos with video-level labels from the existing dataset. Optionally, filter out videos that are too long or too short.

[0062] The definition of an abnormal event here is based on the event to be detected. For example, in this embodiment, it is defined as a dangerous event. However, this application does not limit the abnormal event to a dangerous event. It can also be defined as a flame, an animal, a moving object, a specific color, or a fixed object, etc.

[0063] (2) Further annotate the video, including time-series single-frame annotation and video segment-level text description annotation, specifically including the following steps:

[0064] For each abnormal event in the uncropped abnormal video, a frame is randomly labeled in time sequence; a frame is randomly selected from the image frames of dangerous events that appear in the video and labeled "abnormal".

[0065] The anomaly score of all image frames in the video was evaluated using a pre-trained video anomaly detection network, UR-DMU. Abnormal or normal segments were then cropped from the video based on single-frame annotations and anomaly scores.

[0066] Optionally, consecutive adjacent frames with anomaly scores greater than or equal to a threshold are grouped into anomaly segments, including labeled frames; consecutive adjacent frames with anomaly scores less than a threshold are grouped into normal segments, including labeled frames.

[0067] Optionally, consecutive adjacent frames with an average anomaly score greater than or equal to a threshold are grouped into anomaly segments, including labeled frames; consecutive adjacent frames with an average anomaly score less than a threshold are grouped into normal segments, including labeled frames.

[0068] A pre-trained video description generation network is used to generate detailed text descriptions for cropped abnormal or normal video clips. In this embodiment, the text description is: "A person in the video is holding a gun."

[0069] Specifically, the video description generation network is the multimodal visual language large model Video-LLaVA.

[0070] (3) Construct text pairs in the form of "instruction text - analysis text" in a dialog format. This includes the following steps:

[0071] The large language model is used to generate analysis text based on the judgment (text) of whether the video clip is abnormal and the description text. In this embodiment, based on the tag "abnormal" and "a person in the video is holding a gun", the analysis text "a person in the video is holding a gun, which is abnormal because it shows a violent and potentially dangerous behavior that could cause injury to others" is generated.

[0072] Specifically, the large language model is the large language model Llama3.

[0073] The instruction dataset construction module randomly selects a single preset question and forms a "instruction text-analysis text" text pair with the analysis text; in the text pair, the instruction text is the prompt text of the analysis text.

[0074] In this embodiment, the preset question can be selected from "Are there any abnormal events?" or "Is there any abnormality in the video?", etc., and the preset question is not limited here.

[0075] In the constructed text pair, the instruction text is: "Is there an abnormal event?"

[0076] The analysis states: "The video shows a person holding a gun, which is abnormal because it demonstrates violence and potentially dangerous behavior that could lead to injury to others."

[0077] Optionally, filter out erroneous and low-quality text pairs.

[0078] The training set consists of a set of training data, which includes instruction text, analysis text, and corresponding normal or abnormal video clips.

[0079] Based on the methods in the above embodiments, this application provides an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the methods in the above embodiments.

[0080] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0081] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0082] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0083] It is understood that the processor in the embodiments of this application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor may be a microprocessor or any conventional processor.

[0084] The method steps in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0085] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted through the storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0086] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application.

[0087] The above content is readily understood by those skilled in the art. The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for training a video anomaly detection model with analytical capabilities, characterized in that, include: The initial video anomaly detection network is trained by taking the image frames and frame labels of the video clips in the training set as input and the anomaly scores of the image frames as output, thus obtaining the video anomaly detection network. The video anomaly detection network evaluates the anomaly score of each frame in the video segment and identifies abnormal image frames with anomaly scores higher than a threshold. The image features of the abnormal image frame and the text features of the instruction text corresponding to the video segment where the abnormal image frame is located are used as inputs, and the analysis text corresponding to the video segment where the abnormal image frame is located is used as outputs to train the initial multimodal decoding network and obtain the multimodal decoding network. The video anomaly detection model is composed of the video detection network and the multimodal decoding network. The video anomaly detection model combines the input instruction text to detect the video to be detected and outputs abnormal video segments and corresponding analysis text.

2. The video anomaly detection model training method according to claim 1, characterized in that, The training set was obtained through the following methods: Select videos with video-level tags; Randomly select one image frame from the abnormal events in the abnormal video as a labeled frame; use a pre-trained video anomaly detection network to evaluate the anomaly score of all image frames in the abnormal video, and divide normal or abnormal video segments with frame-level labels based on the anomaly score and the position of the labeled frame. A pre-trained video description generation network is used to generate descriptive text based on the normal or abnormal video clips; A large language model is used to generate analysis text based on whether the fragment is abnormal and the corresponding descriptive text. The training set consists of a set of training data, which includes preset instruction text, analysis text, and corresponding normal or abnormal video clips.

3. The video anomaly detection model training method according to claim 2, characterized in that, Based on the anomaly score and the position of the labeled frame, normal or abnormal video segments with frame-level labels are divided, specifically as follows: Consecutive adjacent frames with anomaly scores greater than or equal to a threshold are grouped into anomaly segments, including labeled frames, wherein the frame labels of the consecutive adjacent frames are all anomaly; consecutive adjacent frames with anomaly scores less than a threshold are grouped into normal segments, including labeled frames, wherein the frame labels of the consecutive adjacent frames are all normal.

4. The video anomaly detection model training method according to claim 2, characterized in that, The video description generation network is specifically the multimodal visual language large model Video-LLaVA.

5. The video anomaly detection model training method according to claim 2, characterized in that, The large language model mentioned is specifically the large language model Llama3.

6. The video anomaly detection model training method according to claim 1, characterized in that, The abnormal image frames are extracted using a video feature extraction network, specifically a video encoder based on the multimodal alignment framework LanguageBind.

7. The video anomaly detection model training method according to claim 1, characterized in that, The text features of the instruction text are extracted by a text feature extraction network, specifically the text encoder of the large language model Vicuna.

8. The video anomaly detection model training method according to claim 1 or 2, characterized in that, The video anomaly detection network is specifically the video anomaly detection model UR-DMU.

9. The video anomaly detection model training method according to claim 1, characterized in that, The multimodal decoding network is specifically the large language model Vicuna.

10. A video anomaly detection method with analytical capabilities, characterized in that, include: The system receives a video to be detected and a command text, and uses a pre-trained video anomaly detection model in conjunction with the command text to detect the video; it outputs anomaly video segments and corresponding analysis text; wherein the video anomaly detection model is trained according to any one of claims 1-9.

Citation Information

Patent Citations

  • Abnormal video identification method, device and equipment and computer readable storage medium

    CN115909116A

  • Real-time detection method and system for abnormal event in monitoring scene

    CN118133198A