Weak supervision video anomaly detection method based on visual language model

By optimizing the video anomaly detection method in the feature space of the visual language model, using text prompt learning and timing module optimization, the problem of data imbalance and time relationship capture is solved, and higher detection accuracy and generalization performance are achieved.

CN120164137APending Publication Date: 2025-06-17SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510060563.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing video anomaly detection schemes have significantly reduced performance when facing high data imbalance, making it difficult to capture typical characteristics of abnormal events, and may cause false positives for uncommon normal activities. At the same time, there are limitations in handling any number of exception segments in exception videos in multi-instance learning, and it is impossible to effectively capture the time relationships within different time steps.

Method used

Weakly supervised video anomaly detection method based on visual language model is adopted, and visual information is optimized in the feature space of visual language model, and text prompt learning is used to optimize visual information to alleviate data imbalance problem. At the same time, the timing module is optimized to adapt to video data of different time lengths, and optimizes the number selection method of abnormal clips in multi-instance learning to project video features onto text.

Benefits of technology

It effectively alleviates the problem of data imbalance, improves the accuracy of distinguishing abnormal videos and normal videos, enhances the ability to capture time relationships within different time steps, reduces the false alarm rate, and improves the accuracy and generalization performance of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164137A_ABST
    Figure CN120164137A_ABST
Patent Text Reader

Abstract

The invention discloses a weak supervision video anomaly detection method based on a visual language model. The method comprises the steps of obtaining a target video; inputting the target video into a detection model subjected to weak supervision training to obtain an abnormal event detection result; wherein the detection model comprises a backbone network and a time sequence processing module, the backbone network is constructed based on a visual language model and comprises an image encoder and a text encoder, the image encoder is used for extracting visual features of an input video frame, and the text encoder is used for extracting text features; the visual features and the text features are projected to a feature space, and the time sequence processing module is used for processing the visual features so as to aggregate time information of different time lengths and learn time information in a video. According to the method, the problem of data imbalance is effectively relieved, and the accuracy of video detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video analysis, and more specifically, to a weakly supervised video anomaly detection method based on a vision-language model. Background Art

[0002] Video action understanding is an active research field with wide applications, covering multiple directions such as action localization, action recognition, video caption generation, etc. Among them, video anomaly detection (VAD) is of great significance in multiple key applications, such as in the fields of monitoring, industrial monitoring, environmental monitoring and disaster response, and behavior analysis and psychological research. Video anomaly detection is the task of automatically identifying activities that deviate from the normal pattern in a video, that is, locating abnormal events in a given video. There are mainly three paradigms in the research of video anomaly detection: fully supervised, unsupervised, and weakly supervised. Currently, the weakly supervised paradigm is mainly used for video anomaly detection, which is called weakly supervised video anomaly detection (WSVAD). And in WSVAD, multi-instance learning (MIL) is basically used to handle the label uncertainty problem in video data. In recent years, vision-language models have developed rapidly, and their powerful zero-shot generalization ability has been favored by more and more researchers. Such generalization ability provides new ideas and methods for weakly supervised abnormal video detection.

[0003] In the work of abnormal video detection, a method of real frame prediction (TFP) has been proposed. By training a generative model to predict the normal content of each frame in a video, based on this prediction, the difference between the current frame and the predicted frame can be judged, so as to identify the anomaly of the video frame. A method of scale-aware spatio-temporal relationship learning (SSRL) has been proposed to solve the problem of spatio-temporal relationship modeling in video anomaly detection, especially focusing on the impact of scale changes on the anomaly detection task, to better capture the multi-scale features in space and time in the video and improve the effect of video anomaly detection. There is also a method that extends the support vector machine (SVM) to the multi-instance learning (MIL) framework, solves the label problem in MIL, and subsequent multiple computer vision tasks indirectly or directly draw on this method.

[0004] In the aspect of visual language, existing solutions have proposed the CLIP (Visual Language Model) model. The CLIP model mainly uses text data described in large-scale natural language to train the visual model, and then generates a general visual representation that can be transferred to various visual tasks. This solution embeds images and texts into a common semantic space through contrastive learning, enabling the model to transfer and achieve good performance on multiple visual tasks. However, CLIP is mainly applied to image-text tasks, while VideoCLIP (Contrastive Pre-training for Zero-shot Video-Text Understanding) is based on the above CLIP model and transfers it from the image domain to the video domain. Its main idea combines the idea of visual-language contrastive learning to achieve zero-shot video understanding through the joint learning of videos and texts.

[0005] After analysis, existing video anomaly detection solutions may significantly reduce performance when dealing with highly imbalanced data. Because in the dataset, normal events usually occupy the vast majority, while abnormal events are relatively rare and sporadic. This imbalance causes the model to often rely on a large amount of normal data to learn features, without enough abnormal data, making it difficult for the model to capture the typical features of abnormal events. And if it encounters an uncommon normal activity, it may trigger false alarms because it is different from the learned normal pattern. Also, in multi-instance learning, there are limitations in dealing with any number of abnormal segments in abnormal videos. Different abnormal events in the video may range from a few seconds to several minutes, and existing models have poor processing capabilities for this kind of video data with an unfixed time length and cannot effectively capture the temporal relationships within different time steps.

[0006] In addition, directly applying a visual language model (such as CLIP) to anomaly detection may face challenges because of the modality differences between images and videos, which makes it difficult for the model to effectively learn the features of videos. And if the visual language model is directly modified so that it can process temporal information and adapt to video data, it may affect its good generalization ability. Also, some anomaly detection methods based on visual language models do not fully utilize the language prompt function but simply optimize the features of the image encoder, which limits their performance. Summary of the Invention

[0007] The object of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a weakly supervised video anomaly detection method based on a visual language model. The method includes the following steps:

[0008] Obtain a target video;

[0009] Input the target video into a detection model trained by weak supervision to obtain an abnormal event detection result;

[0010] Among them, the detection model includes a backbone network and a temporal processing module. The backbone network is constructed based on a vision-language model and includes an image encoder and a text encoder. The image encoder is used to extract visual features of the input video frames, and the text encoder is used to extract text features. Furthermore, the visual features and the text features are projected into a feature space, and the temporal processing module is used to process the visual features to aggregate asynchronous time information and learn the time information in the video.

[0011] Compared with the prior art, the advantages of the present invention are as follows. The provided weakly supervised video anomaly detection method based on a vision-language model optimizes in the feature space of the vision-language model and uses text prompt learning to optimize visual information, effectively alleviating the data imbalance problem. In addition, for the processing of temporal information, by optimizing the temporal module, it can adapt to video data of different time lengths. Moreover, in multi-instance learning, the method for selecting the number of anomaly segments is optimized, and the video features are projected onto the text, solving the problem brought by data imbalance, thereby more accurately selecting abnormal and normal videos.

[0012] Other features and advantages of the present invention will become clear from the following detailed description of exemplary embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings incorporated in the specification and constituting a part of the specification illustrate embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.

[0014] Figure 1 is a flowchart of a weakly supervised video anomaly detection method based on a vision-language model according to an embodiment of the present invention;

[0015] Figure 2 is a feature space feature distribution diagram according to an embodiment of the present invention;

[0016] Figure 3 is a schematic diagram of a feature space with a newly defined origin according to an embodiment of the present invention;

[0017] Figure 4 is a schematic diagram of a multi-scale perceptron according to an embodiment of the present invention;

[0018] Figure 5 is a schematic diagram of the overall process of training an anomaly detection model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.

[0020] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present invention or its application or use.

[0021] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the specification.

[0022] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0023] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, further discussion thereof is not required in subsequent drawings.

[0024] Generally speaking, compared with calculating the similarity between text features and visual features in traditional vision-language models, the present invention mainly uses text prompts as parameters in the feature space to optimize visual features. For example, the visual features are adjusted through text prompts so that the generated video representation can be mapped to the text description of abnormal events. Subsequently, these optimized features are input into the temporal module to aggregate time information at multiple scales and capture the temporal relationships with different time steps in the video (for example, using a recurrent multi-layer perceptron to capture different time steps). In the following, mainly taking the selection of the CLIP (Contrastive Language-Image Pre-Training) vision-language model as the main network of the model for illustration.

[0025] See Figure 1 As shown, the provided weakly supervised video anomaly detection method based on a vision-language model includes the following steps:

[0026] Step S1, select the vision-language model CLIP as the network backbone to extract features, use the image encoder of CLIP to extract the visual features of the input video frames, and use the text encoder to extract text features.

[0027] Assume a set of weakly supervised video data is given which consists of S unclipped videos, and y (k) represents the video-level annotation. For example, y (k)∈ {0, 1}, where 0 indicates a normal video and 1 indicates an abnormal video. According to the multiple instance learning (MIL) principle, the original video is cut into s non-overlapping segments. After MIL processing, the video is where s represents the number of each segment, F is the number of frames in each segment, and D is the feature dimension of each frame. Each frame is fed into the image encoder E (V) , and we get which represents the feature corresponding to one frame. Multiple instance learning calculates the possibility that each frame is abnormal, and based on this, selects the most abnormal frames and maximizes the difference in prediction possibility between normal frames and the frames selected as the most abnormal. In the text encoder, the category t class in each video anomaly and the learnable context vector t ctx are input. They are input into the fine-tuned text encoder E (t) , and the context feature obtained is H = [h1, h2, …, h N ∈ R N×D , where h ∈ R D represents a text context feature dimension.

[0028] f = E (V) F (1)

[0029] h = E (t) (t ctx + t class ) (2)

[0030] where N represents a batch of data.

[0031] Step S2: Project the visual feature and the text feature into the feature space, and use the visual feature and the text feature to re-perform the transformation of center positioning.

[0032] Specifically, after CLIP extracts the features of each frame image, it projects them into the feature space. Assuming that the features with the same class label are clustered in a certain area in the space, a feature projection into the feature space diagram can be constructed to understand this process. See Figure 2 the feature space feature distribution diagram shown. It mainly shows the distribution of projecting the video frame features (such as visual features) into the feature space. The same class will be clustered around its same class (for example, by calculating the text and visual losses, the features with the same class are pulled closer in distance, and those not similar to the class are moved farther away). Different colored dots represent the features of different classes. Based on this concept, the origin can be redefined. The direction starting from the origin can represent different classes, and the distance of these points from the origin can be used to judge the probability that this class belongs to the abnormal.

[0033] Based on the above considerations, constructing the origin in the feature space is a crucial step. Since the origin is used to judge abnormal and normal frames, and there may be many abnormal events, it is not very realistic to select and locate the origin through abnormal events. Therefore, the origin can be located through normal frames. For example, define O1 as the normal origin. O1 extracts the average features of all M frames i included in the videos marked as normal in the dataset through the image encoder. The definition of the origin is as follows:

[0034]

[0035] Among them, M is the total number of frames included in the normal video, that is, a total of M normal frames.

[0036] After defining the origin, text information can be used to determine its direction. The similarity between each video frame in the original feature space and the text encoding is calculated. And through this process, it can be judged which categories in the text information the video frame has a higher similarity with, and then the direction and the magnitude of the direction can be judged with the help of the text information. The formula is defined as follows:

[0037] d i = E (t) [t ctx ,t class -O1 (4)

[0038] Among them, d i represents the i-th text feature, and E (t) represents the text encoder.

[0039] Because the origin is defined, when the features passed through the image encoder are projected into the feature space later, the origin needs to be subtracted. Because the origin is redefined, the projected frames must also use the new origin for similarity calculation, rather than calculating according to the original origin. The original image features After passing through the new origin, the features are The formula is defined as:

[0040] f1 = E (V) (i j ) - O1 (5)

[0041] Project the obtained feature f1 onto the feature space d i The size of the projection indicates the probability of belonging to the abnormal category. The formula is defined as:

[0042]

[0043] Among them, C represents a C-dimensional real vector space. Denotes the projection operation. Since the sizes of certain types of abnormal features may be different from those of other types of abnormal features, when projecting the feature vectors, they will be affected by their scales, resulting in significant differences. Therefore, in order to eliminate the variations brought by their scales, after projection, batch normalization (BN) is performed on them, and the formula is defined as:

[0044]

[0045] Finally, for the video input with this segment, the likelihood is summed for each frame to obtain the type of this segment in the feature space. The formula is defined as:

[0046]

[0047] Among them, denotes the likelihood of segment S, denotes the i-th frame feature after passing through the new origin (refer to formula 5).

[0048] Figure 3 is a schematic diagram of the redefined origin of the feature space, obtained through the average feature of normal videos.

[0049] Step S3: Send the extracted visual features into the temporal processing module to aggregate asynchronous long-term information and learn the temporal information in the video.

[0050] The durations of different abnormal events are often different. Some abnormal events may only last for a few seconds, while some can reach dozens of seconds or more. Therefore, it is very important to consider using local and global temporal dependencies. For example, using the network structure of a multi-layer perceptron to aggregate local and global temporal information, different from the existing single aggregation temporal structure, this structure can better capture the information in different time periods.

[0051] See Figure 4 As shown, in one embodiment, the multi-layer perceptron mainly includes a recurrent fully-connected layer (GC) and a recurrent local fully-connected layer (LC). The process of this stage is that the features after being transformed through the origin of the feature space After layer normalization, the features Send the normalized features into the recurrent fully-connected layer and the recurrent local fully-connected layer Add Gaussian noise to the data after passing through the fully-connected layer to increase the randomness of the data and improve the robustness. Then, add and fuse the processed data. Multiply and fuse the data after addition and fusion processing with the original input features, add it to f1, and then output after layer normalization

[0052]

[0053] Among them, Norm represents normalization processing, represents Gaussian noise, stride represents the loop step size of the local recurrent fully connected layer and the global recurrent fully connected layer, and D and D out represent the input and output dimensions. represents the multiplication of the original feature f1 and the feature obtained through the global and local multi-layer perceptrons multiply, represents the original feature f1 and add.

[0054] In Figure 4 , LN represents layer normalization, LC is the local recurrent fully connected layer, GC is the global recurrent fully connected layer (or simply referred to as the recurrent fully connected layer), and GN represents Gaussian noise.

[0055] Step S4, calculate the loss of the first K most abnormal segments and the first K most normal segments in the feature space together with the time information to train the anomaly detection model.

[0056] For example, according to the first K most abnormal segments projected in the feature space and the defined first K least abnormal segments Maximize the likelihood of predicting anomalies in the feature space by minimizing the loss. The formula is defined as follows:

[0057]

[0058] Among them, represents the loss function related to the abnormal segments in the abnormal video, and K represents the number of the most abnormal (or least abnormal) segments selected from the abnormal video. For example, K = 10 means selecting the first 10 most abnormal segments or the last 10 least abnormal segments. F represents the number of frames contained in each segment. That is, the video is divided into several segments, and each segment has F frames. represents the likelihood of the abnormal segment with category c.

[0059] For f2 obtained by the multi-layer perceptron, obtain the anomaly score through a shallow perceptron, and then obtain the probability through a softmax. Assume that the anomaly probability p A (x) is obtained, then its normal probability p N (x) = 1 - p a (x). Use cross-entropy to maximize the anomaly probability p A (x) contained in each frame of these segments. The formula is defined as follows:

[0060]

[0061] Among them, represents the probability of the most abnormal frame j in the i-th segment, where i is the segment index and j is the frame index.

[0062] For normal videos, cross-entropy is used to maximize p N (x) for each frame in these segments, and its formula is defined as follows:

[0063]

[0064] Among them, represents the probability of the most normal frame j in the i-th segment.

[0065] To utilize the information in normal videos, for each segment of normal videos minimize the likelihood predicted by the feature space, and its formula is defined as follows:

[0066]

[0067] Finally, use the sparse loss loss spa and the smooth term loss loss smo to regularize the training, and its formula is as follows:

[0068]

[0069] Among them, represents the probability of the most normal frame j in the i-th segment, V i,j represents the j-th frame in the i-th segment, p A (V i ) represents the probability that the i-th frame in the video is predicted as abnormal, p A (V i-1 ) represents the probability that the (i - 1)-th frame in the video is predicted as abnormal.

[0070] The total loss of training is expressed as:

[0071]

[0072] The flowchart of the entire model is as Figure 5 shown:

[0073] Step S5, use the trained detection model to identify abnormal events in the target video.

[0074] After the model training is completed, it can be used for detecting abnormal events in actual videos. For example, the application process of the model includes: collecting the target video; inputting the target video into the trained detection model to obtain the abnormal event detection result. From the above training process, it can be seen that the detection model generally includes an image encoder, a text encoder, a feature space mapping process, and a multi-scale temporal processing process, etc. The model application process is basically similar to the training process and will not be elaborated here.

[0075] It should be noted that the model training process involved in the present invention can be carried out offline on a server or in the cloud. Embedding the trained model into an electronic device can achieve real-time video abnormal event detection. The electronic device can be a terminal device or a server. The terminal device includes any terminal device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a point-of-sale terminal (POS), an in-vehicle computer, a smart wearable device (smart watch, virtual reality glasses, virtual reality helmets, etc.). The server includes but is not limited to an application server or a web server, and can be an independent server, a cluster server, or a cloud server, etc.

[0076] To further verify the effect of the present invention, tests were carried out on a public dataset. The test results show that the detection accuracy rate can be comparable to that of a single-modal backbone abnormal detection video model, but the generalization performance and the performance in dealing with complex scenarios of the model of the present invention are better, and better performance and effects can be achieved in actual applications.

[0077] In summary, the present invention applies a vision-language model (LLV) to the field of video anomaly detection. Compared with the prior art, it has the following advantages:

[0078] 1) By using the image encoder and text encoder of the vision-language model, the generated video representation can be mapped to the text description of abnormal events, improving the distinction between abnormal videos and normal videos, and solving the false alarms caused by the highly unbalanced data distribution between abnormal and normal videos.

[0079] 2) Using a multi-scale perceptron to extract the temporal information in the video makes up for the errors caused by the uncertain duration of different abnormal events.

[0080] 3) Through the projection of the feature space, the top K normal and abnormal video segments can be selected more accurately.

[0081] 4) It has been verified that compared with the prior art, using a vision-language model as the backbone network not only makes full use of the good generalization and robustness of the vision-language model, but also improves the accuracy rate. Therefore, it has better performance in reducing errors and dealing with different scenario problems.

[0082] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of the present invention.

[0083] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0084] The computer-readable program instructions described herein may be downloaded to respective computing / processing devices from a computer-readable storage medium or may be downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0085] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.

[0086] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0087] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0088] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0089] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As is well known to those skilled in the art, implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.

[0090] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A weakly supervised video anomaly detection method based on a visual language model, comprising the following steps: Get the target video; Inputting the target video into a detection model trained by weak supervision to obtain an abnormal event detection result; Among them, the detection model includes a backbone network and a timing processing module. The backbone network is built based on a visual language model and includes an image encoder and a text encoder. The image encoder is used to extract visual features of input video frames, and the text encoder is used to extract text features. Then the visual features and the text features are projected into a feature space. The timing processing module is used to process the visual features to aggregate information of different synchronous time periods and learn the time information in the video.

2. The method according to claim 1, characterized in that The visual language model is a CLIP model.

3. The method according to claim 1, characterized in that The visual features and the text features are extracted based on the following formula: f=E (V) F h=E (t) (t ctx +t class ) Among them, f represents the visual feature corresponding to a frame, E (V) represents the image encoder, F is the number of frames in each video segment, and E (t) represents the text encoder, t ctx represents the learned context vector, t class Indicates the category in the video anomaly.

4. The method according to claim 1, characterized in that: Projecting the visual features and the text features into a feature space comprises: For the original image feature f, the feature after passing through the new origin is f1, which is expressed as: f1=E (v) 9i j )-O1 Projecting feature f1 onto the feature space is expressed as: After projection, batch normalization BN is performed, which is expressed as: The type of the input video clip in the feature space is obtained according to the following formula, expressed as: in, represents the likelihood of segment S, represents the i-th frame feature passing through the new origin, C represents a C-dimensional real vector space, represents the projection operation, d i represents i text features, O1 represents the origin, E (v) Represents an image encoder.

5. The method according to claim 4, characterized in that The origin O1 is located through the normal frame, expressed as: The text features are calculated according to the following formula: d i =E (t) [t ctx ,t class ]-O1 Where M is the total number of frames contained in a normal video, t ctx represents the learned context vector, t class Indicates the category of video anomaly, E (t) Represents a text encoder.

6. The method according to claim 4, characterized in that The time series processing module is implemented by a multi-scale multi-layer perceptron, which includes a first normalization layer, a local recurrent fully connected layer, a global recurrent fully connected layer, and a second normalization layer, and performs the following operations: Among them, Norm means normalization processing, represents Gaussian noise, stride represents the cycle step size of the local and global recurrent fully connected layers, D and D out represents the input and output dimensions, and FC means fully connected.

7. The method according to claim 4, characterized in that The total loss function for training the detection model is set to: Among them, loss represents the total loss value, Represents the loss term related to the abnormal segment in the abnormal video, loss spa Represents the sparse loss term, loss smo represents the smoothing loss term, is the cross entropy loss term associated with normal videos, is the cross entropy loss term associated with the abnormal video, and λ1 and λ2 are the coefficients of the related loss terms.

8. The method according to claim 7, characterized in that Each loss item is set as: Where F is the number of frames in each segment, represents the probability that the jth frame in the i-th segment is the most abnormal, i is the segment index, j is the frame index, represents the probability that the jth frame in the i-th segment is the most normal, V i,j represents the jth frame in the i-th segment, p A (V i ) represents the probability that the i-th frame in the video is predicted to be abnormal, p A (V i-1 ) represents the probability that the i-1th frame in the video is predicted to be abnormal, K represents the number of the most abnormal or least abnormal segments selected from the abnormal video, represents the likelihood of an abnormal segment of category c.

9. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Head condition monitoring method based on region best matching feature points

    CN109145684A

  • Method and system for automatically generating subtitles in short video

    CN113159034A

  • Multi-mode self-supervising progressive video abstract model, method and device

    CN115620213A

  • Multi-modal representation learning method based on text guide image block screening

    CN117421591A

  • Progressive multi-scale context learning time sequence target positioning method

    CN117876929A