Weak supervision video anomaly detection method based on text rewriting CLIP model

By adopting a text rewriting-based CLIP model in video anomaly detection, combining diverse text descriptions and enhanced visual features, the problem of imbalance between text and visual features is solved, efficient video anomaly detection is achieved, and the accuracy and robustness of detection is improved.

CN120047872APending Publication Date: 2025-05-27XINJIANG UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510128878.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When the existing video anomaly detection method deals with the problem of imbalance between text and visual features in video, it is difficult to make full use of the semantic information of text labels, resulting in insufficient detection accuracy and robustness.

Method used

Using the CLIP model based on text rewriting, diversified text descriptions are generated through the SwinBERT and LLaMA-7b models, and visual features are extracted and enhanced by the CLIP image encoder and the LGM-Mamba module to perform deep fusion and abnormal detection of classification branches.

Benefits of technology

By combining visual and linguistic information, effective weak supervision and detection of video abnormal events is achieved, the accuracy and robustness of detection are improved, and wide application prospects can be shown in the fields of monitoring video analysis and intelligent security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047872A_ABST
    Figure CN120047872A_ABST
Patent Text Reader

Abstract

The invention discloses a weak supervision video anomaly detection method based on a text rewriting CLIP model. The method comprises the following steps: acquiring a video frame sequence to be detected; a video anomaly detection model is constructed, the video anomaly detection model comprises a visual branch, a text branch, a feature fusion module and two classification branches, dense subtitles of video frames are generated in the text branch through a SwinBERT model, and two different text rewrites are generated through an LLaMA-7b model; after texts corresponding to the dense subtitles are converted into sentences to be embedded, the time relation between the texts is captured through a multi-scale tense network, and text features are obtained; in the visual branch, a CLIP image encoder is used to extract initial visual features of a video frame, and the initial visual features are processed by an LGM-Mama module to enhance the expression ability of the initial visual features. The feature fusion module is used for performing deep fusion on the final visual features and the text features; the two classification branches are respectively used for carrying out coarse-grained and fine-grained anomaly detection on the fused features; and training the video anomaly detection model by using the video frame sequence in combination with the weighted loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video abnormal behavior recognition, and relates to but is not limited to a weakly supervised video anomaly detection method based on a text rewritten CLIP model. Background Art

[0002] In recent years, Weakly-Supervised Video Anomaly Detection (WSVAD) has received increasing attention due to its broad application prospects. The goal of WSVAD is to train a detection model that can predict the confidence of frame-level anomalies (whether a certain frame image is abnormal) using only video-level label information (whether the video is abnormal). This technology is particularly prominent in the field of public safety, where it can detect potential criminal activities such as fighting or theft, ensure timely response, and enhance the sense of security within the community. In addition, in the agricultural field, this technology can analyze surveillance videos to identify crop diseases, pests, or abnormal weather events, thereby improving crop yield and quality. These applications highlight the key role of video anomaly detection in improving the quality and speed of monitoring.

[0003] Most existing works usually adopt the Multiple Instance Learning (MIL) framework. In the MIL framework, the entire video is regarded as a "bag" containing multiple instances (segments). The positive bag contains at least one abnormal segment, while the negative bag does not contain any abnormal segments. The goal is to minimize the anomaly score of each instance in the negative bag while maximizing the anomaly score of each abnormal instance in the positive bag.

[0004] The ambiguity of abnormal semantics and the scarcity of fine-grained annotations lead to the possibility that the MIL detector may wrongly assign high confidence to certain segments or only focus on abnormal segments with simple backgrounds. Since MIL optimizes each instance independently and ignores the temporal correlation between adjacent segments, it is prone to false alarms.

[0005] Recently, the Contrastive Language–Image Pre-training (CLIP) model, a multimodal vision and text learning method, has also attracted great attention in the VAD community. Based on the visual features of CLIP, lv et al. proposed a new MIL framework called Unbiased MIL (UMIL) to learn unbiased abnormal features and thus improve the performance of WSVAD.

[0006] However, this classification paradigm does not fully utilize the semantic information of text tags. Wu et al. proposed the VadCLIP model, which uses the image encoder of CLIP to extract visual features of videos. In addition, it also uses prompt learning and text tags, and they are respectively embedded using the text encoder of CLIP. The visual features enter two branches respectively for binary classification (abnormal or normal) and multi-classification (abnormal types). Binary classification directly uses a binary classifier for binary classification. The multi-classification multiplies with each label to obtain the score of each class, thus realizing multi-classification. However, this method ignores the deep semantic information hidden in the videos.

[0007] Chen et al. proposed a Text-Enhanced Video Anomaly Detection (TEVAD). TEVAD first divides the input video into T segments and feeds them into two separate branches. The text branch calculates text features based on the generated segment dense captions, while the visual branch extracts visual features. Both modalities of features pass through a multi-scale temporal network before being fused together and passed to a binary classifier, which outputs the anomaly score for each video segment, and then propagates this anomaly score to predict the frame-level anomaly score. However, such captions cannot be directly applied to CLIP because CLIP requires a large amount of data for training. During the training process of CLIP, data augmentation is a commonly used strategy to enhance the generalization ability of the model and prevent overfitting. This method mainly includes applying transformations such as rotation, cropping, and flipping to enhance the visual input. However, the text data remains unchanged throughout the training process and lacks any form of augmentation. This leads to some problems: First, the enhanced video frames are always paired with the same text captions, resulting in asymmetry. Second, encountering exactly the same text in each epoch increases the risk of text overfitting.

[0008] State Space Models (SSMs) represented by Mamba have recently become a powerful tool in the field of sequence-based reasoning, showing great research potential and application value in capturing long-term dependencies, achieving linear scalability, and maintaining computational efficiency.

[0009] SSMs have shown great potential in various fields, including images, tabular learning, genomics, and graph data, but they have not been applied to WSVAD. How to effectively apply Mamba to WSVAD is also a significant challenge. Summary of the Invention

[0010] In view of this, an embodiment of the present invention provides a weakly supervised video anomaly detection method based on a text-rewritten CLIP model, aiming to solve the problem of imbalance between text and visual features in video anomaly detection and strengthen visual modeling and feature fusion.

[0011] The technical solution of the embodiment of the present invention is specifically as follows:

[0012] The embodiment of the present invention provides a weakly supervised video anomaly detection method based on a text rewriting CLIP model. The method includes:

[0013] Obtain a video frame sequence to be detected; construct a video anomaly detection model, including a visual branch, a text branch, a feature fusion module, and two classification branches, where: in the text branch, the SwinBERT model is used to generate dense captions for video frames, and two different text rewritings are generated through the LLaMA-7b model; after converting the text corresponding to the dense captions into sentence embeddings, the multi-scale temporal network is used to capture the temporal relationship between texts to obtain text features; in the visual branch, the CLIP image encoder is used to extract the initial visual features of video frames, and the LGM-Mamba module is used to process them to enhance their expressive ability; the LGM-Mamba module includes a local-global temporal adapter, a multi-scale temporal network, and a Mamba module; the feature fusion module is used to deeply fuse the final visual features and text features; the two classification branches are respectively used to perform coarse-grained and fine-grained anomaly detection on the fused features; use the video frame sequence to train the video anomaly detection model in combination with the target loss; where the target loss is a weighted sum of the classification loss, the alignment loss, and the contrast loss.

[0014] In some embodiments, the method further includes: processing the video frame sequence through a data augmentation technique; the data augmentation technique includes random rotation, flipping, and cropping; randomly pair the original text and its two rewritten texts with the video frame sequence after data augmentation processing and input them into the video anomaly detection model.

[0015] In some embodiments, the LGM-Mamba module integrates four Mamba modules and a multi-scale temporal network on the basis of the local-global temporal adapter to process visual features; the four Mamba modules are combined in pairs to form an outer Mamba module pair and an inner Mamba module pair, which are respectively used to capture local and global dependencies in the data.

[0016] In some embodiments, after passing the initial visual features through the local-global temporal adapter, they pass through two embedding layers N1 and N2, and the two embedding layers respectively transform the shape of the input data from the original B×S×F to B×S×n 1 and B×S×n 2 , where B, S, and F are the batch size, sequence length, and initial feature dimension respectively, and n 1 and n 2 respectively represent the output feature dimensions of the two embedding layers and n 1 >n 2After the data transformed by the embedding layers N1 and N2 are respectively input into the Inner Mamba module pair and the Outer Mamba module pair for processing, they are merged through element-wise sum operation and finally concatenated together to form new features; finally, they are processed by the multi-scale temporal network to obtain the final visual features.

[0017] In some embodiments, the two classification branches include: Branch C, which is used to input the fused features into a binary classifier to obtain an anomaly confidence score, and applies the Top-K mechanism to filter out the video frames with the highest anomaly confidence score as video-level predictions, and uses binary cross-entropy to calculate the classification loss to achieve binary classification; Branch A, which is used to input the text label into the text encoder of CLIP to convert it into a corresponding class embedding, calculate the similarity between the class embedding and the final visual feature, and generate an alignment map; fine-grained video anomaly detection is realized through the alignment map to identify specific anomaly classes.

[0018] In some embodiments, the target loss is calculated by the following formula:

[0019]

[0020] L nce =-∑ j y j log(p j );

[0021]

[0022] L = L bce + L nce + λL cts ;

[0023] In the formula, L bce is the classification loss, y i is the true label of the i-th sample, taking values of 0 or 1, is the probability that the model predicts the i-th sample belongs to the positive class, and N is the total number of samples; L nce is the alignment loss, y j is the true label of the j-th category, p j is the probability that the model predicts the j-th category; L cts is the contrast loss, t n and t aj are the maximum anomaly class embeddings in the normal and the j-th category respectively, is the transpose of t n , L is the target loss, and λ is the weight coefficient.

[0024] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include:

[0025] In the embodiments of the present invention, a video anomaly detection model that combines visual and language information is constructed to achieve weakly supervised detection of video anomaly events. It uses the SwinBERT and LLaMA-7b models to generate diverse text descriptions, and extracts and enhances visual features through the CLIP image encoder and the LGM-Mamba module. Finally, coarse-grained and fine-grained anomaly detection is achieved through feature fusion and classification branches. This method has broad application prospects in the field of video anomaly detection, such as in the fields of surveillance video analysis and intelligent security. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:

[0027] Figure 1 It is a schematic flowchart of the weakly supervised video anomaly detection method based on the text rewritten CLIP model provided by the embodiments of the present invention;

[0028] Figure 2 It is a schematic diagram of the working principle of the video anomaly detection model constructed based on the TrCLIP-Vad framework provided by the embodiments of the present invention;

[0029] Figure 3 It is a schematic diagram of the structure of the LGM-Mamba module provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but not to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0031] In the following description, reference is made to "some embodiments", which describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subsets or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0032] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as herein.

[0033] The present invention is an extension based on VadCLIP, and proposes a new weakly-supervised video anomaly detection framework: the TrCLIP-Vad framework. It introduces a text branch to generate rich text features for videos, and adopts a novel text rewriting strategy to enhance the text features and generate richer text descriptions. In addition, the present invention also designs an LGM-Mamba module, aiming to strengthen visual modeling in video anomaly detection and promote the deep fusion of visual and text features. Thereby, it can solve the problem of imbalance between text and visual features in video anomaly detection, and strengthen visual modeling and feature fusion.

[0034] Figure 1 The flow chart of a weakly-supervised video anomaly detection method based on a text-rewritten CLIP model provided by the embodiments of the present invention is as Figure 1 shown, and the method at least includes the following steps:

[0035] Step S110, obtain the video frame sequence to be detected.

[0036] Here, first, a series of video frames are extracted from the video to be detected, and these video frames will be used as the input data for subsequent processing.

[0037] Step S120, construct a video anomaly detection model, including a visual branch, a text branch, a feature fusion module, and two classification branches, where: in the text branch, the SwinBERT model is used to generate dense captions for video frames, and two different text rewritings are generated through the LLaMA-7b model; after converting the text corresponding to the dense captions into sentence embeddings, the temporal relationship between texts is captured through a multi-scale temporal network to obtain text features; in the visual branch, the CLIP image encoder is used to extract the initial visual features of video frames, and is processed by the LGM-Mamba module to enhance its expressive ability; the LGM-Mamba module includes a local-global temporal adapter, a multi-scale temporal network, and a Mamba module; the feature fusion module is used to deeply fuse the final visual features and text features; the two classification branches are respectively used for coarse-grained and fine-grained anomaly detection of the fused features.

[0038] Here, as Figure 2As shown, the video anomaly detection model constructed based on the TrCLIP-Vad framework designed by the present invention includes: a visual branch, a text branch, a feature fusion module, and two classification branches. Visual branch: Use the CLIP image encoder to extract the initial visual features of video frames. Then, the designed LGM-Mamba module is introduced, which includes a local-global temporal adapter, a multi-scale temporal network, and a Mamba module, to further enhance the expressive ability of visual features.

[0039] Text branch: Use the SwinBERT model to generate dense captions for video frames, which describe the content of the video frames, such as "Two people are climbing over the fence"; rewrite the original text in two different ways through the LLaMA-7b model to enhance text features and obtain diverse text descriptions, such as rewritten texts like "Two people, one in black on the left and one in white on the right, are climbing over the fence" or "A tall and thin person and a slightly overweight person are climbing over the fence", etc. Then, convert the text corresponding to the dense captions into sentence embeddings through a pre-trained text embedding network, and capture the temporal relationship between texts through a multi-scale temporal network, thereby obtaining text feature F t 。

[0040] Feature fusion module, responsible for deeply fusing the features extracted by the visual branch and the text branch respectively to make full use of visual and language information. The two classification branches, namely the C-branch and the A-branch, are used for coarse-grained and fine-grained anomaly detection respectively. The coarse-grained branch can quickly identify the abnormal areas in the video, while the fine-grained branch is used to more accurately locate abnormal events.

[0041] Step S130, use the video frame sequence to train the video anomaly detection model in combination with the target loss; wherein, the target loss is a weighted sum of the classification loss, the alignment loss, and the contrast loss.

[0042] Here, use the extracted video frame sequence to train the video anomaly detection model in combination with the target loss function. The target loss function is a weighted sum of the classification loss, the alignment loss, and the contrast loss, and these losses jointly guide the optimization process of the model. The classification loss is used to measure the accuracy of the model in classifying abnormal events. The alignment loss is used to encourage the model to correctly sort abnormal events in time. The contrast loss is used to enhance the model's ability to distinguish abnormal events from normal events.

[0043] In the embodiments of the present invention, a video anomaly detection model that combines visual and language information is constructed to achieve weakly supervised detection of video anomaly events. It uses the SwinBERT and LLaMA-7b models to generate diverse text descriptions, and extracts and enhances visual modeling in video anomaly detection through the CLIP image encoder and the LGM-Mamba module. Finally, coarse-grained and fine-grained anomaly detection is achieved through feature fusion and a classification branch. This method has broad application prospects in the field of video anomaly detection, such as in fields like surveillance video analysis and intelligent security.

[0044] In some embodiments, the method further includes: processing the video frame sequence through data augmentation techniques; the data augmentation techniques include random rotation, flipping, and cropping; randomly pairing the original text and its two rewritten texts with the video frame sequence after data augmentation processing and inputting them into the video anomaly detection model.

[0045] Here, processing the video frame sequence through data augmentation techniques can effectively improve the generalization ability of the model. Among them, random rotation, flipping, and cropping operations can simulate different perspectives and shooting conditions, thereby helping the model learn more robust feature representations.

[0046] Randomly pair the original text and its two rewritten texts with the video frame sequence after data augmentation. This approach has the following benefits: On the one hand, through random pairing, each video frame may be associated with different text descriptions, which increases the diversity of training data and helps the model learn more generalized feature representations. On the other hand, since the correspondence between text descriptions and video frames is random, the model needs to learn to extract useful information from multiple possible text descriptions, thereby enhancing the robustness to changes in text descriptions. Moreover, in the scenario of weakly supervised learning, text descriptions can serve as weak labels or auxiliary information for video content. Through random pairing, the model can learn to utilize these weak labels to improve the performance of anomaly detection.

[0047] The video frame sequence and its corresponding text descriptions after data augmentation processing and text random pairing will be used as inputs and fed into the video anomaly detection model. The model will use these diverse input data to learn how to effectively detect anomaly events in the video. By combining data augmentation techniques and text random pairing strategies, it can help the model better adapt to different shooting conditions and scene changes, thereby improving the accuracy and reliability of anomaly detection. Thus, the performance of the video anomaly detection model is significantly improved.

[0048] In some embodiments, the LGM-Mamba module integrates four Mamba modules and a multi-scale temporal network on the basis of the local-global temporal adapter to process visual features; the four Mamba modules are combined in pairs to form an outer Mamba module pair and an inner Mamba module pair, which are respectively used to capture local and global dependencies in the data.

[0049] Here, the local-global temporal adapter, abbreviated as LGT-Adapter, was initially proposed by vadclip. The local temporal adapter mainly captures short-term dynamic changes in the video frame sequence by introducing a Transformer encoder layer. Different from the traditional Transformer encoder layer, it restricts the calculation range of self-attention within a local window. This method can not only make the model pay more attention to the local patterns and short-term dependencies of the video, but also reduce the computational complexity.

[0050] The global temporal adapter uses a lightweight graph convolutional network (GCN) module to capture global temporal dependencies, and the GCN models the global temporal dependencies from the perspectives of feature similarity and relative distance.

[0051] The Mamba module is a deep learning component based on the state space model (SSM), which has the ability to efficiently process long sequences and capture long-term dependencies. In the LGM-Mamba module, the Mamba module is used to capture local and global features in the data.

[0052] Outer Mamba module pair: mainly used to capture global dependencies in the data. Global dependencies refer to the extensive connections that exist between different time points or different positions in the data. Through the outer Mamba module pair, the LGM-Mamba module can capture these global features, so as to understand the data more comprehensively.

[0053] Inner Mamba module pair: focuses on capturing local dependencies in the data. Local dependencies refer to the close connections that exist between adjacent time points or adjacent positions in the data. The inner Mamba module pair extracts useful local features by carefully analyzing the relationships between these data points.

[0054] The multi-scale temporal network is another important component in the LGM-Mamba module. This network can process temporal information at different scales, so as to capture the dynamic changes of the data more accurately. By combining the multi-scale temporal network, the LGM-Mamba module can analyze the data at different time scales and improve its ability to process complex visual features.

[0055] In some embodiments, the initial visual features are passed through the local-global temporal adapter and then through two embedding layers N1 and N2, which respectively transform the shape of the input data from the original B×S×F to B×S×n 1 and B×S×n 2 , where B, S, and F are the batch size, sequence length, and initial feature dimension respectively, and n 1 and n 2 represent the output feature dimensions of the two embedding layers respectively, and n 1 >n 2 ; After being processed by the Neimamba module pair and the Waimamba module pair respectively, the data transformed by the embedding layers N1 and N2 are merged through element-wise sum operation and finally concatenated together to form new features; Finally, it is processed by the multi-scale temporal network to obtain the final visual features.

[0056] Here, as Figure 3 shown, the initial visual features extracted by the CLIP image encoder are first processed by the local-global temporal adapter, i.e., the LGT adapter. This adapter aims to preliminarily adjust the temporal sequence characteristics of the visual features and lay the foundation for subsequent processing. After passing through the local-global temporal adapter, the visual features are fed into two embedding layers N1 and N2. Embedding layer N1 transforms the data shape from the original [B, S, F] (batch size, sequence length, initial feature dimension) to [B, S, n1], where n1 is the output feature dimension of N1. Embedding layer N2 transforms the data shape to [B, S, n2], where n2 is the output feature dimension of N2, and n1>n2. This means that N1 and N2 may extract feature information at different granularities.

[0057] The transformed data are respectively input into the Neimamba module pair and the Waimamba module pair. Each Mamba module pair consists of two Mamba layers, which work together to capture local and global dependencies in the data. The Neimamba module pair focuses on capturing local dependencies in the data and processes the output of N1. The Waimamba module pair is responsible for capturing global dependencies and processes the output of N2. After the two module pairs extract local and global features respectively, they are merged through element-wise sum operation to integrate the information of these two types of dependencies.

[0058] The merged features are concatenated together to form a new feature representation. This step aims to fuse local and global features into a unified feature vector for subsequent processing. Finally, the new features are fed into the multi-scale temporal network for processing. This network can capture feature information at different time scales, further refining and enhancing the feature representation. After being processed by the multi-scale temporal network, the final visual features are obtained, which can be used for subsequent video anomaly detection or other visual processing tasks.

[0059] In some embodiments, the two classification branches include: Branch C, which inputs the fused features into a binary classifier to obtain an anomaly confidence score, applies the Top-k mechanism to filter out the video frames with the highest anomaly confidence scores as video-level predictions, and uses binary cross-entropy to calculate the classification loss to achieve binary classification; Branch A, which inputs the text labels into the text encoder of CLIP to convert them into corresponding class embeddings, calculates the similarity between the class embeddings and the final visual features, and generates an alignment map; fine-grained video anomaly detection is achieved through the alignment map to identify specific anomaly classes.

[0060] Here, for the C-branch, the present invention follows previous work and the Top-K mechanism. The Top-K mechanism selects the K video frames with the highest anomaly confidence scores, i.e., the frames most likely to contain anomalous behavior, as video-level predictions. Subsequently, binary cross-entropy is used to calculate the classification loss L between the video-level predictions and the ground truth. bce 。

[0061] In Branch A, the present invention adopts the MIL-Align mechanism proposed by VadCLIP. Specifically, MIL-Align focuses on the alignment map M, which captures the similarity between frame-level video features and various types of embeddings. For each row, MIL-Align selects the top k features most similar to the category and calculates their average similarity, which is used as a measure of the alignment degree between the video and the specific category. In this way, MIL-Align generates a vector Y = {y1,......,ym}, which reflects the similarity between the video and various categories. The goal of MIL-Align is to ensure that the video and its corresponding text label have the highest similarity score among all other labels. First, multi-category predictions are calculated, and finally, the alignment loss L can be calculated through cross-entropy. nce 。

[0062] In some embodiments, the target loss is calculated by the following formula:

[0063]

[0064] L nce =-∑ j y j log(p j );

[0065]

[0066] L=L bce +L nce +λL cts ;

[0067] Wherein, L bce is the classification loss, y i is the true label of the i-th sample, taking values of 0 or 1, is the probability that the model predicts the i-th sample belongs to the positive class, and N is the total number of samples; L nce is the alignment loss, y j is the true label of the j-th class, p j is the probability that the model predicts the j-th class; L cts is the contrastive loss, t n and t aj are the maximum anomaly class embeddings in the normal and the j-th class respectively, is the transpose of t n , L is the target loss, and λ is the weight coefficient, which are set to 1×10 -4 and 1×10 -1 respectively in the XD-Violence and UCF-Crime datasets.

[0068] The multi-classification branch will also output a binary classification accuracy. The best one is selected from the values in the multi-classification branch and the values in the binary classification branch to evaluate the frame-level anomaly degree.

[0069] The new weakly supervised video anomaly detection framework (TrCLIP-Vad framework) provided by the embodiments of the present invention realizes innovation in the field of weakly supervised video anomaly detection by introducing a text rewriting branch and an LGM-Mamba module, and adopting multiple loss functions to optimize the performance. This framework not only solves the problem of imbalance between text and visual features, but also strengthens visual modeling and feature fusion, improving the accuracy and robustness of video anomaly detection. Through the collaborative work of the C branch and the A branch, this framework can simultaneously achieve coarse-grained and fine-grained anomaly detection, providing new ideas and methods for the field of video anomaly detection. The TrCLIP-Vad framework is an innovative video anomaly detection system that combines a text rewriting branch, an LGM-Mamba module, and multiple loss functions to optimize the performance.

[0070] This framework achieved the best results on both the XD-Violence and UCF-Crime datasets. VadCLIP achieved 84.51% AP for binary classification and 24.70% AVG for multi-class classification on the XD-Violence dataset, and 88.02% AUC for binary classification and 6.68% AVG for multi-class classification on the UCF-Crime dataset. However, the framework TrCLIP-Vad provided by the present invention far exceeded their results, achieving 86.49% AP for binary classification and 30.49% AVG for multi-class classification on the XD-Violence dataset, and 88.90% AUC for binary classification and 9.17% AVG for multi-class classification on the UCF-Crime dataset.

[0071] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the order numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The serial numbers of the embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0072] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0073] In several embodiments provided by the present invention, it should be understood that the disclosed methods can be implemented in other ways. The methods disclosed in several method embodiments provided by the present invention can be combined arbitrarily without conflict to obtain new method embodiments. The features disclosed in several method embodiments provided by the present invention can be combined arbitrarily without conflict to obtain new method embodiments.

[0074] As described above, it is only the implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims described above.

Claims

1. A weakly supervised video anomaly detection method based on text rewriting CLIP model, characterized in that: include: Obtain a video frame sequence to be detected; A video anomaly detection model is constructed, including a visual branch, a text branch, a feature fusion module and two classification branches, wherein: the SwinBERT model is used in the text branch to generate dense subtitles of video frames, and two different text rewrites are generated through the LLaMA-7b model; after the dense subtitles corresponding text is converted into sentence embeddings, the temporal relationship between texts is captured through a multi-scale temporal network to obtain text features; the CLIP image encoder is used in the visual branch to extract the initial visual features of the video frame, and is processed through the LGM-Mamba module to enhance its expressiveness; the LGM-Mamba module includes a local-global temporal adapter, a multi-scale temporal network and a Mamba module; the feature fusion module is used to deeply fuse the final visual features with the text features; the two classification branches are used to perform coarse-grained and fine-grained anomaly detection on the fused features, respectively; The video anomaly detection model is trained by using the video frame sequence in combination with a target loss; wherein the target loss is a weighted sum of a classification loss, an alignment loss, and a contrast loss.

2. The method according to claim 1, characterized in that The method further comprises: Processing the video frame sequence by data enhancement technology; the data enhancement technology includes random rotation, flipping and cropping; The original text and its two rewritten texts are randomly selected and paired with the video frame sequence after data enhancement processing, and then input into the video anomaly detection model.

3. The method according to claim 1, characterized in that The LGM-Mamba module integrates four Mamba modules and a multi-scale temporal network on the basis of the local-global temporal adapter to process visual features; The four Mamba modules are combined in pairs to form an outer Mamba module pair and an inner Mamba module pair, which are used to capture local and global dependencies in the data respectively.

4. The method according to claim 3, characterized in that The initial visual features are passed through the local-global temporal adapter and then through two embedding layers N1 and N2, which transform the shape of the input data from the original B×S×F to B×S×n1 and B×S×n2, respectively, where B, S, and F are the batch size, sequence length, and initial feature dimension, respectively, and n1 and n2 represent the output feature dimensions of the two embedding layers, respectively, and n1>n2; The data converted by the embedding layers N1 and N2 are respectively input into the inner Mamba module pair and the outer Mamba module pair for processing, merged by element and addition operations, and finally spliced ​​together to form new features; finally, the data is processed by a multi-scale temporal network to obtain the final visual features.

5. The method according to any one of claims 1 to 4, characterized in that: The two classification branches include: Branch C is used to input the fused features into the binary classifier to obtain the anomaly confidence score, and apply the Top-K mechanism to filter out the video frames with the highest anomaly confidence score as the video-level prediction, and use binary cross entropy to calculate the classification loss to achieve binary classification; Branch A is used to input the text label into the text encoder of CLIP to convert it into a corresponding class embedding, calculate the similarity between the class embedding and the final visual feature, and generate an alignment map; fine-grained video anomaly detection is achieved through the alignment map to identify specific anomaly classes.

6. The method according to any one of claims 1 to 4, characterized in that: The target loss is calculated by the following formula: L nce =-∑ j y j log(p j ); L=L bce +L nce +λL cts ; Where, L bce is the classification loss, y i is the true label of the i-th sample, which takes a value of 0 or 1. is the probability that the model predicts that the i-th sample belongs to the positive class, N is the total number of samples; L nce is the alignment loss, y j is the true label of the jth category, p j is the probability of the model predicting the jth category; L cts is the contrast loss, t n and t aj are the maximum abnormal class embeddings in the normal and j-th categories, respectively, t n is the transpose of , L is the target loss, and λ is the weight coefficient.

Citation Information

Cited By

  • Multi-mode prompt memory unsupervised continuous anomaly detection method and system

    CN120707974A