Video anomaly detection method and system based on OPEN-VCLIP model
By combining the features extracted by the OPEN-VCLIP model and the Inflated 3D encoder, the video anomaly detection method with adaptive graph convolution network and cross-batch clustering strategy is used to solve the problem of feature extraction deviation and insufficient detection accuracy in the prior art, and more efficient video anomaly detection is achieved.
Patent Information
- Application Number
- CN202510254546.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-03
AI Technical Summary
The existing video anomaly detection methods have multiple instances to learn sorting losses. Focus on a few instances with strong distinctiveness and ignore other information. The characteristics of the pre-trained model are biased against the original video representation, and the reliability and richness of positive and negative samples are insufficient, which affects the detection accuracy.
The video anomaly detection method based on the OPEN-VCLIP model is adopted, and the spatiotemporal features and the OPEN-VCLIP model are extracted through the Inflated 3D encoder, and the appearance features of the item are extracted. After the fusion is input to the adaptive graph convolution network for abnormal prediction, and the model is optimized through cross-batch clustering strategy and preset loss function.
It significantly improves the accuracy and feature distinction ability of video anomaly detection, enhances intra-class consistency and inter-class differences, and reduces the deviation of the original video representation in the pre-training stage.
Smart Images

Figure CN120088706A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video data processing, and particularly relates to a video anomaly detection method and system based on the OPEN-VCLIP model. Background Art
[0002] Efficient and accurate video anomaly detection algorithms can help maintain social security and ensure social stability. In the field of intelligent security, video anomaly detection is of great significance, aiming to promptly identify various abnormal events and safeguard social security. The popularization of surveillance cameras in public places has made video anomaly detection a key technology for ensuring public safety. It can quickly detect abnormal situations such as traffic accidents, fights, thefts, or explosions in scenarios such as shopping malls, banks, and traffic intersections, providing a basis for taking timely measures, thereby improving the efficiency and accuracy of public safety management. However, abnormal behaviors often hide among normal behaviors, with the characteristics of low occurrence frequency and difficulty in identification, which poses a huge challenge to the reliability and accuracy of detection technologies.
[0003] Existing anomaly detection methods have the following problems: (1) The multi-instance learning ranking loss often only focuses on a few of the most discriminative instances, and the information of a large number of other video segments is ignored, unable to fully mine the abnormal features in the video; (2) Existing anomaly detection methods are easily affected by pre-trained models, and different pre-trained features have biases in the representation of the original video, affecting the accuracy of anomaly detection; (3) The reliability of positive sample pairs and the richness of negative sample pairs are insufficient, thus affecting the performance of anomaly detection. Summary of the Invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a video anomaly detection method and system based on the OPEN-VCLIP model. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0005] 1. A video anomaly detection method based on the OPEN-VCLIP model, comprising:
[0006] S1. Reading a training video, using an Inflated 3D encoder as a feature extraction network to extract the spatio-temporal features of the training video, and using the OPEN-VCLIP model as a feature extraction network to extract the appearance features of the items in the training video;
[0007] S2. Fusing the spatio-temporal features and the appearance features of the items to obtain pre-trained video features that simultaneously contain temporal relationships and object co-occurrence relationships;
[0008] S3. Input the pre-trained video features into the adaptive graph convolutional network in sequence to obtain anomaly prediction scores, and apply a cross-batch clustering strategy to the intermediate layer features to obtain the clustering centers of normal videos and the clustering centers of abnormal videos;
[0009] S4. Substitute the normal clustering center, the abnormal clustering center, and the anomaly prediction scores into a preset loss function, and stop training when the loss value reaches a preset threshold to obtain a trained weakly supervised video anomaly detection model;
[0010] S5. Input the video data to be detected into the trained weakly supervised video anomaly detection model for detection to obtain an anomaly detection result.
[0011] In a specific embodiment, the step S3 includes:
[0012] S31. Construct an adaptive graph convolutional network. The adaptive convolutional network includes three graph convolutional layers. After the first layer and the second layer, a ReLU activation function and a dropout function are connected. After the last layer, a Sigmoid activation function is connected. Construct an adaptive global graph, input the training video features and the adjacency matrix of the adaptive global graph into the adaptive graph convolutional network, and output the anomaly prediction scores of the training videos;
[0013] S32. The first graph convolutional network of the adaptive graph convolutional network outputs intermediate features, which include intermediate normal features and intermediate abnormal features. Through the K-means algorithm, the standardized intermediate normal feature representations and intermediate abnormal feature representations are respectively clustered into two categories; for abnormal videos, the two clustered feature categories represent normal events and abnormal events, and the loss based on batch clustering is used to iteratively push the centers of the two clusters farther apart; for normal videos, the two clustered feature categories both represent normal events, and the loss based on batch clustering is used to iteratively pull the centers of the two clusters closer, that is,
[0014]
[0015] where d = ∥c 1 - c 2 ∥ 2 is the distance between the two clustering centers, c 1 and c 2 are the two cluster centers obtained after clustering all the normal videos or abnormal videos in the batch (i.e., ), u is an upper bound to help the model be robust to different video data, and b is the batch size;
[0016] S33. During the iteration process of step S32, construct an abnormal clustering memory bank C a and a normal clustering memory bank C n, after the training of each batch is completed, all abnormal videos and normal videos in the current batch are respectively batch-clustered to obtain abnormal clustering centers and normal clustering centers, which are added to the abnormal clustering memory bank C a and the normal clustering memory bank C n ;
[0017] During the training process of the nth batch, the clustering centers stored in C a 、C n are respectively binary-clustered to obtain and where, represents the result after the binary clustering of the abnormal clustering memory bank, and is used as the two initial clustering centers for the abnormal feature clustering during the training process of the nth batch; represents the result after the binary clustering of the normal clustering memory bank, and is used as the two initial clustering centers for the normal feature clustering during the training process of the nth batch. After the training of the current batch is completed, the clustering centers obtained after training are added to the clustering memory bank, that is, m is the current iteration number.
[0018] In a specific embodiment, the preset loss function includes a feature contrast loss function and a center contrast loss function, and the step S4 includes:
[0019] S41. Select segments with the highest and lowest abnormal prediction scores and the same quantity, and respectively determine them as candidate normal event features, candidate abnormal event features, and candidate background features;
[0020] S42. Calculate the feature contrast loss function and the center contrast loss function respectively to obtain the preset loss function;
[0021] Use the candidate normal event features and the clustering centers of the corresponding categories as feature positive sample pairs, and use the candidate abnormal event features and the clustering centers of the opposite categories as feature negative sample pairs to obtain the feature contrast loss function. Among them, the feature positive sample pairs include: candidate normal event features and normal clustering centers, candidate background features and normal clustering centers, candidate abnormal event features and abnormal clustering centers, and the feature negative sample pairs include: candidate normal event features and abnormal clustering centers, candidate background features and abnormal clustering centers, candidate abnormal event features and normal clustering centers;
[0022] S43. Use the cluster center and the candidate features of the corresponding category as the central positive sample pairs, and use the cluster center and the candidate features of the opposite category as the central negative samples to obtain the central contrast loss function. Among them, the central positive sample pairs include: the normal cluster center and the candidate normal event features, the normal cluster center and the candidate background features, and the abnormal cluster center and the candidate abnormal event features. The central negative sample pairs include: the normal cluster center and the candidate abnormal event features, the abnormal cluster center and the candidate normal event features, and the abnormal cluster center and the candidate background features.
[0023] In a specific embodiment, the preset loss function is: Among them, represents the feature contrast loss function, represents the feature contrast loss function.
[0024] In a specific embodiment, the feature contrast loss function is:
[0025]
[0026] Among them, X v represents the candidate normal event features, abnormal event features or background features, T represents transpose, represents the reliable features of the category opposite to X v candidate features, C v represents the normal cluster center or the abnormal cluster center, s (·,·) represents the cosine similarity function, which is used to calculate the similarity between two vectors; τ is the temperature parameter of the contrast loss, representing the expected degree of distinction between features.
[0027] In a specific embodiment, the central contrast loss function is:
[0028]
[0029] Among them, represents the candidate features of the category opposite to the cluster center C v , s (·,·) represents the cosine similarity function, which calculates the similarity between two vectors; τ is the temperature parameter of the contrast loss, representing the degree of distinction between features.
[0030] The present invention also discloses a video anomaly detection system based on the OPEN-VCLIP model, including:
[0031] A feature extraction module, configured to read a training video, and use an Inflated 3D encoder as a feature extraction network to extract the spatio-temporal features of the training video and use the OPEN-VCLIP model as a feature extraction network to extract the appearance features of the items in the training video;
[0032] A feature fusion module, configured to fuse the spatio-temporal features and the item appearance features to obtain pre-trained video features that simultaneously contain temporal relationships and object co-occurrence relationships;
[0033] A clustering module, configured to sequentially input the pre-trained video features into an adaptive graph convolutional network to obtain anomaly prediction scores, and apply a cross-batch clustering strategy to the intermediate layer features to obtain the clustering centers of normal videos and the clustering centers of abnormal videos;
[0034] A training module, configured to substitute the normal clustering center, the abnormal clustering center, and the anomaly prediction scores into a preset loss function, and stop training when the loss value reaches a preset threshold to obtain a trained weak supervision video anomaly detection model;
[0035] A detection module, configured to input the video data to be detected into the trained weak supervision video anomaly detection model for detection to obtain an anomaly detection result.
[0036] Advantages of the present invention:
[0037] The video anomaly detection method based on the OPEN-VCLIP model of the present invention realizes more accurate normal and abnormal classification guidance through the clustering center, and through the contrast loss mechanism, narrows the distance between the same-class clustering center and the video features, while keeping away from different-class features, significantly improving the discrimination ability of the features, enhancing the intra-class consistency and inter-class difference. In addition, by combining three complementary pre-trained features, the spatio-temporal information and the object co-occurrence information are more perfectly fused, reducing the deviation of the original video representation in the pre-training stage.
[0038] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings
[0039] Figure 1 is a schematic flowchart of a video anomaly detection method based on the OPEN-VCLIP model provided by an embodiment of the present invention;
[0040] Figure 2 is a block diagram of a video anomaly detection system module based on the OPEN-VCLIP model provided by an embodiment of the present invention. Detailed Embodiments
[0041] The following further describes the present invention in detail in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0042] Embodiment 1
[0043] Please refer to Figure 1 , Figure 1It is a schematic flowchart of a video anomaly detection method based on the OPEN-VCLIP model provided by an embodiment of the present invention, including:
[0044] S1. Read the training video, use the Inflated 3D encoder as the feature extraction network to extract the spatio-temporal features of the training video, and use the OPEN-VCLIP model as the feature extraction network to extract the object appearance features of the training video;
[0045] Specifically, for the training video input into the network, it passes through two pre-trained models respectively to extract the spatio-temporal features and object appearance features of the video dataset, and fuse them to obtain more complete and general pre-trained features. The original video contains N normal videos and N abnormal videos and their corresponding video-level labels Y n = 0, Y a = 1. The abnormal video is composed of abnormal video segments and normal video segments. Each abnormal video contains at least one abnormal video segment. The normal video is composed of normal video segments and does not contain abnormal video segments;
[0046] Select the Inflated 3D (I3D) model pre-trained on the large-scale video dataset Kinetics as the feature extractor to extract the spatio-temporal features of the video. The spatio-temporal features of video V i of can be expressed as:
[0047]
[0048] Among them, is the spatio-temporal feature of the j-th video segment of the training video V i , The dimension of is i dimensions, T i is the number of video segments contained in the training video V represents the dimension of the spatio-temporal feature;
[0049] Select the pre-trained OPEN-VCLIP model (based on the ViT-B / 16 architecture) as the feature extraction network to extract the features of the video. These features contain rich object co-occurrence relationships and appearance features, and can effectively capture the spatio-temporal dynamic information in the video. For video V i , its appearance features The extraction process is as follows:
[0050] OPEN-VCLIP improves the original CLIP model by constructing the VCLIP model to better adapt to video tasks. It has a special design in the self-attention layer, expanding the temporal attention view, enabling each video segment to obtain information from adjacent frames, thereby enhancing the context relevance of features. Specifically, the new implementation of its self-attention layer is as follows:
[0051]
[0052] where d is the vector dimension, and q s,t refers to the query vector of the s-th token in the t-th frame, and [K (t-1)~(t+1) and [V (t-1)~(t+1) respectively represent the matrices composed of the key vectors and value vectors of the t-th frame and its adjacent frames. This improvement enables the model to better capture the changes in actions and events in the video in the temporal dimension.
[0053] In this way, the appearance features i of video V can integrate the dynamic information of adjacent frames and are no longer limited to the feature representation of individual frames. Its dimension is still dimensions, where T i is the number of video segments contained in the training video V i , represents the dimension of the item appearance feature, but the features at this time contain richer spatio-temporal dynamic information, which can provide a more discriminative feature representation for subsequent video anomaly detection.
[0054] S2. Fuse the spatio-temporal features and the item appearance features to obtain pre-trained video features that simultaneously contain temporal relationships and object co-occurrence relationships;
[0055] Specifically, by fusing the spatio-temporal features and appearance features extracted by the two pre-trained models, more comprehensive original video features that simultaneously contain temporal relationships and object co-occurrence relationships are obtained. Finally, the feature X i of video V i can be expressed as:
[0056]
[0057] where X i represents the training video feature, the dimension of X i is T i ×D i , T i is the number of video segments contained in the training video V i , and D i represents the dimension of the fused feature.
[0058] S3. Input the pre-trained video features into the adaptive graph convolutional network in sequence to obtain anomaly prediction scores, and apply a cross-batch clustering strategy to the intermediate layer features to obtain the clustering centers of normal videos and the clustering centers of abnormal videos;
[0059] In a specific embodiment, the step S3 includes:
[0060] S31. Construct an adaptive graph convolutional network. The adaptive convolutional network contains three graph convolutional layers. After the first layer and the second layer, ReLU activation functions and dropout functions are connected. After the last layer, a Sigmoid activation function is connected. Construct an adaptive global graph, input the training video features and the adjacency matrix of the adaptive global graph into the adaptive graph convolutional network, and output the anomaly prediction scores of the training videos;
[0061] S32. The first graph convolutional network of the adaptive graph convolutional network outputs intermediate features. The intermediate features include intermediate normal features and intermediate abnormal features. Through the K-means algorithm, the standardized intermediate normal features and intermediate abnormal feature representations are respectively clustered into two categories; for abnormal videos, the two categories of features after clustering represent normal events and abnormal events respectively, and the loss based on batch clustering is used to iteratively push the centers of the two clusters farther apart; for normal videos, the two categories of features after clustering both represent normal events, and the loss based on batch clustering is used to iteratively pull the centers of the two clusters closer, that is,
[0062]
[0063] where d = ∥c 1 - c 2 ∥ 2 is the distance between the two clustering centers, c 1 , c 2 are the centers of the two clusters obtained after clustering all normal videos or abnormal videos in the batch (that is, ), u is an upper bound to help the model be robust to different video data, and b is the batch size;
[0064] The effect of the clustering algorithm is usually affected by the selection of the initial clustering centers. When performing batch clustering based on the K-means algorithm, randomly selecting two segments from all video segments as the initial clustering centers is very likely to lead to significantly different clustering results due to different initial center selections. To better select appropriate initial clustering centers, a cross-batch learning module is used to provide guidance for the current batch clustering by introducing the clustering results of the previous batch, making the convergence speed of the clustering algorithm faster and obtaining more accurate clustering centers and clustering results.
[0065] Due to the limitation of GPU memory and the need to ensure accurate guidance for clustering centers, and since the batch clustering algorithm has a high dependence on clustering centers, a cross-batch clustering strategy is used. During the iterative process, an abnormal clustering memory bank C a and a normal clustering memory bank C n are constructed respectively to store the learned knowledge, and the stored information is used to provide guidance for the current batch of clustering to ensure the stability and accuracy of clustering.
[0066] S33. During the iterative process of step S32, an abnormal clustering memory bank C a and a normal clustering memory bank C n are constructed respectively. After the training of each batch is completed, all abnormal videos and normal videos in the current batch are respectively subjected to batch clustering to obtain abnormal clustering centers and normal clustering centers, which are added to the abnormal clustering memory bank C a and the normal clustering memory bank C n ;
[0067] During the training process of the nth batch, the clustering centers stored in C a and C n are respectively subjected to binary clustering to obtain and Among them, represents the result after binary clustering of the abnormal clustering memory bank and serves as the two initial clustering centers for abnormal feature clustering during the training process of the nth batch; represents the result after binary clustering of the normal clustering memory bank and serves as the two initial clustering centers for normal feature clustering during the training process of the nth batch. After the training of the current batch is completed, the trained clustering centers are added to the clustering memory bank, that is, m is the current number of iterations.
[0068] Then, during the training process of the n + 1th batch, the above operations are repeated, and the clustering centers obtained by binary clustering of the clustering centers stored in C a and C n are used as the initial training centers for the new training.
[0069] S4. Substitute the normal clustering center, the abnormal clustering center, and the abnormal prediction score into a preset loss function, and stop training when the loss value reaches a preset threshold to obtain a trained weak-supervised video anomaly detection model;
[0070] In a specific embodiment, the preset loss function includes a feature contrast loss function and a center contrast loss function, and step S4 includes:
[0071] S41. Select segments with the highest and lowest anomaly prediction scores and the same quantity, and determine them as candidate normal event features, candidate abnormal event features, and candidate background features respectively;
[0072] Specifically, k segments with the highest scores in the normal videos and abnormal videos can be selected as candidate normal event features and abnormal event features, and at the same time, k segments with the lowest scores in the normal videos and abnormal videos can be selected as candidate background features. The normal clustering center and the abnormal clustering center represent reliable normal features and abnormal features respectively.
[0073] S42. Calculate the feature contrast loss function and the center contrast loss function respectively to obtain a preset loss function;
[0074] Take the candidate normal event features and the clustering centers of the corresponding categories as feature positive sample pairs, and take the candidate abnormal event features and the clustering centers of the opposite categories as feature negative sample pairs to obtain the feature contrast loss function. Among them, the feature positive sample pairs include: candidate normal event features and normal clustering centers, candidate background features and normal clustering centers, candidate abnormal event features and abnormal clustering centers. The feature negative sample pairs include: candidate normal event features and abnormal clustering centers, candidate background features and abnormal clustering centers, candidate abnormal event features and normal clustering centers;
[0075] S43. Take the clustering centers and the candidate features of the corresponding categories as center positive sample pairs, and take the clustering centers and the candidate features of the opposite categories as center negative samples to obtain the center contrast loss function. Among them, the center positive sample pairs include: normal clustering centers and candidate normal event features, normal clustering centers and candidate background features, abnormal clustering centers and candidate abnormal event features. The center negative sample pairs include: normal clustering centers and candidate abnormal event features, abnormal clustering centers and candidate normal event features, abnormal clustering centers and candidate background features.
[0076] In a specific embodiment, the preset loss function is: Among them, represents the feature contrast loss function, represents the feature contrast loss function.
[0077] In a specific embodiment, the feature contrast loss function is:
[0078]
[0079] Among them, X v represents the candidate normal event features, abnormal event features or background features, T represents transpose, represents the reliable features of the opposite category to X v candidate features, C vIndicates a normal clustering center or an abnormal clustering center. s (·,·) represents the cosine similarity function, which is used to calculate the similarity between two vectors; τ is the temperature parameter of the contrast loss, representing the degree of discrimination between expected features.
[0080] The feature contrast loss is the n+1-class cross-entropy loss for candidate positive samples, expecting the candidate positive samples to be correctly classified into the clustering center C v 's category, rather than other categories where negative samples are located. Therefore, the feature contrast loss optimizes the candidate event features and background features through the precise guidance of the clustering center, helping the normal / abnormal event features and background features to be closer to their distributions.
[0081] In a specific embodiment, the center contrast loss function is:
[0082]
[0083] Wherein, represents the candidate feature of the category opposite to the clustering center C v (·,·) represents the cosine similarity function, calculating the similarity between two vectors; τ is the temperature parameter of the contrast loss, representing the degree of discrimination between features. s (·,·) represents the cosine similarity function, calculating the similarity between two vectors; τ is the temperature parameter of the contrast loss, representing the degree of discrimination between features.
[0084] The center contrast loss is the n+1-class cross-entropy loss for the clustering center, expecting the clustering center to be correctly classified into the candidate positive sample x v , rather than the n negative sample categories
[0085] The difference between the center contrast loss and the feature contrast loss lies in whether the sample to be classified is the clustering center C v , or the candidate positive sample x v . When the sample to be classified is the clustering center, the contrast loss hopes to pull the clustering center closer to the candidate positive sample and push it away from the negative samples at the same time, optimizing the clustering center through the reliable candidate normal / abnormal video segments in this batch to ensure the correctness and reliability of the clustering center; while when the sample to be classified is the candidate sample, the contrast loss hopes that the candidate positive sample approaches the clustering center and is pushed away from the remaining negative samples; it is expected to help the normal / abnormal event features and background features to be closer to their corresponding feature distributions through the normal / abnormal clustering centers with statistical information. Therefore, the total loss function simultaneously updates the clustering center in the current round to ensure its precise guidance; and optimizes the normal / abnormal feature distribution in the current round according to the statistical information possessed by the clustering center.
[0086] S5. Input the video data to be detected into the trained weakly supervised video anomaly detection model for detection to obtain the anomaly detection result.
[0087] The present invention also discloses a video anomaly detection system based on the OPEN-VCLIP model, including:
[0088] A feature extraction module, configured to use an Inflated 3D encoder as a feature extraction network to extract spatio-temporal features of a training video, and use the OPEN-VCLIP model as a feature extraction network to extract appearance features of items in the training video;
[0089] A feature fusion module, configured to fuse the spatio-temporal features and the appearance features of the items to obtain pre-trained video features that simultaneously include temporal relationships and object co-occurrence relationships;
[0090] A clustering module, configured to sequentially input the pre-trained video features into an adaptive graph convolutional network to obtain anomaly prediction scores, and apply a cross-batch clustering strategy to the intermediate layer features to obtain the clustering centers of normal videos and the clustering centers of abnormal videos;
[0091] A training module, configured to substitute the normal clustering center, the abnormal clustering center, and the anomaly prediction scores into a preset loss function, and stop training when the loss value reaches a preset threshold to obtain a trained weakly supervised video anomaly detection model;
[0092] A detection module, configured to input the video data to be detected into the trained weakly supervised video anomaly detection model for detection to obtain an anomaly detection result.
[0093] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0094] Although the present application has been described in conjunction with various embodiments, those skilled in the art will understand and realize other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit may implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce a favorable effect.
[0095] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system (device), or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects, which are collectively referred to herein as "modules" or "systems". Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The computer program is stored / distributed in a suitable medium, provided together with other hardware or as part of the hardware, or may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.
[0096] The present application is described with reference to the flowcharts and / or block diagrams of the methods, systems (devices), and computer program products of the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a system for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0097] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufacture including an instruction system that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 one or more processes and / or blocks Figure 1 steps for the functions specified in one block or a plurality of blocks.
[0099] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A video anomaly detection method based on OPEN-VCLIP model, characterized in that: include: S1, read the training video, use the Inflated 3D encoder as the feature extraction network to extract the spatiotemporal features of the training video, and use the OPEN-VCLIP model as the feature extraction network to extract the appearance features of the objects in the training video; S2, fusing the spatiotemporal features with the object appearance features to obtain pre-trained video features that simultaneously include temporal relationships and object co-occurrence relationships; S3, sequentially inputting the pre-trained video features into the adaptive graph convolutional network to obtain anomaly prediction scores, and applying a cross-batch clustering strategy to the intermediate layer features to obtain the clustering centers of normal videos and abnormal videos; S4, substituting the normal cluster center, the abnormal cluster center and the abnormal prediction score into a preset loss function so that the training is stopped when the loss value reaches a preset threshold to obtain a trained weakly supervised video anomaly detection model; S5. Input the video data to be detected into a trained weakly supervised video anomaly detection model to perform detection and obtain an anomaly detection result.
2. The video anomaly detection method based on the OPEN-VCLIP model according to claim 1, characterized in that: The step S3 comprises: S31. Construct an adaptive graph convolution network. The adaptive convolution network includes three graph convolution layers. The first and second layers are connected to the ReLU activation function and the dropout function. The last layer is connected to the Sigmoid activation function. An adaptive global graph is constructed. The adjacency matrix of the training video features and the adaptive global graph is input into the adaptive graph convolution network. The abnormality prediction score of the training video is output. S32. The first layer of the adaptive graph convolutional network outputs intermediate features, which include intermediate normal features and intermediate abnormal features. The standardized intermediate normal features and intermediate abnormal features are clustered into two categories through the K-means algorithm. For abnormal videos, the two types of clustered features represent normal events and abnormal events, respectively, and the centers of the two clusters are pushed away by the loss based on batch clustering iteration. For normal videos, the two types of clustered features both represent normal events, and the centers of the two clusters are pulled closer by the loss based on batch clustering iteration, that is, Among them, d = ||c1-c2||2 is the distance between the two cluster centers, c1 and c2 are all normal videos or abnormal videos in the batch (i.e. ) The two cluster centers obtained after clustering, u is an upper bound that helps the model to be robust to different video data, and b is the batch size; S33, in the iterative process of step S32, respectively construct an abnormal cluster memory library C a and the normal cluster memory C n After the training of each batch is completed, all abnormal videos and normal videos in the current batch are batch clustered separately to obtain the abnormal cluster center and the normal cluster center, and add them to the abnormal cluster memory library C a and the normal cluster memory C n middle; During the nth batch training, C a , C n The cluster centers stored in are clustered by binary values, and we get as well as in, Represents the result of binary clustering of the abnormal cluster memory library, which serves as the two initial clustering centers of abnormal feature clustering during the nth batch training process; It represents the result of binary clustering of the normal cluster memory library, which is used as the two initial cluster centers of normal feature clustering in the nth batch training process. After the training of the current batch is completed, the cluster centers obtained after training are added to the cluster memory library, that is, m is the current iteration number.
3. The video anomaly detection method based on the OPEN-VCLIP model according to claim 1, characterized in that: The preset loss function includes a feature contrast loss function and a center contrast loss function, and the step S4 includes: S41, selecting the segments with the highest and lowest abnormal prediction scores and the same number, and determining them as candidate normal event features, candidate abnormal event features, and candidate background features, respectively; S42, respectively calculating a feature contrast loss function and a center contrast loss function to obtain a preset loss function; The candidate normal event features and the cluster centers of the corresponding categories are taken as feature positive sample pairs, and the candidate abnormal event features and the cluster centers of the opposite categories are taken as feature negative sample pairs to obtain a feature contrast loss function, wherein the feature positive sample pairs include: the candidate normal event features and the normal cluster centers, the candidate background features and the normal cluster centers, the candidate abnormal event features and the abnormal cluster centers, and the feature negative sample pairs include: the candidate normal event features and the abnormal cluster centers, the candidate background features and the abnormal cluster centers, and the candidate abnormal event features and the normal cluster centers; S43. The cluster center and the candidate features of the corresponding category are taken as the center positive sample pair, and the cluster center and the candidate features of the opposite category are taken as the center negative sample to obtain the center contrast loss function, wherein the center positive sample pair includes: the normal cluster center and the candidate normal event features, the normal cluster center and the candidate background features, the abnormal cluster center and the candidate abnormal event features, and the center negative sample pair includes: the normal cluster center and the candidate abnormal event features, the abnormal cluster center and the candidate normal event features, and the abnormal cluster center and the candidate background features.
4. The video anomaly detection method based on the OPEN-VCLIP model according to claim 3, characterized in that: The preset loss function is: in, represents the feature contrast loss function, represents the feature contrast loss function.
5. The video anomaly detection method based on the OPEN-VCLIP model according to claim 4, characterized in that: The feature contrast loss function is: Among them, X v represents candidate normal event features, abnormal event features or background features, T represents transposition, Represents and X v Reliable features of the opposite category of candidate features, C v represents the normal cluster center or the abnormal cluster center, s (·,·) represents the cosine similarity function, which is used to calculate the similarity between two vectors; τ is the temperature parameter of the contrast loss, which represents the degree of distinction between the desired features.
6. The video anomaly detection method based on the OPEN-VCLIP model according to claim 4, characterized in that: The center contrast loss function is: in, Representative and cluster center C v Candidate features of the opposite category, s (·,·) represents the cosine similarity function, which calculates the similarity between two vectors; τ is the temperature parameter of the contrast loss, which represents the degree of distinction between features.
7. A video anomaly detection system based on OPEN-VCLIP model, characterized in that: include: A feature extraction module is used to read the training video, use the Inflated 3D encoder as a feature extraction network to extract the spatiotemporal features of the training video, and use the OPEN-VCLIP model as a feature extraction network to extract the appearance features of the objects in the training video; A feature fusion module, used to fuse the spatiotemporal features with the object appearance features to obtain pre-trained video features that simultaneously include temporal relationships and object co-occurrence relationships; A clustering module, used to sequentially input the pre-trained video features into the adaptive graph convolutional network to obtain anomaly prediction scores, and apply a cross-batch clustering strategy to the intermediate layer features to obtain the clustering centers of normal videos and the clustering centers of abnormal videos; A training module, used for substituting the normal cluster center, the abnormal cluster center and the abnormal prediction score into a preset loss function so as to stop the training when the loss value reaches a preset threshold value to obtain a trained weakly supervised video anomaly detection model; The detection module is used to input the video data to be detected into a trained weakly supervised video anomaly detection model to perform detection and obtain an anomaly detection result.