Video event feature clustering method and system based on text description

Through the video event feature clustering method based on text description, combined with adaptive graph convolution network and cross-batch clustering strategy, the problems of high false alarm rate of video anomaly detection and bias of pre-trained model in the prior art are solved, and more accurate video anomaly detection is achieved.

CN120107862APending Publication Date: 2025-06-06NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510254954.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When existing video anomaly detection technology deals with complex scenarios, the model false alarm rate is high, which is difficult to meet actual needs. The pre-trained model has deviations in video representation, making it difficult to fully reflect the video content.

Method used

The video event feature clustering method based on text description is adopted, and the training data set is obtained and the spatiotemporal and object appearance features are extracted using two pretrained models, and the pretrained video features are obtained after fusion. Then, using an adaptive graph convolution network and cross-batch clustering strategy, the clustering centers of normal and anomaly videos are obtained, and the weakly supervised video anomaly detection model is trained through the preset loss function.

Benefits of technology

It significantly improves the ability to distinguish features, enhances intra-class consistency and inter-class differences, reduces the deviation from the original video representation in the pre-training stage, and achieves more accurate normal and abnormal classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107862A_ABST
    Figure CN120107862A_ABST
Patent Text Reader

Abstract

The invention discloses a video event feature clustering method based on text description, and the method comprises the steps: obtaining a training data set, extracting spatial-temporal features and object appearance features of a video data set from the training data set through two pre-training models, and fusing the spatial-temporal features and the object appearance features to obtain pre-training video features; for a batch of input pre-training video features, obtaining an abnormal prediction score through an adaptive graph convolutional network, and applying a cross-batch clustering strategy to intermediate layer features to obtain a clustering center of normal videos and a clustering center of abnormal videos; substituting the normal clustering center, the abnormal clustering center and the abnormal prediction score into a preset loss function so as to stop training when a loss value reaches a preset threshold value to obtain a trained weak supervision video anomaly detection model; and inputting the to-be-detected video data into a trained weak supervision video anomaly detection model for detection to obtain an anomaly detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of video data processing, and in particular relates to a video event feature clustering method and system based on text description. Background Art

[0002] With the development of society, surveillance cameras are widely used in public places such as shopping malls, banks, and transportation hubs. Video anomaly detection technology is crucial in ensuring social security. It can promptly identify abnormal events such as traffic accidents, violent conflicts, and thefts, provide key support for law enforcement agencies, and improve the efficiency and accuracy of public safety management. However, the proportion of abnormal behaviors in video data is extremely low, and they are hidden in a large number of normal behaviors. The patterns are complex and changeable, which brings great challenges to anomaly detection and puts higher requirements on the reliability and accuracy of the technology.

[0003] Existing solutions usually identify anomalies by learning normal patterns, but due to the complexity and diversity of scenarios, they cannot cover all normal behavior patterns, resulting in a high false alarm rate of the model and difficulty in meeting actual needs. In addition, existing methods rely heavily on pre-trained models, and different pre-trained features have deviations when representing the original video, making it difficult to fully reflect the video content. In particular, when using graph convolutional networks to model temporal contextual relationships, there are problems with single modeling or insufficient dynamic updates, making it difficult to capture complex spatiotemporal dependencies, further affecting the accuracy of anomaly detection. Summary of the invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a method and system for clustering video event features based on text description. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0005] A video event feature clustering method based on text description, comprising:

[0006] S1. Obtain a training data set, and respectively extract the spatiotemporal features and object appearance features of a video data set from the training data set through two pre-training models, and obtain pre-trained video features after fusing the spatiotemporal features with the object appearance features, wherein the training data set includes normal videos and abnormal videos, and the abnormal videos include at least one abnormal segment;

[0007] S2. For a batch of pre-trained video features, anomaly prediction scores are obtained through an adaptive graph convolutional network, and a cross-batch clustering strategy is applied to the intermediate layer features to obtain the cluster centers of normal videos and abnormal videos;

[0008] S3, substituting the normal cluster center, the abnormal cluster center and the abnormal prediction score into a preset loss function so that the training is stopped when the loss value reaches a preset threshold to obtain a trained weakly supervised video anomaly detection model;

[0009] S4: input the video data to be detected into a trained weakly supervised video anomaly detection model to perform detection and obtain an anomaly detection result.

[0010] In a specific embodiment, the step S1 includes:

[0011] S11, using the Inflated 3D encoder as a feature extraction network to extract the spatiotemporal features of the training video;

[0012] S12, using the CLIP encoder as a feature extraction network to extract the appearance features of the object to be detected in the training video;

[0013] S13, the video frame judged as abnormal is used to generate a block-level feature map x through a sliding window scheme P , the window size is P×P, and the step size is s. The block-level feature map is input into the image encoder of CLIP to obtain the block-level feature representation, and the normal text description and the abnormal text description are encoded through the text encoder of CLIP to obtain the corresponding text embedding, and the block-level feature map x is calculated. P Similarity with the text query, block-level retrieval to obtain semantic features, for each block x P [i,j], calculate its similarity score with the abnormal text description:

[0014]

[0015] Among them, q T [r] is the anomaly text description and τ is the temperature parameter.

[0016] S14, integrating the spatiotemporal features, object appearance features and semantic features to obtain pre-trained video features that include temporal relationships, object co-occurrence relationships and semantic features.

[0017] In a specific embodiment, the step S2 includes:

[0018] S21. Adaptive graph convolution network simultaneously models the global dynamic association and temporal adjacency between video clips to obtain an adaptive global graph, which is used as the input of the graph convolution network together with the fusion features. The graph convolution network contains three layers of graph convolution layers. Except for the last layer, each layer is followed by a ReLU activation function and a dropout function. The last layer is followed by a Sigmoid activation function. For each training video V i , the input is the pre-trained feature X extracted by the feature extraction module i The adjacency matrix with the global graph is output as the anomaly score vector of the video clip in, is the i-th training video V i The anomaly prediction score of It is a training video V i The anomaly score of the jth segment in ;

[0019] S22. The output of the first-layer graph convolutional network is obtained as the intermediate feature representation of the video. The standardized intermediate normal and abnormal feature representations are clustered into two categories through the K-means clustering algorithm; for abnormal videos, the two types of features obtained by clustering represent normal events and abnormal events respectively, and the centers of the two clusters are pushed away by the loss based on batch clustering; for normal videos in a batch, the two types of features obtained by clustering both represent normal events, and the centers of the two clusters are brought closer by the loss based on batch clustering;

[0020] S23, add the cluster centers obtained by batch clustering all abnormal / normal video clips to C a and C n In the nth batch training process, C a , C n The cluster centers stored in are clustered by binary values, and we get as well as in, Represents the result of binary clustering of the abnormal cluster memory library, which serves as the two initial clustering centers of abnormal feature clustering during the nth batch training process; It represents the result of binary clustering of the normal cluster memory library, which is used as the two initial cluster centers of normal feature clustering in the nth batch training process. After the training of the current batch is completed, the cluster centers obtained after training are added to the cluster memory library, that is, m is the number of iterations, and then, in the training process of n+1 batches, the above operation is repeated through C a , C n The cluster centers stored in the binary clustering are used as the initial training centers for a new training.

[0021] In a specific implementation, the preset loss function includes a feature contrast loss function and a center contrast loss function, and the step S3 includes:

[0022] S31, selecting the segments with the highest and lowest abnormal prediction scores and the same number, and determining them as candidate normal event features, candidate abnormal event features, and candidate background features, respectively;

[0023] S32, respectively calculating the feature contrast loss function and the center contrast loss function to obtain a preset loss function;

[0024] The candidate normal event features and the cluster centers of the corresponding categories are taken as feature positive sample pairs, and the candidate abnormal event features and the cluster centers of the opposite categories are taken as feature negative sample pairs to obtain a feature contrast loss function, wherein the feature positive sample pairs include: the candidate normal event features and the normal cluster centers, the candidate background features and the normal cluster centers, the candidate abnormal event features and the abnormal cluster centers, and the feature negative sample pairs include: the candidate normal event features and the abnormal cluster centers, the candidate background features and the abnormal cluster centers, and the candidate abnormal event features and the normal cluster centers;

[0025] S33. The cluster center and the candidate features of the corresponding category are taken as the central positive sample pair, and the cluster center and the candidate features of the opposite category are taken as the central negative sample to obtain the central contrast loss function, wherein the central positive sample pair includes: the normal cluster center and the candidate normal event features, the normal cluster center and the candidate background features, the abnormal cluster center and the candidate abnormal event features, and the central negative sample pair includes: the normal cluster center and the candidate abnormal event features, the abnormal cluster center and the candidate normal event features, and the abnormal cluster center and the candidate background features.

[0026] In a specific implementation, the preset loss function is: in, represents the feature contrast loss function, represents the feature contrast loss function.

[0027] In a specific implementation, the feature contrast loss function is:

[0028]

[0029] Among them, X v represents candidate normal event features, abnormal event features or background features, T represents transposition, Represents and X v Reliable features of the opposite category of candidate features, C v represents the normal cluster center or the abnormal cluster center, s (·,·) represents the cosine similarity function, which is used to calculate the similarity between two vectors; τ is the temperature parameter of the contrast loss, which represents the degree of distinction between the desired features.

[0030] In a specific implementation, the center contrast loss function is:

[0031]

[0032] in, Representative and cluster center C v Candidate features of the opposite category, s(·,·) represents the cosine similarity function, which calculates the similarity between two vectors; τ is the temperature parameter of the contrast loss, which represents the degree of distinction between features.

[0033] The present invention also discloses a video event feature clustering system based on text description, comprising:

[0034] A training set acquisition module is used to acquire a training data set, and respectively extract the spatiotemporal features and object appearance features of a video data set from the training data set through two pre-training models, and obtain pre-trained video features after fusing the spatiotemporal features with the object appearance features, wherein the training data set includes normal videos and abnormal videos, and the abnormal videos include at least one abnormal segment;

[0035] The cluster center calculation module is used to obtain the abnormal prediction score of a batch of pre-trained video features through an adaptive graph convolutional network, and apply the cross-batch clustering strategy to the intermediate layer features to obtain the cluster centers of normal videos and abnormal videos;

[0036] A training module, used for substituting the normal cluster center, the abnormal cluster center and the abnormal prediction score into a preset loss function so as to stop the training when the loss value reaches a preset threshold value to obtain a trained weakly supervised video anomaly detection model;

[0037] The detection module is used to input the video data to be detected into a trained weakly supervised video anomaly detection model to perform detection and obtain an anomaly detection result.

[0038] Beneficial effects of the present invention:

[0039] The video event feature clustering method based on text description of the present invention realizes more accurate normal and abnormal classification guidance through cluster centers, and through the contrast loss mechanism, shortens the distance between the cluster centers of the same type and the video features, while keeping away from different types of features, significantly improving the ability to distinguish features, enhancing intra-class consistency and inter-class differences. In addition, by combining three complementary pre-training features, the spatiotemporal information and object co-occurrence information are more perfectly integrated, reducing the deviation of the original video representation in the pre-training stage.

[0040] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flow chart of a method for clustering video event features based on text description provided by an embodiment of the present invention;

[0042] Figure 2 The present invention provides a module block diagram of a video event feature clustering system based on text description. DETAILED DESCRIPTION

[0043] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0044] Embodiment 1

[0045] See also Figure 1 , Figure 1 The present invention provides a flow chart of a method for clustering video event features based on text description, including:

[0046] S1. Obtain a training data set, and respectively extract the spatiotemporal features and object appearance features of a video data set from the training data set through two pre-training models, and obtain pre-trained video features after fusing the spatiotemporal features with the object appearance features, wherein the training data set includes normal videos and abnormal videos, and the abnormal videos include at least one abnormal segment;

[0047] Specifically, the training videos input into the network are respectively subjected to two pre-training models to extract the spatiotemporal features and object appearance features of the video dataset and fuse them to obtain more complete and universal pre-training features. The original video contains N normal videos. With N abnormal videos And its corresponding video-level label Y n =0,Y a =1, the abnormal video consists of abnormal video segments and normal video segments, each abnormal video contains at least one abnormal video segment, and the normal video consists of normal video segments and does not contain abnormal video segments.

[0048] S2. For a batch of pre-trained video features, anomaly prediction scores are obtained through an adaptive graph convolutional network, and a cross-batch clustering strategy is applied to the intermediate layer features to obtain the cluster centers of normal videos and abnormal videos;

[0049] S3, substituting the normal cluster center, the abnormal cluster center and the abnormal prediction score into a preset loss function so that the training is stopped when the loss value reaches a preset threshold to obtain a trained weakly supervised video anomaly detection model;

[0050] S4: input the video data to be detected into a trained weakly supervised video anomaly detection model to perform detection and obtain an anomaly detection result.

[0051] In a specific embodiment, the step S1 includes:

[0052] S11, using the Inflated 3D encoder as a feature extraction network to extract the spatiotemporal features of the training video;

[0053] Specifically, the Inflated 3D (I3D) model pre-trained on the large-scale video dataset Kinetics is selected as the feature extractor to extract the spatiotemporal features of the video. i The spatiotemporal characteristics of It can be expressed as:

[0054]

[0055] in, It is a training video V i The spatiotemporal features of the jth video clip, The dimension is Wei, T i It is a training video V i The number of video clips included, The dimension representing the spatiotemporal characteristics.

[0056] S12, using the CLIP encoder as a feature extraction network to extract the appearance features of the object to be detected in the training video;

[0057] Specifically, the pre-trained CLIP (ViT-B / 16) model is selected as the feature extraction network to extract the object co-occurrence relationship and appearance features of the video. i Appearance features of X i c It can be expressed as:

[0058]

[0059] in, It is a training video V i The spatiotemporal features of the jth video clip, The dimension is Wei, T i It is a training video V i The number of video clips included, The dimension representing the spatiotemporal characteristics.

[0060] S13, the video frame judged as abnormal is used to generate a block-level feature map x through a sliding window scheme P , the window size is P×P, and the step size is s. These image blocks are input into the image encoder of CLIP to obtain block-level feature representation, and the normal text description and abnormal text description are encoded through the text encoder of CLIP to obtain the corresponding text embedding, and the block-level feature map x is calculated. P Similarity with the text query, block-level retrieval to obtain semantic features, for each block x P [i,j], calculate its similarity score with the abnormal text description:

[0061]

[0062] Among them, q T [r] is the anomaly text description and τ is the temperature parameter.

[0063] S14, integrating the spatiotemporal features, object appearance features and semantic features to obtain pre-trained video features that include temporal relationships, object co-occurrence relationships and semantic features.

[0064] Specifically, the spatiotemporal features, object appearance features, and semantic features are integrated to obtain video features that contain both temporal relationships and object co-occurrence relationships, and incorporate the semantic information brought by text descriptions. The final video V i Features of X i It can be expressed as:

[0065]

[0066] Among them, X i represents the training video features, X i The dimension is T i ×D i , T i It is a training video V i The number of video clips included, D i Represents the dimension of the fused features

[0067] In a specific embodiment, the step S2 includes:

[0068] S21. Adaptive graph convolution network simultaneously models the global dynamic association and temporal adjacency between video clips to obtain an adaptive global graph, which is used as the input of the graph convolution network together with the fusion features. The graph convolution network contains three layers of graph convolution layers. Except for the last layer, each layer is followed by a ReLU activation function and a dropout function. The last layer is followed by a Sigmoid activation function. For each training video V i , the input is the pre-trained feature X extracted by the feature extraction module i The adjacency matrix with the global graph is output as the anomaly score vector of the video clip in, is the i-th training video V i The anomaly prediction score of It is a training video V i The anomaly score of the jth segment in ;

[0069] S22. The output of the first-layer graph convolutional network is obtained as the intermediate feature representation of the video. The standardized intermediate normal and abnormal feature representations are clustered into two categories through the K-means clustering algorithm; for abnormal videos, the two types of features obtained by clustering represent normal events and abnormal events respectively, and the centers of the two clusters are pushed away by the loss based on batch clustering; for normal videos in a batch, the two types of features obtained by clustering both represent normal events, and the centers of the two clusters are brought closer by the loss based on batch clustering;

[0070] Specifically, a batch clustering module is used to cluster normal video features and abnormal video features respectively, and normal cluster centers and abnormal cluster centers are obtained respectively. The cluster centers provide supervision to enhance the discriminative ability of features.

[0071] The output of the first-layer graph convolutional network is obtained as the intermediate feature representation of the video. Through the K-means clustering algorithm, the standardized intermediate normal and abnormal feature representations are clustered into two categories respectively. For abnormal videos, the two types of features obtained by clustering represent normal events and abnormal events respectively, and the centers of the two clusters are pushed away by the loss based on batch clustering; for normal videos in a batch, the two types of features obtained by clustering both represent normal events, and the centers of the two clusters are brought closer by the loss based on batch clustering. In the K-means clustering process, the normal cluster centers are close to each other, and the abnormal cluster centers are pushed away from each other, thereby enhancing the discriminative power of the features. The calculation formula is:

[0072]

[0073] Where d = ∥ c 1 -c 2 ∥ 2 is the distance between two cluster centers, c 1 、c 2 are all normal videos or abnormal videos in the batch (i.e. ) The two cluster centers obtained after clustering, u is an upper bound that helps the model to be robust to different video data, and b is the batch size.

[0074] S23, add the cluster centers obtained by batch clustering all abnormal / normal video clips to C a and C n In the nth batch training process, C a , C n The cluster centers stored in are clustered by binary values, and we get as well as in, Represents the result of binary clustering of the abnormal cluster memory library, which serves as the two initial clustering centers of abnormal feature clustering during the nth batch training process; It represents the result of binary clustering of the normal cluster memory library, which is used as the two initial cluster centers of normal feature clustering in the nth batch training process. After the training of the current batch is completed, the cluster centers obtained after training are added to the cluster memory library, that is, m is the number of iterations, and then, in the training process of n+1 batches, the above operation is repeated through C a , C n The cluster centers stored in the binary clustering are used as the initial training centers for a new training.

[0075] Specifically, the effect of clustering algorithms is usually affected by the selection of initial cluster centers. When performing batch clustering based on the K-means algorithm, two clips are randomly selected from all video clips as initial cluster centers, which can easily lead to significantly different clustering results for different initial center selections. In order to better select appropriate initial cluster centers, a cross-batch learning module is used to introduce the clustering results of the previous batch to provide guidance for the current batch clustering, so that the clustering algorithm converges faster and obtains more accurate cluster centers and clustering results.

[0076] Due to the limitation of GPU memory and the need to ensure accurate guidance of cluster centers, and the high dependence of batch clustering algorithm on cluster centers, a cross-batch clustering strategy is used to construct abnormal cluster memory libraries C in the iteration process. a and the normal cluster memory C n ,Store the learned knowledge and use the stored information to provide guidance for the current batch clustering to ensure the stability and accuracy of clustering.

[0077] In a specific implementation, the preset loss function includes a feature contrast loss function and a center contrast loss function, and the step S3 includes:

[0078] S31, selecting the segments with the highest and lowest abnormal prediction scores and the same number, and determining them as candidate normal event features, candidate abnormal event features, and candidate background features, respectively;

[0079] Specifically, the k clips with the highest scores in normal videos and abnormal videos are selected as candidate normal event features and abnormal event features. At the same time, the k clips with the lowest scores in normal videos and abnormal videos are selected as candidate background features. The normal cluster center and the abnormal cluster center represent reliable normal features and abnormal features, respectively.

[0080] S32, respectively calculating the feature contrast loss function and the center contrast loss function to obtain a preset loss function;

[0081] The candidate normal event features and the cluster centers of the corresponding categories are taken as feature positive sample pairs, and the candidate abnormal event features and the cluster centers of the opposite categories are taken as feature negative sample pairs to obtain a feature contrast loss function, wherein the feature positive sample pairs include: the candidate normal event features and the normal cluster centers, the candidate background features and the normal cluster centers, the candidate abnormal event features and the abnormal cluster centers, and the feature negative sample pairs include: the candidate normal event features and the abnormal cluster centers, the candidate background features and the abnormal cluster centers, and the candidate abnormal event features and the normal cluster centers;

[0082] S33. The cluster center and the candidate features of the corresponding category are taken as the central positive sample pair, and the cluster center and the candidate features of the opposite category are taken as the central negative sample to obtain the central contrast loss function, wherein the central positive sample pair includes: the normal cluster center and the candidate normal event features, the normal cluster center and the candidate background features, the abnormal cluster center and the candidate abnormal event features, and the central negative sample pair includes: the normal cluster center and the candidate abnormal event features, the abnormal cluster center and the candidate normal event features, and the abnormal cluster center and the candidate background features.

[0083] In a specific implementation, the preset loss function is: in, represents the feature contrast loss function, represents the feature contrast loss function.

[0084] In a specific implementation, the feature contrast loss function is:

[0085]

[0086] Among them, X v represents candidate normal event features, abnormal event features or background features, T represents transposition, Represents and X v Reliable features of the opposite category of candidate features, C v represents the normal cluster center or the abnormal cluster center, s (·,·) represents the cosine similarity function, which is used to calculate the similarity between two vectors; τ is the temperature parameter of the contrast loss, which represents the degree of distinction between the desired features.

[0087] In a specific implementation, the center contrast loss function is:

[0088]

[0089] in, Representative and cluster center C v Candidate features of the opposite category, s(·,·) represents the cosine similarity function, which calculates the similarity between two vectors; τ is the temperature parameter of the contrast loss, which represents the degree of distinction between features.

[0090] The video event feature clustering method based on text description of the present invention realizes more accurate normal and abnormal classification guidance through cluster centers, and through the contrast loss mechanism, shortens the distance between the cluster centers of the same type and the video features, while keeping away from different types of features, significantly improving the ability to distinguish features, enhancing intra-class consistency and inter-class differences. In addition, by combining three complementary pre-training features, the spatiotemporal information and object co-occurrence information are more perfectly integrated, reducing the deviation of the original video representation in the pre-training stage.

[0091] Please continue to see Figure 2 , Figure 2 The module block diagram of a video event feature clustering system based on text description provided by an embodiment of the present invention includes:

[0092] A training set acquisition module is used to acquire a training data set, and respectively extract the spatiotemporal features and object appearance features of a video data set from the training data set through two pre-training models, and obtain pre-trained video features after fusing the spatiotemporal features with the object appearance features, wherein the training data set includes normal videos and abnormal videos, and the abnormal videos include at least one abnormal segment;

[0093] The cluster center calculation module is used to obtain the abnormal prediction score of a batch of pre-trained video features through an adaptive graph convolutional network, and apply the cross-batch clustering strategy to the intermediate layer features to obtain the cluster centers of normal videos and abnormal videos;

[0094] A training module, used for substituting the normal cluster center, the abnormal cluster center and the abnormal prediction score into a preset loss function so as to stop the training when the loss value reaches a preset threshold value to obtain a trained weakly supervised video anomaly detection model;

[0095] The detection module is used to input the video data to be detected into a trained weakly supervised video anomaly detection model to perform detection and obtain an anomaly detection result.

[0096] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification.

[0097] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "one" or "an" does not exclude multiple situations. A single processor or other unit may implement several functions listed in a claim. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0098] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems (equipment), or computer program products. Therefore, the present application may adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware, which are collectively referred to as "modules" or "systems" herein. Moreover, the present application may adopt the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program codes. The computer program is stored / distributed in a suitable medium, provided together with other hardware or as a part of hardware, or may adopt other distribution forms, such as by Internet or other wired or wireless telecommunication systems.

[0099] The present application is described with reference to the flowcharts and / or block diagrams of the methods, systems (devices) and computer program products of the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A system that specifies the functions of a box or multiple boxes.

[0100] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction system, which is implemented in the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0102] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A video event feature clustering method based on text description, characterized in that: include: S1. Obtain a training data set, and respectively extract the spatiotemporal features and object appearance features of a video data set from the training data set through two pre-training models, and obtain pre-trained video features after fusing the spatiotemporal features with the object appearance features, wherein the training data set includes normal videos and abnormal videos, and the abnormal videos include at least one abnormal segment; S2. For a batch of pre-trained video features, anomaly prediction scores are obtained through an adaptive graph convolutional network, and a cross-batch clustering strategy is applied to the intermediate layer features to obtain the cluster centers of normal videos and abnormal videos; S3, substituting the normal cluster center, the abnormal cluster center and the abnormal prediction score into a preset loss function so that the training is stopped when the loss value reaches a preset threshold to obtain a trained weakly supervised video anomaly detection model; S4: input the video data to be detected into a trained weakly supervised video anomaly detection model to perform detection and obtain an anomaly detection result.

2. The video event feature clustering method based on text description according to claim 1 is characterized in that: The step S1 comprises: S11, using the Inflated 3D encoder as a feature extraction network to extract the spatiotemporal features of the training video; S12, using the CLIP encoder as a feature extraction network to extract the appearance features of the object to be detected in the training video; S13, the video frame judged as abnormal is used to generate a block-level feature map x through a sliding window scheme P , the window size is P×P, the step size is s, the block-level feature map is input into the image encoder of CLIP, the block-level feature representation is obtained, the normal text description and the abnormal text description are encoded through the text encoder of CLIP, the corresponding text embedding is obtained, and the block-level feature map x is calculated. P Similarity with the text query, block-level retrieval to obtain semantic features, for each block x P [i,j], calculate its similarity score with the abnormal text description: Among them, q T [r] is the anomaly text description and τ is the temperature parameter. S14, integrating the spatiotemporal features, object appearance features and semantic features to obtain pre-trained video features that include temporal relationships, object co-occurrence relationships and semantic features.

3. The video event feature clustering method based on text description according to claim 2 is characterized in that: The step S2 comprises: S21. Adaptive graph convolution network simultaneously models the global dynamic association and temporal adjacency between video clips to obtain an adaptive global graph, which is used as the input of the graph convolution network together with the fusion features. The graph convolution network contains three layers of graph convolution layers. Except for the last layer, each layer is followed by a ReLU activation function and a dropout function. The last layer is followed by a Sigmoid activation function. For each training video V i , the input is the pre-trained feature X extracted by the feature extraction module i The adjacency matrix with the global graph is output as the anomaly score vector of the video clip in, is the i-th training video V i The anomaly prediction score of It is a training video V i The anomaly score of the jth segment in ; S22. The output of the first-layer graph convolutional network is obtained as the intermediate feature representation of the video. The standardized intermediate normal and abnormal feature representations are clustered into two categories through the K-means clustering algorithm; for abnormal videos, the two types of features obtained by clustering represent normal events and abnormal events respectively, and the centers of the two clusters are pushed away by the loss based on batch clustering; for normal videos in a batch, the two types of features obtained by clustering both represent normal events, and the centers of the two clusters are brought closer by the loss based on batch clustering; S23, add the cluster centers obtained by batch clustering all abnormal / normal video clips to C a and C n In the nth batch training process, C a , C n The cluster centers stored in are clustered by binary values, and we get as well as in, Represents the result of binary clustering of the abnormal cluster memory library, which serves as the two initial clustering centers of abnormal feature clustering during the nth batch training process; It represents the result of binary clustering of the normal cluster memory library, which is used as the two initial cluster centers of normal feature clustering in the nth batch training process. After the training of the current batch is completed, the cluster centers obtained after training are added to the cluster memory library, that is, m is the number of iterations, and then, in the training process of n+1 batches, the above operation is repeated through C a , C n The cluster centers stored in the binary clustering are used as the initial training centers for a new training.

4. The video event feature clustering method based on text description according to claim 2 is characterized in that: The preset loss function includes a feature contrast loss function and a center contrast loss function, and the step S3 includes: S31, selecting the segments with the highest and lowest abnormal prediction scores and the same number, and determining them as candidate normal event features, candidate abnormal event features, and candidate background features, respectively; S32, respectively calculating the feature contrast loss function and the center contrast loss function to obtain a preset loss function; The candidate normal event features and the cluster centers of the corresponding categories are taken as feature positive sample pairs, and the candidate abnormal event features and the cluster centers of the opposite categories are taken as feature negative sample pairs to obtain a feature contrast loss function, wherein the feature positive sample pairs include: the candidate normal event features and the normal cluster centers, the candidate background features and the normal cluster centers, the candidate abnormal event features and the abnormal cluster centers, and the feature negative sample pairs include: the candidate normal event features and the abnormal cluster centers, the candidate background features and the abnormal cluster centers, and the candidate abnormal event features and the normal cluster centers; S33. The cluster center and the candidate features of the corresponding category are taken as the central positive sample pair, and the cluster center and the candidate features of the opposite category are taken as the central negative sample to obtain the central contrast loss function, wherein the central positive sample pair includes: the normal cluster center and the candidate normal event features, the normal cluster center and the candidate background features, the abnormal cluster center and the candidate abnormal event features, and the central negative sample pair includes: the normal cluster center and the candidate abnormal event features, the abnormal cluster center and the candidate normal event features, and the abnormal cluster center and the candidate background features.

5. The video event feature clustering method based on text description according to claim 4 is characterized in that: The preset loss function is: ,in, represents the feature contrast loss function, represents the feature contrast loss function.

6. The video event feature clustering method based on text description according to claim 5, characterized in that: The feature contrast loss function is: Among them, X v represents candidate normal event features, abnormal event features or background features, T represents transposition, Represents and X v Reliable features of the opposite category of candidate features, C v represents the normal cluster center or the abnormal cluster center, s (·,·) represents the cosine similarity function, which is used to calculate the similarity between two vectors; τ is the temperature parameter of the contrast loss, which represents the degree of distinction between the desired features.

7. The video event feature clustering method based on text description according to claim 5, characterized in that: The center contrast loss function is: in, Representative and cluster center C v Candidate features of the opposite category, s (·,·) represents the cosine similarity function, which calculates the similarity between two vectors; τ is the temperature parameter of the contrast loss, which represents the degree of distinction between features.

8. A video event feature clustering system based on text description, characterized in that: include: A training set acquisition module is used to acquire a training data set, and respectively extract the spatiotemporal features and object appearance features of a video data set from the training data set through two pre-training models, and obtain pre-trained video features after fusing the spatiotemporal features with the object appearance features, wherein the training data set includes normal videos and abnormal videos, and the abnormal videos include at least one abnormal segment; The cluster center calculation module is used to obtain the abnormal prediction score of a batch of pre-trained video features through an adaptive graph convolutional network, and apply the cross-batch clustering strategy to the intermediate layer features to obtain the cluster centers of normal videos and abnormal videos; A training module, used for substituting the normal cluster center, the abnormal cluster center and the abnormal prediction score into a preset loss function so as to stop the training when the loss value reaches a preset threshold value to obtain a trained weakly supervised video anomaly detection model; The detection module is used to input the video data to be detected into a trained weakly supervised video anomaly detection model to perform detection and obtain an anomaly detection result.