Prototype pseudo instance enhanced video anomaly detection method and system
By introducing learnable normal prototype sets and attention mechanisms into the video anomaly detection method, the feature context vectors are dynamically generated, and combined with the general supervision comparison loss and multi-example learning loss, the problems of high model complexity and limited generalization ability in the existing technology are solved, significantly improving the performance and positioning accuracy of video anomaly detection.
Patent Information
- Application Number
- CN202510580074.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
When handling abnormal videos, existing video anomaly detection methods are difficult to learn fine-grained abnormal features, and lack explicit modeling of normal modes, resulting in high model complexity, limited generalization capabilities, and insufficient optimization of boundary instances, which affects positioning accuracy.
A prototype pseudo-instance enhanced video anomaly detection method is proposed. The feature context vectors are dynamically generated through the learnable normality prototype set and attention mechanism, a robust normality benchmark is established, and the model is trained through the general supervision comparison loss and multi-example learning loss to enhance the abnormality detection ability.
Through the combined use of dynamically generated feature context vectors and general supervision comparison losses, the model's ability to perceive normal modes and abnormal deviation detection is significantly improved, and the performance and positioning accuracy of video anomaly detection are improved.
Smart Images

Figure CN120088712A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly relates to a method and system for prototype pseudo-instance enhanced video anomaly detection. Background Art
[0002] As a core branch field of artificial intelligence, computer vision is dedicated to realizing the perception, analysis, and understanding of image and video data through algorithms. Its technological development has experienced a transformation from traditional image processing to deep learning-driven intelligence. In the early days, computer vision mainly relied on manually designed features (such as SIFT, HOG) and statistical learning methods to complete basic tasks such as target recognition and image classification through means such as edge detection and feature matching.
[0003] Existing methods such as RTFM based on MIL and CLAWS based on contrastive learning, although they improve performance through ranking loss or clustering, do not solve the problems of normal prototype structured embedding and extreme instance contrast enhancement, resulting in high model complexity and limited generalization ability. Specifically, in abnormal videos, normal segments dominate, making it difficult for the model to learn fine-grained abnormal features; there is a lack of explicit modeling of normal patterns, leading to fuzzy anomaly detection benchmarks; and boundary instances (such as the highest / lowest scoring segments) lack targeted optimization, affecting the positioning accuracy. Summary of the Invention
[0004] In view of the above situation, the main purpose of the present invention is to propose a method and system for prototype pseudo-instance enhanced video anomaly detection to solve the above technical problems.
[0005] The present invention proposes a method for prototype pseudo-instance enhanced video anomaly detection, and the method includes the following steps: Step 1: Input the video into a video anomaly detection model, perform temporal segmentation and feature extraction on the video in sequence to obtain deep features, and obtain a feature sequence based on the deep features; Step 2: Input the feature sequence into a prototype interaction layer, and calculate the cosine similarity through a set of normal prototypes; Generate attention weights based on the cosine similarity; Obtain a normal context vector through the attention weights; Perform a linear transformation on the normal context vector and perform a residual connection with the deep features to obtain enhanced features; Step 3: Based on the enhanced features, calculate an anomaly score through a classifier; Screen the video through the anomaly score to obtain an abnormal set and a normal set respectively, and merge the abnormal set and the normal set to obtain an index set of extreme instances; Step 4: Based on the index set of extreme instances, perform low-dimensional mapping on the enhanced features corresponding to the extreme instances to obtain a deep feature embedding representation; Construct a total supervision contrast loss based on the deep feature embedding representation; Train the video anomaly detection model by combining the total supervision contrast loss and the multi-instance learning loss to obtain the trained video anomaly detection model; Step 5: Input the video into the trained video anomaly detection model, and output the video-level anomaly classification result and the instance-level anomaly localization score.
[0006] The present invention also proposes a prototype pseudo-instance enhanced video anomaly detection system, and the system includes: A feature extraction module, which is used for: Input the video into the video anomaly detection model, perform temporal segmentation and feature extraction on the video in sequence to obtain deep features, and obtain a feature sequence based on the deep features; An enhanced feature module, which is used for: Input the feature sequence into the prototype interaction layer, and calculate the cosine similarity through the normal prototype set; Generate attention weights based on the cosine similarity; Obtain the normal context vector through the attention weights; Perform a linear transformation on the normal context vector and then perform a residual connection with the deep features to obtain enhanced features; A screening module, which is used for: Based on the enhanced features, calculate the anomaly score through a classifier; Screen the video through the anomaly score, respectively obtain the anomaly set and the normal set, and combine the anomaly set and the normal set to obtain the index set of extreme instances; A training module, which is used for: Based on the index set of extreme instances, perform low-dimensional mapping on the enhanced features corresponding to the extreme instances to obtain the deep feature embedding representation; Construct a total supervision contrast loss based on the deep feature embedding representation; Train the video anomaly detection model by combining the total supervision contrast loss and the multi-instance learning loss to obtain the trained video anomaly detection model; A result output module, which is used for: Input the video into the trained video anomaly detection model, and output the video-level anomaly classification result and the instance-level anomaly localization score.
[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention dynamically generates a feature context vector through a learnable normal prototype set and an attention mechanism, embeds the structured normal knowledge into the feature stay, and establishes a robust normal benchmark; 2. Select pseudo-positive / negative samples with extreme scores based on the model prediction scores, and force the separation of difficult instance features through the targeted contrast loss to sharpen the decision boundary. The two modules work together to enhance the normal mode perception and abnormal deviation detection capabilities simultaneously during the feature learning stage.
[0008] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is a flowchart of a prototype pseudo-instance enhanced video anomaly detection method proposed by the present invention; Figure 2 is a framework diagram of a prototype pseudo-instance enhanced video anomaly detection method proposed by the present invention; Figure 3 is a schematic diagram of the overall framework of a prototype pseudo-instance enhanced video anomaly detection system proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0010] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0011] These and other aspects of the embodiments of the present invention will be clear from the following description and the accompanying drawings. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0012] Please refer to Figure 1 and Figure 2 , an embodiment of the present invention proposes a prototype pseudo-instance enhanced video anomaly detection method, which includes the following steps: Step 1: Input the video into the video anomaly detection model, perform temporal segmentation and feature extraction on the video in sequence to obtain deep features, and obtain a feature sequence based on the deep features.
[0013] Specifically, in this step, after inputting the video into the video anomaly detection model, the video is evenly divided into a certain number of non-overlapping segments, each segment contains 16 frames, and the time span is 1 second; Perform 10-crop spatial enhancement on each segment: Crop 5 224×224 pixel regions from the four corners and the central region of the video frame, and perform horizontal flipping to generate a total of 10 enhanced samples.
[0014] Step 2: Input the feature sequence into the prototype interaction layer, and calculate the cosine similarity through the normal prototype set; Generate the attention weights based on the cosine similarity; Obtain the normal context vector through the attention weights; Perform a linear transformation on the normal context vector and perform a residual connection with the depth feature to obtain the enhanced feature; In Step 2, when the feature sequence is input into the prototype interaction layer and the cosine similarity is calculated through the normal prototype set, the relational expression existing in the corresponding process is: ; where, represents the cosine similarity, represents the th depth feature, represents the feature of the th normal prototype, represents the L2 norm, represents the transpose; The attention weights are generated based on the cosine similarity, and the relational expression existing in the corresponding process is: ; where, represents the attention weight of the th depth feature to the feature of the th normal prototype, represents the temperature coefficient, represents the total number of features of the normal prototype; The normal context vector is obtained through the attention weights, and the relational expression existing in the corresponding process is: ; where, represents the normal context vector; After performing a linear transformation on the normal context vector and performing a residual connection with the depth feature, the enhanced feature is obtained, and the relational expression existing in the corresponding process is: ; where, represents the enhanced feature; represents the learnable weight matrix, and ; represents the learnable bias vector, and .
[0015] Specifically, in this step, the normal prototype set is composed of five normal prototype vectors, and the dimension of each normal prototype vector is 512 and follows a normal distribution (0, 0.02).
[0016] Step 3: Based on the enhanced features, the anomaly score is calculated by the classifier; The videos are filtered by anomaly scores to obtain anomaly sets and normal sets respectively. The anomaly sets and normal sets are merged to obtain an index set of extreme instances. In step 3, based on the enhanced features, the anomaly score is calculated by the classifier, and the relationship between the corresponding process is: ; in, represents the anomaly score, Indicates that it is processed by the activation function. Indicates processing through two layers of multi-layer perceptron classifiers.
[0017] Specifically, in this step, extreme instances are selected based on the anomaly score: Select a certain number of indexes of the highest-scoring instances to form an anomaly set; Select a certain number of indexes of the lowest-scoring instances to form a normal set; The abnormal set and the normal set are aggregated into an index set of extreme instances.
[0018] Step 4: Based on the index set of extreme instances, the enhanced features corresponding to the extreme instances are low-dimensionally mapped to obtain a deep feature embedding representation; Constructing a total supervised contrastive loss based on deep feature embedding representation; The video anomaly detection model is trained by combining the total supervision contrast loss and the multi-instance learning loss to obtain a trained video anomaly detection model; In step 4, based on the index set of extreme instances, the enhanced features corresponding to the extreme instances are low-dimensionally mapped to obtain a deep feature embedding representation. The relationship between the corresponding process is: ; in, Indicates The embedded representation of deep features, Indicates that it is processed by the projection head. represents the learnable projection weight matrix, represents the learnable projection bias vector; Constructing the total supervised contrast loss based on the deep feature embedding representation, the specific steps are as follows: The index set of other instances is obtained through deep feature embedding representation. The relationship between the corresponding process is: ; in, Represents the index set of other instances, Indicates the index of the anchor instance in the set of extreme instances, Indicates the indices of other instances in the set of extreme instances except the current anchor instance, Indicates the index of the current anchor instance, Indicates the operation of extracting the pseudo - label of the instance; Obtain the dynamic weights of negative samples through deep feature embedding representation. The relational formula for the corresponding process is: ; Among them, Indicates the dynamic weight of the negative sample, Indicates the th negative sample, Indicates the th negative sample, Indicates the temporary index of the negative sample, Indicates the weight temperature coefficient; Obtain the supervised contrast loss through the index set of other instances and the dynamic weights of negative samples. The relational formula for the process is: ; Among them, Indicates the supervised contrast loss of the th instance, Indicates the feature decoupling regularization, Indicates the dynamic temperature coefficient for adaptively adjusting the feature difference according to the sample pair, and both indicate the results of the structural decoupling of the feature, Indicates the projected feature of other extreme instances, Indicates the weight coefficient of the feature decoupling regularization term; Sum up all the supervised contrast losses to obtain the total supervised contrast loss. The relational formula for the corresponding process is: ; Among them, Indicates the total supervised contrast loss, Indicates the set of extreme instances; Train the video anomaly detection model by combining the total supervised contrast loss and the multi - instance learning loss to obtain the trained video anomaly detection model. Among them, the total loss is: ; Among them, Indicates the total loss, Indicates the uncertainty parameter related to the multi - instance learning task, Indicates the uncertainty parameter of the total supervised contrast task, Indicates the multi - instance learning loss, represents the base mixing coefficient, represents the regularization coefficient, represents the prototype regularization term.
[0019] Step 5: Input the video into the trained video anomaly detection model, and output the video-level anomaly classification result and the instance-level anomaly localization score.
[0020] To verify the effectiveness of the present invention, comprehensive tests are carried out on two major mainstream datasets, SH-AUC and UCF-ADC. SH-AUC contains 437 videos of 13 types of scenarios, covering abnormal events of general degree; UCF-ADC contains 1900 long videos, covering more complex abnormal types, ensuring the reliability of the evaluation of the model generalization ability; Compare with the prior art: Experiments are carried out on two commonly used weakly supervised video anomaly detection datasets, SH-AUC and UCF-ADC. Use AUC (Area Under the ROC Curve) as the evaluation index to measure the anomaly detection performance of the model. Compared with a variety of SOTA methods published in recent years, these methods use different features (C3D, I3D, ViT) and technologies. It can be seen from Table 1 and Table 2 that the present invention reaches or approaches the current state-of-the-art level in terms of performance, and has great advantages in terms of lightweight.
[0021] Table 1 Comparison with the prior art
[0022] Table 2 Lightweight comparison
[0023] To verify the effectiveness of the prototype interaction layer and the supervised contrast (PIDE) loss of the present invention (SND-VAD) and their synergistic effect, four model configurations are also compared: Baseline (only ViT features), Baseline + Protolnteract, Baseline + PIDE, and the complete SND-VAD; It can be seen from Table 3 and Table 4 that adding the prototype interaction layer or the PIDE loss alone can significantly improve the performance. The complete SND-VAD model achieves the best performance, indicating that there is a synergistic effect between these two components.
[0024] Table 3 Ablation comparison
[0025] Table 4 Number k of normal prototypes
[0026] Please refer toFigure 3 , an embodiment of the present invention further provides a prototype pseudo-instance enhanced video anomaly detection system, and the system includes: A feature extraction module, configured to: Input the video into a video anomaly detection model, perform temporal segmentation and feature extraction on the video in sequence to obtain deep features, and obtain a feature sequence based on the deep features; An enhanced feature module, configured to: Input the feature sequence into the prototype interaction layer, and calculate the cosine similarity through the normal prototype set; Generate attention weights based on the cosine similarity; Obtain a normal context vector through the attention weights; Perform a linear transformation on the normal context vector and perform a residual connection with the deep features to obtain enhanced features; A screening module, configured to: Based on the enhanced features, calculate an anomaly score through a classifier; Screen the video through the anomaly score, respectively obtain an anomaly set and a normal set, and merge the anomaly set and the normal set to obtain an index set of extreme instances; A training module, configured to: Based on the index set of extreme instances, perform a low-dimensional mapping on the enhanced features corresponding to the extreme instances to obtain a deep feature embedding representation; Construct a total supervision contrast loss based on the deep feature embedding representation; Train the video anomaly detection model jointly with the total supervision contrast loss and the multi-instance learning loss to obtain a trained video anomaly detection model; A result output module, configured to: Input the video into the trained video anomaly detection model, and output a video-level anomaly classification result and an instance-level anomaly localization score.
[0027] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiment, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0028] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0029] The above-described embodiments merely represent several implementation manners of the present invention. The descriptions thereof are relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A prototype pseudo-instance enhanced video anomaly detection method, characterized in that: The method comprises the following steps: Step 1: Input the video into the video anomaly detection model, perform time segmentation and feature extraction on the video in sequence to obtain deep features, and obtain a feature sequence based on the deep features; Step 2: Input the feature sequence into the prototype interaction layer and calculate the cosine similarity through the normal prototype set; Attention weights are generated based on cosine similarity; Get the normal context vector through the attention weight; After linear transformation of the normal context vector, residual connection is performed with the deep feature to obtain enhanced features; Step 3: Based on the enhanced features, the anomaly score is calculated by the classifier; The videos are filtered by anomaly scores to obtain anomaly sets and normal sets respectively. The anomaly sets and normal sets are merged to obtain an index set of extreme instances. Step 4: Based on the index set of extreme instances, the enhanced features corresponding to the extreme instances are low-dimensionally mapped to obtain a deep feature embedding representation; Constructing a total supervised contrastive loss based on deep feature embedding representation; The video anomaly detection model is trained by combining the total supervision contrast loss and the multi-instance learning loss to obtain a trained video anomaly detection model; Step 5: Input the video into the trained video anomaly detection model and output the video-level anomaly classification results and instance-level anomaly location scores.
2. The prototype pseudo-instance enhanced video anomaly detection method according to claim 1, characterized in that: In step 2, the feature sequence is input into the prototype interaction layer, and the cosine similarity is calculated by the normal prototype set. The relationship between the corresponding process is: ; in, represents the cosine similarity, Indicates deep features, Indicates The characteristics of a normal archetype, represents the L2 norm, Indicates transpose.
3. The prototype pseudo-instance enhanced video anomaly detection method according to claim 2, characterized in that: In step 2, the attention weight is generated based on the cosine similarity, and the relationship between the corresponding process is: ; in, Indicates The deep feature pair The feature attention weights of the normal prototypes, represents the temperature coefficient, The total number of features representing the normality prototype.
4. The prototype pseudo-instance enhanced video anomaly detection method according to claim 3, characterized in that: In step 2, the normal context vector is obtained by the attention weight, and the relationship between the corresponding process is: ; in, Represents the normality context vector.
5. The prototype pseudo-instance enhanced video anomaly detection method according to claim 4, characterized in that: In step 2, the normal context vector is linearly transformed and then residually connected with the deep feature to obtain the enhanced feature. The relationship between the corresponding process is: ; in, Represents enhanced features, represents the learnable weight matrix, represents the learnable bias vector.
6. The prototype pseudo-instance enhanced video anomaly detection method according to claim 5, characterized in that: In step 3, based on the enhanced features, the anomaly score is calculated by the classifier, and the relationship between the corresponding process is: ; in, represents the anomaly score, Indicates that it is processed by the activation function. Indicates processing through two layers of multi-layer perceptron classifiers.
7. The prototype pseudo-instance enhanced video anomaly detection method according to claim 6, characterized in that: In step 4, based on the index set of the extreme instance, the enhanced features corresponding to the extreme instance are low-dimensionally mapped to obtain a deep feature embedding representation. The relationship between the corresponding process is: ; in, Indicates The embedded representation of deep features, Indicates that it is processed by the projection head. represents the learnable projection weight matrix, represents the learnable projection bias vector.
8. The prototype pseudo-instance enhanced video anomaly detection method according to claim 7, characterized in that: In step 4, the total supervised contrast loss is constructed based on the deep feature embedding representation. The specific steps are as follows: The index set of other instances is obtained through deep feature embedding representation. The relationship between the corresponding process is: ; in, Represents the index set of other instances, represents the index of the anchor instance in the set of extreme instances, Represents the index of other instances in the extreme instance set except the current anchor instance, Represents the index of the current anchor instance, represents the pseudo-labeling operation of extracting instances; The dynamic weight of negative samples is obtained through deep feature embedding representation, and the relationship between the corresponding process is: ; in, represents the dynamic weight of negative samples, Indicates negative samples, Indicates negative samples, represents the temporary index of negative samples, represents the weight temperature coefficient; Through the index set of other instances and the dynamic weights of negative samples, the supervised contrast loss is obtained, and the relationship between the process is: ; in, Indicates The supervised contrast loss for each instance, represents feature decoupling regularization, represents the dynamic temperature coefficient for adaptive adjustment of characteristic differences according to sample pairs, and Both represent the results of structural decoupling of features. Represents the projection features of other extreme instances, Represents the weight coefficient of the feature decoupling regularization term; All supervised contrast losses are summed up to get the total supervised contrast loss. The relationship between the corresponding process is: ; in, represents the total supervised contrast loss, Represents a set of extreme instances.
9. The prototype pseudo-instance enhanced video anomaly detection method according to claim 8, characterized in that: In step 4, the video anomaly detection model is trained by combining the total supervision contrast loss and the multi-instance learning loss to obtain a trained video anomaly detection model, wherein the total loss is: ; in, represents the total loss, represents the uncertainty parameter related to the multi-instance learning task, represents the uncertainty parameter of the total supervision comparison task, represents the multi-instance learning loss, represents the basic mixing coefficient, represents the regularization coefficient, represents the prototype regularization term.
10. A prototype pseudo-instance enhanced video anomaly detection system, characterized in that The system applies any one of the prototype pseudo-instance enhanced video anomaly detection methods of claims 1 to 9, and the system comprises: Feature extraction module for: The video is input into the video anomaly detection model, and the video is sequentially segmented and feature extracted to obtain deep features, and a feature sequence is obtained based on the deep features; Enhanced feature modules for: The feature sequence is input into the prototype interaction layer, and the cosine similarity is calculated through the normal prototype set; Attention weights are generated based on cosine similarity; Get the normal context vector through the attention weight; After linear transformation of the normal context vector, residual connection is performed with the deep feature to obtain enhanced features; Screening modules for: Based on the enhanced features, the anomaly score is calculated by the classifier; The videos are filtered by anomaly scores to obtain anomaly sets and normal sets respectively. The anomaly sets and normal sets are merged to obtain an index set of extreme instances. Training modules for: Based on the index set of extreme instances, the enhanced features corresponding to the extreme instances are mapped to low dimensions to obtain deep feature embedding representation; Constructing a total supervised contrastive loss based on deep feature embedding representation; The video anomaly detection model is trained by combining the total supervision contrast loss and the multi-instance learning loss to obtain a trained video anomaly detection model; Result output module, used for: The video is input into the trained video anomaly detection model, which outputs the video-level anomaly classification results and instance-level anomaly location scores.
Citation Information
Patent Citations
Attention mechanism guided weak supervision visual anomaly event detection method
CN116883885A
Video anomaly detection method based on space-time enhanced associated memory
CN116958878A
Feature enhancement and fusion-based weak supervision video anomaly detection method and system
CN118470608A
Video anomaly detection model network structure, training method and detection method
CN118865202A
Video abnormal event detection method based on prompt learning and multi-scale time sequence fusion
CN118918506A
Cited By
Weak supervision video anomaly detection method and system based on prototype orthogonality
CN121640198A