Video anomaly detection method and system based on time interaction entropy
Through the video anomaly detection method based on temporal interaction entropy, the temporal structural entropy and global volatility prior descriptor optimization model are used to solve the shortcomings of the existing technology in the detection of temporal anomalies and dynamic pattern mutations, and achieve efficient and accurate video anomaly detection, which is suitable for scenarios such as security monitoring and intelligent transportation.
Patent Information
- Application Number
- CN202511224320.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing video anomaly detection technology is not effective in detecting timing anomalies, dynamic pattern mutations, and low-energy anomalies. It relies on frame-level annotation data and has low computational efficiency, making it difficult to adapt to complex scenarios.
A method based on temporal interaction entropy is adopted to build a video anomaly detection model through pre-trained visual encoder, sliding window algorithm, fully connected layer and asymmetric attention mechanism. Temporal structural entropy and global volatility prior descriptor are used to perform anomaly scoring, and the model is optimized by combining multi-example learning and global structure contrast loss function.
It improves the accuracy and computational efficiency of video anomaly detection, reduces data annotation costs, and is suitable for scenarios such as security monitoring and intelligent transportation. It has cross-scenario generalization capabilities and robustness.
Smart Images

Figure CN120726546A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and artificial intelligence technology, and in particular to a video anomaly detection method and system based on temporal interaction entropy. Background Art
[0002] Current video anomaly detection technologies are mainly divided into three categories: 1. Reconstruction-based methods (such as Autoencoder and GAN): detect anomalies through reconstruction errors, assuming that abnormal events are difficult to reconstruct from normal patterns.
[0003] Typical examples include: ①Conv-AE (2016), which uses convolutional autoencoders to learn normal patterns; and ②FutureFrame Prediction (2018), which detects anomalies by predicting the next frame. Its advantage is that it does not require training with abnormal samples.
[0004] 2. Classification-based methods (such as 3D CNN, Two-Stream): directly model the normal / abnormal classification boundary.
[0005] Typical methods include: ① C3D (2015), which uses 3D convolution to extract spatiotemporal features; and ② TSN (2016), which uses segmented sampling and a two-stream network. Its advantage lies in its high detection rate for overt anomalies (such as violent behavior).
[0006] 3. Weakly supervised methods (such as MIL and attention mechanism): trained only using video-level labels.
[0007] Mainstream technologies include: ①RTFM (2021): based on feature amplitude sorting; ②MGFN (2022): multi-granularity feature fusion. Its advantage is to reduce annotation costs.
[0008] The above video anomaly detection technology has the following defects: 1. Reconstruction-based method: The core defect is that it is insensitive to timing anomalies and has difficulty detecting dynamic pattern mutations.
[0009] 2. Classification-based method: The core flaw is that it relies heavily on frame-level annotation data, and the actual application cost is too high.
[0010] 3. Weakly supervised method: The core flaw is that it over-relies on feature amplitude and has difficulty detecting low-energy anomalies.
[0011] 4. Common Problems: The fundamental bottleneck is that existing methods are all based on the reasoning logic of "semantics → anomaly". However, actual anomalies may be highly similar to normal events in visual appearance (such as stealth theft), or only appear as subtle violations of temporal dynamics (such as sudden changes in mechanical vibration frequency).
[0012] Current technology is limited by semantic ambiguity, strong annotation dependence and low computational efficiency, and there is an urgent need for a new method that can directly quantify the stability of temporal structure. Summary of the Invention
[0013] In view of the above situation, the main purpose of the present invention is to propose a video anomaly detection method and system based on temporal interaction entropy to solve the above technical problems.
[0014] The present invention proposes a video anomaly detection method based on temporal interaction entropy, which includes the following steps: Step 1: Build a feature extraction module based on the pre-trained visual encoder, build a temporal structure entropy calculation module based on the sliding window algorithm, build an attention modulation module based on the fully connected layer and the asymmetric attention mechanism, and build an anomaly scoring module based on the multi-layer perceptron and the linear transformation layer. The video anomaly detection model is constructed by combining the feature extraction module, the temporal structure entropy calculation module, the attention modulation module, and the anomaly scoring module. Step 2: Obtain the original video stream, divide the original video stream into video segments, generate time-perturbed samples from the video segments, and use the feature extraction module to extract features from the time-perturbed samples generated by each video segment and calculate the average to obtain a video feature sequence; Step 3: Process the video feature sequence using the time structure entropy calculation module to obtain a time structure entropy sequence; Step 4: Use the fully connected layer to map the video feature sequence to a low-dimensional semantic space to obtain semantic features; use the attention modulation module to perform asymmetric attention calculation on the semantic features to obtain structural entropy modulation attention output features; Step 5: Extract the global temporal volatility prior descriptor based on the temporal structural entropy sequence, and use the anomaly scoring module to process the global temporal volatility prior descriptor and the structural entropy modulated attention output features to obtain an anomaly score; Step 6: Construct a multi-example learning loss function based on the anomaly score, and construct a global structure contrast loss function based on the global structure descriptor. Use the multi-example learning loss function and the global structure contrast loss function to optimize the video anomaly detection model to obtain the optimized video anomaly detection model. Input the original video stream into the optimized video anomaly detection model for processing to obtain the final anomaly score.
[0015] The present invention further proposes a video anomaly detection system based on temporal interaction entropy, wherein the system applies the video anomaly detection method based on temporal interaction entropy as described above, and the system includes: Building blocks for: A feature extraction module is built based on the pre-trained visual encoder, a temporal structure entropy calculation module is built based on the sliding window algorithm, an attention modulation module is built based on the fully connected layer and the asymmetric attention mechanism, and an anomaly scoring module is built based on the multi-layer perceptron and the linear transformation layer. The video anomaly detection model is composed of the feature extraction module, the temporal structure entropy calculation module, the attention modulation module, and the anomaly scoring module. Extraction module for: Obtain the original video stream, divide the original video stream into video segments, generate time-perturbed samples from the video segments, and use the feature extraction module to extract features from the time-perturbed samples generated by each video segment and calculate the average to obtain a video feature sequence; Temporal structure entropy module, used for: The video feature sequence is processed using the time structure entropy calculation module to obtain the time structure entropy sequence; Modulated attention module for: The video feature sequence is mapped to a low-dimensional semantic space using a fully connected layer to obtain semantic features. The attention modulation module is used to perform asymmetric attention calculation on the semantic features to obtain structural entropy modulated attention output features. Scoring module for: Based on the temporal structure entropy sequence, a prior descriptor of global temporal volatility is extracted. The anomaly scoring module is used to process the prior descriptor of global temporal volatility and the structural entropy modulated attention output features to obtain an anomaly score. Based on the anomaly score, a multi-example learning loss function is constructed. The global temporal volatility prior descriptor is used as the global structure descriptor. Based on the global structure descriptor, a global structure contrast loss function is constructed. The video anomaly detection model is optimized using the multi-example learning loss function and the global structure contrast loss function to obtain the optimized video anomaly detection model. The original video stream is input into the optimized video anomaly detection model for processing to obtain the final anomaly score.
[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. Superior detection performance: On mainstream public datasets such as UCF-Crime (complex scene anomaly detection), XD-Violence (violent behavior detection), and ShanghaiTech (multi-scene anomaly detection), this method's frame-level AUC (area under the curve) and AP (average precision) reach or exceed the current state-of-the-art (SOTA) levels. By quantifying local dynamic instability through temporal structural entropy and combining it with structural entropy to modulate attention, it can accurately locate abnormal segments and reduce false positives and missed detections. This makes it suitable for scenarios such as security monitoring and intelligent transportation, significantly improving the accuracy of automated detection of abnormal events and reducing manual review costs. 2. High computational efficiency: The lightweight design incorporates a parameter-free temporal structure entropy calculation module. This module performs instant feature centering and temporal autocorrelation matrix calculation within a sliding window, eliminating the need for trainable parameters and significantly reducing computational overhead. This makes the system suitable for deployment on embedded devices and can be widely used in resource-constrained scenarios, enabling real-time anomaly detection with high frame rates and low latency. 3. Strong weak supervision capabilities: Only video-level labels are required: Through multi-instance learning loss and global structural contrast loss, training relies solely on video-level abnormal / normal labels, eliminating the need for expensive frame-level annotation. It automatically focuses on potentially abnormal segments using temporal structural entropy and generates video-level predictions through maximum aggregation, effectively addressing label sparsity and significantly reducing data annotation costs. This is particularly suitable for long video surveillance scenarios and can quickly adapt to the anomaly detection needs of new scenarios. 4. Robustness: Structural stability is prioritized. Temporal structural entropy is used to quantify temporal structural instability, rather than relying on semantic content or feature amplitude, to avoid misjudgments caused by illumination changes, background noise, etc. Global-local collaborative calibration combines global volatility priors to perform macroscopic calibration of local predictions, suppressing local noise interference and improving cross-scenario generalization capabilities.
[0017] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flowchart of the steps of a video anomaly detection method based on temporal interaction entropy proposed by the present invention; Figure 2 This is the overall framework diagram of a video anomaly detection method based on temporal interaction entropy proposed in the present invention; Figure 3 This is a visualization diagram of the anomaly detection results of a video anomaly detection method based on temporal interaction entropy proposed in the present invention; Figure 4 This is a system structure diagram of a video anomaly detection system based on temporal interaction entropy proposed in this invention. DETAILED DESCRIPTION
[0019] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0020] These and other aspects of the embodiments of the present invention will become clear with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0021] See also Figure 1 This embodiment provides a video anomaly detection method based on temporal interaction entropy, the method comprising the following steps: Step 1. Build a feature extraction module based on the pre-trained visual encoder, build a temporal structure entropy calculation module based on the sliding window algorithm, build an attention modulation module based on the fully connected layer and asymmetric attention mechanism, and build an anomaly scoring module based on the multi-layer perceptron and linear transformation layer. The video anomaly detection model is constructed by the feature extraction module, temporal structure entropy calculation module, attention modulation module and anomaly scoring module.
[0022] Step 2: Obtain the original video stream, divide the original video stream into video segments, generate time perturbation samples through the video segments, and use the feature extraction module to extract features and calculate the average of the time perturbation samples generated by each video segment to obtain a video feature sequence.
[0023] See also Figure 2 In step 2, the original video stream is obtained and divided into video segments. Time perturbation samples are generated from the video segments. The feature extraction module is used to extract features from the time perturbation samples generated by each video segment and average them to obtain a video feature sequence. The specific steps include the following: Obtain an original video stream, and divide the original video stream into segments of equal length to obtain a video segment sequence; Based on the video clip sequence, three time perturbation samples are generated for each video clip; The pre-trained visual encoder is used to extract features from the three temporal perturbation samples generated for each video clip and average them to obtain the feature representation of each video clip. The following relationship exists in the corresponding process: ; in, Indicates the The feature representation of a video clip, Indicates the index of the video segment, Indicates the Video clip of Time perturbation sample characteristics; Arrange the feature representation of each video clip in chronological order to obtain a video feature sequence ; in, represents the video feature sequence, Indicates the number of video clips.
[0024] It should be noted that each segment contains 16 frames and the time span is 1 second. If the total number of video frames is not a multiple of 16, the missing part at the end will be filled with black frames. Three time perturbation samples are generated for each segment, specifically: sample 1 is the original 16 frames (frames 1 to 16), sample 2 is right-shifted by 1 frame (frames 2 to 17), and sample 3 is left-shifted by 1 frame (frames 0 to 15, the missing part is filled with black frames). Figure 2 In , TSE represents the temporal structure entropy and TVP represents the global volatility prior.
[0025] Step 3: Use the time structure entropy calculation module to process the video feature sequence to obtain the time structure entropy sequence.
[0026] In step 3, the video feature sequence is processed using a temporal structure entropy calculation module to obtain a temporal structure entropy sequence, which specifically includes the following sub-steps: The video feature sequence is segmented using the sliding window algorithm, and the feature representation of each video clip in the window is immediately centralized to obtain a centralized feature vector. The following relationship exists in the corresponding process: ; in, Indicates the The centralized feature vector of the video clips, Indicates that the mean value along the feature dimension is calculated; All the centralized feature vectors in the window are stacked to obtain the centralized feature matrix in the window. The following relationship exists in the corresponding process: ; in, represents the centralized feature matrix within the window, represents the sliding window size, and ; The time autocorrelation matrix of the centralized feature matrix in the window is calculated to obtain the time autocorrelation matrix in the window. The following relationship exists in the corresponding process: ; in, represents the temporal autocorrelation matrix within the window, represents transpose; Perform eigenvalue decomposition on the time autocorrelation matrix within the window to obtain the eigenvalue spectrum; Based on the eigenvalue spectrum, the eigenvalue is normalized into a probability distribution, and the entropy value is calculated to obtain the structural entropy value. The following relationship exists in the corresponding process: ; in, represents the normalized probability distribution of eigenvalues, Representation matrix No. eigenvalues, Representation matrix No. eigenvalues, Indicates the The temporal structure entropy of a video clip, Representation matrix The number of nonzero eigenvalues of Indicates taking the logarithm; The structural entropy values of each video clip are spliced together to obtain the temporal structural entropy sequence ; in, Represents the time structure entropy sequence.
[0027] Step 4: Use the fully connected layer to map the video feature sequence to the low-dimensional semantic space to obtain semantic features; use the attention modulation module to perform asymmetric attention calculation on the semantic features to obtain the structural entropy modulated attention output features.
[0028] In step 4, the video feature sequence is mapped to a low-dimensional semantic space using a fully connected layer to obtain semantic features. The attention modulation module is used to perform asymmetric attention calculation on the semantic features to obtain structural entropy modulated attention output features. The specific steps include the following: The video feature sequence is mapped to a low-dimensional semantic space using a fully connected layer to obtain semantic features. The following relationship exists in the corresponding process: ; in, Represents semantic features, Indicates that it has been processed by the fully connected layer; Project the semantic features to obtain the query matrix and the key matrix generated by relying solely on the semantic feature flow. The following relationship exists in the corresponding process: ; in, represents the query matrix generated by relying only on the semantic feature flow, represents the trainable projection matrix used to generate the query matrix, represents the key matrix generated by relying only on the semantic feature flow, represents the trainable projection matrix used to generate the key matrix; The Sigmoid function is used to process the time structure entropy sequence to obtain the dynamic weight. The following relationship exists in the corresponding process: ; in, represents the dynamic weight, Indicates that it has been processed by the Sigmoid function. Represents the arithmetic mean of the time structure entropy series; The dynamic weights are multiplied element-by-element by the semantic features and projected to obtain the value matrix generated by the structural value stream. The following relationship exists in the corresponding process: ; in, represents the value matrix generated by the structural value stream, represents element-wise multiplication, represents the trainable value projection matrix; Attention calculation is performed based on the query matrix, key matrix and value matrix to obtain the structural entropy modulated attention output feature. The following relationship exists in the corresponding process: ; in, represents the structural entropy modulation attention output feature, represents the attention mechanism, Indicates that it has been processed by the softmax function. Indicates the key vector dimension.
[0029] Step 5: Extract the global temporal volatility prior descriptor based on the temporal structural entropy sequence, and use the anomaly scoring module to process the global temporal volatility prior descriptor and the structural entropy modulated attention output features to obtain the anomaly score.
[0030] In step 5, a global temporal volatility prior descriptor is extracted based on the temporal structural entropy sequence. The global temporal volatility prior descriptor and the structural entropy modulated attention output features are processed using the anomaly scoring module to obtain an anomaly score. The specific steps include the following: The global temporal volatility prior descriptor is extracted based on the temporal structure entropy sequence. The following relationship exists in the corresponding process: ; in, represents the global temporal volatility prior descriptor, represents the standard deviation of the time structure entropy series, Indicates taking the minimum value, Indicates taking the maximum value; The global temporal volatility prior descriptor is mapped to the feature space using a multi-layer perceptron and multiplied element-by-element with the structural entropy modulated attention output feature to obtain the attention feature modulated by the global temporal volatility prior. The following relationship exists in the corresponding process: ; in, represents the attention feature modulated by the global temporal volatility prior, represents the scaling factor, Indicates that it has been processed by a multi-layer perceptron; After subtracting the semantic features from the structural entropy modulated attention output features and performing layer normalization, the residual features are obtained. The following relationship exists in the corresponding process: ; in, represents the residual feature, Indicates that it has been processed by layer normalization; The attention feature modulated by the global temporal volatility prior is added to the residual feature to obtain the final fused feature. The following relationship exists in the corresponding process: ; in, Represents the final feature after fusion; The final fused features are processed by the linear transformation layer and the Sigmoid function in sequence to obtain the anomaly score. The following relationship exists in the corresponding process: ; in, represents the anomaly score, Indicates that it has been processed by the linear transformation layer.
[0031] Step 6. Based on the anomaly score, a multi-example learning loss function is constructed. The global temporal volatility prior descriptor is used as the global structure descriptor. Based on the global structure descriptor, a global structure contrast loss function is constructed. The video anomaly detection model is optimized using the multi-example learning loss function and the global structure contrast loss function to obtain an optimized video anomaly detection model. The original video stream is input into the optimized video anomaly detection model for processing to obtain the final anomaly score.
[0032] In step 6, a multi-example learning loss function is constructed based on the anomaly score, and the global temporal volatility prior descriptor is used as the global structure descriptor. A global structure contrast loss function is constructed based on the global structure descriptor. The video anomaly detection model is optimized using the multi-example learning loss function and the global structure contrast loss function to obtain an optimized video anomaly detection model. The original video stream is input into the optimized video anomaly detection model for processing to obtain the final anomaly score. The multi-example learning loss function is constructed based on the anomaly score, and the corresponding process has the following relationship: ; in, represents the multi-example learning loss, represents the number of training samples, Indicates the Tags for videos, Indicates the The highest segment score of the videos; Among them, the global structure contrast loss function is constructed based on the global structure descriptor, and the following relationship exists in the corresponding process: ; in, represents the centroid of normal samples, represents a normal sample set, represents the global structure contrast loss, Indicates that it has been processed by L2 regularization term. represents a set of abnormal samples, represents the boundary hyperparameter; It should be noted that the global structure descriptor in the global structure contrast loss function directly reuses the global temporal volatility prior descriptor output in step 4, and its statistics characterize the overall dynamic stability of the video.
[0033] The video anomaly detection model is optimized using the multi-example learning loss function and the global structure contrast loss function to obtain the optimized video anomaly detection model. The original video stream is input into the optimized video anomaly detection model for processing to obtain the final anomaly score.
[0034] Furthermore, the multi-example learning loss and the global structure contrast loss are weighted and fused to obtain the total loss. The following relationship exists in the corresponding process: ; in, represents the total loss, Represents all learnable parameters in the model.
[0035] To validate the effectiveness of this invention, we conducted comprehensive tests on three major datasets: ShanghaiTech, UCF-Crime, and XD-Violence. ShanghaiTech contains 437 videos from 13 different scenarios, covering common abnormal events; UCF-Crime contains 1,900 long videos covering complex abnormal events; and XD-Violence focuses on violent behavior detection to ensure the reliability of the model's generalization ability assessment.
[0036] 1. Comparison with the State-of-the-Art (SOTA) In a weakly supervised setting, we use frame-level AUC (Area Under the ROC Curve) as the main evaluation metric to compare with current mainstream methods. The experimental results are shown in Table 1: Table 1 Performance comparison with SOTA methods (AUC / %)
[0037] Key conclusions: When using only the RGB modality, the proposed TIENet achieves an AUC of 98.03% and 88.34% on ShanghaiTech and UCF-Crime, respectively, significantly outperforming the multi-modal VadCLIP (88.02%).
[0038] On the XD-Violence dataset, the AP (87.23%) and APA (87.95%) of the proposed TIENet surpassed existing methods, verifying its ability to detect complex violence scenes.
[0039] Lightweight comparison: As shown in Table 2, TIENet has much fewer parameters (0.46M) and a smaller model size (1.76MB) than other methods while maintaining the highest performance, reflecting its “efficient and lightweight” design advantage.
[0040] Table 2 Model efficiency comparison
[0041] 2. Ablation Experiment To verify the contributions of Temporal Structure Entropy (TSE), Structure Entropy Modulated Attention (SEMA), and Global Volatility Prior (TVP), we gradually stack modules for testing (UCF-Crime dataset): Table 3 Ablation experiment results (AUC / %)
[0042] Key conclusions: TSE module: After independent introduction, the AUC increased by 0.73%, demonstrating the ability of temporal structure entropy to quantify dynamic instability.
[0043] SEMA mechanism: By modulating attention through structural entropy, the performance is further improved by 0.93%, verifying the effectiveness of "structure-guided semantics".
[0044] TVP and contrastive loss: Global calibration achieves an AUC of 88.34%, indicating the necessity of integrating multi-level information.
[0045] 3. Hyperparameter sensitivity analysis Explore the impact of sliding window size W on the TSE module (UCF-Crime dataset): Table 4 Effect of window size W (AUC / %)
[0046] Key conclusions: The optimal window size is W=5, which can balance local dynamic capture and computational efficiency.
[0047] The performance fluctuation is less than 0.2% when W∈[3,7], showing the robustness of the method to hyperparameters.
[0048] 4. Visualization of abnormal scenarios See also Figure 3 Through keyframe and TSE curve visualization, TIENet performs well in the following scenarios: Traffic accidents: The TSE value suddenly increases at the moment of collision (the peak value is 300% higher than the baseline), but traditional methods miss it due to background noise.
[0049] Stealth behavior: The amplitude of semantic features changes slightly, but TSE accurately marks abnormal periods through "dynamic instability".
[0050] Crowd gathering: TSE remains high, reflecting the formation of a new structure, which is clearly distinguishable from short-term energy fluctuations.
[0051] See also Figure 4 This embodiment further provides a video anomaly detection system based on temporal interaction entropy, wherein the system applies the video anomaly detection method based on temporal interaction entropy as described above, and the system includes: Building blocks for: A feature extraction module is built based on the pre-trained visual encoder, a temporal structure entropy calculation module is built based on the sliding window algorithm, an attention modulation module is built based on the fully connected layer and the asymmetric attention mechanism, and an anomaly scoring module is built based on the multi-layer perceptron and the linear transformation layer. The video anomaly detection model is composed of the feature extraction module, the temporal structure entropy calculation module, the attention modulation module, and the anomaly scoring module. Extraction module for: Obtain the original video stream, divide the original video stream into video segments, generate time-perturbed samples from the video segments, and use the feature extraction module to extract features from the time-perturbed samples generated by each video segment and calculate the average to obtain a video feature sequence; Temporal structure entropy module, used for: The video feature sequence is processed using the time structure entropy calculation module to obtain the time structure entropy sequence; Modulated attention module for: The video feature sequence is mapped to a low-dimensional semantic space using a fully connected layer to obtain semantic features. The attention modulation module is used to perform asymmetric attention calculation on the semantic features to obtain structural entropy modulated attention output features. Scoring module for: Based on the temporal structure entropy sequence, a prior descriptor of global temporal volatility is extracted. The anomaly scoring module is used to process the prior descriptor of global temporal volatility and the structural entropy modulated attention output features to obtain an anomaly score. Based on the anomaly score, a multi-example learning loss function is constructed. The global temporal volatility prior descriptor is used as the global structure descriptor. Based on the global structure descriptor, a global structure contrast loss function is constructed. The video anomaly detection model is optimized using the multi-example learning loss function and the global structure contrast loss function to obtain the optimized video anomaly detection model. The original video stream is input into the optimized video anomaly detection model for processing to obtain the final anomaly score.
[0052] It should be understood that, although the various steps in the flow chart of each embodiment of the present invention are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0053] It should be understood that various components of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0054] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0055] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A video anomaly detection method based on temporal interaction entropy, characterized in that: The method comprises the following steps: Step 1: Build a feature extraction module based on the pre-trained visual encoder, build a temporal structure entropy calculation module based on the sliding window algorithm, build an attention modulation module based on the fully connected layer and the asymmetric attention mechanism, and build an anomaly scoring module based on the multi-layer perceptron and the linear transformation layer. The video anomaly detection model is constructed by combining the feature extraction module, the temporal structure entropy calculation module, the attention modulation module, and the anomaly scoring module. Step 2: Obtain the original video stream, divide the original video stream into video segments, generate time-perturbed samples from the video segments, and use the feature extraction module to extract features from the time-perturbed samples generated by each video segment and calculate the average to obtain a video feature sequence; Step 3: Process the video feature sequence using the time structure entropy calculation module to obtain a time structure entropy sequence; Step 4: Use the fully connected layer to map the video feature sequence to a low-dimensional semantic space to obtain semantic features; The attention modulation module is used to perform asymmetric attention calculation on the semantic features to obtain the structural entropy modulated attention output features; Step 5: Extract the global temporal volatility prior descriptor based on the temporal structural entropy sequence, and use the anomaly scoring module to process the global temporal volatility prior descriptor and the structural entropy modulated attention output features to obtain an anomaly score; Step 6. Based on the anomaly score, a multi-example learning loss function is constructed. The global temporal volatility prior descriptor is used as the global structure descriptor. Based on the global structure descriptor, a global structure contrast loss function is constructed. The video anomaly detection model is optimized using the multi-example learning loss function and the global structure contrast loss function to obtain an optimized video anomaly detection model. The original video stream is input into the optimized video anomaly detection model for processing to obtain the final anomaly score.
2. The video anomaly detection method based on temporal interaction entropy according to claim 1 is characterized in that: In step 2, the original video stream is obtained, the original video stream is divided into video segments, time-perturbed samples are generated from the video segments, and a feature extraction module is used to extract features from the time-perturbed samples generated by each video segment and average them to obtain a video feature sequence. The specific steps include the following: Obtain an original video stream, and divide the original video stream into segments of equal length to obtain a video segment sequence; Based on the video clip sequence, three time perturbation samples are generated for each video clip; The pre-trained visual encoder is used to extract features from the three temporal perturbation samples generated for each video clip and average them to obtain the feature representation of each video clip. The following relationship exists in the corresponding process: ; in, Indicates the The feature representation of a video clip, Indicates the index of the video segment, Indicates the Video clip of Time perturbation sample characteristics; Arrange the feature representation of each video clip in chronological order to obtain a video feature sequence ; in, represents the video feature sequence, Indicates the number of video clips.
3. The video anomaly detection method based on temporal interaction entropy according to claim 2 is characterized in that: In step 3, the video feature sequence is processed using a temporal structure entropy calculation module to obtain a temporal structure entropy sequence, which specifically includes the following sub-steps: The video feature sequence is segmented using a sliding window algorithm, and the feature representation of each video segment within the window is immediately centralized to obtain a centralized feature vector. Stack all the centralized feature vectors in the window to obtain the centralized feature matrix in the window; Performing time autocorrelation matrix calculation on the centralized feature matrix within the window to obtain the time autocorrelation matrix within the window; Perform eigenvalue decomposition on the time autocorrelation matrix within the window to obtain the eigenvalue spectrum; Based on the eigenvalue spectrum, the eigenvalues are normalized into probability distribution and the entropy value is calculated to obtain the structural entropy value; The structural entropy values of each video clip are spliced together to obtain the temporal structural entropy sequence ; in, represents the time structure entropy sequence, Indicates the The temporal structure entropy of a video clip.
4. The video anomaly detection method based on temporal interaction entropy according to claim 3 is characterized in that: In the step of using the sliding window algorithm to segment the video feature sequence and perform real-time feature centering on the feature representation of each video segment within the window to obtain a centralized feature vector, the following relationship exists: ; in, Indicates the The centralized feature vector of the video clips, Indicates that the mean value along the feature dimension is calculated; In the step of stacking all the centralized feature vectors in the window to obtain the centralized feature matrix in the window, the following relationship exists: ; in, represents the centralized feature matrix within the window, Indicates the sliding window size; In the step of calculating the time autocorrelation matrix of the centered feature matrix in the window to obtain the time autocorrelation matrix in the window, the following relationship exists: ; in, represents the temporal autocorrelation matrix within the window, represents transpose; In the step of normalizing the eigenvalues to a probability distribution based on the eigenvalue spectrum and calculating the entropy value to obtain the structural entropy value, the following relationship exists: ; in, represents the normalized probability distribution of eigenvalues, Representation matrix No. eigenvalues, Representation matrix No. eigenvalues, Representation matrix The number of nonzero eigenvalues of Indicates taking the logarithm.
5. The video anomaly detection method based on temporal interaction entropy according to claim 4 is characterized in that: In step 4, the video feature sequence is mapped to a low-dimensional semantic space using a fully connected layer to obtain semantic features; and the attention modulation module is used to perform asymmetric attention calculation on the semantic features to obtain structural entropy modulated attention output features, which specifically includes the following sub-steps: Use the fully connected layer to map the video feature sequence into a low-dimensional semantic space to obtain semantic features; Performing a projection operation on the semantic features to obtain a query matrix generated solely by the semantic feature flow and a key matrix generated solely by the semantic feature flow; The time structure entropy sequence is processed using the Sigmoid function to obtain dynamic weights; Multiply the dynamic weights by the semantic features element-wise and perform a projection operation to obtain the value matrix generated by the structural value stream; Attention calculation is performed based on the query matrix, key matrix and value matrix to obtain the structural entropy modulated attention output features.
6. The video anomaly detection method based on temporal interaction entropy according to claim 5 is characterized in that: In the step of using the fully connected layer to map the video feature sequence to a low-dimensional semantic space to obtain semantic features, the following relationship exists: ; in, Represents semantic features, Indicates that it has been processed by the fully connected layer; In the step of performing a projection operation on the semantic features to obtain a query matrix generated solely by the semantic feature stream and a key matrix generated solely by the semantic feature stream, the following relationship exists: ; in, represents the query matrix generated by relying only on the semantic feature flow, represents the trainable projection matrix used to generate the query matrix, represents the key matrix generated by relying only on the semantic feature flow, represents the trainable projection matrix used to generate the key matrix; In the step of using the Sigmoid function to process the time structure entropy sequence to obtain the dynamic weight, the following relationship exists: ; in, represents the dynamic weight, Indicates that it has been processed by the Sigmoid function. Represents the arithmetic mean of the time structure entropy series; In the step of multiplying the dynamic weights by the semantic features element-wise and performing a projection operation to obtain the value matrix generated by the structural value stream, the following relationship exists: ; in, represents the value matrix generated by the structural value stream, represents element-wise multiplication, represents the trainable value projection matrix; In the step of performing attention calculation based on the query matrix, key matrix, and value matrix to obtain the structural entropy modulated attention output feature, the following relationship exists: ; in, represents the structural entropy modulation attention output feature, represents the attention mechanism, Indicates that it has been processed by the softmax function. Indicates the key vector dimension.
7. The video anomaly detection method based on temporal interaction entropy according to claim 6 is characterized in that: In step 5, a global temporal volatility prior descriptor is extracted based on the temporal structural entropy sequence, and the global temporal volatility prior descriptor and the structural entropy modulated attention output feature are processed using an anomaly scoring module to obtain an anomaly score, which specifically includes the following sub-steps: Extracting a priori descriptors of global temporal volatility based on temporal structure entropy sequences; The global temporal volatility prior descriptor is mapped to the feature space using a multi-layer perceptron and is element-wise multiplied with the structured entropy modulated attention output feature to obtain the attention feature modulated by the global temporal volatility prior. After subtracting the semantic features from the structural entropy modulated attention output features and performing layer normalization, the residual features are obtained. The attention feature modulated by the global temporal volatility prior is added to the residual feature to obtain the final fused feature; The final fused features are processed by the linear transformation layer and the Sigmoid function in sequence to obtain the anomaly score.
8. The video anomaly detection method based on temporal interaction entropy according to claim 7 is characterized in that: In the step of extracting the global temporal volatility prior descriptor based on the temporal structure entropy sequence, the following relationship exists: ; in, represents the global temporal volatility prior descriptor, represents the standard deviation of the time structure entropy series, Indicates taking the minimum value, Indicates taking the maximum value; In the step of mapping the global temporal volatility prior descriptor to the feature space using a multilayer perceptron and performing element-wise multiplication with the structural entropy modulated attention output feature to obtain the attention feature modulated by the global temporal volatility prior, the following relationship exists: ; in, represents the attention feature modulated by the global temporal volatility prior, represents the scaling factor, Indicates that it has been processed by a multi-layer perceptron; After subtracting the semantic features from the structural entropy modulated attention output features and performing layer normalization to obtain the residual features, the following relationship exists: ; in, represents the residual feature, Indicates that it has been processed by layer normalization; In the step of adding the attention feature modulated by the global temporal volatility prior to the residual feature to obtain the final fused feature, the following relationship exists: ; in, Represents the final feature after fusion; In the step of processing the fused final features through the linear transformation layer and the Sigmoid function in sequence to obtain the anomaly score, the following relationship exists: ; in, represents the anomaly score, Indicates that it has been processed by the linear transformation layer.
9. The video anomaly detection method based on temporal interaction entropy according to claim 8, characterized in that: In step 6, a multi-example learning loss function is constructed based on the anomaly score, a global temporal volatility prior descriptor is used as a global structure descriptor, and a global structure contrast loss function is constructed based on the global structure descriptor. The video anomaly detection model is optimized using the multi-example learning loss function and the global structure contrast loss function to obtain an optimized video anomaly detection model. The original video stream is input into the optimized video anomaly detection model for processing to obtain a final anomaly score. The multi-example learning loss function is constructed based on the anomaly score, and the following relationship exists in the corresponding process: ; in, represents the multi-example learning loss, represents the number of training samples, Indicates the Tags for videos, Indicates the The highest segment score of the videos; Among them, the global structure contrast loss function is constructed based on the global structure descriptor, and the following relationship exists in the corresponding process: ; in, represents the centroid of normal samples, represents a normal sample set, represents the global structure contrast loss, Indicates that it has been processed by L2 regularization term. represents a set of abnormal samples, represents the boundary hyperparameter.
10. A video anomaly detection system based on temporal interaction entropy, characterized in that: The system applies the video anomaly detection method based on temporal interaction entropy according to any one of claims 1 to 9, and the system includes: Building blocks for: A feature extraction module is built based on the pre-trained visual encoder, a temporal structure entropy calculation module is built based on the sliding window algorithm, an attention modulation module is built based on the fully connected layer and the asymmetric attention mechanism, and an anomaly scoring module is built based on the multi-layer perceptron and the linear transformation layer. The video anomaly detection model is composed of the feature extraction module, the temporal structure entropy calculation module, the attention modulation module, and the anomaly scoring module. Extraction module for: Obtain the original video stream, divide the original video stream into video segments, generate time-perturbed samples from the video segments, and use the feature extraction module to extract features from the time-perturbed samples generated by each video segment and calculate the average to obtain a video feature sequence; Temporal structure entropy module, used for: The video feature sequence is processed using the time structure entropy calculation module to obtain the time structure entropy sequence; Modulated attention module for: The video feature sequence is mapped to a low-dimensional semantic space using a fully connected layer to obtain semantic features. The attention modulation module is used to perform asymmetric attention calculation on the semantic features to obtain structural entropy modulated attention output features. Scoring module for: Based on the temporal structure entropy sequence, a prior descriptor of global temporal volatility is extracted. The anomaly scoring module is used to process the prior descriptor of global temporal volatility and the structural entropy modulated attention output features to obtain an anomaly score. Based on the anomaly score, a multi-example learning loss function is constructed. The global temporal volatility prior descriptor is used as the global structure descriptor. Based on the global structure descriptor, a global structure contrast loss function is constructed. The video anomaly detection model is optimized using the multi-example learning loss function and the global structure contrast loss function to obtain the optimized video anomaly detection model. The original video stream is input into the optimized video anomaly detection model for processing to obtain the final anomaly score.
Citation Information
Patent Citations
A traffic incident detection method based on depth learning and entropy model
CN108345894A
Weak supervision video anomaly detection method based on multiple text prompts
CN119206563A
DATA-DRIVEN ANOMALITY DETECTION AND PERFORMANCE OF SENSOR DATA
DE102021200344A1
KR20230095845A
Cited By
Associated imaging impurity detection method and system for drug production
CN120971448A
Abnormality monitoring processing method and system based on artificial intelligence
CN121259750A
Weak supervision video anomaly detection method and system based on multi-head spectrum residual gating
CN121963060A
Weakly supervised video anomaly detection method and system based on multi-head spectral residual gating
CN121963060B