Video abnormal event identification method and device, computer equipment and storage medium

By aligning and fusing visual, audio, and textual features through a multi-scale temporal model, the method enhances anomaly detection accuracy and adaptability to complex scenarios, addressing the limitations of existing systems in accurately identifying diverse abnormal events.

CN120316682APending Publication Date: 2025-07-15HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510491870.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing video abnormality detection system has problems of missed and false alarms in complex scenarios, and it is difficult to effectively identify multimodal information, resulting in low accuracy of abnormal event recognition.

Method used

By performing modal consistency processing of video, audio and text features, combining multi-scale timing networks and multi-head attention mechanisms, visual, audio and text information are integrated to generate temporal consistency features to improve detection accuracy.

Benefits of technology

It significantly improves the accuracy and robustness of abnormal detection, enhances the model's adaptability to complex scenarios and noise interference, and improves the generalization performance and applicability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316682A_ABST
    Figure CN120316682A_ABST
Patent Text Reader

Abstract

The invention discloses a video abnormal event identification method and device, computer equipment and a storage medium, and the method comprises the steps: carrying out the feature extraction of a target video, and obtaining a visual feature, an audio feature and a text feature; performing modal consistency processing on the visual features, the audio features and the text features to obtain a visual consistency feature, an audio consistency feature and a text consistency feature; and calling a preset mature student model to perform multi-scale time sequence processing on the visual consistency feature, the audio consistency feature and the text consistency feature to obtain at least one video fusion time sequence feature and a global weight value thereof, and obtaining an abnormal value according to the video fusion time sequence feature and the global weight value thereof. According to the invention, the comprehensiveness and accuracy of abnormal event recognition are ensured, the adaptability of the model to complex scenes and noise interference is improved, and the accuracy of anomaly detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method, device, computer device and storage medium for identifying video abnormal events. Background Art

[0002] Video anomaly detection is a core issue in the development of computer vision technology and a basic requirement for realizing intelligent monitoring and public safety. Affected by device performance, environmental conditions, algorithm defects and other factors, existing anomaly detection systems often have problems such as missed alarms and false alarms, seriously reducing the reliability and practicality of the monitoring system. The task of video anomaly detection is to quickly identify abnormal events occurring in the monitoring video and achieve real-time detection while ensuring the accuracy rate to meet the actual application requirements. Therefore, it is of great significance to study efficient and reliable video anomaly detection methods.

[0003] However, the inventor found that due to the complexity and diversity of current abnormal events, there are significant differences in the definitions and manifestation forms of abnormal behaviors in different scenarios. Therefore, the existing video anomaly detection models that gradually analyze the video in chronological order result in low accuracy in identifying abnormal events. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, device, computer device and storage medium for identifying video abnormal events, which are used to solve the problem of low accuracy in identifying abnormal events in videos existing in the prior art.

[0005] To achieve the above purpose, the present invention provides a method for identifying video abnormal events, including:

[0006] Performing feature extraction on the target video to obtain a visual feature, an audio feature and a text feature;

[0007] Performing modal consistency processing on the visual feature, the audio feature and the text feature to obtain a visual consistency feature, an audio consistency feature and a text consistency feature.

[0008] Invoking a pre-set mature student model to perform multi-scale temporal processing on the visual consistency feature, the audio consistency feature and the text consistency feature to obtain at least one video fusion temporal feature and its global weight value, and obtaining an anomaly value according to the video fusion temporal feature and its global weight value; wherein, the anomaly value reflects the probability of abnormal events occurring in the visual feature, audio feature and text feature of the target video.

[0009] In the above solution, performing feature extraction on the target video to obtain a visual feature, an audio feature and a text feature includes:

[0010] Extract the visual information in the target video, perform feature extraction on the visual information at a preset visual sampling rate to obtain at least one visual temporal feature, and integrate at least one of the visual temporal features to obtain a visual feature;

[0011] Extract the audio information in the target video, perform feature extraction on the audio information at a preset audio sampling rate to obtain at least one audio temporal feature, and integrate at least one of the audio temporal features to obtain an audio feature;

[0012] Convert the visual information into visual text information through a preset video subtitle generation network, and / or perform text conversion on the audio information to obtain audio text information, and integrate the visual text information and / or the audio text information to obtain the text information of the target video;

[0013] Perform feature extraction on the text information at a preset text sampling rate to obtain at least one text temporal feature, and integrate at least one of the text temporal features to obtain a text feature.

[0014] In the above solution, perform modality consistency processing on the visual feature, the audio feature, and the text feature to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature, including:

[0015] Project the visual feature, the audio feature, and the text feature into a preset common feature space respectively to obtain a visual projection feature, an audio projection feature, and a text projection feature; wherein, there is at least one visual temporal projection feature in the visual projection feature; there is at least one audio temporal projection feature in the audio projection feature; there is at least one text temporal projection feature in the text projection feature;

[0016] Calculate the differences between the visual temporal projection features, the audio temporal projection features, and the text temporal projections at each time step in sequence to obtain consistency loss information; wherein, the time step is a preset time span, and there is at least one visual temporal projection feature, at least one audio temporal projection feature, and at least one text temporal projection at one time step; the consistency loss information includes first difference data and second difference data; the first difference data represents the degree of difference between the visual temporal projection feature and the audio temporal projection feature; the second difference data represents the degree of difference between the audio temporal projection feature and the text temporal projection feature;

[0017] With the goal of minimizing the consistency loss information, the visual temporal projection features, audio temporal projection features, and text temporal projection features at each time step are adjusted respectively to obtain the visual temporal consistency features, audio temporal consistency features, and text temporal consistency features at each time step;

[0018] The visual temporal consistency features, audio temporal consistency features, and text temporal consistency features at each time step are aggregated respectively to obtain a visual consistency feature, an audio consistency feature, and a visual consistency feature.

[0019] In the above solution, a pre-set mature student model is called to perform multi-scale temporal processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain at least one video fusion temporal feature and its global weight value, including:

[0020] The mature student model performs enhancement processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain a visual enhancement feature, an audio enhancement feature, and a text enhancement feature; wherein, the enhancement processing includes one-dimensional convolution operation and non-local module processing; the one-dimensional convolution operation is used to identify the short-term dependence relationships between the visual temporal consistency features in the visual consistency feature and other visual temporal consistency features, the short-term dependence relationships between the audio temporal consistency features in the audio consistency feature and other audio temporal consistency features, and the short-term dependence relationships between the text temporal consistency features in the text consistency feature and other text temporal consistency features; the non-local module processing is used to identify the long-term dependence relationships between the visual temporal consistency features in the visual consistency feature and other visual temporal consistency features, the long-term dependence relationships between the audio temporal consistency features in the audio consistency feature and other audio temporal consistency features, and the long-term dependence relationships between the text temporal consistency features in the text consistency feature and other text temporal consistency features;

[0021] The mature student model fuses the visual consistency feature, the audio consistency feature, and the text consistency feature into a video fusion feature according to the visual enhancement feature, the audio enhancement feature, and the text enhancement feature; wherein, there is at least one video fusion temporal feature in the video fusion feature, and one video fusion temporal feature represents the visual content, sound features, and text description of the target video at a time step.

[0022] The mature student model identifies the correlation between each video fusion temporal feature and other video fusion temporal features in the video fusion features based on the multi-head attention mechanism, and obtains the global weight value of each video fusion temporal feature; wherein, the global weight value reflects the degree of content association and context dependence between a video fusion temporal feature and all video fusion temporal features.

[0023] In the above solution, obtaining the outlier according to the video fusion temporal feature and its global weight value includes:

[0024] The mature student model outputs the video fusion temporal feature and its global weight value to a preset fully connected layer;

[0025] The mature student model calls the fully connected layer for the video fusion temporal feature and its global weight value to obtain the outlier.

[0026] In the above solution, before extracting a visual feature, an audio feature, and a text feature from the target video, the method further includes:

[0027] Training a preset initial student model through at least one preset training video and a preset loss function to obtain a mature student model; wherein, there is a training outlier in the training video, and the training outlier characterizes one or more of visual content, sound features, and text descriptions, and records at least one abnormal event; the loss function includes classification loss, contrast learning loss, and modality consistency loss.

[0028] In the above solution, the initial student model belongs to an anomaly aggregation model, and the anomaly aggregation model further has a mature teacher model for training the initial student model.

[0029] To achieve the above object, the present invention also provides a video abnormal event recognition device, including:

[0030] A feature extraction module for extracting features from the target video to obtain a visual feature, an audio feature, and a text feature;

[0031] A modality consistency module for performing modality consistency processing on the visual feature, the audio feature, and the text feature to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature;

[0032] A multi-scale time series module is used to call a pre-set mature student model to perform multi-scale time series processing on the visual consistency feature, the audio consistency feature, and the text consistency feature, obtain at least one video fusion time series feature and its global weight value, and obtain an outlier according to the video fusion time series feature and its global weight value; wherein, the outlier reflects the probability of abnormal events occurring in the visual feature, audio feature, and text feature of the target video.

[0033] To achieve the above object, the present invention also provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor of the computer device executes the computer program, the steps of the above video anomaly event recognition method are implemented.

[0034] To achieve the above object, the present invention also provides a computer-readable storage medium. A computer program is stored on the readable storage medium. When the computer program stored on the readable storage medium is executed by a processor, the steps of the above video anomaly event recognition method are implemented.

[0035] A video anomaly event recognition method, device, computer device, and storage medium provided by the present invention extract a visual feature, an audio feature, and a text feature by performing feature extraction on the target video, so as to extract information of multiple modalities in the target video, facilitating subsequent analysis of visual content, sound features, and text descriptions respectively to ensure the comprehensiveness and accuracy of anomaly event recognition.

[0036] Through modal consistency processing, the distributions of modal features are aligned, the consistency features are matched at the semantic level (such as the association between visual actions and text descriptions), the time series features of different modalities are synchronized, and time consistency features are generated to ensure that information participates in subsequent calculations equally; at the same time, the semantic differences between different modalities are effectively processed, improving the model's adaptability to complex scenarios and noise interference. At the same time, the complementarity of multi-modal features also enhances the generalization performance of the model.

[0037] Comprehensively understand the monitoring scenario from three dimensions of vision, audio, and text. Ensure the semantic correspondence relationship of feature representations through modal consistency processing, and use multi-scale time series processing and the multi-head attention mechanism added in the multi-scale time series processing to enhance global correlation, significantly improving the accuracy of anomaly detection; moreover, the design of the feature alignment based on modal consistency processing and the fusion mechanism of visual features, audio features, and text features also ensures the applicability of the method in the detection of various anomaly events. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flowchart of a multi-modal video anomaly detection method driven by subtitles in an embodiment of the present invention;

[0039] Figure 2 This is a flowchart of a method for identifying video abnormal events according to an embodiment of the present invention;

[0040] Figure 3 It is a specific method flowchart of the method for identifying video abnormal events in the second embodiment of the present invention;

[0041] Figure 4 This is a schematic diagram of program modules of the video abnormal event identification device of the present invention;

[0042] Figure 5 This is a schematic diagram of the hardware structure of a computer device according to the fourth embodiment of the computer device of the present invention. Detailed implementation manners

[0043] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Video anomaly detection is the core issue in the development of computer vision technology and a basic requirement for realizing intelligent monitoring and public safety. Affected by device performance, environmental conditions, algorithm defects and other factors, existing anomaly detection systems often have problems such as missed reports and false reports, seriously reducing the reliability and practicality of the monitoring system. The task of video anomaly detection is to quickly identify abnormal events that occur in surveillance videos and achieve real-time detection while ensuring accuracy to meet the actual application requirements. Therefore, it is of great significance to study efficient and reliable video anomaly detection methods.

[0045] The difficulties of video anomaly detection are mainly reflected in three aspects.

[0046] First, the complexity and diversity of abnormal events. There are significant differences in the definitions and manifestation forms of abnormal behaviors in different scenarios, and a single detection model is difficult to adapt to various actual application scenarios;

[0047] Second, the expression and fusion problem of multi-modal features. In addition to visual information, video data also contains multi-modal information such as audio and text. How to effectively utilize these complementary information to improve the detection performance is a challenge;

[0048] Third, the temporal dependence of feature representation. Abnormal events often have a complex temporal evolution process, and it is necessary to consider both local and global temporal relationships.

[0049] Regarding the first difficulty, most existing methods only rely on visual features for detection, such as using deep networks like I3D or C3D to extract spatio-temporal features. These methods perform well in dealing with anomalies with obvious visual patterns, but have limited ability to recognize abnormal behaviors in complex scenarios. Some methods attempt to introduce high-level semantic features, but still have difficulty in comprehensively understanding the scene content and behavior semantics.

[0050] Regarding the second difficulty, some studies have started to explore multi-modal learning methods, mainly focusing on the interactive learning between the video and audio modalities. However, these methods often ignore other modality information such as text, and there is a semantic gap between different modality features, resulting in unsatisfactory feature fusion effects.

[0051] Regarding the third difficulty, most existing methods adopt a fixed-scale temporal modeling approach, such as feature extraction based on sliding windows or sequence modeling of recurrent neural networks. These methods are difficult to capture abnormal patterns at different time scales simultaneously and have low computational efficiency.

[0052] The latest research shows that the multi-scale temporal modeling method based on the attention mechanism can better handle complex temporal dependencies and improve the accuracy and efficiency of detection.

[0053] In summary, there are still many deficiencies in the existing technologies when dealing with video anomaly detection in complex scenarios, and there is an urgent need to develop new technical solutions to solve these problems. The present invention proposes a video anomaly detection method based on a multi-modal temporal network. By fusing video, audio, and text modality information and introducing innovative feature alignment and multi-scale temporal modeling mechanisms, accurate detection of complex abnormal events is achieved.

[0054] Aiming at the above defects or improvement requirements of the existing technologies, the present invention provides a video anomaly detection method based on a multi-modal temporal network, aiming to quickly and accurately detect abnormal events in surveillance videos. This anomaly detection method converts visual content into text descriptions through a video caption generation network, introduces a multi-modal temporal scale network to adaptively fuse different modality information, and uses a modality consistency module to dynamically adjust the contributions and feature alignment of each modality, thereby solving the technical problems of video anomaly detection in complex scenarios.

[0055] Generally speaking, compared with the existing technologies, the above technical solutions conceived by the present invention have the following beneficial effects:

[0056] High detection accuracy. The invention comprehensively understands the surveillance scene from three dimensions: vision, audio, and text, ensures the semantic correspondence relationship of feature representations through the modality consistency module, and enhances the global correlation using the multi-head attention mechanism, significantly improving the accuracy of anomaly detection.

[0057] Strong robustness. The introduced modal consistency module in this invention effectively addresses the semantic differences between different modalities, enhancing the model's adaptability to complex scenarios and noise interference. Meanwhile, the complementarity of multi-modal features also improves the generalization performance of the model.

[0058] Good generality. The video caption generation network and multi-modal feature extraction module adopted in this invention have good transfer capabilities and are applicable to anomaly detection tasks in different scenarios. The design of the feature alignment and fusion mechanism also ensures the applicability of the method in detecting various types of abnormal events.

[0059] Strong scalability. The framework structure proposed in this invention has good modular characteristics, and each functional module can be independently optimized and extended according to specific application requirements, facilitating further improvement of system performance.

[0060] Figure 1 It is the flowchart of the subtitle-driven multi-modal video anomaly detection method in the embodiments of this invention;

[0061] The following first explains and describes the technical terms of this invention:

[0062] Visual Features: Feature vectors representing visual content extracted from video frames through a deep neural network, used to capture the appearance and motion information of images.

[0063] Audio Features: Acoustic features extracted from the audio track of a video, including information such as spectrum, pitch, and volume, used to describe various attributes of sound.

[0064] Textual Features: Semantic feature vectors transformed from the text descriptions obtained through the video caption generation network, used to provide high-level semantic understanding of the scene.

[0065] Modal Consistency: An indicator measuring the semantic correspondence relationship between different modal features, calculated through metrics such as cosine similarity and Euclidean distance.

[0066] Multi-head Attention: An attention mechanism that maps queries, key-value pairs to outputs, and captures different types of dependencies through parallel calculations of multiple attention heads.

[0067] As Figure 1 shown, the video anomaly detection method based on a multi-modal temporal network in this invention includes the following steps:

[0068] (1) Feature extraction module

[0069] (1.1) Visual feature extraction: The I3D network pre-trained on the Kinetics-400 dataset is used to extract feature vectors at a sampling rate of 24 frames per second and a sliding window of 16 frames.

[0070] (1.2) Audio feature extraction: The VGGish network is used to process 960ms long audio segments, and feature vectors are extracted after calculating the log mel spectrogram.

[0071] (1.3) Text feature extraction: The SwinBERT model is used to generate video segment descriptions, and then converted into feature vectors through the SimCSE model.

[0072] (2) Modal consistency module

[0073] (2.1) Feature alignment: Project different modal features into a common feature space through learnable projection matrices:

[0074]

[0075] Among them, W v , W a , W t are projection matrices, and b v , b a , b t are bias terms.

[0076] (2.2) Consistency loss calculation:

[0077]

[0078] Among them, T is the sequence length, x i and y i are feature vectors at the i-th time step.

[0079] (3) Multi-scale temporal network

[0080] (3.1) Temporal feature enhancement: For each modality m∈{v,a,t}, process the temporal information through one-dimensional convolution operations and non-local modules:

[0081] α(c out ) = softmax((θ(c out )) T ·φ(c out ))

[0082] n(c out ) = c out + Conv1D(α(c out )·g(c out ))

[0083] (3.2) Feature fusion:

[0084]

[0085] (3.3) Global correlation modeling: Applying the multi-head self-attention mechanism

[0086]

[0087] f fusion = att(h f , h f , h f )

[0088] (4) Training and optimization

[0089] (4.1) Loss function: Combining classification loss, modality consistency loss, and contrastive learning loss:

[0090] L = L cls + λ MCL · L MCL + λ consist · L MC

[0091] (4.2) Parameter update: Using the Adam optimizer with an initial learning rate of 4e-4 and adjusting the learning rate using cosine annealing.

[0092] The method of the present invention has been verified on four standard datasets (XD-Violence, UCF-Crime, ShanghaiTech, and UCSD-Ped2). The experimental results show that the method is superior to the existing technologies in terms of detection accuracy and generalization performance.

[0093] Now, an embodiment is provided as follows:

[0094] Embodiment 1:

[0095] Please refer to Figure 2 , a method for identifying video abnormal events in this embodiment, including:

[0096] S101: Extract features from the target video to obtain a visual feature, an audio feature, and a text feature.

[0097] S102: Perform modality consistency processing on the visual feature, the audio feature, and the text feature to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature.

[0098] S103: Invoke a pre-set mature student model to perform multi-scale temporal processing on the visual consistency feature, the audio consistency feature, and the text consistency feature, obtain at least one video fusion temporal feature and its global weight value, and obtain an outlier according to the video fusion temporal feature and its global weight value; wherein, the outlier reflects the probability of abnormal events occurring in the visual feature, audio feature, and text feature of the target video.

[0099] In this example, by extracting features from the target video, a visual feature, an audio feature, and a text feature are obtained, realizing the extraction of information in multiple modalities in the target video, so as to facilitate subsequent analysis of visual content, sound features, and text descriptions respectively, to ensure the comprehensiveness and accuracy of abnormal event recognition.

[0100] Through modal consistency processing, the distributions of modal features are aligned, the consistency features are matched at the semantic level (such as the association between visual actions and text descriptions), the temporal features of different modalities are synchronized, and temporal consistency features are generated to ensure that information equally participates in subsequent calculations; at the same time, the semantic differences between different modalities are effectively processed, improving the model's adaptability to complex scenarios and noise interference. At the same time, the complementarity of multi-modal features also enhances the generalization performance of the model.

[0101] Comprehensively understand the monitoring scenario from three dimensions of vision, audio, and text, ensure the semantic correspondence relationship of feature representations through modal consistency processing, and enhance the global correlation by using multi-scale temporal processing and the multi-head attention mechanism added in the multi-scale temporal processing, significantly improving the accuracy of anomaly detection; moreover, the design of the feature alignment based on modal consistency processing and the fusion mechanism of visual features, audio features, and text features also ensures the applicability of the method in the detection of various abnormal events.

[0102] Embodiment Two:

[0103] This embodiment is a specific application scenario of the above Embodiment One. Through this embodiment, the method provided by the present invention can be more clearly and specifically elaborated.

[0104] Figure 3 It is a specific method flowchart of a video abnormal event recognition method provided by an embodiment of the present invention. The method specifically includes steps S200 to S203. A video abnormal event recognition method includes:

[0105] S200: Train a preset initial student model through at least one preset training video and a preset loss function to obtain a mature student model; wherein, there is one training outlier in the training video, and the training outlier characterizes one or several of visual content, sound features, and text descriptions, and records at least one abnormal event; the loss function includes classification loss, contrastive learning loss, and modality consistency loss.

[0106] In this example, collect training videos containing abnormal events to ensure that the videos contain visual content, sound features, and text descriptions (such as video titles, subtitles, etc.). Each training video should be labeled with at least one training outlier to clarify the location and type of the abnormal event in the video.

[0107] Visual content: Use computer vision techniques (such as OpenCV, YOLO, etc.) to extract video frames and key visual features (such as object detection, action recognition, etc.).

[0108] Sound features: Use audio processing techniques (such as Librosa, MFCC, etc.) to extract audio features (such as pitch, volume, speech rate, etc.).

[0109] Text description: Extract text features from text information such as video titles and subtitles (such as using models like BERT, Word2Vec, etc. for text embedding).

[0110] Annotate the extracted multi-modal data to ensure that each outlier has an annotation of the corresponding abnormal event in the visual, sound, and text modalities, so that the initial student model should be able to process inputs in the three modalities of vision, sound, and text and output abnormal recognition results.

[0111] Optionally, adopt the batch training method, and use a batch of training data for each update. Set the number of training epochs and the number of batches per epoch. Use the validation set to monitor the model performance to prevent overfitting. Use the test set to evaluate the abnormal recognition performance of the model, and calculate the AUC (area under the ROC curve) and average precision; the AUC is the area under the Receiver Operating Characteristic Curve (ROC curve), which is used to measure the overall performance of the classification model at different thresholds; the average precision is the area under the Precision-Recall (PR) curve, which measures the performance of the model in the ranking task. Analyze the performance of the model in different modalities to ensure that the model can make full use of multi-modal information. Adjust the model parameters, loss function weights, or optimizer configuration according to the evaluation results. Try different model architectures or fusion strategies to improve the model performance.

[0112] Meanwhile, the framework structure of the mature aggregation network proposed in this application has good modular characteristics, and each functional module can be independently optimized and expanded according to specific application requirements, which is convenient for further improving the system performance.

[0113] Specifically, the initial student model belongs to an anomaly aggregation model, and the anomaly aggregation model also has a mature teacher model, which is used to train the initial student model.

[0114] The anomaly recognition model is a knowledge-enhanced complementary aggregation network, which aims to enhance the model's ability by introducing external knowledge, and through the knowledge transfer between the mature teacher model and the initial student model, to train the initial student model and improve the training efficiency and model performance of the initial student model.

[0115] The mature teacher model is a large, complex and superior pre-trained model trained with a large number of training videos. Its role is to provide knowledge or guidance to help the student model learn. In self-distillation, the teacher model guides the student model through its prediction results (soft labels) or intermediate layer features. The initial student model is a smaller and simpler model. Its goal is to learn the knowledge of the teacher model to achieve or approach the performance of the teacher model while maintaining a small scale. In the training stage, pre-trained weights or randomly initialize the model parameters. Ensure that the model architecture can support multi-modal input and the calculation of composite loss functions. The teacher model processes the input data and generates prediction results or intermediate layer features, which are used as the training targets of the student model. During self-distillation, the student model not only learns the true labels (hard labels), but also learns the soft labels (i.e., prediction probability distributions) provided by the teacher model. This approach helps the student model capture the complex patterns and generalization ability in the teacher model.

[0116] Furthermore, training the preset initial student model through at least one preset training video and a preset loss function to obtain a mature student model, including:

[0117] S01: Extract the training visual information in the training video, perform feature extraction on the training visual information according to the preset visual sampling rate to obtain at least one training visual temporal feature, and integrate at least one of the training visual temporal features to obtain a training visual feature;

[0118] S02: Extract the training audio information in the training video, perform feature extraction on the training audio information according to the preset audio sampling rate to obtain at least one training audio temporal feature, and integrate at least one of the training audio temporal features to obtain a training audio feature;

[0119] S03: Obtain training visual text information by performing visual description on the training visual information, and / or obtain training audio text information by performing text conversion on the training audio information, and integrate the training visual text information and / or the training audio text information to obtain the training text information of the training video;

[0120] S04: Extract features from the training text information according to a preset text sampling rate to obtain at least one training text time series feature, and integrate at least one of the training text time series features to obtain a training text feature.

[0121] S05: Enable the training process, where the training process includes:

[0122] Input the training visual feature, the training audio feature, and the training text feature into the initial student model in the anomaly recognition model, and input the intermediate data generated by the mature teacher model into the initial student model;

[0123] Call a preset optimizer to assign a learning rate to each parameter in the initial student model, and take reducing the loss function as the training objective, and train the initial student model through the training visual feature, the training audio feature, the training text feature, and the intermediate data to obtain an intermediate model; where the intermediate data is the outlier generated by the mature teacher model according to the training visual feature, or the intermediate layer data in the process of the mature teacher model generating the outlier.

[0124] Exemplarily, use the Adam optimizer, with an initial learning rate of 4e-4, and use cosine annealing to adjust the learning rate. The Adam (Adaptive Moment Estimation) optimizer is an adaptive learning rate optimization algorithm based on gradient descent, which is widely used in deep learning. It combines the advantages of the Momentum and RMSProp optimization algorithms, and can dynamically adjust the learning rate of each parameter, so as to converge to the minimum value of the loss function more effectively. Momentum is the exponential weighted average of the previous gradients, which accelerates the gradient descent and reduces the oscillation. RMSProp: By adaptively adjusting the learning rate of each parameter, different parameters can have different update steps during the training process.

[0125] The initial learning rate refers to the learning rate value set at the beginning of the training. The learning rate determines the step size of parameter update and is an important hyperparameter that affects the convergence speed and stability of the model. 4e-4 is 0.0004.

[0126] Cosine Annealing is a learning rate scheduling strategy used to dynamically adjust the learning rate during training. It is based on the shape of the cosine function, causing the learning rate to gradually decrease from an initial value to a minimum value within a certain period, and then potentially increase again (if multiple periods are performed).

[0127] The loss function includes classification loss, contrastive learning loss, and modality consistency loss;

[0128] L = L cls + λ MCL ·L MCL + λ consist ·L MC

[0129] Where L is the loss function, L cls is the classification loss function; L MCL is the contrastive learning loss function, and λ MCL is the contrastive learning parameter, which can be set as needed; L MC is the modality consistency loss function, and λ consist is the modality consistency parameter, which can be set as needed.

[0130] L cls is the classification loss function, which is the difference between the outliers output by the initial teacher model and the outliers output by the initial student model;

[0131] L MCL is the contrastive learning loss, which is the difference between the weight value vector in the fully connected layer called by the initial teacher model to output outliers and the weight value vector in the fully connected layer called by the initial student model to output outliers;

[0132] L MC is the modality consistency loss function, including: the degree of difference between the training visual temporal projection features and the training audio temporal projection features, and the degree of difference between the training audio temporal projection features and the training text temporal projection features.

[0133] S06: Repeat the training process. When the loss result of the loss function reaches below a preset loss threshold, set the obtained intermediate model as the anomaly recognition model.

[0134] In this step, the loss threshold is a parameter set to avoid excessive model training time or model overfitting. When the loss result is less than the loss threshold, it indicates that the intermediate model is a qualified model, and set the intermediate model as the anomaly recognition model. The anomaly recognition model includes a mature teacher model and a mature student model.

[0135] It should be noted that the anomaly recognition model obtained in this embodiment was verified on standard datasets (XD-Violence, UCF-Crime, ShanghaiTech, and UCSD-Ped2), and the experimental results show that this method is superior to the prior art in terms of detection accuracy and generalization performance.

[0136] Optionally, for the classification loss function: use Cross-Entropy Loss or Binary Cross-Entropy Loss to measure the difference between the model's prediction results and the true labels. The classification loss function should act on the final output layer of the model to optimize the anomaly recognition task.

[0137] Modal consistency loss: Design a modal consistency loss function to ensure that the feature representations between different modalities (vision, sound, text) are consistent under abnormal events. For example, cosine similarity or Euclidean distance can be used to measure the similarity between different modality features and be used as part of the loss function. The modal consistency loss function should act on the intermediate layer of the model to enhance the correlation between multi-modal features.

[0138] Contrastive learning loss function: Use a contrastive learning loss function (such as InfoNCE Loss) to enhance the model's ability to distinguish abnormal events. By maximizing the feature differences between abnormal events and normal situations, the anomaly recognition performance of the model is improved. The contrastive learning loss function should act on the feature extraction layer of the model to learn more discriminative feature representations.

[0139] S201: Extract features from the target video to obtain a visual feature, an audio feature, and a textual feature.

[0140] In this step, the visual feature (VisualFeatures): is a feature vector that represents visual content extracted from video frames through a deep neural network, used to capture visual appearance and motion information.

[0141] The audio feature (Audio Features): is an acoustic feature extracted from the audio track of the video, including information such as spectrum, pitch, and volume, used to describe various attributes of the sound.

[0142] The textual feature (Textual Features): is a semantic feature vector transformed from the text description obtained through video captions and an audio conversion generation network, used to provide high-level semantic understanding of the scene.

[0143] In this step, by extracting features from the target video, a visual feature, an audio feature, and a text feature are obtained, realizing the extraction of information in multiple modalities in the target video, so as to facilitate subsequent analysis of visual content, sound features, and text descriptions respectively, to ensure the comprehensiveness and accuracy of abnormal event recognition.

[0144] In a preferred embodiment, extracting features from the target video to obtain a visual feature, an audio feature, and a text feature includes:

[0145] S11: Extract the visual information in the target video, extract features from the visual information according to a preset visual sampling rate to obtain at least one visual temporal feature, and integrate at least one of the visual temporal features to obtain a visual feature;

[0146] S12: Extract the audio information in the target video, extract features from the audio information according to a preset audio sampling rate to obtain at least one audio temporal feature, and integrate at least one of the audio temporal features to obtain an audio feature;

[0147] S13: Convert the visual information into visual text information through a preset video caption generation network, and / or perform text conversion on the audio information to obtain audio text information, and integrate the visual text information and / or the audio text information to obtain the text information of the target video;

[0148] S14: Extract features from the text information according to a preset text sampling rate to obtain at least one text temporal feature, and integrate at least one of the text temporal features to obtain a text feature.

[0149] Exemplarily, for visual feature extraction: Use the I3D network pre-trained on the Kinetics-400 dataset to extract feature vectors at a sampling rate of 24 frames per second and a sliding window of 16 frames.

[0150] For audio feature extraction: Use the VGGish network to process audio segments 960ms long, calculate the log mel spectrogram, and then extract feature vectors.

[0151] For text feature extraction: Use the SwinBERT model to generate video segment descriptions, and then convert them into feature vectors through the SimCSE model.

[0152] Among them, Kinetics-400 is a large-scale and high-quality dataset of YouTube video URLs, containing various human-centered actions. This dataset contains 400 human action categories, with at least 400 video clips for each action. Each clip lasts about 10 seconds and is taken from different YouTube videos. These actions are human-centered and cover a wide range of categories, including human-object interactions such as playing musical instruments, and human-human interactions such as handshakes.

[0153] The I3D (Inflated 3D ConvNet) network is a network often used in video processing to extract spatio-temporal features from videos and perform classification. The I3D network expands a 2D convolutional network (such as Inception-v1) into a 3D convolutional network by "inflating" the convolutional kernels and pooling kernels in the temporal dimension. This structure can capture both spatial and temporal information in videos.

[0154] VGGish is an audio feature extractor based on convolutional neural networks (CNNs), mainly used to convert audio signals into feature vectors with semantic meanings. VGGish adopts a similar structure of convolutional layers and pooling layers. It contains multiple convolutional layers and fully connected layers, and can extract high-level features in audio signals.

[0155] The Log-Mel Spectrogram is a method to convert audio signals into visual representations, often used in fields such as speech recognition, music information retrieval, and audio analysis. It first frames the audio signal and performs a fast Fourier transform (FFT) on each frame signal to obtain a spectrogram. Then, a set of Mel filters is used to perform weighted averaging on the spectrogram to simulate the non-linear perception of sound frequencies by the human ear. Finally, the logarithm of the values output by the filter bank is taken to obtain the Log-Mel Spectrogram.

[0156] The SwinBERT model is an end-to-end Transformer-based video captioning model for directly generating natural language descriptions from video frames. SwinBERT uses a video Transformer to encode spatio-temporal representations, which can adapt to variable-length video inputs. Different from previous methods, it does not require special design for different frame rates.

[0157] The SimCSE (Similarity Contrastive Estimation) model is a technique for enhancing the semantic understanding ability of pre-trained language models (such as BERT, RoBERTa, etc.). SimCSE performs unsupervised training by generating positive and negative sample pairs and uses contrastive learning to improve the performance of text similarity tasks. Specifically, it randomly perturbs the input sentences to generate positive sample pairs and maximizes the similarity between them by calculating the similarity score between two positive samples, while minimizing the similarity scores with all negative samples.

[0158] In this embodiment, the visual text information is the text obtained according to the captions generated in the target video and / or the text content describing the content shown in the visual image; the audio text information is used to display the audio information in the target video in text form. The adopted video caption generation network and multi-modal feature extraction module have good transfer capabilities and are applicable to anomaly detection tasks in different scenarios. Rapid and accurate detection of abnormal events in surveillance videos is achieved. This anomaly detection method converts visual content into text descriptions through a video caption generation network and introduces a multi-modal temporal scale network to adaptively fuse different modal information, thereby solving the technical problem of video anomaly detection in complex scenarios.

[0159] S202: Perform modal consistency processing on the visual feature, the audio feature, and the text feature to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature.

[0160] In this step, through modal consistency processing, the distributions of each modal feature are aligned, the consistency features are matched at the semantic level (such as the association between visual actions and text descriptions), the temporal features of different modalities are synchronized to generate temporal consistency features, ensuring that information equally participates in subsequent calculations; at the same time, the semantic differences between different modalities are effectively processed, improving the model's adaptability to complex scenarios and noise interference. At the same time, the complementarity of multi-modal features also enhances the generalization performance of the model.

[0161] Modal Consistency is an index used to measure the semantic correspondence between different modal features, which is calculated through metrics such as cosine similarity and Euclidean distance. Modal consistency processing is an index that dynamically measures the speech correspondence between visual features, audio features, and text features based on the Modal Consistency module, and adjusts the contributions and feature alignments of visual features, audio features, and text features to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature, effectively dealing with the semantic differences between different modalities, improving the model's adaptability to complex scenarios and noise interference. At the same time, the complementarity of multi-modal features also enhances the generalization performance of the model. Among them, the index for measuring the semantic correspondence between different modal features is obtained by calculating through metrics such as cosine similarity and Euclidean distance.

[0162] Therefore, features of different modalities (such as CNN features for vision, MFCC for audio, and BERT embeddings for text) usually have different numerical ranges, dimensions, and statistical distributions. Direct fusion will cause certain modalities to dominate the model's decision-making; and in multi-modal data, visual, audio, and text descriptions of the same event may be out of sync or semantically misaligned (such as the "barking dog" action in the video corresponding to the barking sound in the audio and the "dog barking" in the text); at the same time, in video or audio streams, there may be time offsets between visual frames, audio segments, and text descriptions (such as the text description lagging behind the visual event).

[0163] In response to this, in this step, through modal consistency processing, the distributions of each modal feature are aligned, the consistency features are matched at the semantic level (such as associating visual actions with text descriptions), the temporal features of different modalities are synchronized, and temporal consistency features are generated to ensure that information equally participates in subsequent calculations.

[0164] In a preferred embodiment, performing modal consistency processing on the visual feature, the audio feature, and the text feature to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature includes:

[0165] S21: Project the visual feature, the audio feature, and the text feature into a preset common feature space respectively to obtain a visual projection feature, an audio projection feature, and a text projection feature; wherein, there is at least one visual temporal projection feature in the visual projection feature; there is at least one audio temporal projection feature in the audio projection feature; and there is at least one text temporal projection feature in the text projection feature;

[0166] Exemplarily, the visual features, the audio features, and the text features are projected into a common feature space through a learnable visual matrix, audio matrix, and text matrix to obtain visual projection features, audio projection features, and text projection features:

[0167]

[0168] where, W v is the visual matrix, W a is the audio matrix, W t is the text matrix, b v is the visual bias term, b a is the audio bias term, b t is the text bias term; is the visual projection feature, f v is the visual feature; is the audio projection feature, f a is the audio feature; is the text projection feature, f t is the text feature.

[0169] S22: Calculate the differences between the visual temporal projection features, the audio temporal projection features, and the text temporal projections at each time step in sequence to obtain consistency loss information; wherein, the time step is a preset time span, and there is at least one visual temporal projection feature, at least one audio temporal projection feature, and at least one text temporal projection at one time step; the consistency loss information includes first difference data and second difference data; the first difference data represents the degree of difference between the visual temporal projection feature and the audio temporal projection feature; the second difference data represents the degree of difference between the audio temporal projection feature and the text temporal projection feature.

[0170] Exemplarily, the calculation formula of the consistency loss information is:

[0171]

[0172] where, T is the sequence length of the visual projection feature, the audio projection feature, or the text projection feature; x is the first feature, y is the second feature; x i and y i are the visual temporal projection feature, the audio temporal projection feature, or the text temporal projection feature at the i-th time step. When the first feature is the visual temporal projection feature and the second feature is the audio temporal projection feature, L consist (x, y) is the first difference data; when the first feature is the audio temporal projection feature and the second feature is the text temporal projection feature, L consist (x, y) is the second difference data.

[0173] S23: With the goal of minimizing the consistency loss information, adjust the visual temporal projection features, audio temporal projection features, and text temporal projection features at each of the time steps respectively to obtain the visual temporal consistency features, audio temporal consistency features, and text temporal consistency features at each of the time steps.

[0174] In this step, by taking the goal of minimizing the consistency loss information and adjusting the visual temporal projection features, audio temporal projection features, and text temporal projection features at each of the time steps respectively, the first difference data and the second difference data are minimized, so that the information conveyed by the visual temporal consistency features, audio temporal consistency features, and text temporal consistency features obtained at each time step is kept consistent, facilitating the subsequent accurate identification of abnormal events.

[0175] S24: Aggregate the visual temporal consistency features, audio temporal consistency features, and text temporal consistency features at each of the time steps respectively to obtain a visual consistency feature, an audio consistency feature, and a visual consistency feature respectively.

[0176] In this step, the visual temporal consistency features at each time step are concatenated to obtain a visual consistency feature; the audio temporal consistency features at each time step are concatenated to obtain an audio consistency feature; the text temporal consistency features at each time step are concatenated to obtain a text consistency feature.

[0177] S203: Invoke a pre - set mature student model to perform multi - scale temporal processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain at least one video fusion temporal feature and its global weight value, and obtain an outlier value according to the video fusion temporal feature and its global weight value; wherein, the outlier value reflects the probability of abnormal events occurring in the visual features, audio features, and text features of the target video.

[0178] In this step, comprehensively understand the monitoring scenario from three dimensions of vision, audio, and text, ensure the semantic correspondence relationship of feature representations through modal consistency processing, and enhance the global correlation by using multi - scale temporal processing and the multi - head attention mechanism added in the multi - scale temporal processing, significantly improving the accuracy of anomaly detection; moreover, the design of the feature alignment based on modal consistency processing and the fusion mechanism of visual features, audio features, and text features also ensures the applicability of the method in the detection of various abnormal events.

[0179] In a preferred embodiment, invoking a pre - set mature student model to perform multi - scale temporal processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain at least one video fusion temporal feature and its global weight value includes:

[0180] S31: The mature student model performs enhancement processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain a visual enhancement feature, an audio enhancement feature, and a text enhancement feature. Among them, the enhancement processing includes one-dimensional convolution operation and non-local module processing. The one-dimensional convolution operation is used to identify the short-term dependence relationships between each visual temporal consistency feature and other visual temporal consistency features in the visual consistency feature, between each audio temporal consistency feature and other audio temporal consistency features in the audio consistency feature, and between each text temporal consistency feature and other text temporal consistency features in the text consistency feature. The non-local module processing is used to identify the long-term dependence relationships between each visual temporal consistency feature and other visual temporal consistency features in the visual consistency feature, between each audio temporal consistency feature and other audio temporal consistency features in the audio consistency feature, and between each text temporal consistency feature and other text temporal consistency features in the text consistency feature.

[0181] The visual enhancement feature records the long-term and short-term dependence relationships between each visual temporal consistency feature and other visual temporal consistency features in the visual consistency feature. The audio enhancement feature records the long-term and short-term dependence relationships between each audio temporal consistency feature and other audio temporal consistency features in the audio consistency feature. The text enhancement feature records the long-term and short-term dependence relationships between each text temporal consistency feature and other text temporal consistency features in the text consistency feature.

[0182] In this step, the long-term dependence relationship reflects the correlation degree between a visual temporal consistency feature and all visual temporal consistency features, between an audio temporal consistency feature and all audio temporal consistency features, and between a text temporal consistency feature and all text temporal consistency features.

[0183] The short-term dependence relationship reflects the correlation degree between a visual temporal consistency feature and adjacent visual temporal consistency features, between an audio temporal consistency feature and adjacent audio temporal consistency features, and between a text temporal consistency feature and adjacent text temporal consistency features.

[0184] Specifically, the mature student model performs enhancement processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain a visual enhancement feature, an audio enhancement feature, and a text enhancement feature, including:

[0185] S311: The mature student model performs non-local module processing on the visual consistency feature, the audio consistency feature, and the text consistency feature respectively to obtain a visual weight, an audio weight, and a text weight. Among them, the visual weight includes the visual correlation weight values between each visual temporal consistency feature in the visual consistency feature and all visual temporal consistency features; the audio weight includes the audio correlation weight values between each audio temporal consistency feature in the audio consistency feature and all audio temporal consistency features; the text weight includes the text correlation weight values between each text temporal consistency feature in the text consistency feature and all text temporal consistency features. The visual correlation weight value reflects the long-term dependence relationship between a visual temporal consistency feature and all visual temporal consistency features; the audio correlation weight value reflects the long-term dependence relationship between an audio temporal consistency feature and all audio temporal consistency features; the text correlation weight value reflects the long-term dependence relationship between a text temporal consistency feature and all text temporal consistency features.

[0186] Exemplarily, for each modality m ∈ {v, a, t}, the visual correlation weight value, the audio correlation weight value, and the text correlation weight value are calculated through the following formula, and the visual weight, the audio weight, and the text weight are obtained.

[0187] α(c out ) = softmax((θ(c out )) T ·φ(c out ));

[0188] Among them, c out is any one of the visual consistency feature, the audio consistency feature, and the text consistency feature; θ(cout) and φ(cout) are two independent linear transformations that map the input to different feature spaces; θ(cout)T·φ(cout) is to calculate the dot product (similarity) of two vectors to measure the correlation degree between each visual sequence consistency feature in the visual consistency feature and each visual sequence consistency feature in the visual consistency feature, the correlation degree between each audio sequence consistency feature in the audio consistency feature and each audio sequence consistency feature in the audio consistency feature, and the correlation degree between each text sequence consistency feature in the text consistency feature and each text sequence consistency feature in the text consistency feature; α(c out ) is any one of the visual weight, the audio weight, and the text weight; softmax maps each visual correlation weight value, audio correlation weight value, and text correlation weight value of the input vector to the interval (0, 1) and ensures that the sum of all output values is 1, thereby forming a probability distribution.

[0189] S312: The mature student model performs one-dimensional convolutional operations on the visual consistency feature, the audio consistency feature, and the text consistency feature according to the visual weight, the audio weight, and the text weight respectively, to obtain a visual dependence feature, an audio dependence feature, and a text dependence feature; wherein, the one-dimensional convolutional operation is used to determine the short-term dependence relationships of the visual temporal consistency features in the visual consistency feature, the audio temporal consistency features in the audio consistency feature, and the text temporal consistency features in the text consistency feature; the visual dependence feature records the visual consistency features of the long-term dependence relationship and the short-term dependence relationship of the visual temporal consistency features in the visual consistency feature; the audio dependence feature records the long-term dependence relationship and the short-term dependence relationship of the audio temporal consistency features in the audio consistency feature; the text dependence feature records the text consistency features of the long-term dependence relationship and the short-term dependence relationship of the text temporal consistency features in the text consistency feature.

[0190] S313: The mature student model obtains a visual enhancement feature according to the visual consistency feature and the visual weight feature, obtains an audio enhancement feature according to the audio consistency feature and the audio weight feature, and obtains a text enhancement feature according to the text consistency feature and the text weight feature.

[0191] Exemplarily, the visual enhancement feature, the audio enhancement feature, and the text enhancement feature are obtained through the following formula.

[0192] n(c out ) = c out +Conv1D(α(c out )·g(c out ))

[0193] wherein, c out is any one of the visual consistency feature, the audio consistency feature, and the text consistency feature; Conv1D is the one-dimensional convolutional operation; α(c out ) is any one of the visual weight, the audio weight, and the text weight; g(c out ) is one of the visual transformation feature, the audio transformation feature, and the text transformation feature obtained by performing a non-linear transformation on the visual consistency feature, the audio consistency feature, and the text consistency feature; the non-linear transformation is one of the power function transformation, the logarithmic function transformation, the exponential function transformation, and the activation function transformation; n(c out) is any one of a visual enhancement feature, an audio enhancement feature, and a text enhancement feature; among them, power function transformation: such as y = ax2+bx+c, fitting a curve relationship through high-order terms. Logarithmic function transformation: such as y = log(x), used to compress the dynamic range of data. Exponential function transformation: such as y = ex, used to enhance the contrast of data. Activation function: In a neural network, activation functions such as ReLU, Sigmoid, and tanh are all examples of non-linear transformations.

[0194] S32: The mature student model fuses the visual consistency feature, the audio consistency feature, and the text consistency feature into a video fusion feature according to the visual enhancement feature, the audio enhancement feature, and the text enhancement feature; among them, there is at least one video fusion timing feature in the video fusion feature, and one video fusion timing feature characterizes the visual content, sound feature, and text description of the target video at a time step.

[0195] Specifically, the mature student model fuses the visual consistency feature, the audio consistency feature, and the text consistency feature into a video fusion feature according to the visual enhancement feature, the audio enhancement feature, and the text enhancement feature, including:

[0196] S321: Perform residual connection on the visual consistency feature according to the visual enhancement feature to obtain a visual residual feature; perform residual connection on the audio consistency feature according to the audio enhancement feature to obtain an audio residual feature; perform residual connection on the text consistency feature according to the text enhancement feature to obtain a text residual feature.

[0197] Exemplarily, perform residual connection through the following formula, and obtain a video residual feature, an audio residual feature, and a text residual feature.

[0198]

[0199] Among them, is any one of a visual residual feature, an audio residual feature, and a text residual feature; f m is any one of a visual consistency feature, an audio consistency feature, and a text consistency feature; n(c out ) is any one of a visual enhancement feature, an audio enhancement feature, and a text enhancement feature; Conv1D5 represents a one-dimensional convolution operation, and extracts local features among the visual enhancement feature, the audio enhancement feature, and the text enhancement feature when extracting in the 5th convolution layer;

[0200] S322: Perform fusion processing on the visual residual feature, the audio residual feature, and the text residual feature to obtain a video fusion feature.

[0201] Exemplarily, the following formula is used for the fusion process to obtain video fusion features.

[0202]

[0203] where h f is the video fusion feature; is the visual residual feature; is the audio residual feature; is the text residual feature; w f is the fusion weight value, which can be set as needed; b f is the fusion bias value, which can be set as needed; Concat represents the "concatenation" operation, which is used to fuse multiple features into one.

[0204] S33: The mature student model identifies the correlation between each video fusion temporal feature and other video fusion temporal features in the video fusion features based on the multi-head attention mechanism, and obtains the global weight value of each video fusion temporal feature; wherein, the global weight value reflects the degree of content association and context dependence between a video fusion temporal feature and all video fusion temporal features.

[0205] In this step, the monitoring scenario is comprehensively understood from three dimensions of vision, audio, and text. The semantic correspondence relationship of the feature representation is ensured through the modality consistency module, and the global correlation is enhanced by using the multi-head attention mechanism, significantly improving the accuracy of anomaly detection. Multi-head Attention is an attention mechanism that maps queries, key-value pairs to outputs, and captures different types of correlations through parallel calculations of multiple attention heads.

[0206] The degree of content association refers to the direct connection or similarity in semantics, theme, or information content between elements at different positions in the input sequence. This degree of association reflects the tightness of the relationship between elements in terms of meaning and is an important aspect of understanding the relationship between elements in sequence data (such as text, images, etc.); it includes:

[0207] Semantic association, which refers to the direct connection between elements in meaning or concept. For example, in text, there is a semantic association between words, phrases, or sentences with the same or similar meanings. Example: In the sentence "Cats like to eat fish", there is a semantic association between "cat" and "fish" because "fish" is the food that "cats" like. Function: In the multi-head attention mechanism, capturing semantic associations helps the model understand the internal relationships between elements in the sequence, thereby more accurately generating outputs or performing tasks such as classification and prediction.

[0208] Thematic relevance. Thematic relevance refers to the correlation of elements in terms of themes or topics. For example, in an article about environmental protection, there is thematic relevance among all sentences, paragraphs, or words related to environmental protection. Example: There is thematic relevance between the sentences "Reducing plastic use is an important measure for environmental protection" and "Promoting degradable materials contributes to environmental protection" in the article because they both revolve around the theme of environmental protection. Capturing thematic relevance helps the model grasp the overall theme and core idea of sequential data, improving the model's understanding and processing capabilities of sequential data.

[0209] Information content relevance. Information content relevance refers to the direct connection of elements in terms of conveying information or expressing viewpoints. For example, in news reports, there is information content relevance among multiple reports or comments about the same event. Example: There is information content relevance among the news report, expert analysis, and public reaction regarding an earthquake because they all revolve around this earthquake. Function: Capturing information content relevance helps the model integrate relevant information in sequential data, improving the model's comprehensive analysis and processing capabilities of sequential data.

[0210] The degree of dependence in context refers to the fact that the meaning or importance of an element in the input sequence depends on its surrounding elements or the context environment of the entire sequence. This degree of dependence reflects the influence of the position of the element in the sequence and its surrounding environment on its meaning; it includes:

[0211] Order dependence. Order dependence refers to the influence of the position of an element in the sequence on its meaning. For example, in text, the order of words determines the grammatical structure and meaning of a sentence. Example: The sentences "I like cats" and "Cats like me" contain the same words, but due to the different word orders, the meanings of the sentences are completely different. Function: Capturing order dependence helps the model understand the grammatical structure and logical order of sequential data, thus generating more accurate outputs or performing understanding tasks more accurately.

[0212] Context dependence. Context dependence refers to the fact that the meaning of an element depends on the context or context environment in which it is located. For example, in text, the meaning of a word may vary depending on its position in the sentence and the content of the surrounding text. Example: The word "bank" refers to a financial institution in the sentence "I withdraw money from the bank", while it may refer to the riverbank in "I take a walk on the riverbank". This change in meaning reflects context dependence. Function: Capturing context dependence helps the model accurately understand the meaning of elements based on the context environment, improving the model's semantic understanding capabilities of sequential data.

[0213] Structural Dependency, which refers to the influence of the structural relationship of elements in a sequence on their meaning. For example, in a text, the hierarchical relationship between sentences and the logical relationship between paragraphs all reflect structural dependency. Example: There is a structural dependency relationship among the introduction, body, and conclusion of an article, which together constitute the overall framework and logical structure of the article. Function: Capturing structural dependency helps the model understand the overall structure and logical relationship of sequential data, and improve the model's comprehensive analysis and processing ability of sequential data.

[0214] Exemplarily, the dependency relationship of each video fusion time sequence is identified through the following formula to obtain the global weight value of each video fusion time sequence.

[0215] f fusion = att(h f , h f , h f );

[0216] Where h f is the video fusion feature; att is the multi-head attention mechanism; f fusion records the global weight value of each said video fusion time sequence;

[0217] The function of att as the multi-head attention mechanism is as follows.

[0218]

[0219] q is the query vector, and the query vector q represents the information that needs to be focused on or queried currently. It usually comes from the result of the input data after linear transformation.

[0220] k is the key vector, and the key vector k represents the features or information at each position in the input sequence. It is also obtained from the input data through linear transformation.

[0221] v is the value vector, and the value vector v represents the actual information or feature representation at each position in the input sequence. In the attention calculation, it is the object to be weighted and summed.

[0222] qk T represents the dot product (or matrix multiplication, depending on the specific implementation) of the query vector q and the key vector k. This operation calculates the similarity between q and k.

[0223] is the scaling factor, d m is the dimension of the key vector or query vector (usually d_k in multi-head attention). The scaling factor is used to prevent the dot product result from being too large, which may lead to gradient disappearance or explosion. By dividing by the dot product result can be within a more reasonable range, thereby improving the stability and effect of training.

[0224] The softmax function is a normalization function used to convert input values into a probability distribution. In the attention mechanism, the softmax function converts the similarity scores into attention weights, which represent the importance of each position to the current query position.

[0225] The scaled dot - product similarity is calculated and converted into a probability distribution to generate attention weights, which are used to weight - sum the value vectors.

[0226] It is to weight - sum the value vector v using the attention weights. The attention output is obtained, which is the result of the current query position fused according to the information of other positions in the input sequence.

[0227] Therefore, by separately converting the video fusion features into query vectors, key vectors, and value vectors, and jointly applying them to the attention calculation process, it helps the model capture the important information and relationships in the input sequence.

[0228] The query vector represents the information that the current video fusion temporal feature needs to focus on or query. In the video fusion feature, the query vector of the video fusion temporal feature at each time step is calculated based on the input features at that time step. The query vector is used to calculate the similarity with the key vector to determine which position information is important for the current video fusion temporal feature.

[0229] The key vector represents the video fusion temporal features at each time step in the video fusion feature. The dot - product or similarity calculation is performed between the key vector and the query vector to generate attention scores. These scores reflect the similarity or correlation between the query vector and the key vector.

[0230] The value vector represents the video fusion temporal features at each time step in the video fusion feature. According to the attention scores, the value vectors are weighted - summed to obtain the attention output. This output is the result of the current video fusion temporal feature fused according to the information of other positions in the input sequence, characterizing the correlation degree between each video fusion temporal feature and each video fusion temporal feature in the video fusion feature.

[0231] Therefore, in this embodiment, by fusing three - modality information of video, audio, and text, and introducing an innovative feature alignment and multi - scale temporal modeling mechanism, accurate detection of complex abnormal events is achieved.

[0232] In a preferred embodiment, obtaining the outlier according to the video fusion temporal feature and its global weight value includes:

[0233] The mature student model outputs the video fusion temporal feature and its global weight value to a preset fully - connected layer;

[0234] The mature student model calls the fully connected layer video fusion temporal features and their global weight values to obtain outliers.

[0235] Specifically, in the fully connected layer, the input features are weighted and summed: y = Wx + b; where y is the outlier, W is the weight matrix, x is the input feature (fusion temporal feature and weight value), and b is the bias term. The weighted sum result is passed through an activation function to obtain the final outlier.

[0236] Embodiment Three:

[0237] Please refer to Figure 4 , a video anomaly event recognition device 3 in this embodiment includes:

[0238] A feature extraction module 31, configured to extract features from the target video to obtain a visual feature, an audio feature, and a text feature;

[0239] A modality consistency module 32, configured to perform modality consistency processing on the visual feature, the audio feature, and the text feature to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature;

[0240] A multi-scale temporal module 33, configured to call a preset mature student model to perform multi-scale temporal processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain at least one video fusion temporal feature and its global weight value, and obtain an outlier according to the video fusion temporal feature and its global weight value; wherein, the outlier reflects the probability of an anomaly event occurring in the visual feature, audio feature, and text feature of the target video.

[0241] Optionally, the video anomaly event recognition device 1 further includes:

[0242] A training module 30, which trains a preset initial student model through at least one preset training video and a preset loss function to obtain a mature student model; wherein, there is a training outlier in the training video, and the training outlier characterizes that there is at least one anomaly event in one or several of visual content, sound features, and text descriptions; the loss function includes classification loss, contrastive learning loss, and modality consistency loss.

[0243] Embodiment Four:

[0244] To achieve the above object, the present invention further provides a computer device 5. The components of the video anomaly event recognition device in Embodiment 3 can be distributed in different computer devices. The computer device 5 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple application servers) that executes a program, etc. The computer device in this embodiment at least includes, but is not limited to, a memory 51 and a processor 52 that can be communicatively connected to each other through a system bus, as Figure 5 shown. It should be noted that Figure 5 only a computer device with components - is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0245] In this embodiment, the memory 51 (i.e., the readable storage medium) includes flash memory, a hard disk, a multimedia card, a card - type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read - only memory (ROM), an electrically erasable programmable read - only memory (EEPROM), a programmable read - only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 51 can be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 51 can also be an external storage device of the computer device, such as a plug - in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Of course, the memory 51 can also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory 51 is generally used to store the operating system and various application software installed in the computer device, such as the program code of the video anomaly event recognition device in Embodiment 3. In addition, the memory 51 can also be used to temporarily store various data that have been output or will be output.

[0246] In some embodiments, the processor 52 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data - processing chips. The processor 52 is generally used to control the overall operation of the computer device. In this embodiment, the processor 52 is used to run the program code stored in the memory 51 or process data, such as running the video anomaly event recognition device to implement the video anomaly event recognition methods in Embodiment 1 and Embodiment 2.

[0247] Embodiment 5:

[0248] To achieve the above object, the present invention further provides a computer-readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, a server, an App application store, etc., on which a computer program is stored, and when the program is executed by the processor 52, the corresponding functions are implemented. The computer-readable storage medium of this embodiment is used to store the computer program for implementing the video anomaly event recognition method, and when executed by the processor 52, it implements the video anomaly event recognition methods of Embodiment 1 and Embodiment 2.

[0249] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0250] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0251] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A method for identifying video abnormal events, characterized in that, Including: Performing feature extraction on the target video to obtain a visual feature, an audio feature, and a text feature; Performing modality consistency processing on the visual feature, the audio feature, and the text feature to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature; Invoking a pre-set mature student model to perform multi-scale temporal processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain at least one video fusion temporal feature and its global weight value, and obtaining an outlier according to the video fusion temporal feature and its global weight value; wherein, the outlier reflects the probability of abnormal events occurring in the visual feature, audio feature, and text feature of the target video.

2. The video anomaly event recognition method according to claim 1, wherein Performing feature extraction on the target video to obtain a visual feature, an audio feature, and a text feature, including: Extracting visual information in the target video, performing feature extraction on the visual information according to a pre-set visual sampling rate to obtain at least one visual temporal feature, and integrating at least one of the visual temporal features to obtain a visual feature; Extracting audio information in the target video, performing feature extraction on the audio information according to a pre-set audio sampling rate to obtain at least one audio temporal feature, and integrating at least one of the audio temporal features to obtain an audio feature; Converting the visual information into visual text information through a pre-set video subtitle generation network, and integrating the visual text information to obtain the text information of the target video; Performing feature extraction on the text information according to a pre-set text sampling rate to obtain at least one text temporal feature, and integrating at least one of the text temporal features to obtain a text feature.

3. The video abnormal event recognition method according to claim 1, wherein, Performing modality consistency processing on the visual feature, the audio feature, and the text feature to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature, including: Projecting the visual feature, the audio feature, and the text feature into a pre-set common feature space respectively to obtain a visual projection feature, an audio projection feature, and a text projection feature; wherein, there is at least one visual temporal projection feature in the visual projection feature; there is at least one audio temporal projection feature in the audio projection feature; there is at least one text temporal projection feature in the text projection feature; Calculating the differences between the visual temporal projection features, audio temporal projection features, and text temporal projections at each time step in sequence to obtain consistency loss information; wherein, the time step is a pre-set time span, and there is at least one visual temporal projection feature, at least one audio temporal projection feature, and at least one text temporal projection at one time step; the consistency loss information includes first difference data and second difference data; the first difference data represents the degree of difference between the visual temporal projection feature and the audio temporal projection feature; the second difference data represents the degree of difference between the audio temporal projection feature and the text temporal projection feature; With the goal of minimizing the consistency loss information, the visual temporal projection features, audio temporal projection features, and text temporal projection features at each time step are adjusted respectively to obtain the visual temporal consistency features, audio temporal consistency features, and text temporal consistency features at each time step; The visual temporal consistency features, audio temporal consistency features, and text temporal consistency features at each time step are aggregated respectively to obtain a visual consistency feature, an audio consistency feature, and a visual consistency feature.

4. The video anomaly event recognition method according to claim 1, characterized in that A pre-set mature student model is called to perform multi-scale temporal processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain at least one video fusion temporal feature and its global weight value, including: The mature student model performs enhancement processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain a visual enhancement feature, an audio enhancement feature, and a text enhancement feature; wherein, the enhancement processing includes one-dimensional convolution operation and non-local module processing; the one-dimensional convolution operation is used to identify the short-term dependence relationships between the visual temporal consistency features in the visual consistency feature and other visual temporal consistency features, the short-term dependence relationships between the audio temporal consistency features in the audio consistency feature and other audio temporal consistency features, and the short-term dependence relationships between the text temporal consistency features in the text consistency feature and other text temporal consistency features; the non-local module processing is used to identify the long-term dependence relationships between the visual temporal consistency features in the visual consistency feature and other visual temporal consistency features, the long-term dependence relationships between the audio temporal consistency features in the audio consistency feature and other audio temporal consistency features, and the long-term dependence relationships between the text temporal consistency features in the text consistency feature and other text temporal consistency features; The mature student model fuses the visual consistency feature, the audio consistency feature, and the text consistency feature into a video fusion feature according to the visual enhancement feature, the audio enhancement feature, and the text enhancement feature; wherein, the video fusion feature has at least one video fusion temporal feature, and one video fusion temporal feature represents the visual content, sound features, and text description of the target video at a time step; The mature student model identifies the correlation between each video fusion temporal feature in the video fusion feature and other video fusion temporal features based on the multi-head attention mechanism to obtain the global weight value of each video fusion temporal feature; wherein, the global weight value reflects the degree of content association and context dependence between a video fusion temporal feature and all video fusion temporal features.

5. The video abnormal event recognition method according to claim 1, wherein Outliers are obtained according to the video fusion temporal feature and its global weight value, including: The mature student model outputs the video fusion temporal feature and its global weight value to a pre-set fully connected layer; The mature student model calls the fully connected layer video fusion temporal features and their global weight values to obtain outliers.

6. The video abnormal event recognition method according to claim 1, wherein Before performing feature extraction on the target video to obtain a visual feature, an audio feature, and a text feature, the method further includes: Training a preset initial student model with at least one preset training video and a preset loss function to obtain a mature student model; wherein, there is a training outlier in the training video, and the training outlier characterizes that at least one abnormal event is recorded in one or several of visual content, sound features, and text descriptions; the loss function includes classification loss, contrastive learning loss, and modality consistency loss.

7. The video anomaly event recognition method according to claim 6, wherein The initial student model belongs to an anomaly aggregation model, and the anomaly aggregation model further has a mature teacher model, and the mature teacher model is used to train the initial student model.

8. A video abnormal event recognition device, characterized in that, It includes: A feature extraction module for performing feature extraction on the target video to obtain a visual feature, an audio feature, and a text feature; A modality consistency module for performing modality consistency processing on the visual feature, the audio feature, and the text feature to obtain a visual consistency feature, an audio consistency feature, and a text consistency feature; A multi-scale temporal module for calling the preset mature student model to perform multi-scale temporal processing on the visual consistency feature, the audio consistency feature, and the text consistency feature to obtain at least one video fusion temporal feature and its global weight value, and obtaining an outlier according to the video fusion temporal feature and its global weight value; wherein, the outlier reflects the probability of an abnormal event occurring in the visual feature, audio feature, and text feature of the target video.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor of the computer device executes the computer program, it implements the steps of the video abnormal event recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program stored in the readable storage medium is executed by the processor, it implements the steps of the video abnormal event recognition method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Illegal content auditing method and device based on multi-modal data, equipment and medium

    CN121278126A

  • Visual feature anomaly detection method and device based on audio frequency guidance

    CN121438211A