Video data behavior compliance detection method based on multi-modal analysis

Through the multimodal analysis method, SwinTransformer, DeepSpeech and BERT extract video data features, combined with generative adversarial networks and spatiotemporal causal reasoning, the problems of high error detection rate, low multimodal fusion efficiency and insufficient behavioral interpretation in video data behavior compliance detection are solved, and efficient and interpretable behavioral detection is achieved.

CN120451854AActive Publication Date: 2025-08-08DATA SPACE RES INST

Patent Information

Application Number
CN202510376721.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-08
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The existing video data behavior compliance detection technology has problems such as high error detection rate, low multimodal information fusion efficiency, insufficient long-tail abnormal behavior detection ability, and insufficient behavior interpretation ability.

Method used

The multimodal analysis method is adopted to extract image features through SwinTransformer with multi-scale feature fusion, combine the improved DeepSpeech and domain pre-trained BERT to extract audio and text features, and use the cross-modal self-attention mechanism to fusion, introduce the generative adversarial network to generate long-tail anomaly behavior samples, combine the spatiotemporal perturbation enhancement data, and finally perform interpretability analysis through multimodal adversarial scoring and spatiotemporal causal reasoning.

Benefits of technology

Significantly reduce the error detection rate, improve the efficiency of multimodal information fusion, enhance the ability to detect long-tail abnormal behaviors, provide detailed behavior explanations, and improve the transparency and trust of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451854A_ABST
    Figure CN120451854A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video data analysis, in particular to a video data behavior compliance detection method based on multi-modal analysis. According to the method, the false detection rate can be reduced, the classification model is optimized through a dynamic weight distribution mechanism, the detection accuracy is improved in a complex environment, the probability that a normal behavior is misjudged to be abnormal is remarkably reduced, meanwhile, the fusion efficiency of multi-modal information can be improved, and by improving the multi-modal feature extraction and deep fusion algorithm, the detection accuracy is improved. Potential relations among modalities such as images, audios and texts are fully mined, detection performance is improved, long-tail abnormal behavior detection capability can be enhanced, a generative adversarial network is designed to generate rare behavior samples, data insufficiency is made up, low-frequency and high-risk behavior recognition capability of the model is improved, behavior interpretation capability can be enhanced, and detection efficiency is improved. And an interpretable analysis framework is introduced, so that the model can explain the detection result in detail, and the transparency and the credibility in application are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video data analysis, and in particular to a method for detecting behavior compliance of video data based on multimodal analysis. Background Art

[0002] With the widespread use of video data, particularly in areas such as content review, public safety, and corporate regulation, the potential for non-compliant behavior within video data has posed serious social and legal challenges. These challenges place higher demands on automated detection technologies. Furthermore, the widespread application of deep learning models in video data behavioral compliance detection in recent years has significantly improved the processing efficiency of massive amounts of video data. These models can automatically extract multi-level feature representations from video frame sequences, providing more precise analytical tools for behavioral detection.

[0003] Existing technological advancements in video-based behavioral compliance detection focus on the integration and analysis of multimodal information. Video data not only contains visual information but also incorporates audio and textual information. The interaction between these modalities provides key clues for behavioral recognition. For example, image data can be used to analyze specific behaviors within the footage, audio can capture verbal and non-verbal cues, and textual information can supplement the semantics of the scene. By effectively integrating this multimodal information, technological advancements have laid the foundation for behavioral detection in complex scenarios.

[0004] However, existing technologies have some limitations and shortcomings in video data behavior compliance detection:

[0005] 1. High false positive rate: Current mainstream behavior detection models are prone to misjudging normal behavior as abnormal in complex environments (such as insufficient lighting, background interference, or multi-person interaction), which limits the reliability and practical application value of the detection results.

[0006] 2. Low efficiency of multimodal information fusion: Video data contains multiple modalities such as images, audio, and text. Existing methods perform poorly in extracting and fusing cross-modal features and fail to fully explore the potential correlations between modalities, thereby limiting detection performance.

[0007] 3. Inadequate detection of long-tail abnormal behaviors: Many abnormal behaviors occur very infrequently but often carry high risk. Existing models are poorly equipped to detect these long-tail behaviors, potentially leading to key risks being overlooked.

[0008] 4. Insufficient behavioral explanation capabilities: Existing methods typically only provide detection results but fail to offer interpretable analysis. This deficiency makes it difficult for detection results to serve as a direct basis for actual decision-making, affecting transparency and trust in applications.

[0009] This invention takes the fusion and analysis of multimodal information as its technical basis, and proposes an innovative solution based on full consideration of the advantages and disadvantages of existing technologies, providing important technical support for future video data behavior compliance detection. Summary of the Invention

[0010] The purpose of the present invention is to provide a video data behavior compliance detection method based on multimodal analysis to solve the problems of high false detection rate, low efficiency of multimodal information fusion, insufficient long-tail abnormal behavior detection capability and insufficient behavior interpretation capability in the existing technology proposed in the above background technology in the field of video data behavior compliance detection.

[0011] To achieve the above objectives, the present invention provides a method for detecting behavior compliance of video data based on multimodal analysis, comprising the following steps:

[0012] S1. Extract multimodal features from video data, including:

[0013] Image modality feature extraction: Generate image representation vectors based on SwinTransformer with multi-scale feature fusion;

[0014] Audio modal feature extraction: The audio is processed by short-time Fourier transform and denoising network and then fed into the improved DeepSpeech network;

[0015] Text modality feature extraction: Use the domain-pretrained BERT model to extract text features;

[0016] S2: Image, audio, and text features are spliced into a cross-modal sequence, fused through a cross-modal self-attention mechanism, and the weights of each modality are dynamically adjusted;

[0017] S3. Generate long-tail abnormal behavior samples using a generative adversarial network with target domain constraints and enhance the data with spatiotemporal perturbations.

[0018] S4. Perform behavioral compliance detection based on the fused multimodal features and output abnormality probability or category;

[0019] S5. Perform interpretable analysis of detection results through multimodal adversarial scoring and spatiotemporal causal reasoning.

[0020] As a further improvement of the present technical solution, the specific implementation of the image modality feature extraction in step S1 includes:

[0021] Extract multi-scale features from the output of each layer of SwinTransformer;

[0022] Through the trainable mapping matrix W fusion Splicing and dimensionality reduction are performed to generate the image representation vector I.

[0023] As a further improvement of the present technical solution, the audio modal feature extraction in step S1 includes:

[0024] Use short-time Fourier transform to generate multi-channel spectrum S c ;

[0025] Correct the spectrum noise through the denoising network and output the enhanced spectrum S' c ;

[0026] The audio feature vector A is extracted using the improved DeepSpeech network.

[0027] As a further improvement of the present technical solution, the text modality feature extraction in step S1 includes:

[0028] Pre-train the BERT model based on domain data and use masked language model loss optimization;

[0029] Extract the text feature vector T by re-pretraining BERT.

[0030] As a further improvement of this technical solution, the specific implementation of the cross-modal self-attention mechanism in step S2 includes: using the image representation vector I as the query, the text feature vector T as the key, and the audio feature vector A as the value; generating the fusion feature H through multi-head attention calculation, and dynamically assigning the weight α of each modality m .

[0031] As a further improvement of this technical solution, the constraints of the generative adversarial network (GAN) in step S3 include: adversarial loss Used to distinguish real from generated samples;

[0032] Target domain loss Constrain the Euclidean distance between generated samples and real abnormal data in the high-level feature space;

[0033] The joint optimization goal is Where λ is the weight coefficient.

[0034] As a further improvement of this technical solution, step S3 further includes:

[0035] Applying spatiotemporal perturbations to generated samples, including random frame insertion, rotation, lighting adjustment, and audio band occlusion;

[0036] The model is trained jointly with the augmented synthetic data and real data.

[0037] As a further improvement of this technical solution, the implementation of the multimodal adversarial scoring in step S5 includes:

[0038] Map each modal feature to the CLIP common vector space;

[0039] Add a small perturbation δ and calculate the cross-modal similarity change △sim i→j ;

[0040] The weighted combination similarity variation generates an adversarial score to evaluate the model's sensitivity to a specific modality.

[0041] As a further improvement of this technical solution, the implementation of spatiotemporal causal reasoning in step S5 includes:

[0042] Divide the video into time-series segments τ and construct temporal edges and semantic edges;

[0043] Aggregate segment features h through spatiotemporal graph networks τ ;

[0044] Matching fragment features and compliance rule vector r k , a regular edge is generated when the similarity exceeds a threshold.

[0045] As a further improvement of this technical solution, the compliance rule vector r k Through the text encoder E text (·) The mapping is generated and compared with the segment feature h τ Calculate the cosine similarity. The specific formula is:

[0046] cos_sim(h τ ,r k )≥ε

[0047] Among them, ε is the preset threshold that triggers the rule violation judgment.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] 1. The present invention can reduce the false detection rate, optimize the classification model through a dynamic weight allocation mechanism, improve detection accuracy in complex environments, and significantly reduce the probability of normal behavior being misjudged as abnormal. At the same time, it can improve the fusion efficiency of multimodal information. By improving multimodal feature extraction and deep fusion algorithms, it can fully explore the potential relationship between modalities such as images, audio and text, improve detection performance, and enhance the ability to detect long-tail abnormal behaviors. It can also design a generative adversarial network to generate rare behavior samples to make up for data shortages and enhance the model's ability to recognize low-frequency and high-risk behaviors. In addition, the present invention can also enhance the ability to explain behavior and introduce an explainable analysis framework so that the model can provide detailed explanations of the detection results, thereby enhancing transparency and trust in applications.

[0050] 2. In this invention, the fusion of image, audio and text features can fully mine video scene information. The model is still robust in the case of modal loss or severe noise. At the same time, with the help of target domain constrained GAN, scarce abnormal data is generated and diversified enhancement is performed, thereby greatly improving the detection ability of low-frequency, high-risk behaviors.

[0051] 3. In this invention, CLIP-based multimodal adversarial scoring is used to perturb and measure the similarity of audio, image, and text, enabling the model to have a more fine-grained vulnerability discovery capability after compliance detection. At the same time, audio and video evidence and text descriptions are generated through spatiotemporal causal reasoning and compliance rule mapping, thereby greatly improving the interpretability of the detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 Schematic diagram of the steps of the video data behavior compliance detection method based on multimodal analysis of the present invention. DETAILED DESCRIPTION

[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0054] In a specific embodiment, Figure 1 As shown, the present invention provides a method for detecting behavior compliance of video data based on multimodal analysis, which specifically includes the following steps:

[0055] 1. Data preprocessing and feature extraction:

[0056] First, in view of the current situation of difficulty in video behavior recognition in complex scenes, we make full use of multi-source information to extract video features. For the image modality of the video, we use SwinTransformer improved by multi-scale feature fusion to extract the image vector of the video frame. The method is to extract the image vector of the video frame from the feature F output of each layer. (l) Extract information of different scales from the dataset and then concatenate them to obtain a combined representation F merged =Concat(F (1) ,...,F (l) )W fusion , where Concat(·) represents the concatenation operation on the dimension, W fusion It is a trainable mapping matrix that maps the concatenated multi-scale features back to the target dimension to generate the final representation vector I of the image modality. In this way, local details are preserved without losing the global context.

[0057] For the audio mode of the video, first perform short-time Fourier transform (STFT) on the original signal to obtain the multi-channel spectrum representation S c , and then use the denoising network to correct the noise and get the enhanced S' c (f,τ)=M c (f,τ)S c (f,τ), where M c (f,τ) is the denoising network. Finally, these spectra are fed into the DeepSpeech backbone network to obtain a stable audio feature vector A=E audio ({S' c (f,τ)}), where E audio (·) represents the feature extraction function of the improved DeepSpeech.

[0058] For the text mode of the video, we use BERT pre-trained on domain data and the loss of MaskedLanguageModel (MLM) Make the model more familiar with industry terms and context, where w i is the word to be masked, w \i represents the context, θ is the BERT model parameter, and finally the text feature T=E is obtained, which is more sensitive to industry semantics. text ({w i}), E text (·) is the feature extraction function for pre-training BERT.

[0059] 2. Multimodal fusion:

[0060] After completing the above modal feature extraction, the image representation vector I, audio feature vector A, and text feature vector T are spliced into a sequence Let cross-modal self-attention capture the interaction between the modalities. When calculating the multi-head attention, it will first project Q = XW Q , K=XW K 、V=XW V , and then through Combine them together. Because this is a cross-modal scenario, the image feature I is used as the query, the text feature T is used as the key, and the audio feature A is used as the value. Finally, in order to cope with the situation of missing modalities or out-of-control noise, a learnable weight α is given to each attention output. m , then by Get the final fused representation, where M is the number of multi-head attention. In this way, if the quality of a certain modality is not good, α m It will automatically become smaller, and the weights of other modes will be relatively increased.

[0061] 3. Generation of long-tail abnormal behavior samples:

[0062] When dealing with the problem of long-tail abnormal behavior, relying solely on the multimodal fusion feature H cannot completely solve the dilemma of data scarcity, so a generative adversarial network (GAN) with target domain constraints is used to help fill in the samples. This approach is based on StyleGAN2, in addition to the conventional adversarial discriminator D adv In addition, a domain discriminator D is added domain In this way, the generator G passes the conventional adversarial discriminator D adv and domain discriminator D domain The joint constraint of , generates synthetic samples that are close to the distribution of real abnormal data. The adversarial loss is defined as:

[0063]

[0064] Among them, p data represents the real data distribution, p z is the prior distribution of the generator input noise. The target domain loss stipulates that the generated sample G(z) is different from the real abnormal data X in the high-level feature space φ(·) target Keeping it close, the domain constraint loss is defined as:

[0065]

[0066] Among them, φ(·) is a feature mapping function compatible with the multimodal fusion network, which is used to measure the distance between the generated sample and the real abnormal data in the high-level semantic space, so that the GAN can be closer to the multimodal distribution of the real abnormal samples. dist(,) takes the Euclidean distance, X target Represents the feature set of the target abnormal domain data. The overall form of the joint loss is:

[0067]

[0068] Here, λ controls the weight of the target domain constraint in the overall loss. With this domain constraint in place, the network generates synthetic data of rare abnormal behavior types for a specific abnormal behavior l using the conditional formula (G(z, l), D(x, l)). Subsequently, the synthetic data is further subjected to spatiotemporal perturbations (such as random frame insertion, rotation, lighting adjustments, and audio band occlusion) and used together with real data for training to enhance the model's adaptability to long-tail scenarios.

[0069] After the training is completed, a compliance detection model with both multimodal fusion and long-tail anomaly adaptability will be obtained. Send it to the final detection head and use the formula:

[0070]

[0071] Output behavior category or abnormal probability. is the trainable weight, d is the dimension of the fused feature, c is the number of classifications, is the bias vector, and σ(·) is the softmax function.

[0072] 4. Explanation of behavioral compliance test results:

[0073] In actual use, it is necessary to explain the detection results to users to improve the credibility of the model. Here, we use CLIP-based multimodal adversarial scoring and spatiotemporal causal reasoning to enhance interpretability.

[0074] 5. Perform interpretability analysis on the detection results:

[0075] CLIP-based multimodal adversarial scoring: Map the image features I, audio features A, and text features T in the cross-modal network to the common vector space of CLIP:

[0076] v I =E img (I)

[0077] v A =E audio (A)

[0078] v T =E text (T)

[0079] Among them, E img (·), E audio (·), E text (·) is the image, audio and text encoder of CLIP. At this time, a small perturbation δ is added to the image to obtain the slightly affected image vector v in the CLIP common vector space I' =E img (I+δ), this perturbation δ is also applicable when added to audio or text. At this time, the change in similarity can be measured:

[0080] △sim I→T =|cos_sim(v I ,v T )-cos_sim(v I' ,v T )|

[0081] Similarly, we can get △sim I→A , △sim A→T , where cos_sim(,) is the cosine similarity function of the vector. Then the similarity changes of different modalities are weighted and combined to obtain the multimodal adversarial score:

[0082]

[0083] where β ij Indicates the attention of each mode to the disturbance, satisfying ∑ (i,j)∈((I,T)(I,A)(T,A)) β ij = 1. This allows us to understand the model's sensitivity to perturbations across the three modalities of image, audio, and text based on the score, and to identify any potential over-reliance on certain modal features (such as pornographic images, alarm sounds, and inappropriate text). This allows the model to interpretably determine whether features from each modality match, and also allows us to further supplement relevant training data and optimize the model network structure based on the score results.

[0084] Spatiotemporal causal reasoning: The video sequence after anomaly detection is divided into several segments τ, each of which contains image frames, corresponding audio and text information. In order to capture temporal causality and semantic association, two types of edges are defined between segments: if a segment τ ’ If the video features (image + audio + text) of the two videos are similar in time sequence, a temporal edge is connected; if the features of the two videos (image + audio + text) are highly similar in the multimodal common space, a semantic edge is connected. At this time, the spatiotemporal graph network is used to define the aggregated features of each segment τ:

[0085]

[0086] Where N(τ') represents the adjacent segments of segment τ, W g 、b g is a trainable parameter, σ(·) is the activation function, α τ,τ’ The attention weight is determined by the segment similarity.

[0087] The compliance rules are then passed through the same encoder E as the text modality text (·) is mapped to the embedding vector r k , when the aggregation feature h of a certain fragment is found τ and the compliance rule vector r k When the similarity reaches a certain threshold:

[0088] cos_sim(h τ ,r k )≥ε

[0089] Among them, ε is the similarity threshold, and a rule edge containing this rule is added to the segment τ, which is convenient for users to use the detection results. Tracing back to a certain fragment τ, it was detected that the rule r was violated k , which improves the interpretability of the model and enhances user trust.

[0090] In summary, the present invention first uses the Swin Transformer and DeepSpeech (including multi-channel spectrum enhancement) improved by multi-scale fusion, and BERT pre-trained for the compliance field to extract deep representations of images, audio and text of video data respectively; then, with the support of cross-modal attention and dynamic weighting mechanism, the features of the three modalities are mapped into a unified representation space, and information complementarity and global dependency modeling are achieved through cross-attention between different modalities; then, the target domain constrained StyleGAN2 is introduced in the long-tail anomaly detection part, which combines the synthesis of scarce behavior scenarios with spatiotemporal perturbation enhancement to alleviate the problem of insufficient real data, and on this basis, the anomaly detection head is trained to perform compliance / anomaly judgment on the fused multimodal features; finally, through the interpretable analysis framework of multimodal adversarial scoring and spatiotemporal causal reasoning, the multimodal analysis explanation link from cause to consequence is presented to users, thereby achieving a combination of high accuracy and traceability in practical scenarios.

[0091] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for detecting behavior compliance of video data based on multimodal analysis, characterized in that: The following steps are involved: S1. Extract multimodal features from video data, including: Image modality feature extraction: Generate image representation vectors based on SwinTransformer with multi-scale feature fusion; Audio modal feature extraction: The audio is processed by short-time Fourier transform and denoising network and then fed into the improved DeepSpeech network; Text modality feature extraction: Use the domain-pretrained BERT model to extract text features; S2: Image, audio, and text features are spliced into a cross-modal sequence, fused through a cross-modal self-attention mechanism, and the weights of each modality are dynamically adjusted; S3. Generate long-tail abnormal behavior samples using a generative adversarial network with target domain constraints and enhance the data with spatiotemporal perturbations. S4. Perform behavioral compliance detection based on the fused multimodal features and output abnormality probability or category; S5. Perform interpretable analysis of detection results through multimodal adversarial scoring and spatiotemporal causal reasoning.

2. The method for detecting compliance of video data behavior based on multimodal analysis according to claim 1, characterized in that: The specific implementation of image modality feature extraction in step S1 includes: Extract multi-scale features from the output of each layer of SwinTransformer; Through the trainable mapping matrix W fusion Splicing and dimensionality reduction are performed to generate the image representation vector I.

3. The method for detecting compliance of video data behavior based on multimodal analysis according to claim 1, characterized in that: The audio modal feature extraction in step S1 includes: Use short-time Fourier transform to generate multi-channel spectrum S c ; Correct the spectrum noise through the denoising network and output the enhanced spectrum S' c ; The audio feature vector A is extracted using the improved DeepSpeech network.

4. The method for detecting compliance of video data behavior based on multimodal analysis according to claim 1, characterized in that: The text modality feature extraction in step S1 includes: Pre-train the BERT model based on domain data and use masked language model loss optimization; Extract the text feature vector T by re-pretraining BERT.

5. The method for detecting compliance of video data behavior based on multimodal analysis according to claim 1, characterized in that: The specific implementation of the cross-modal self-attention mechanism in step S2 includes: The image representation vector I is used as the query, the text feature vector T is used as the key, and the audio feature vector A is used as the value; Generate fusion feature H through multi-head attention calculation and dynamically assign weight α to each modality m .

6. The method for detecting compliance of video data behavior based on multimodal analysis according to claim 1, characterized in that: The constraints of the Generative Adversarial Network (GAN) in step S3 include: Fighting Losses Used to distinguish real from generated samples; Target domain loss Constrain the Euclidean distance between generated samples and real abnormal data in the high-level feature space; The joint optimization goal is Where λ is the weight coefficient.

7. The method for detecting compliance of video data behavior based on multimodal analysis according to claim 1, characterized in that: The step S3 further comprises: Applying spatiotemporal perturbations to generated samples, including random frame insertion, rotation, lighting adjustment, and audio band occlusion; The model is trained jointly with the augmented synthetic data and real data.

8. The method for detecting compliance of video data behavior based on multimodal analysis according to claim 1, characterized in that: The implementation of the multimodal adversarial scoring in step S5 includes: Map each modal feature to the CLIP common vector space; Add a small perturbation δ and calculate the cross-modal similarity change △sim i→j ; The weighted combination similarity variation generates an adversarial score to evaluate the model's sensitivity to a specific modality.

9. The method for detecting compliance of video data behavior based on multimodal analysis according to claim 1, characterized in that: The implementation of spatiotemporal causal reasoning in step S5 includes: Divide the video into time-series segments τ and construct temporal edges and semantic edges; Aggregate segment features h through spatiotemporal graph networks τ ; Matching fragment features and compliance rule vector r k , a regular edge is generated when the similarity exceeds a threshold.

10. The method for detecting compliance of video data behavior based on multimodal analysis according to claim 1, characterized in that: The compliance rule vector r k Through the text encoder E text (·) The mapping is generated and compared with the segment feature h τ Calculate the cosine similarity. The specific formula is: cos_sim(h τ ,r k )≥ε Among them, ε is the preset threshold that triggers the rule violation judgment.

Citation Information

Patent Citations

  • Multi-mode multi-disease long-tail distribution ophthalmic disease classification model training method and device

    CN113011485A

  • Multi-modal sentiment analysis method based on multi-task learning and stacked cross-modal fusion

    CN114694076A

Cited By

  • Video content security understanding method, system and device and storage medium

    CN120823549A

  • Long-tail target detection method and device based on visual union scene and computer equipment

    CN120953898A

  • Motion feature vector optimization method, device and equipment based on multi-modal fusion

    CN121214534A