Double-flow gating violence detection method and system based on expansion 3D convolutional network and Transform

By combining the expanded 3D convolutional network and the Transformer encoder architecture, the dual-stream gated brute force detection method is adopted to solve the limitations of the existing model in dealing with the space-time dependence of videos, and to achieve more efficient and accurate brute force detection, suitable for real-time applications.

CN119942407AActive Publication Date: 2025-05-06CHONGQING UNIV OF TECH

Patent Information

Application Number
CN202510035710.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

The existing brute force detection model has limitations in processing the spatiotemporal dependence of video data, which is difficult to adapt to dynamic changes in video, and has high computational complexity, which limits the application of real-time detection.

Method used

A dual-stream gated brute force detection method based on expanded 3D convolutional network and Transformer is used to generate temporal feature maps for violent behavior detection by acquiring video streams, calculating dense optical streams, performing background suppression, spatiotemporal modeling, time modeling and classification.

Benefits of technology

It significantly improves the accuracy and efficiency of brute force detection, can capture the spatial and temporal features in the video more accurately, reduces the computational complexity, and makes the model more suitable for real-time detection applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942407A_ABST
    Figure CN119942407A_ABST
Patent Text Reader

Abstract

The invention discloses a double-flow gating violence detection method and system based on an expansion 3D convolutional network and Transform. Classification is realized through the following steps; the method comprises the following steps: S1, acquiring a video stream, segmenting the video stream into frames, and calculating dense optical streams among continuous frames of the video stream; s2, segmenting video stream segments through a sliding window, and inhibiting non-motion elements in an RGB mode by using an optical stream mode; s3, space-time modeling is carried out on the dense optical flow and the SRGB mode through a 3D convolutional neural network, feature fusion is carried out, and a feature map is created; s4, performing time modeling on the spatial-temporal feature sequence of the video stream segment through a Transform encoder to generate a time feature map; and S5, average pooling is carried out through a 1D convolutional neural network, classification is carried out through a multi-layer perceptron, and videos are classified as violent videos or non-violent videos. According to the method, the defects of a violence detection model in processing the spatial-temporal characteristics are overcome, the accuracy and efficiency of violence detection are improved, and the application effect in a dynamic video monitoring scene is particularly achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent video surveillance technology, and in particular to a dual-stream gated violence detection method and system based on an expanded 3D convolutional network and a Transformer. Background Art

[0002] Accurate violence detection is crucial for intelligent surveillance systems, which enable proactive safety prevention, timely response to emergencies, and support public safety management decisions. With the increasing public safety needs and the growth of surveillance video data, real-time detection of violent behavior faces significant challenges, which mainly stems from the inherent spatiotemporal dependencies and dynamics in videos. Traditional single-modality input methods often have difficulty in effectively capturing key motion features in videos.

[0003] In recent years, advances in deep learning have promoted the application of neural network architectures in modeling spatial and temporal dependencies. Two-dimensional convolutional neural networks (2DCNNs) and three-dimensional convolutional neural networks (3DCNNs) have been widely used in video content understanding. However, models based on a single modality have difficulties in capturing spatial and temporal features in videos, and are difficult to effectively model complex motion dependencies in videos due to the large number of parameters and high training cost. Existing violence detection models have limitations in handling the spatiotemporal dependencies of video data. Traditional methods often rely on fixed video frames and have difficulty adapting to dynamic changes in videos. Although some advanced models, such as Transformer-based models, perform well in modeling long-range dependencies, their high computational complexity limits their application in real-time detection. Therefore, a new model is needed to improve the accuracy of violence detection while reducing the computational complexity.

[0004] In order to solve the problem of spatiotemporal dependency, a two-stream network structure is introduced to effectively model the relational structure in the video. FGITNet combines the dilated 3D convolutional network (I3D) and the Transformer encoder architecture to process spatiotemporal data. In addition, the adaptability of the model is further enhanced by dynamically updating the spatial relationship between frames. However, these methods usually rely on predefined or static video frames, which limits their adaptability to the ever-changing patterns of violent behavior.

[0005] Transformer-based models have attracted widespread attention due to their ability to effectively capture long-term dependencies using self-attention mechanisms. FGITNet uses a multi-head attention mechanism to model spatiotemporal correlations. Although these models perform well on complex video data, they are accompanied by high computational complexity, making them less suitable for real-time detection applications. Although the FGITNet approach aims to reduce this complexity, it still has difficulties in achieving a balance between efficiency and accuracy.

[0006] State-space models have been re-examined in sequence modeling because they can efficiently capture long-term dependencies. This application introduces a method of combining state-space models with neural networks through FGITNet, which can effectively model complex temporal dynamics without the computational overhead of traditional recursive units. Summary of the invention

[0007] The present invention provides a dual-stream gated violence detection method based on an expanded 3D convolutional network and a Transformer, comprising the following steps:

[0008] S1: Get the video stream, divide the video stream into frames, and calculate the dense optical flow between consecutive frames of the video stream;

[0009] S2: Segment the video stream segments by sliding windows and use the optical flow modality to suppress non-motion elements in the RGB modality;

[0010] S3: Use 3D convolutional neural network to perform spatiotemporal modeling of dense optical flow and SRGB modalities, perform feature fusion, and create feature maps;

[0011] S4: Temporal modeling of the spatiotemporal feature sequence of the video stream segment is performed through the Transformer encoder to generate a temporal feature map;

[0012] S5: Videos are classified as violent or non-violent using a 1D convolutional neural network with average pooling and a multi-layer perceptron.

[0013] Preferably, the continuous frames of the video stream in S1 are normalized by Z-score, and after the normalization, the dense optical flow between the continuous frames of the video stream is calculated by the Farneback method, and the formula is: DOF t =Farneback(RGB t , RGB t+1 ), where RGB t is a continuous frame at time t, RGB t+1 is the continuous frame at time t+1, DOF t is the dense optical flow at time t; by iterating the above formula for each frame of the continuous frame of the video stream, n-1 dense optical flow sequences are obtained, where n represents the length of the video.

[0014] Preferably, in S2, the dense optical flow is divided into m segments by a sliding window method, and the suppression effect is enhanced by accumulating the optical flow vector amplitudes of all frames. The formula is: Where |DOF t | is the DOF frame at time t∈[i, j], i and j are indices, indicating the start frame and end frame in the kth DOF segment, A kTo accumulate all frames; the optical flow vector amplitude of all accumulated frames is obtained by normalizing the accumulated optical flow amplitude through the minimum-maximum normalization MMN, scaling it to a value between 0 and 1, and using the normalized accumulated value as a scaling factor to perform element-by-element multiplication and adjustment of the RGB fragment to suppress non-motion elements. The formula is: in is the i-th frame of the k-th RGB fragment, is the suppressed RGB frame, Get the suppressed SRGB fragment for element-wise multiplication over all spatial dimensions and color channels.

[0015] Preferably, in S3, spatiotemporal features are extracted from the SRGB modality and the dense optical flow modality through a pre-trained I3D model to obtain fragment spatiotemporal features, and the fragment spatiotemporal features are fused by element-by-element multiplication to generate a comprehensive fragment spatiotemporal feature map, and the fragment spatiotemporal feature map is connected along the time dimension to form two spatiotemporal feature sequences of shape l×1024, where l is the sum of all fragment output logical values ​​determined by the fragment size and the number of fragments.

[0016] Preferably, the Transformer encoder in S4 receives the output X of the segment spatiotemporal modeling; the output X is mapped to query Q, key K and value V respectively, and the mapping relationship is as follows: Q = XW Q , K = XW K , V = XW V , where W Q , W K , W V are parameter matrices used to map the input X to the specific space of query, key, and value respectively; the Transformer encoder is used to perform temporal modeling on the spatiotemporal feature sequence of each video clip to generate a temporal feature map.

[0017] Preferably, the video clips of violent activities are selected by the multi-head self-attention mechanism of the Transformer encoder; the multi-head self-attention mechanism is transformed by the input X, and the formula is: MultiHead(Q, K, V) = {Concat(head}1, ..., head h )W o , where h represents the number of attention heads, W o The parameter matrix of the multi-head self-attention mechanism, where is a linear transformation of the attention heads of the query Q, key K and value V; the multi-head self-attention mechanism is transformed by the input query Q, key K and value V, and the formula is: where d k is the dimension of each key and softmax is the activation function.

[0018] Preferably, the multi-head self-attention mechanism of the Transformer encoder is normalized by outputting X and transmitted through a residual connection, the formula is: Z=LayerNorm(X+MultiHead(Q, K, V)), each position in Z is independently processed by a feedforward network FFN, the formula is: FFN(z)=max(0, zW1+b1)W2+b2, wherein W1 and W2 are feedforward network weight parameters, and b1 and b2 are bias coefficients; the output of the feedforward network FFN is normalized and added back to Z through another residual connection, the formula is: Y=LayerNorm(Z+FFN(Z)), wherein Y is the output value.

[0019] Preferably, in S5, average pooling is performed through a 1D convolutional neural network, the input data is sent to a 1D convolutional neural network architecture in the time dimension, and a sliding window operation is performed on the time series. For each data point in the window, the average value is calculated to realize the downsampling process of the time feature; the dimension of the input data is compressed and a tensor is output.

[0020] Preferably, after completing the average pooling of the 1D convolutional neural network, the data enters a multi-layer perceptron for classification, and the output tensor obtained by the average pooling is flattened into a vector, and the vector is sent to the 1D convolutional neural network as the input of the multi-layer perceptron; the multi-layer perceptron is composed of multiple fully connected layers, and the input vector and the neurons of the multi-layer perceptron are linearly transformed through a weight matrix, and nonlinearly transformed through an activation function. The transformed result is continuously passed to the next layer, and the above-mentioned linear transformation and nonlinear transformation process are repeated. After multiple layers of processing, the data reaches the output layer, corresponding to the classification output of non-violent or violent behavior in the video, and the video classification is determined as violent or non-violent category through the activation values ​​of two neurons in the output layer.

[0021] A dual-stream gated violence detection system based on an expanded 3D convolutional network and a Transformer, comprising a dual-stream gated network FGITNet model; the dual-stream gated network FGITNet model comprises a data preprocessing module, a background suppression module, a fragment spatiotemporal modeling module, a fragment time modeling module and a classification module;

[0022] The data preprocessing module performs standardization processing on the input surveillance video, adjusts the resolution pixels of the image, normalizes the image data by Z-score, and extracts a fixed number of frames by uniform sampling to generate standardized video frame data;

[0023] The background suppression module suppresses non-moving elements in the video, extracts segments from the RGB frame and the dense optical flow frame using a sliding window method, calculates the amplitude of motion information in the dense optical flow segment, adjusts the optical flow amplitude to a value between 0 and 1 through minimum-maximum normalization, and adjusts the RGB segment by element-by-element multiplication;

[0024] The fragment spatiotemporal modeling module uses the pre-trained I3D model to extract spatiotemporal features from the suppressed RGB modality and dense optical flow modality, independently models spatiotemporal features on each fragment of length clip_size, and fuses the two modal features by element-by-element multiplication to form a comprehensive spatiotemporal feature map of the fragment;

[0025] The fragment temporal modeling module adopts the Transformer encoder structure, introduces a multi-head self-attention mechanism, takes the spatiotemporal feature sequence output by the fragment spatiotemporal modeling module as input, and through linear transformation, multi-head self-attention calculation, residual connection and feedforward network processing operations, enables the FGITNet model to selectively observe fragments of violent activities and generate a temporal feature map;

[0026] The classification module downsamples the temporal features through 1D average pooling, compresses the dimension of the input data, outputs a tensor, and then flattens it into a vector, processes it through a multi-layer perceptron containing multiple fully connected layers, and outputs a classification result of whether each video clip is non-violent or violent behavior;

[0027] The video frame data processed by the data preprocessing module enters the background suppression module, and the data processed by the background suppression module is sent to the fragment spatiotemporal modeling module. The spatiotemporal feature sequence generated by the fragment spatiotemporal modeling module is passed as input to the fragment time modeling module, and the time feature map output by the fragment time modeling module is then sent to the classification module, and the video classification result is output after being processed by the classification module.

[0028] Compared with the prior art, the technical solution of this application has the following technical effects:

[0029] By combining the expanded 3D convolutional network with the Transformer encoder architecture, the present invention innovatively solves the shortcomings of the traditional violence detection model in processing spatiotemporal features, significantly improving the accuracy and efficiency of violence detection, especially in dynamic video surveillance scenarios.

[0030] The FGITNet model proposed in this paper has the ability to process video structure data and capture the complex relationship between spatial and temporal dimensions. The FGITNet model not only uses optical flow as the unique feature of each frame, but also uses RGB data as multiple features to capture video relationships through the spatiotemporal FGITNet layer, which is then processed by the classification layer to make the final prediction. By combining the expanded 3D convolutional network with the Transformer encoder architecture, the shortcomings of traditional violence detection models in dealing with spatiotemporal dependencies are innovatively solved.

[0031] The present invention improves the feature expression ability of the model through multi-dimensional feature learning of the data preprocessing module, can deeply capture the rich features of the input video, and enhance the learning ability of the model; through the proposed background suppression method, the RGB frame is effectively filtered through the optical flow modality, non-motion elements are suppressed, and moving objects in the video are highlighted, thereby improving the model's ability to capture motion features while maintaining a low computational overhead.

[0032] The fragment spatiotemporal modeling module of the present invention optimizes the video sequence analysis capability. By expanding the 3D convolutional network and the Transformer encoder mechanism, the model can more accurately capture the dynamic dependencies between time points and spatial positions. The FGITNet layer structure of the present invention supports flexible adjustment. Users can adjust the depth of the model according to specific task requirements to adapt to video data detection tasks of different complexities, and has good adaptability.

[0033] The present invention excels in computational efficiency and provides an optimized computational method through the integration of Transformer, which significantly reduces the computational complexity and makes the model more suitable for real-time detection. At the same time, it can dynamically adapt to changes in video surveillance and improve detection accuracy.

[0034] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application so that it can be implemented in accordance with the contents of the specification, and to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following is a detailed description of the preferred embodiments of the present application in conjunction with the accompanying drawings as follows.

[0035] Based on the detailed description of the specific embodiments of the present application in combination with the accompanying drawings below, those skilled in the art will become more aware of the above and other objects, advantages and features of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings without creative work. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual scale.

[0037] Figure 1 , a schematic diagram of the dual-stream gating network FGITNet model of the present invention;

[0038] Figure 2 , flow chart of the dual-stream gated violence detection method of the present invention;

[0039] Figure 3 , a schematic diagram of the spatiotemporal modeling process of editing of the present invention;

[0040] Figure 4 , Schematic diagram of the FGITNet classification stage of the present invention. DETAILED DESCRIPTION

[0041] To make the purpose, technical scheme and advantages of the embodiment of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiment of the present application. Obviously, the described embodiment is a part of the embodiment of the present application, rather than all of the embodiments. In the following description, specific details such as specific configuration and components are provided only to help fully understand the embodiments of the present application. Therefore, it should be clear to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. In addition, for clarity and brevity, the description of known functions and structures is omitted in the embodiment.

[0042] It should be understood that the references to "one embodiment" or "this embodiment" throughout the specification mean that the specific features, structures, or characteristics associated with the embodiment are included in at least one embodiment of the present application. Therefore, the references to "one embodiment" or "this embodiment" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0043] In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplicity and clarity, and does not in itself indicate the relationship between the various embodiments and / or settings discussed.

[0044] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist at the same time. The term " / and" in this article describes another type of association object relationship, indicating that there can be two relationships. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the previous and next associated objects are in an "or" relationship.

[0045] The term "at least one" in this article is merely a description of the association relationship of associated objects, indicating that there may be three relationships. For example, at least one of A and B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0046] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusions.

[0047] Example 1

[0048] This embodiment mainly describes a dual-stream gated violence detection method based on an expanded 3D convolutional network and a Transformer. Figure 2 As shown, the following steps are included:

[0049] S1: Get the video stream, divide the video stream into frames, and calculate the dense optical flow between consecutive frames of the video stream;

[0050] S2: Segment the video stream segments by sliding windows and use the optical flow modality to suppress non-motion elements in the RGB modality;

[0051] S3: Use 3D convolutional neural network to perform spatiotemporal modeling of dense optical flow and SRGB modalities, perform feature fusion, and create feature maps;

[0052] S4: Temporal modeling of the spatiotemporal feature sequence of the video stream segment is performed through the Transformer encoder to generate a temporal feature map;

[0053] S5: Videos are classified as violent or non-violent using a 1D convolutional neural network with average pooling and a multi-layer perceptron.

[0054] Furthermore, the continuous frames of the video stream in S1 are normalized by Z-score, and after processing, the dense optical flow between the continuous frames of the video stream is calculated by the Farneback method. The formula is: DOF t=Farneback(RGB t , RGB t+1 ), where RGB t is a continuous frame at time t, RGB t+1 is the continuous frame at time t+1, DOF t is the dense optical flow at time t; by iterating the above formula for each frame of the continuous frame of the video stream, n-1 dense optical flow sequences are obtained, where n represents the length of the video.

[0055] Furthermore, as shown in Table 1, in S2, the dense optical flow is divided into m segments by the sliding window method, and the suppression effect is enhanced by accumulating the optical flow vector amplitude of all frames. The formula is: Where |DOF t | is the DOF frame at time t∈[i, j], i and j are indices, indicating the start frame and end frame in the kth DOF segment, A k To accumulate all frames; obtain the accumulated optical flow vector amplitude of all frames, normalize the accumulated optical flow amplitude through the minimum-maximum normalization MMN, scale it to a value between 0 and 1, and use the normalized accumulated value as the scaling factor to perform element-by-element multiplication and adjustment of the RGB fragment to suppress non-motion elements. The formula is: in is the i-th frame of the k-th RGB fragment, is the suppressed RGB frame, Get the suppressed SRGB fragment for element-wise multiplication over all spatial dimensions and color channels.

[0056] Table 1 Background suppression

[0057]

[0058]

[0059] Further, if Figure 3 As shown in FIG. 1 , in S3, the pre-trained I3D model is used to extract spatiotemporal features from the SRGB modality and the dense optical flow modality to obtain the fragment spatiotemporal features, and the fragment spatiotemporal features are fused by element-by-element multiplication to generate a comprehensive fragment spatiotemporal feature map. The fragment spatiotemporal feature map is connected along the time dimension to form two spatiotemporal feature sequences of shape l×1024, where l is the sum of the logical values ​​of all fragment outputs determined by the fragment size and the number of fragments.

[0060] Furthermore, the Transformer encoder in S4 receives the output X of the spatiotemporal modeling of the fragment, and X is converted into a query Q, a key K, and a value V, as follows: Q = XW Q , K = XW K , V = XW V , where WQ , W K , W V It is a parameter matrix that maps the input X to the respective spaces of query, key, and value, models the spatiotemporal feature sequence of each fragment, and generates a temporal feature map through the Transformer encoder.

[0061] Furthermore, the video clips of violent activities are selected through the multi-head self-attention mechanism of the Transformer encoder; the multi-head self-attention mechanism transforms the input X, and the formula is: MultiHead(Q, K, V) = {Concat(head}1, ..., head h )W O , where h represents the number of attention heads, W O The parameter matrix of the multi-head self-attention mechanism, where is a linear transformation of the attention head of the query Q, key K and value V; the multi-head self-attention mechanism transforms the input query Q, key K and value V, and the formula is: where d k is the dimension of each key and softmax is the activation function.

[0062] Furthermore, the multi-head self-attention mechanism of the Transformer encoder is normalized by outputting X and transmitted through the residual connection, as follows: Z = LayerNorm(X + MultiHead(Q, K, V)), and each position in Z is processed independently through the feedforward network FFN, as follows: FFN(z) = max(0, zW1 + b1)W2 + b2, where W1 and W2 are the weight parameters of the feedforward network, and b1 and b2 are bias coefficients; the output of the feedforward network FFN is normalized and added back to Z through another residual connection, as follows: Y = LayerNorm(Z + FFN(Z)), where Y is the output value.

[0063] Furthermore, in S5, average pooling is performed through a 1D convolutional neural network. The input data is sent to the 1D convolutional neural network architecture in the time dimension, and a sliding window operation is performed on the time series. For each data point in the window, the average value is calculated to realize the downsampling process of the time feature; the dimension of the input data is compressed and the tensor is output.

[0064] Furthermore, after completing the average pooling of the 1D convolutional neural network, the data enters the multi-layer perceptron for classification. The output tensor obtained after average pooling will be flattened into a vector, and the vector is sent to the 1D convolutional neural network as the input of the multi-layer perceptron; the multi-layer perceptron is composed of multiple fully connected layers. The input vector and the neurons of the multi-layer perceptron are linearly transformed through the weight matrix, and nonlinearly transformed through the activation function. The transformed result is passed on to the next layer, and the above linear transformation and nonlinear transformation process are repeated. After multiple layers of processing, the data reaches the output layer, corresponding to the classification output of non-violent or violent behavior in the video. The activation values ​​of the two neurons in the output layer determine whether the video is classified as violent or non-violent.

[0065] This embodiment captures the complex relationship between spatial and temporal dimensions by processing video structure data, using optical flow as the unique feature of each frame, and also using RGB data as multiple features. The video relationship is captured by the spatiotemporal FGITNet layer, which is then processed by the classification layer to make the final prediction. By combining the expanded 3D convolutional network with the Transformer encoder architecture, the shortcomings of traditional violence detection models in dealing with spatiotemporal dependencies are innovatively solved.

[0066] Example 2

[0067] This embodiment describes in detail a dual-stream gated violence detection system based on an expanded 3D convolutional network and a Transformer. The specific connection method is as follows: Figure 1 As shown, the dual-stream gated violence detection system includes a dual-stream gated network FGITNet model; the dual-stream gated network FGITNet model includes a data preprocessing module, a background suppression module, a fragment spatiotemporal modeling module, a fragment time modeling module and a classification module;

[0068] The data preprocessing module standardizes the input surveillance video and adjusts the resolution pixels of the image. For example, the resolution of each image is standardized to 224×224 pixels. The image data is normalized by Z-score to conform to the standard format, and a fixed number of frames are extracted by uniform sampling to generate standardized video frame data.

[0069] The background suppression module suppresses non-moving elements in the video, extracts segments from RGB frames and dense optical flow frames using a sliding window method, calculates the amplitude of motion information in dense optical flow segments, adjusts the optical flow amplitude to a value between 0 and 1 through minimum-maximum normalization, and adjusts the RGB segments by element-by-element multiplication;

[0070] The background suppression module uses the optical flow modality to suppress non-moving elements, which are denoted as And SRGBclips, for each time step, the optical flow magnitude is retrieved from DOFclips, which are normalized and broadcasted to form a motion suppressed representation SRGBclips of the video sequence. The background suppression module effectively encodes dynamic patterns in the video and enhances the model's ability to capture repetitive motion in the video.

[0071] The fragment spatiotemporal modeling module uses the pre-trained I3D model to extract spatiotemporal features from the suppressed RGB modality and dense optical flow modality, independently models spatiotemporal features on each clip of length clip_size, and fuses the two modal features by element-by-element multiplication to form a comprehensive spatiotemporal feature map of the fragment;

[0072] The fragment spatiotemporal modeling module uses I3D to perform spatiotemporal modeling on DOF and RGB modalities for each fragment in the video, where DOF is dense optical flow; it combines the spatiotemporal features of DOF and RGB in each fragment, performs feature fusion, and creates l feature maps.

[0073] The present invention uses the I3D model as the backbone network to extract spatiotemporal features from the SRGB and DOF modalities. The SRGB modality consists of three channels, and the DOF modality contains two channels in the vertical and horizontal directions, which are processed using the I3D model respectively. The FGITNet model is configured to the corresponding number of input channels, SRGB is set to 3, DOF is set to 2, and the output layer dimension is set to 1024. Each video clip with a length of clip_size is independently spatiotemporally modeled to construct a spatiotemporal feature sequence of each clip. By using the I3D model on these clips respectively, the features of each clip are captured. These features are fused using element-by-element multiplication to generate a comprehensive spatiotemporal feature map of the clips, and the spatiotemporal feature map of each clip is connected along the time dimension to form a spatiotemporal feature sequence of shape l×1024, where r is determined by the size of the clip and the number of clips. The sequence integrates the spatiotemporal features of the SRGB and DOF modalities of each clip, and each logit in the sequence will be used as an input token for the subsequent model to perform temporal modeling between clips, providing support for the analysis of spatiotemporal features.

[0074] The fragment temporal modeling module adopts the Transformer encoder structure and introduces a multi-head self-attention mechanism. It takes the spatiotemporal feature sequence output by the fragment spatiotemporal modeling module as input. Through linear transformation, multi-head self-attention calculation, residual connection and feedforward network processing operations, the FGITNet model selectively observes fragments of violent activities and generates a temporal feature map.

[0075] The classification module downsamples the temporal features through 1D average pooling, compresses the dimensions of the input data, outputs a tensor, and then flattens it into a vector. It is processed by a multi-layer perceptron containing multiple fully connected layers, and outputs the classification result of whether each video clip is non-violent or violent.

[0076] like Figure 4 As shown in the figure, the input data in the classification module is downsampled by time features, and 1D average pooling (AvgPooling) is performed on the time dimension to reduce the dimension and retain the time information. The output of the pooling layer is a tensor with a shape of 8×1024, which is then flattened and normalized into a vector of size 8192. The vector is processed by an MLP containing fully connected layers of 2048, 512, and finally 2 dimensions, corresponding to the classification output of non-violent or violent behavior in the video.

[0077] The video frame data processed by the data preprocessing module enters the background suppression module, and the data processed by the background suppression module is sent to the fragment spatiotemporal modeling module. The spatiotemporal feature sequence generated by the fragment spatiotemporal modeling module is passed as input to the fragment time modeling module, and the time feature map output by the fragment time modeling module is then sent to the classification module, and the video classification result is output after being processed by the classification module.

[0078] This embodiment describes in detail a dual-stream gated violence detection system, which optimizes the video sequence analysis capability for the fragment spatiotemporal modeling module. By expanding the 3D convolutional network and the Transformer encoder mechanism, it can more accurately capture the dynamic dependencies between time points and spatial positions. The depth of the model can be adjusted according to specific task requirements to adapt to video data detection tasks of different complexities, and has good adaptability.

[0079] Example 3

[0080] This embodiment is based on Embodiment 1 and describes in detail the technical effects of the present application, including:

[0081] This application evaluates the effectiveness of our proposed model using three main metrics: accuracy, precision, and recall. Accuracy quantifies the proportion of correct predictions, including true positives and true negatives among all predictions. Precision is a measure of the accuracy of positive predictions, which indicates the proportion of predicted positive examples that are true positives, which can minimize unnecessary alerts and maintain trust in system notifications. Recall determines the model's ability to identify all relevant instances, indicating the proportion of actual positive examples that are correctly predicted as positive. This is critical to detecting as many real violent incidents as possible, thereby enhancing public safety and reducing the risk of missing incidents. Our model is fully evaluated using these metrics and compared with other models using accuracy.

[0082] The RWF-2000 dataset is used, which is the largest and most extensive benchmark dataset for the VD task, containing 2000 real-world surveillance videos from YouTube. These videos are divided into two groups: 1000 videos labeled as “fight” and the other 1000 videos labeled as “non-fight”. For comparison, the Hockey dataset is used, covering 1000 ice hockey game videos capturing National Hockey League (NHL) games. Similar to RWF-2000, the Hockey dataset is also divided into two groups: 500 videos labeled as “fight” and the other 500 videos labeled as “non-fight”. Both datasets present an equal distribution between violent and non-violent videos, demonstrating the good property of no class imbalance, which is beneficial for binary classification tasks. See Table 2 for details on these datasets;

[0083] Table 2 Details of the dataset used in the invention

[0084]

[0085] Experimental results:

[0086] The performance of FGITNet is evaluated on the RWF-2000 and Hockey datasets to assess its ability to detect violence in real-world fight videos. In the experiments, we use 80% of each dataset for training and validation, and retain the remaining 20% ​​for testing. Before training, all videos are preprocessed to ensure the consistency of frame resolution and length, as follows: each frame is resized to 224×224 pixels, normalized using Z-score, and uniformly sampled to a fixed number of frames to ensure consistency between datasets. The FGITNet model is compared with existing models, as shown in Table 3;

[0087] Table 3 Comparison of binary classification accuracy (%) on benchmark datasets

[0088]

[0089] It can be seen from Table 3 that the technical solution of the present application has a high accuracy rate;

[0090] By comparing the evaluation of the model on the benchmark datasets in Table 4, we can further obtain that the accuracy of this application has achieved the highest accuracy of 93.75% and 99.50% on the RWF-2000 and Hockey datasets, respectively, as shown in Table 4;

[0091] Table 4 Evaluation of the model on the benchmark dataset

[0092]

[0093] To further evaluate the effectiveness of the background suppression mechanism in FGITNet, an ablation study was performed on two variants of the model on the RWF-2000 dataset; the baseline model is FGITNet without the background suppression module, while the experimental model includes the aforementioned background suppression mechanism; the background suppression mechanism improves the performance of the model by weakening non-motion elements and emphasizing motion-related features, specifically by calculating the cumulative optical flow amplitude of each fragment k, using minimum-maximum normalization MMN, and adjusting the corresponding RGB fragments by element-wise multiplication.

[0094] By integrating this mechanism, the model effectively suppresses static background regions, allowing it to focus on dynamic regions that are more indicative of violent activities. This targeted attention enhances the extraction of motion-related features, thereby improving classification performance, as shown in Table 5;

[0095] Table 5 Ablation study of FGITNet on the RWF-2000 dataset

[0096]

[0097] Table 5 shows the comparative analysis of the baseline model and the experimental model in terms of various performance indicators such as accuracy, precision and recall. The experimental model shows significant improvement in various indicators compared with the baseline model, highlighting the effectiveness of the background suppression mechanism.

[0098] This example verifies the effectiveness of the background suppression mechanism through significant improvements observed in the experimental model, which achieves superior classification results on the RWF-2000 dataset by focusing on motion center features and reducing background interference. The background suppression mechanism plays a key role in improving the performance of FGITNet, enabling the model to focus on important motion features while mitigating the impact of non-informative background elements, thereby improving the accuracy and reliability of violence detection tasks.

[0099] The above are only preferred embodiments of the present invention, which do not limit the scope of protection of the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any changes, modifications, replacements, integrations and parameter changes to these embodiments within the spirit and principles of the present invention through conventional substitutions or without departing from the principles and spirit of the present invention fall within the scope of protection of the present invention.

Claims

1. A dual-stream gated violence detection method based on an expanded 3D convolutional network and Transformer, characterized in that: The following steps are involved: S1: Get the video stream, divide the video stream into frames, and calculate the dense optical flow between consecutive frames of the video stream; S2: Segment the video stream segments by sliding windows and use the optical flow modality to suppress non-motion elements in the RGB modality; S3: Use 3D convolutional neural network to perform spatiotemporal modeling of dense optical flow and SRGB modalities, perform feature fusion, and create feature maps; S4: Temporal modeling of the spatiotemporal feature sequence of the video stream segment is performed through the Transformer encoder to generate a temporal feature map; S5: Videos are classified as violent or non-violent using a 1D convolutional neural network with average pooling and a multi-layer perceptron.

2. The dual-stream gated violence detection method based on dilated 3D convolutional network and Transformer according to claim 1 is characterized in that: The continuous frames of the video stream in S1 are normalized by Z-score, and after the processing, the dense optical flow between the continuous frames of the video stream is calculated by the Farneback method, and the formula is: DOF t =Farneback(RGB t ,RGB t+1 ), where RGB t is a continuous frame at time t, RGB t+1 is the continuous frame at time t+1, DOF t is the dense optical flow at time t; by iterating the above formula for each frame of the continuous frame of the video stream, n-1 dense optical flow sequences are obtained, where n represents the length of the video.

3. The dual-stream gated violence detection method based on dilated 3D convolutional network and Transformer according to claim 2 is characterized in that: In S2, the dense optical flow is divided into m segments by a sliding window method, and the suppression effect is enhanced by accumulating the optical flow vector amplitude of all frames. The formula is: Where |DOF t | is the DOF frame at time t∈[i,j], i and j are indices, indicating the start frame and end frame in the kth DOF segment, A k To accumulate all frames; the optical flow vector amplitude of all accumulated frames is obtained by normalizing the accumulated optical flow amplitude through the minimum-maximum normalization MMN, scaling it to a value between 0 and 1, and using the normalized accumulated value as a scaling factor to perform element-by-element multiplication and adjustment of the RGB fragment to suppress non-motion elements. The formula is: in is the i-th frame of the k-th RGB fragment, is the suppressed RGB frame, Get the suppressed SRGB fragment for element-wise multiplication over all spatial dimensions and color channels.

4. The dual-stream gated violence detection method based on dilated 3D convolutional network and Transformer according to claim 1, characterized in that: In the S3, the spatiotemporal features are extracted from the SRGB modality and the dense optical flow modality through the pre-trained I3D model to obtain the fragment spatiotemporal features, and the fragment spatiotemporal features are fused by element-by-element multiplication to generate a comprehensive fragment spatiotemporal feature map. The fragment spatiotemporal feature map is connected along the time dimension to form two spatiotemporal feature sequences with a shape of l×1024, where l is the sum of all fragment output logical values ​​determined by the fragment size and the number of fragments.

5. The dual-stream gated violence detection method based on dilated 3D convolutional network and Transformer according to claim 1 or 4, characterized in that: The Transformer encoder in S4 receives the output X of the spatiotemporal modeling of the fragment; the output X is mapped to query Q, key K and value V respectively, and the mapping relationship is as follows: Q = XW Q ,k=XW K ,V=XW V , where W Q ,W K ,W V are parameter matrices used to map the input X to the specific space of query, key, and value respectively; the Transformer encoder is used to perform temporal modeling on the spatiotemporal feature sequence of each video clip to generate a temporal feature map.

6. The dual-stream gated violence detection method based on dilated 3D convolutional network and Transformer according to claim 5, characterized in that: The multi-head self-attention mechanism of the Transformer encoder is used to select video clips of violent activities; the multi-head self-attention mechanism is transformed by the input X, and the formula is: MultiHead(Q,K,V)={Concat(head}1,…,head h )W O , where h represents the number of attention heads, W O The parameter matrix of the multi-head self-attention mechanism, where is a linear transformation of the attention heads of the query Q, key K and value V; the multi-head self-attention mechanism is transformed by the input query Q, key K and value V, and the formula is: where d k is the dimension of each key and softmax is the activation function.

7. The dual-stream gated violence detection method based on dilated 3D convolutional network and Transformer according to claim 6, characterized in that: The multi-head self-attention mechanism of the Transformer encoder is normalized by outputting X and transmitted through the residual connection, and the formula is: Z = LayerNorm(X+MultiHead(Q,K,V)), where Z is the output result of the multi-head self-attention mechanism, the normalized intermediate representation, the intermediate state in the Transformer encoder, and carries the global feature information of the current layer for the input X; each position in Z is independently processed by the feedforward network FFN, and the formula is: FFN(z) = max(0,zW1+b1)W2+b2, where W1 and W2 are the weight parameters of the feedforward network, and b1 and b2 are bias coefficients; the output of the feedforward network FFN is normalized and added back to Z through another residual connection, and the formula is: Y = LayerNorm(Z+FFN(Z)), where Y is the output value, which combines the global features in Z and retains part of the information of the input Z through the residual connection.

8. The dual-stream gated violence detection method based on dilated 3D convolutional network and Transformer according to claim 1, characterized in that: In S5, average pooling is performed through a 1D convolutional neural network. The input data is sent to the 1D convolutional neural network architecture in the time dimension, and a sliding window operation is performed on the time series. For the data points in each window, the average value is calculated to realize the downsampling process of the time feature; the dimension of the input data is compressed and a tensor is output.

9. The dual-stream gated violence detection method based on dilated 3D convolutional network and Transformer according to claim 8, characterized in that: After the 1D convolutional neural network is completed with average pooling, the data enters a multi-layer perceptron for classification, and the output tensor obtained by average pooling is flattened into a vector, and the vector is sent to the 1D convolutional neural network as the input of the multi-layer perceptron; the multi-layer perceptron is composed of multiple fully connected layers, and the input vector and the neurons of the multi-layer perceptron are linearly transformed through a weight matrix, and nonlinearly transformed through an activation function. The transformed result is continuously passed to the next layer, and the above linear transformation and nonlinear transformation processes are repeated. After multiple layers of processing, the data reaches the output layer, corresponding to the classification output of non-violent or violent behavior in the video, and the video classification is determined as violent or non-violent category through the activation values ​​of two neurons in the output layer.

10. A dual-stream gated violence detection system based on an expanded 3D convolutional network and Transformer, characterized in that: It includes a dual-stream gating network FGITNet model; the dual-stream gating network FGITNet model includes a data preprocessing module, a background suppression module, a fragment spatiotemporal modeling module, a fragment time modeling module and a classification module; The data preprocessing module performs standardization processing on the input surveillance video, adjusts the resolution pixels of the image, normalizes the image data by Z-score, and extracts a fixed number of frames by uniform sampling to generate standardized video frame data; The background suppression module suppresses non-moving elements in the video, extracts segments from the RGB frame and the dense optical flow frame using a sliding window method, calculates the amplitude of motion information in the dense optical flow segment, adjusts the optical flow amplitude to a value between 0 and 1 through minimum-maximum normalization, and adjusts the RGB segment by element-by-element multiplication; The fragment spatiotemporal modeling module uses the pre-trained I3D model to extract spatiotemporal features from the suppressed RGB modality and dense optical flow modality, independently models spatiotemporal features on each fragment of length clip_size, and fuses the two modal features by element-by-element multiplication to form a comprehensive spatiotemporal feature map of the fragment; The fragment temporal modeling module adopts the Transformer encoder structure, introduces a multi-head self-attention mechanism, takes the spatiotemporal feature sequence output by the fragment spatiotemporal modeling module as input, and through linear transformation, multi-head self-attention calculation, residual connection and feedforward network processing operations, enables the FGITNet model to selectively observe fragments of violent activities and generate a temporal feature map; The classification module downsamples the temporal features through 1D average pooling, compresses the dimension of the input data, outputs a tensor, and then flattens it into a vector, processes it through a multi-layer perceptron containing multiple fully connected layers, and outputs a classification result of whether each video clip is non-violent or violent behavior; The video frame data processed by the data preprocessing module enters the background suppression module, and the data processed by the background suppression module is sent to the fragment spatiotemporal modeling module. The spatiotemporal feature sequence generated by the fragment spatiotemporal modeling module is passed as input to the fragment time modeling module, and the time feature map output by the fragment time modeling module is then sent to the classification module, and the video classification result is output after being processed by the classification module.

Citation Information

Patent Citations

  • Violence detection system based on time-space information credible fusion

    CN117876746A

  • Long duration structured video action segmentation

    US20240104915A1

  • Deep learning-based intelligent recognition method and system for locomotive signboard information

    WO2024131380A1

Cited By

  • Two-photon photoetching part quality inspection method based on three-dimensional shift window multi-head self-attention

    CN120219394A

  • Double-flow circulation attention dynamic gesture recognition method based on motion guidance

    CN120220253A