Behavior recognition method and system based on adaptive sliding window multi-head self-attention

Through the combination of adaptive sliding window multi-head self-attention and residual Transformer-Block, the problem that 3D convolutional neural networks are difficult to capture long-term dependencies in behavior recognition is solved, achieving higher accuracy and robustness, and improving the effect of behavior recognition.

CN120340111APending Publication Date: 2025-07-18SHANDONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510217113.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing 3D convolutional neural networks are difficult to capture long-term dependencies in videos in behavior recognition, and the fixed sliding window position leads to target splitting, which is less robust.

Method used

The Video-Aswin-Residual-Transformer behavior recognition model of adaptive sliding window multi-head self-attention and residual Transformer-Block is adopted. Through the adaptive sliding window offset and residual attention connection module, feature multiplexing and robustness are enhanced, and important features are transmitted across layers.

Benefits of technology

It improves the accuracy of behavior detection and the generalization ability of the model, can better capture the long-term dependence and global features in the video, and improves the efficiency and accuracy of behavior recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340111A_ABST
    Figure CN120340111A_ABST
Patent Text Reader

Abstract

The invention relates to a behavior recognition method and system based on adaptive sliding window multi-head self-attention, and the method comprises the steps: preprocessing a training set of a human behavior video data set, inputting the preprocessed training set into a neural network model for training, preprocessing a test set, and inputting the preprocessed test set into the trained neural network model for behavior detection, obtaining the detection accuracy of the neural network model; and inputting to-be-recognized human body behavior video data into the tested neural network model for behavior recognition. According to the method, the long-time dependency and global features in the video can be captured, so that the extracted features are more comprehensive, the problems that the receptive field of the convolutional neural network is limited and the long-time dependency relationship in the video is difficult to capture can be effectively solved, and the recognition accuracy of various behaviors is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a behavior recognition method and system based on adaptive sliding window multi-head self-attention, belonging to the technical fields of deep learning and behavior recognition. Background Art

[0002] With the continuous development of artificial intelligence technology, behavior recognition has become an important research direction in the field of computer vision. By analyzing the actions of people in videos, the behaviors they perform can be inferred. In the field of video surveillance, behavior recognition can detect illegal behaviors of people, help security personnel discover abnormal behaviors in a timely manner, and eliminate potential risks; in the field of autonomous driving, behavior recognition can identify the traffic conditions around vehicles, make corresponding driving decisions, and improve driving safety and efficiency; in the field of health monitoring, behavior recognition can monitor the daily activities of the elderly or patients, and assist medical staff to provide more user-friendly health care services. Therefore, automatically capturing and analyzing human behaviors in a non-intrusive manner based on surveillance videos to achieve all-day continuous behavior detection has broad application prospects.

[0003] In recent years, with the rapid development of deep learning and neural networks, human behavior recognition algorithms based on 3D convolutional neural networks have been widely used. Based on 2D convolution, 3D convolution takes into account the temporal continuity, and the convolutional kernel contains convolutional operations in both temporal and spatial dimensions, which can extract spatio-temporal features to improve the accuracy of behavior detection. There are some drawbacks in applying convolutional networks in the field of behavior recognition. For example, the receptive field of 3D convolution is limited, and it is difficult to capture long-term dependencies in videos. 3D convolutional neural networks gradually extract features of different scales through multiple layers of convolutional kernels, but each layer still mainly focuses on local information. With the rise of the attention mechanism, the Video-Swin-Transformer model based on the self-attention mechanism can naturally capture long-term dependencies and global features in videos, can effectively address the problems of 3D convolutional neural networks, and has better performance in behavior recognition.

[0004] Video-Swin-Transformer is an extension of Swin-Transformer. This model is based on the self-attention mechanism and can effectively process spatio-temporal information. In order to capture cross-window information and perform information interaction between window boundaries, the Video-Swin-Transformer introduces a sliding window multi-head self-attention mechanism in the network layer. This sliding window mechanism enables the model to capture the global information of images at multiple scales. However, the fixed position of the sliding window is likely to cause target fragmentation and has low robustness for extracting different video features. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a behavior recognition method based on adaptive sliding window multi-head self-attention;

[0006] The present invention also provides a behavior recognition system based on adaptive sliding window multi-head self-attention;

[0007] The present invention proposes a Video-Aswin-Residual-Transformer behavior recognition model based on adaptive sliding window multi-head self-attention and residual Transformer-Block. Compared with the fixed window position, the present invention can adaptively offset the sliding window. In order to enhance feature reuse and prevent the loss of low-level key feature information, a residual attention connection module is introduced in the network layer, and important features are transmitted across layers through residual connections, effectively transmitting feature information across layers, improving the model's expression ability and generalization ability, and thus improving the accuracy of behavior detection.

[0008] By analyzing video datasets containing various human behaviors, the present invention detects and recognizes various behaviors, which helps to promote the implementation of monitoring video behavior detection applications in key places, improve the detection efficiency and accuracy, timely discover abnormal behaviors, eliminate potential risks in a timely manner, and safeguard the personal safety of personnel and the stability of the place.

[0009] In order to more flexibly explore the correlation between different windows in the key areas of the video, focus more on the key positions to calculate local attention weights, and enable the network to better utilize multi-level features, the present invention constructs a Video-Aswin-Residual-Transformer behavior recognition model based on adaptive sliding window multi-head self-attention, which adaptively offsets the sliding window. Compared with the fixed window position, the adaptive sliding window can perform sliding processing operations according to the position changes of the people in the video. The shallow key features are directly transmitted to the deep layer through skip connections, improving the generalization ability and robustness of the model, and further enhancing the model accuracy.

[0010] Term Explanation:

[0011] 1. Swin-Transformer, that is, Hierarchical Vision Transformer using ShiftedWindows (Hierarchical Vision Transformer based on sliding windows) realizes cross-window information interaction through sliding windows, avoiding the limitations of local modeling.

[0012] 2. Video-Swin-Transformer is a neural network designed specifically for video understanding based on Swin-Transformer, which realizes the expansion of 2D windows into 3D spatio-temporal windows, thereby expanding into spatio-temporal modeling to capture spatial action details and temporal dependencies (such as action continuity).

[0013] 3. Video-Aswin-Residual-Transformer is a behavior recognition model proposed in the present invention, which is constructed based on adaptive sliding window multi-head self-attention and residual Transformer-Block. The adaptive sliding window multi-head self-attention can adaptively offset the sliding window to ensure the integrity of the action information inside the window. The residual attention connection module in the residual Transformer-Block can cross-layer transfer the important features assigned with channel weights by the 3D-SE attention mechanism, thereby realizing feature reuse and preventing the loss of key feature information at the lower layer.

[0014] The technical solution of the present invention is as follows:

[0015] A behavior recognition method based on adaptive sliding window multi-head self-attention, including:

[0016] Preprocess the training set of the human behavior video dataset and input it into the neural network model for training. Preprocess the test set and input it into the trained neural network model for behavior detection to obtain the detection accuracy of the neural network model;

[0017] Input the human behavior video data to be recognized into the tested neural network model for behavior recognition.

[0018] According to the preference of the present invention, the neural network model includes a 3DPatch-Partition layer connected in sequence, a stage1 composed of a Linear-Embe ding layer and a Video-Aswin-Residual-Transformer-Block module, a stage2, a stage3, and a stage4 composed of a Patch-Merging layer and a Video-Aswin-Residual-Transformer-Block module, an average pooling layer, a fully connected layer, and a softmax function layer;

[0019] The 3DPatch-Partition layer cuts the input video frame picture into several patches to realize the adjustment of size and the increase of channels;

[0020] The Linear-Embeding layer linearly maps the input image to a 128-dimensional vector, which serves as the input vector for calculating the self-attention weights of local windows in subsequent computations;

[0021] The Patch-Merging layer merges adjacent 2×2 patches into one patch, implementing a downsampling operation similar to that of a convolutional neural network;

[0022] The Video-Aswin-Residual-Transformer-Block module is used to calculate the self-attention weights within different windows and model the dependencies related to action information.

[0023] According to the preference of the present invention, the Video-Aswin-Residual-Transformer-Block module includes a first normalization layer, a window multi-head attention mechanism layer, a second normalization layer, a first multi-layer perceptron, a third normalization layer, an adaptive sliding window multi-head attention mechanism, a fourth normalization layer, a second multi-layer perceptron, and a residual attention branch with cross-layer connection connected in sequence;

[0024] The first normalization layer normalizes the feature information;

[0025] The window multi-head attention mechanism layer divides the input feature map into non-overlapping 3D windows and calculates the self-attention weights within each window to extract local spatio-temporal features;

[0026] The adaptive sliding window multi-head attention mechanism divides the input feature map into non-overlapping 3D windows, and on this basis, the windows slide adaptively according to the input feature map, and then calculate the self-attention weights within each window to extract global spatio-temporal features;

[0027] The multi-layer perceptron includes two fully connected layers, which perform non-linear transformation and enhancement on the features after the self-attention mechanism operation;

[0028] The calculation formula of the Video-Aswin-Residual-Transformer-Block module is shown in formulas (1), (2), (3), and (4):

[0029]

[0030] where z l is the input layer, is the intermediate layer, is the output layer, LN is the first normalization layer, MLP is the multi-layer perceptron, and SE is the attention mechanism; z l-1 and z l+1 are the input and output of the adaptive sliding window multi-head attention mechanism respectively.

[0031] Preferably according to the present invention, the window multi-head attention mechanism layer is used to calculate local self-attention after dividing the input feature map into windows of a fixed size.

[0032] Preferably according to the present invention, the adaptive sliding window includes a fixed sliding window and a sliding window offset calculation module. The sliding window offset calculation module includes a convolutional layer, a fully connected layer, and a Sigmoid function layer. The offset of the sliding window is calculated according to different video inputs to offset the sliding window.

[0033] Preferably according to the present invention, the window multi-head attention mechanism layer calculates local self-attention, which means that each patch in the same window is calculated with other patches, and is calculated through the Q vector, K vector, and V vector. All patches within the same window share the Q vector, K vector, and V vector; it includes:

[0034] ① Create three Q vectors, K vectors, and V vectors based on the input vector. The input vector is the vector obtained by linearly mapping the feature map to the channel dimension; the Q vector, K vector, and V vector are calculated as shown in equations (5), (6), and (7):

[0035]

[0036] where W Q 、W K 、W V are weights, represents the output of the previous layer, and LN represents the Layer Normalization layer;

[0037] ② Generate self-attention weights α through the Q vector and K vector; as shown in equation (8):

[0038]

[0039] In equation (8), represents the square root of the K dimension, B is the relative position offset, and softmax is the normalization exponential function;

[0040] ③ Multiply the self-attention weights α by the V vector to obtain the output vector, that is, complete the self-attention mechanism calculation.

[0041] Preferably according to the present invention, the residual attention branch assigns weights to the features by the SE module. The SE module calculates the importance of each channel in the feature map through a series of operations and assigns corresponding attention weights to each channel; the SE module includes a compression operation and an excitation operation:

[0042] During the compression operation, the input feature map is compressed into a vector through global average pooling, and this vector contains the global spatial information of each channel;

[0043] During the excitation operation, the complex dependencies between channels are modeled through a bottleneck structure including two fully connected layers. The first fully connected layer reduces the channel dimension and introduces non-linearity through the ReLU activation function; the second fully connected layer restores the channel dimension to its original size and generates the weight value for each channel through the Sigmoid activation function; finally, the output weight values are multiplied with the input feature map channel by channel to obtain the weighted feature map, realizing the dynamic adjustment of different channels.

[0044] Preferably according to the present invention, the calculation of the residual attention branch is shown in formula (9):

[0045] x t+1 = SE(h(x t )) + F(x t ) #(9);

[0046] In formula (9), x t represents the input of the residual attention unit, and x t+1 represents the output of the residual attention unit; h(x t ) = x t represents the identity mapping relationship; SE represents the attention function; F is the non-linear residual function.

[0047] Preferably according to the present invention, the specific implementation process of the preprocessing includes the following steps:

[0048] Dataset division: The public dataset is divided into a training set and a test set; each video segment is labeled;

[0049] Frame processing: Each video in the dataset is frame-processed to extract each frame image of the video;

[0050] Segment sampling: Segment sampling is performed on the frame-processed video, and the sampled segments are used as input data and fed into the neural network model for training.

[0051] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of a behavior recognition method based on adaptive sliding window multi-head self-attention.

[0052] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of a behavior recognition method based on adaptive sliding window multi-head self-attention.

[0053] A behavior recognition system based on adaptive sliding window multi-head self-attention, comprising:

[0054] Neural network model construction and training module: The training set of the human behavior video dataset is preprocessed and then input into the neural network model for training. The test set is preprocessed and then input into the trained neural network model for behavior detection to obtain the detection accuracy of the neural network model.

[0055] Behavior recognition module: The human behavior video data to be recognized is input into the neural network model after training and testing for behavior recognition, and the recognition result is output.

[0056] The beneficial effects of the present invention are as follows:

[0057] 1. The present invention provides 3D-ASW-MSA (Adaptive Sliding Window Multi-Head Attention Mechanism), which can adaptively adjust the offset of window sliding according to different video inputs. Compared with the traditional 3D-SW-MSA (Sliding Window Multi-Head Attention Mechanism) that slides the window at a fixed position, it has a certain degree of flexibility in video modeling. It can focus more on key positions to calculate local attention weights according to the position changes of people in the video, improving the generalization ability of the model.

[0058] 2. The present invention provides a residual attention branch module. By introducing the attention mechanism, the model can more effectively focus on the important parts of the input data, suppress irrelevant information, thereby improving the accuracy of feature extraction. It directly transmits shallow features to the deep layer, enabling the model to reuse these features, avoiding information loss, and being able to integrate features at different levels, enhancing the model's ability to process multi-scale information.

[0059] 3. The present invention provides a Video-Aswin-Residual-Transformer behavior recognition method based on the adaptive sliding window multi-head self-attention module and the residual attention module, which can capture long-term dependencies and global features in the video, making the extracted features more comprehensive. It can effectively address the problem that the receptive field of the convolutional neural network is limited and it is difficult to capture long-term dependency relationships in the video, significantly improving the accuracy of various behavior recognitions. Description of the Drawings

[0060] Figure 1 It is a schematic structural diagram of the behavior recognition model constructed based on the adaptive sliding window multi-head self-attention and the residual Transformer-Block of the present invention

[0061] Figure 2 Based on the adaptive sliding window multi-head self-attention and the residual attention branch block of the present invention

[0062] Figure 3Schematic diagram of the window division operation of the 3D-W-MSA (Window Multi-Head Attention Mechanism) and 3D-SW-MSA (Sliding Window Multi-Head Attention Mechanism) of the present invention;

[0063] Figure 4 Schematic diagram of the adaptive sliding window multi-head self-attention mechanism of the present invention;

[0064] Figure 5 Schematic diagram of the recognition accuracy rate of the present invention on the HMDB51 validation set;

[0065] Figure 6 Schematic diagram of the recognition accuracy rate of the present invention on the UCF101 validation set;

[0066] Figure 7 Schematic diagram of the confusion matrix (the top 20 categories with the highest recognition accuracy rate) of the present invention on HMDB51;

[0067] Figure 8 Schematic diagram of the confusion matrix (the top 20 categories with the highest recognition accuracy rate) of the present invention on UCF101. Detailed implementation manners

[0068] The present invention will be further limited below in conjunction with the accompanying drawings of the specification and embodiments, but not limited thereto.

[0069] Embodiment 1

[0070] A behavior recognition method based on adaptive sliding window multi-head self-attention, comprising:

[0071] Preprocess the training set of the human behavior video dataset and input it into the neural network model for training, preprocess the test set and input it into the trained neural network model for behavior detection to obtain the detection accuracy rate of the neural network model;

[0072] Input the human behavior video data to be recognized into the tested neural network model for behavior recognition.

[0073] Embodiment 2

[0074] A behavior recognition method based on adaptive sliding window multi-head self-attention according to Embodiment 1, wherein the difference lies in:

[0075] As Figure 1As shown in the figure, the neural network model includes a 3DPatch-Partition layer (block segmentation layer) connected in sequence, stage1 composed of a Linear-Embeding layer (linear embedding layer) and a Video-Aswin-Residual-Transformer-Block module, stage2, stage3, and stage4 composed of a Patch-Merging layer (block merging layer) and a Video-Aswin-Residual-Transformer-Block module (based on an adaptive sliding window multi-head self-attention and residual attention branch block), an average pooling layer, a fully connected layer, and a softmax function layer;

[0076] The 3DPatch-Partition layer cuts the input video frame picture into several patches, realizing the adjustment of size and the increase of channels;

[0077] The Linear-Embeding layer linearly maps the input image to a 128-dimensional vector, serving as the input vector for calculating the local window self-attention weights in subsequent calculations;

[0078] The Patch-Merging layer merges adjacent 2×2 patches into one patch, realizing a downsampling operation similar to that of a convolutional neural network;

[0079] The Video-Aswin-Residual-Transformer-Block module is used to calculate the self-attention weights within different windows and model the dependency relationships related to action information.

[0080] As shown in Table 1, it is the size change of the feature map in each layer of the network.

[0081] Table 1

[0082] Stage Time dimension Spatial dimension Number of channels Input 32 224×224 3 3D-Patch-Partition 16 56×56 128 Stage1 16 56×56 128 Stage2 16 28×28 256 Stage3 16 14×14 512 Stage4 16 7×7 1024

[0083] As Figure 2 shown, the Video-Aswin-Residual-Transformer-Block module includes a first normalization layer (Layer-Normalization), a window multi-head attention mechanism layer (3D-W-MSA), a second normalization layer (Layer-Normalization), a first multi-layer perceptron (MLP), a third normalization layer (Layer-Normalization), an adaptive sliding window multi-head attention mechanism (3D-ASW-MSA), a fourth normalization layer (Layer-Normalization), a second multi-layer perceptron (MLP), and a residual attention branch connected across layers;

[0084] The first layer normalization layer (Layer-Normalization) normalizes the feature information to prevent gradient vanishing or explosion;

[0085] The window multi-head attention mechanism layer (3D-W-MSA) divides the input feature map into non-overlapping 3D windows and calculates the self-attention weights within each window to extract local spatio-temporal features;

[0086] The adaptive sliding window multi-head attention mechanism (3D-ASW-MSA) divides the input feature map into non-overlapping 3D windows, and on this basis, the windows slide adaptively according to the input feature map, and then calculate the self-attention weights within each window to extract global spatio-temporal features;

[0087] The multi-layer perceptron (MLP) includes two fully connected layers to perform non-linear transformation and enhancement on the features after the self-attention mechanism operation; improving the expressive ability of the model.

[0088] The calculation formulas of the Video-Aswin-Residual-Transformer-Block module are shown in Formulas (1), (2), (3), and (4):

[0089]

[0090] Formulas (1), (2), (3), and (4) are respectively Figure 2 the calculation formulas of the input, output, and intermediate layers, and specific references can be made to Figure 2 .

[0091] Among them, z l is the input layer, is the intermediate layer, is the output layer, LN is the first layer normalization layer (normalization layer), MLP is the multi-layer perceptron, and SE is the attention mechanism; z l-1 and z l+1 are respectively the input and output of the adaptive sliding window multi-head attention mechanism.

[0092] The window multi-head attention mechanism layer is used to divide the input feature map into windows of a fixed size (such as Figure 3After calculating the local self-attention as shown in layer1. With a window size of 8×7×7, the feature map is divided into multiple non-overlapping small windows, and the pixels or features within each window are adjacent. For the features within each window, local self-attention is calculated. This means that the features at each position within the window interact with the features at other positions within the same window to generate attention weights. This local self-attention calculation can capture the local context information within the window. 3D-W-MSA (Window Multi-Head Attention Mechanism) significantly reduces the computational complexity compared to 3D MSA (Multi-Head Attention Mechanism). Their computational complexities are shown as follows:

[0093] Ω(MSA) = 4hwtc 2 + 2(hwt) 2 C;

[0094] Ω(W-MSA) = 4hwt 2 C 2 + 2M 2 hwtC;

[0095] Among them, h represents the height of the feature map, w represents the width of the feature map, t represents the depth of the feature map, C represents the number of channels of the feature map, and M represents the size of each window.

[0096] 3D-ASW-MSA (Adaptive Sliding Window Multi-Head Attention Mechanism) is an improvement based on 3D-SW-MSA (Sliding Window Multi-Head Attention Mechanism). The window partitioning methods of 3D-W-MSA (Window Multi-Head Attention Mechanism) and 3D-SW-MSA (Sliding Window Multi-Head Attention Mechanism) are as Figure 3 shown. By comparing the left and right figures, it can be found that the window has shifted (it can be understood that the window has shifted pixels to the right and down respectively from the upper left corner). The fixed position partitioning of the window in 3D-SW-MSA (Sliding Window Multi-Head Attention Mechanism) easily leads to target fragmentation and incomplete action information within the window. To solve this problem, the present invention proposes 3D-ASW-MSA (Sliding Window Multi-Head Attention Mechanism), that is, the sliding window can be adaptively offset. The adaptive sliding window multi-head attention mechanism is as Figure 4 shown. The adaptive sliding window includes a fixed sliding window and a sliding window offset calculation module. The sliding window offset calculation module includes a convolutional layer, a fully connected layer, and a Sigmoid function layer. The offset of the sliding window is calculated according to different video inputs to offset the sliding window. For the convenience of calculation, the part obtained after processing by the sliding window is moved, and then the redundant information is filtered out through a masking operation for multi-head attention calculation.

[0097] The window multi-head attention mechanism layer calculates local self-attention, which means that each patch in the same window is calculated with other patches, and is calculated through the Q (query) vector, K (key) vector, and V (value) vector. All patches within the same window share the Q vector, K vector, and V vector. It includes:

[0098] ① Create three Q vectors, K vectors, and V vectors based on the input vector. The input vector is the vector obtained by linearly mapping the feature map to the channel dimension. The calculation of the Q vector, K vector, and V vector is shown in equations (5), (6), and (7):

[0099]

[0100]

[0101] Among them, W Q 、W K 、W V are weights, which are parameters trained in the neural network, represents the output of the previous layer, and LN represents the Layer Normalization layer;

[0102] ② Generate self-attention weight α through the Q vector and K vector; as shown in equation (8):

[0103]

[0104] In equation (8), represents the square root of the K dimension, B is the relative position offset, and the role of B is to add a value to each element in the feature map. Its essence is to hope that the feature map can be further biased, because the lower a value in the feature map, after softmax, this value will be lower, and ultimately the contribution to the feature will be lower. Softmax is the normalized exponential function; it can convert the output values of multi-classification into numerical values within the range of [0,1], and the probability distribution with a sum of 1, aiming to show the results of multi-classification in the form of probability.

[0105] ③ Multiply the self-attention weight α by the V vector to obtain the output vector, that is, complete the self-attention mechanism calculation.

[0106] The residual attention branch assigns weights to the features by the SE module. The SE module calculates the importance of each channel in the feature map through a series of operations and assigns corresponding attention weights to each channel; thus enabling the convolutional network to pay more attention to the feature channels useful for the current task while suppressing those channels that contribute less to the task. The SE module includes a squeeze operation and an excitation operation:

[0107] In the compression operation, the input feature map is compressed into a vector through global average pooling. This vector includes the global spatial information of each channel, reflecting the overall feature distribution of each channel.

[0108] In the excitation operation, a bottleneck structure consisting of two fully connected layers is used to model the complex dependencies between channels. This design introduces more non-linearity and can better capture the correlations between channels. Specifically, the first fully connected layer reduces the channel dimension and introduces non-linearity through the ReLU activation function; the second fully connected layer restores the channel dimension to its original size and generates the weight value for each channel through the Sigmoid activation function; finally, the output weight value is multiplied by the input feature map channel by channel to obtain the weighted feature map, realizing the dynamic adjustment of different channels.

[0109] The calculation of the residual attention branch is shown in Equation (9):

[0110] x t+1 = SE(h(x t )) + F(x t ) #(9);

[0111] In Equation (9), x t represents the input of the residual attention unit, and x t+1 represents the output of the residual attention unit; h(x t ) = x t represents the identity mapping relationship; SE represents the attention function; F is the non-linear residual function.

[0112] The training process of the neural network model includes the following steps:

[0113] Step 1: Data Download and Preprocessing; Download the HMDB51 dataset and UCF101 dataset from official channels and preprocess the data. The HMDB51 dataset was released by Brown University in 2011, containing 51 action categories with a total of 6,849 videos, and each action category contains at least 100 videos. The action categories cover the following: common facial actions (such as smiling, chewing, speaking); complex facial actions (such as smoking, eating, drinking); common limb actions (such as climbing, diving, jumping); complex limb actions (such as combing hair, grasping, drawing a sword); multi-person interaction actions (such as hugging, kissing, shaking hands). The UCF101 dataset was released by the Computer Vision Research Center of the University of Central Florida in 2012, containing 101 action categories with a total of 13,320 videos, and each action category contains at least 100 videos. The action categories cover the following: human-object interaction (such as applying eye makeup, brushing teeth); simple limb actions (such as hammering, doing push-ups, walking); human-human interaction (such as sumo wrestling, military parade, wrestling); playing musical instruments (such as playing the guitar, playing the piano, playing the violin); sports (such as skiing, shot put, skydiving). After downloading the dataset, divide the dataset into a training set and a test set according to the provided partition file, and the sample quantity ratio of the training set to the test set is 7:3; perform frame splitting on all video clips and add corresponding action labels to each frame image to complete the data annotation work;

[0114] Step 2: Training of the Neural Network Model: Input the preprocessed training set in Step 1 into the neural network model for supervised training, and finally obtain a trained neural network model; the specific implementation process is as follows:

[0115] Initialization of the Neural Network Model: Load some network parameters of the neural network model using the pre-trained parameters of the Video-Swin-Transformer model; for the parameters without pre-training, use the Kaiming initialization method for initialization; initialize the learning rate to 0.0001 and adopt the warmup strategy to warm up the learning rate to stabilize the convergence process in the initial stage of training;

[0116] Training Process: The neural network model is trained for a total of 30 epochs. In each epoch of training, all samples in the training set are evenly divided into S subsets, and each subset contains the same number of samples (each subset contains 2 samples). In the forward propagation process, the neural network model calculates the output value according to the input sample and compares it with the label value of the sample to calculate the loss function value. Take the average of the loss function values of all samples in each subset to obtain the average loss value of this subset. Adopt the backpropagation algorithm and combine the AdamW optimizer to dynamically adjust the model parameters to make the loss function value gradually converge to the optimal value.

[0117] Neural Network Model Evaluation and Tuning: After each round of training, use the test set to evaluate the action detection accuracy of the neural network model. According to the accuracy results of the test set, adjust the hyperparameters of the model (such as learning rate, regularization parameter, etc.) to further optimize the model performance. Through multiple rounds of training and tuning, finally obtain a neural network model with the optimal detection accuracy.

[0118] The specific implementation process of preprocessing includes the following steps:

[0119] Dataset Partitioning: According to the provided partitioning files by the official, partition the public datasets (such as HMDB51 and UCF101) into a training set and a test set; the sample quantity ratio of the training set to the test set is 7:3. Annotate each video clip; ensure that its action category label is consistent with the original video for use in supervised training.

[0120] Frame Processing: Perform frame processing on each video in the dataset to extract each frame image of the video;

[0121] Segment Sampling: Conduct segment sampling from the framed video. Take 1 frame every 2 frames and continuously intercept 32 frames each time as a video segment. Use the sampled segments as input data and feed them into the neural network model for training.

[0122] Through the above preprocessing steps, convert the original video data into a format suitable for neural network model training, providing high-quality data input for subsequent neural network model training.

[0123] The experimental results of this embodiment are shown in Table 2:

[0124] Table 2

[0125] Common models HMDB51 UCF101 VidTr-L 74.4% 96.7% AMD 79.6% 97.1% D3D+D3D 80.5% 97.6% Multi-stream I3D 80.92% 97.2% MARS+RGB+FLow 80.9% 97.8% Two-stream I3D 80.7% 97.8% The neural network model of the present invention HMDB51 UCF101 Video-Swin-Transformer 80.21% 98.15% Video-Aswin-Transformer 80.76% 98.48% Video-Swin-Residual-Transformer 80.85% 98.56% Video-Aswin-Residual-Transformer 81.02% 98.73%

[0126] For parameters without pre-training, use the Kaiming initialization method for initialization, initialize the learning rate to 0.0001, and at the same time use the warmup strategy to warm up the learning rate to stabilize the convergence process in the initial stage of training. The model is trained for a total of 30 epochs. In each training round, the training set samples of the two datasets (HMDB51 and UCF101) are evenly divided into S subsets (S = total number of training set samples / 2), and each subset contains 2 samples. Finally, the Video-Aswin-Residual-Transformer action recognition model proposed by the present invention based on the adaptive sliding window multi-head self-attention module is significantly superior to the Video-Swin-Transformer model in terms of performance. Specifically, as Figure 5 shown, the accuracy on the HMDB51 dataset has increased by 0.81%, asFigure 6 As shown, the accuracy rate on the UCF101 dataset has increased by 0.58%. Figure 5 and Figure 6 In [figure], the abscissa is the number of training rounds, and the ordinate is the recognition accuracy rate of the verification machine; as Figure 7 shown, it is the confusion matrix of the recognition accuracy rate of the Video-Aswin-Residual-Transformer behavior recognition model on the HMDB51 verification set (the top 20 categories with the highest recognition accuracy rate), and the top 20 categories with the highest accuracy rate are all 100%; as Figure 8 shown, it is the confusion matrix of the recognition accuracy rate of the Video-Aswin-Residual-Transformer behavior recognition model on the HMDB51 verification set (the top 20 categories with the highest recognition accuracy rate), and the top 8 categories with the highest accuracy rate are all 100%. Figure 7 and Figure 8 In [figure], the abscissa is the true action category, and the ordinate is the action category recognized by the model.

[0127] Example 3

[0128] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of a behavior recognition method based on adaptive sliding window multi-head self-attention described in Example 1 or 2.

[0129] Example 4

[0130] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of a behavior recognition method based on adaptive sliding window multi-head self-attention described in Example 1 or 2.

[0131] Example 5

[0132] A behavior recognition system based on adaptive sliding window multi-head self-attention includes:

[0133] A neural network model construction and training module: preprocess the training set of the human behavior video dataset and input it into the neural network model for training, preprocess the test set and input it into the trained neural network model for behavior detection, and obtain the detection accuracy rate of the neural network model;

[0134] A behavior recognition module: input the human behavior video data to be recognized into the neural network model after training and testing for behavior recognition, and output the recognition result.

Claims

1. A behavior recognition method based on adaptive sliding window multi-head self-attention, characterized in that, Including: Preprocess the training set of the human behavior video dataset and input it into the neural network model for training. Preprocess the test set and input it into the trained neural network model for behavior detection to obtain the detection accuracy of the neural network model; Input the human behavior video data to be recognized into the tested neural network model for behavior recognition.

2. The behavior recognition method based on adaptive sliding window multi-head self-attention according to claim 1, characterized in that The neural network model includes a 3DPatch-Partition layer connected in sequence, a stage1 composed of a Linear-Embeding layer and a Video-Aswin-Residual-Transformer-Block module, a stage2, stage3, and stage4 composed of a Patch-Merging layer and a Video-Aswin-Residual-Transformer-Block module, an average pooling layer, a fully connected layer, and a softmax function layer; The 3DPatch-Partition layer cuts the input video frame image into several patches to adjust the size and increase the channels; The Linear-Embeding layer linearly maps the input image to a 128-dimensional vector as the input vector for calculating the local window self-attention weights in subsequent calculations; The Patch-Merging layer merges adjacent 2×2 patches into one patch to achieve a downsampling operation similar to that of a convolutional neural network; The Video-Aswin-Residual-Transformer-Block module is used to calculate the self-attention weights within different windows and model the dependencies related to action information.

3. The behavioral recognition method based on adaptive sliding window multi-head self-attention according to claim 2, wherein The Video-Aswin-Residual-Transformer-Block module includes a first normalization layer, a window multi-head attention mechanism layer, a second normalization layer, a first multi-layer perceptron, a third normalization layer, an adaptive sliding window multi-head attention mechanism, a fourth normalization layer, a second multi-layer perceptron, and a residual attention branch with cross-layer connection; The first normalization layer normalizes the feature information; The window multi-head attention mechanism layer divides the input feature map into non-overlapping 3D windows and calculates the self-attention weights within each window to extract local spatio-temporal features; The adaptive sliding window multi-head attention mechanism divides the input feature map into non-overlapping 3D windows, and on this basis, the window slides adaptively according to the input feature map, and then calculates the self-attention weights within each window to extract global spatio-temporal features; The multi-layer perceptron includes two fully connected layers to perform non-linear transformation and enhancement on the features after the self-attention mechanism operation; The calculation formula of the Video-Aswin-Residual-Transformer-Block module is shown in formulas (1), (2), (3), and (4): Among them, z l is the input layer, is the intermediate layer, is the output layer, LN is the first normalization layer, MLP is the multi-layer perceptron, and SE is the attention mechanism; z l-1 and z l+1 are the input and output of the adaptive sliding window multi-head attention mechanism respectively.

4. The behavior recognition method based on adaptive sliding window multi-head self-attention according to claim 1, characterized in that, The window multi-head attention mechanism layer is used to calculate the local self-attention after dividing the input feature map according to windows of a fixed size; The adaptive sliding window includes a fixed sliding window and a sliding window offset calculation module. The sliding window offset calculation module includes a convolutional layer, a fully connected layer, and a Sigmoid function layer, which calculates the offset of the sliding window according to different video inputs and offsets the sliding window.

5. A behavior recognition method based on adaptive sliding window multi-head self-attention according to claim 4, characterized in that The window multi-head attention mechanism layer calculates local self-attention, which means that each patch in the same window is calculated with other patches, and is calculated through the Q vector, K vector, and V vector. All patches within the same window share the Q vector, K vector, and V vector. It includes: ① Create three Q vectors, K vectors, and V vectors based on the input vector. The input vector is the vector obtained by linearly mapping the feature map to the channel dimension. The calculation of the Q vector, K vector, and V vector is shown in equations (5), (6), and (7): Q = W Q · LN(z l-1 )#(5); K = W K · LN(z l-1 )#(6); V = W V ·LN(z l-1 )#(7); Among them, W Q , W K , W V are weights, z l-1 represents the output of the previous layer, and LN represents the Layer Normalization layer; ② Generate the self-attention weight α through the Q vector and the K vector. As shown in equation (8): In formula (8), represents the square root of the K dimension, B is the relative position offset, and softmax is the normalized exponential function; ③ Multiply the self-attention weight α by the V vector to obtain the output vector, that is, complete the self-attention mechanism calculation.

6. A behavior recognition method based on adaptive sliding window multi-head self-attention according to claim 3, characterized in that, The residual attention branch assigns weights to the features by the SE module. The SE module calculates the importance of each channel in the feature map through a series of operations and assigns corresponding attention weights to each channel. The SE module includes a squeezing operation and an excitation operation: In the squeezing operation, the input feature map is compressed into a vector through global average pooling, and this vector includes the global spatial information of each channel. In the excitation operation, a bottleneck structure including two fully connected layers is used to model the complex dependencies between channels. The first fully connected layer reduces the channel dimension and introduces non-linearity through the ReLU activation function. The second fully connected layer restores the channel dimension to the original size and generates the weight value of each channel through the Sigmoid activation function. Finally, the output weight value is multiplied by the input feature map channel by channel to obtain the weighted feature map, realizing the dynamic adjustment of different channels. Further preferably, the calculation of the residual attention branch is shown in formula (9): x t+1 = SE(h(x t )) + F(x t ) #(9); In formula (9), x t represents the input of the residual attention unit, and x t+1 represents the output of the residual attention unit; h(x t ) = x t represents the identity mapping relationship; SE represents the attention function; F is a non-linear residual function.

7. A behavior recognition method based on adaptive sliding window multi-head self-attention according to any one of claims 1-6, characterized in that The specific implementation process of the preprocessing includes the following steps: Dataset division: The public dataset is divided into a training set and a test set; each video segment is labeled. Frame processing: Each video in the dataset is frame-processed to extract each frame image of the video. Segment sampling: Segment sampling is performed on the frame-processed video, and the sampled segments are used as input data and fed into the neural network model for training.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of a behavior recognition method based on adaptive sliding window multi-head self-attention.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of a behavior recognition method based on adaptive sliding window multi-head self-attention.

10. A behavior recognition system based on adaptive sliding window multi-head self-attention, characterized in that, It includes: Neural network model construction and training module: The training set of the human behavior video dataset is preprocessed and input into the neural network model for training. The test set is preprocessed and input into the trained neural network model for behavior detection to obtain the detection accuracy of the neural network model. Behavior recognition module: Input the video data of the human behavior to be recognized into the neural network model after training and testing for behavior recognition, and output the recognition result.

Citation Information

Cited By

  • Smart classroom interaction analysis method based on double-layer architecture voice segmentation

    CN120783757A

  • A Smart Classroom Interaction Analysis Method Based on Two-Layer Architecture Speech Segmentation

    CN120783757B