Human body action recognition method and system based on double-flow self-attention mechanism

By adopting the human body movement recognition method with the dual-flow self-attention mechanism in human body movement recognition, the problem of accuracy and calculation amount of human body movement recognition in complex environments is solved, and the balance of high accuracy and appropriate speed is achieved, with good practical application prospects.

CN120047999AActive Publication Date: 2025-05-27UNIV OF CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510114883.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

When recognizing human movements in complex environments, it is difficult to accurately distinguish the characteristics of human bodies and background areas in videos, and the calculation amount required to improve recognition accuracy is huge, which affects the practicality of the model.

Method used

The human body movement recognition method based on the dual-stream self-attention mechanism is adopted. Multiple modal features of video data are extracted through the basic network of the dual-stream self-attention mechanism model, and fuse them with the time flow. The space-time memory network is input to the space-time memory network for spatial-time features to interact at different scales, and preliminary action classification results are generated, and the action boundary and center offset positions are estimated through the action prediction head to generate the final human body movement recognition results.

Benefits of technology

It realizes high accuracy of human movement recognition in complex environments, and the model inference speed and accuracy have a certain balance, and has good practical application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047999A_ABST
    Figure CN120047999A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision, and relates to a human body action recognition method and system based on a double-flow self-attention mechanism, and the method comprises the steps: collecting video data of human body actions; extracting a plurality of modal features of the video data through the basic network of the double-flow self-attention mechanism model, and fusing the modal features with the time flow; inputting the fused spatio-temporal features into a spatio-temporal memory network of a double-flow self-attention mechanism model, performing spatio-temporal feature interaction of different scales, and generating a preliminary action classification result; and estimating an action boundary and an action center offset position of the preliminary action classification result through an action prediction head, and generating a final human body action recognition result. According to the method, the human body action video features can be effectively extracted, and human body feature extraction can be carried out from two levels of a macroscopic instance level and a microscopic fine-grained level, so that the interference of an action video background environment on an action recognition result is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a human action recognition method and system based on a two-stream self-attention mechanism, belonging to the technical field of computer vision. Background Art

[0002] Human action recognition, as an important research direction in the field of computer vision, has a profound impact on people's daily lives. At present, human action recognition methods based on deep learning have broad application prospects in the directions of human-computer interaction, security monitoring, health evaluation, sports analysis, etc. Human action recognition methods need to use video data samples collected in daily life, and these data samples usually have multi-modal information, such as RGB image data, optical flow information, and audio information, etc. When performing human action recognition in an actual complex environment, not only the above-mentioned complex multi-modal data needs to be processed, but also the characteristics of the human body and background regions in the video need to be accurately distinguished. In addition, the action recognition method needs to have a certain robustness to changes in the illumination and shooting perspectives of the data background; and the computational cost required to improve the accuracy of human action recognition in a complex environment will increase significantly. In actual production applications, there are usually real-time requirements for the action recognition model algorithm, and too large a computational cost will affect the practicability of the recognition model. Summary of the Invention

[0003] Aiming at the above problems, the purpose of the present invention is to provide a human action recognition method and system based on a two-stream self-attention mechanism that balances accuracy and speed.

[0004] To achieve the above purpose, the present invention proposes the following technical solutions: A human action recognition method based on a two-stream self-attention mechanism, including the following steps: collecting video data of human actions; extracting multiple modal features of the video data through the basic network of the two-stream self-attention mechanism model and fusing them with the time stream; inputting the fused spatio-temporal features into the spatio-temporal memory network of the two-stream self-attention mechanism model to perform spatio-temporal feature interaction at different scales and generate a preliminary action classification result; through an action prediction head, estimating the action boundary and the offset position of the action center of the preliminary action classification result to generate the final human action recognition result.

[0005] Further, in the spatio-temporal memory network, the fused spatio-temporal features are enhanced in local context features through several convolutional blocks. A part of the features output by the convolutional blocks is encoded through a cross-channel interaction encoder, and the encoded data is combined with the features output by another part of the convolutional blocks, and the combined features are dimension-reduced. The dimension-reduced features are input into several spatio-temporal memory convolutional blocks for feature encoding, and a preliminary action classification result is output after feature encoding.

[0006] Further, the dimension reduction of the features is realized through a two-stream Manhattan self-attention layer.

[0007] Furthermore, the dual-stream Manhattan self-attention layer includes an instance-level feature extraction stream, a window-level feature extraction stream, and an original path branch; the instance-level feature extraction stream differentiates action and non-action features by expanding the feature distance of video features; the window-level feature extraction stream uses a receptive field larger than the original video data to understand the high-dimensional semantic information features of the dual-stream self-attention mechanism model; the original path branch performs an identity mapping on the original features and conducts feature transmission of the deep network through the residual connection method.

[0008] Furthermore, a part of the data in the instance-level feature extraction stream is input into the fully connected layer, and another part of the data enters the Manhattan self-attention layer. The data passing through the fully connected layer is multiplied by the data passing through the Manhattan self-attention layer to generate an output result; the window-level feature extraction stream is divided into three branches. The first branch, after convolution, is added to the second branch that has undergone k groups of convolutions; the data stream after addition is multiplied by the third branch that has undergone convolution to generate an output result, and the output results of the instance-level feature extraction stream and the window-level feature extraction stream are added to the output result of the original path branch to generate the output result of the dual-stream Manhattan self-attention layer.

[0009] Furthermore, the basic network performs x , y feature extraction in two dimensions on the video data, performs local feature enhancement in two dimensions on the video data, and the expression of the features extracted from the video data is:

[0010]

[0011] where is the Manhattan self-attention mechanism, X is the input feature, QKV The matrix is three matrices used for calculating weights in the self-attention mechanism, Q is the query vector matrix, K is the key vector matrix, V is the value vector matrix, K T is K the transpose matrix of the matrix, D is the causal mask and exponential decay matrix, D The matrix is a real matrix and has the same dimension as the input feature X , stores the relative distances in the one-dimensional sequence of the matrix, and provides explicit temporal prior information. D 2d is a two-dimensional D matrix, n and mare the two dimensions of this matrix, is QKV the exponential decay factor of the matrix.

[0012] Furthermore, the base network performs sequence modeling on x , y the features of the two dimensions, and performs temporal modeling on the action sequence of the image data through the Transformers layer with a retention mechanism. The calculation formula in the self-attention stage of the Transformers layer is: . Among them, and are the Q query vector matrix and the K key vector matrix decomposed along the y axis direction, and are the Q query vector matrix and the K key vector matrix decomposed along the y axis direction. B is the batch size, L is the length after the original spatial decay matrix is flattened into a one-dimensional tensor, C is the number of matrix channels, and W and H are the width and height after the set decay matrix is decomposed. is the spatial self-attention decomposed by H, is the spatial self-attention decomposed by W, is the spatial decay matrix decomposed by H, is the spatial decay matrix decomposed by W. is the Manhattan self-attention for the original features. On this basis, the basic topological structure of the feature extraction part is formed, and its expression is , is the local feature enhancement, and depthwise separable convolution is specifically used for local feature enhancement.

[0013] Furthermore, the action prediction head includes an action start boundary branch, an action end boundary branch, and an intermediate offset branch. The action start boundary branch is used to predict the response intensity of the start boundary of each action branch; the action end boundary branch is used to predict the response intensity of the end boundary of each action branch; the intermediate offset branch takes a certain instance as a reference standard, uses the start or end point of the two adjacent local time sets before and after it as the response intensity, calculates the response expectation value through a local window, and obtains the boundary prediction value of each instance.

[0014] Furthermore, the overall loss function of the two-stream self-attention mechanism model is as follows:

[0015] Among them, is the loss function, is the number of positive samples, lis the length of the spatial feature pyramid layer, t is the length of the read feature time, is to calculate the L1 norm; is the classification label; is the IoU distance in the time dimension between the predicted sample and the Ground Truth, is the classification loss function, is the regression loss, is the number of negative samples.

[0016] The present invention also discloses a human action recognition system based on a two-stream self-attention mechanism, including: a data acquisition module for acquiring video data of human actions; a basic network module for extracting multiple modal features of the video data through the basic network of the two-stream self-attention mechanism model and fusing them with the time stream; a spatio-temporal memory network module for inputting the fused spatio-temporal features into the spatio-temporal memory network of the two-stream self-attention mechanism model to perform spatio-temporal feature interaction at different scales and generate a preliminary action classification result; an action recognition module for estimating the action boundary and the offset position of the action center of the preliminary action classification result through an action prediction head to generate a final human action recognition result.

[0017] The technical solution of the present invention has at least the following technical effects or advantages: 1. The present invention adopts a two-stream self-attention structure to build a lightweight network for effectively extracting human action video features; 2. The present invention can extract human features from two levels: the macroscopic instance level and the microscopic fine-grained level, so as to reduce the interference of the action video background environment on the action recognition result; 3. The present invention has a high accuracy rate for recognizing human actions in complex environments, has a certain balance between the model inference speed and accuracy, and has good practical application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is the overall flowchart of the two-stream self-attention mechanism network in an embodiment of the present invention; Figure 2 is the structural schematic diagram of the spatio-temporal memory network in an embodiment of the present invention; Figure 3 is the action recognition schematic diagram of the action prediction head in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention is described in detail through specific embodiments. However, it should be understood that the provision of specific embodiments is only for better understanding of the present invention, and they should not be construed as limitations on the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.

[0020] To solve the problems in the prior art such as the inability to accurately distinguish the features of the human body and the background area in the video, and the huge increase in the amount of calculation required to improve the accuracy of human action recognition in complex environments, the present invention builds a network structure based on the two-stream self-attention mechanism to fuse and learn the spatio-temporal features of video data, and proposes a human action recognition method and system based on the two-stream self-attention mechanism to achieve a balance between accuracy and speed. It collects human action data to construct and train a two-stream self-attention network model; uses the spatio-temporal signs extracted by the two-stream self-attention network model; and uses the extracted spatio-temporal features for action classification and time localization. The present invention adopts a two-stream self-attention structure to build a lightweight network for effectively extracting the features of human action videos. At the same time, this method can extract human features from two levels: the macroscopic instance level and the microscopic fine-grained level, thereby reducing the interference of the action video background environment on the action recognition result. The present invention has a high accuracy for human action recognition in complex environments, and there is a certain balance between the model inference speed and accuracy, and has practical application prospects. The solution of the present invention is elaborated in detail below through embodiments with reference to the accompanying drawings.

[0021] Embodiment 1 This embodiment discloses a human action recognition method based on the two-stream self-attention mechanism, as Figure 1 shown, including the following steps: S1 Collect video data of human actions.

[0022] Collect video data of human actions to make a data set. The data set can be selected such as UCF101, HMDB51, Kinetics, etc. These data sets contain various action categories and rich video samples. Use channels such as camera video acquisition equipment, video website downloads, and common data set collections to obtain human action videos.

[0023] Before inputting the data into the two-stream self-attention mechanism model, it is necessary to preprocess the video data. The preprocessing process includes: First, video frame extraction is performed to decompose the video into consecutive frames. Usually, a certain number of frames are extracted per second. In this embodiment, it is preferably 25 frames, but it can also be determined according to actual needs. Then, the extracted image data is enhanced. The methods for image data enhancement for each frame include rotation, scaling, cropping, and color transformation, etc. Image data enhancement is performed to improve the robustness of the two-stream self-attention mechanism model. Finally, label processing is carried out to count and record the representative action categories and make action labels to ensure that each video frame or sequence has a corresponding action label.

[0024] S2 extracts multiple modal features of the video data through the basic network of the two-stream self-attention mechanism model and fuses them with the time stream.

[0025] The two-stream self-attention mechanism network includes a basic network and a spatio-temporal memory network.

[0026] In order to process 2D image data, the basic network performs a 1D to 2D operation on the existing Retention layer. It extracts features of the video data through RetNet x , y in two dimensions, enhances the local features of the video data in two dimensions. The performances of 1D and 2D are different. 1D language tasks are one-dimensional unidirectional data, but the video data processed in this embodiment requires feature extraction in two dimensions. Therefore, the n-m part of the Retention layer needs to take the absolute value, which can be expressed as:

[0027]

[0028]

[0029] Among them, is the output matrix corresponding to the n direction, is the query vector matrix corresponding to the n direction, is the rotation position embedding angle of the transformer, is the conjugate transpose, is the value vector matrix corresponding to the m direction, is the two-dimensional bidirectional retention mechanism, X is the input feature, T is the matrix transpose, D is the spatial decay matrix, V is the value vector matrix, n is x the number of columns of the dimension direction matrix, m is y the number of rows of the direction matrix, is the exponential decay factor.

[0030] The calculation process of the D matrix in the Retention layer is similar. The absolute values need to be taken for the two dimensions of x and y:

[0031] where d is the matrix dimension, is the x coordinate of the token after decomposition in the n - dimensional direction, is the y coordinate of the token after decomposition in the m - dimensional direction, is the y coordinate of the token after decomposition in the n - dimensional direction, is the y - coordinate of the token after decomposition in the m - dimensional direction.

[0032] The expression of the features extracted from the video data is:

[0033]

[0034] where, is the Manhattan self - attention mechanism, X is the input feature, Q is the query vector matrix, K is the key vector matrix, T is the matrix transpose, D is the spatial decay matrix, d is the matrix dimension, V is the value vector matrix, n is the position of the token after flattening the matrix in the n - dimensional direction, m is the position of the token after flattening the matrix in the m - direction, is the exponential decay factor.

[0035] The basic module follows the feature extraction form of the Transformer layer. The time series modeling of the action sequence of the image data is carried out through the Transformers layer with a retention mechanism. The time decay brought by the retention mechanism is not available in the existing Transformers methods. In this embodiment, this structure is used to carry out the time series modeling of the action sequence of the image data. It solves the problem that the number of Tokens is too large in the self - attention stage, resulting in a huge increase in the computational amount. The image is decomposed into x , y two axes to reduce the computational consumption. The calculation formula in the self - attention stage is: . where, is the query vector matrix after decomposition in the H direction, is the key vector matrix after decomposition in the H direction, B is the batch size, L is the length after flattening the matrix, C is the number of channels, W represents the width direction of the original feature matrix, H is the height direction of the original feature matrix, is the query vector matrix after decomposition in the W direction, is the key vector matrix after decomposition in the W direction, Perform self-attention mechanism on the H direction, Perform self-attention mechanism on the W direction, Is the spatial attenuation matrix decomposed in the H direction, Is the spatial attenuation matrix decomposed in the W direction, Is the Manhattan self-attention mechanism. On this basis, the basic topological structure of the feature extraction part is formed, and its expression is , Is local feature enhancement, performing local feature enhancement on the original features in two dimensions.

[0036] This embodiment is for the temporal action recognition task (TAD), and the video stream to be recognized , where n represents the number of video frames, V i Represents the video frame at a certain moment. If the video data input to the network contains multi-modal features, such as RBG channel image features, optical flow features, sound features, etc., then the feature group input to the network is denoted as X i , Extracted from V i The original video data is also related to the time sequence T Corresponding. For the action label of the video frame K i , record the start frame of this action as s k , the end frame is e k , and the action category is c k , then there is an action label table . The basic network of the two-stream self-attention mechanism is used for the feature extraction basic module. Process actions of different time lengths through a feature pyramid.

[0037] S3 inputs the fused spatio-temporal features into the spatio-temporal memory network of the two-stream self-attention mechanism model to perform spatio-temporal feature interaction at different scales and generate a preliminary action classification result.

[0038] At the backend of the basic network, perform multiple downsamplings with a maximum pooling stride of 2, and build a spatio-temporal memory network similar to the Transformer layer at the backend to enhance the spatio-temporal dimensions of the pyramid features at each scale. Effective cross-scale perception of feature interactions at different granularities and different time ranges can be achieved through feature iterative downsampling.

[0039] In the spatio-temporal memory network, the fused spatio-temporal features are enhanced for local context features through several convolutional blocks. A part of the features output by the convolutional blocks is encoded through a cross-channel interaction encoder. The encoded data is combined with the features output by another part of the convolutional blocks, and the combined features are dimensionally reduced. The dimensionally reduced features are input into several spatio-temporal memory convolutional blocks for feature encoding, and the preliminary action classification results are output after feature encoding.

[0040] Perform two-stream self-attention processing on the enhanced context features. As follows:

[0041]

[0042]

[0043] Among them, is the two-stream self-attention mechanism, is the video-level average feature, is the fully convolutional operation on the input feature X, is the local feature branch, is the convolutional operation, is the convolution with the window size of w, is the convolution of the input feature x with the window size of w according to the expansion factor k, and the k factor is used to measure the large-grained time information.

[0044] The biggest problem with the spatio-temporal feature enhancement mechanism is that it is easy to introduce complex spatio-temporal features, resulting in a high computational burden. In this embodiment, the decomposition of the two-stream Manhattan self-attention layer MaSA is adopted to ensure that the receptive fields between Tokens are the same, so as to reduce the computational complexity. The two-stream Manhattan self-attention layer is mainly used to achieve feature dimensionality reduction. As Figure 2 shown, a two-stream self-attention feature pyramid is built in the two-stream Manhattan self-attention layer, which includes an instance-level feature extraction stream, a window-level feature extraction stream, and an original path branch; the instance-level feature extraction stream distinguishes action and non-action features by expanding the feature distance of video features; the window-level feature extraction stream uses a receptive field larger than the original video data to understand the high-dimensional semantic information features of the two-stream self-attention mechanism model; the original path branch performs an identity mapping on the original features and conducts feature transmission of the deep network through the residual connection method. This network structure can improve the discrimination between action and non-action features, enhance the model's understanding ability of high-dimensional semantic information features, enhance the feature transmission ability when the deep network is enhanced, and improve the learning effect.

[0045] In a part of the instance-level feature extraction stream, some data is input into a fully connected layer, and the other part of the data enters the Manhattan self-attention layer. The data passing through the fully connected layer is multiplied by the data passing through the Manhattan self-attention layer to generate an output result. The window-level feature extraction stream is divided into three branches. The first branch, after convolution, is added to the second branch that has undergone k groups of convolutions. The resulting data stream is multiplied by the third branch that has undergone convolution to generate an output result. The output results of the instance-level feature extraction stream and the window-level feature extraction stream are added to the output result of the original path branch to generate the output result of the dual-stream Manhattan self-attention layer.

[0046] To handle action features of different time lengths, in the window-level feature extraction stream, a window scaling coefficient w is introduced to expand the distance between actions and non-actions and the instance average features, thereby distinguishing action boundaries. The window-level feature extraction stream can not only introduce semantic feature information with different receptive fields, increase the multi-scale feature recognition ability, but also perform feature interaction on actions at different times.

[0047] S4 estimates the action boundaries and the offset positions of the action centers of the preliminary action classification results through an action prediction head to generate the final human action recognition results.

[0048] The action prediction head includes an action start boundary branch, an action end boundary branch, and an intermediate offset branch. The action start boundary branch is used to predict the response intensity of the start boundary of each action branch; the action end boundary branch is used to predict the response intensity of the end boundary of each action branch; the intermediate offset branch takes a certain instance as a reference standard, uses the start or end point of the two adjacent local time sets before and after it as the response intensity, calculates the response expectation value through a local window, and obtains the boundary prediction value of each instance. For example, the action start distance of the t th instance is d st , and The recognition start boundary prediction formula is as follows:

[0049]

[0050]

[0051]

[0052] Similarly, the end boundary prediction formula is as follows:

[0053]

[0054] Among them, is the probability of the start boundary of an action within an action set at a certain moment, is the value of instance t within the action start set, is the value of instance t within the action offset set, and b is the number of boundary prediction sets, is the expectation of the start boundary prediction for instance t, is the probability of an action occurring within the b-th prediction set, is the distance of the action relative to the start moment, l is the feature scaling factor, and t is the motion feature instance, is the distance of the action relative to the end moment; B is the number of frames of the boundary prediction F s and F e are the response values as the start and end points of an action at each moment respectively, is the probability of the end boundary of an action within an action set at a certain moment; is the t action end distance of the th instance.

[0055] Three action prediction heads are modeled using ordinary convolutional layers, sharing parameters at all feature pyramid scales, aiming to reduce the number of parameters.

[0056] Each layer of the feature pyramid in the two-stream self-attention mechanism model outputs temporal features , and its action prediction head is used for action classification. The feature pyramid of each instance t is represented as: . The overall loss function of the two-stream self-attention mechanism model is as follows:

[0057] Among them, is the overall loss function, is the number of positive samples, l is the length of the spatial feature pyramid, t is the time length of reading features, is to find the first norm; is the classification label; is the IoU distance in the time dimension between the predicted sample and GT, is the classification loss function, is the regression loss function, is the number of negative samples.

[0058] In this embodiment, the hardware configuration for executing the dual-stream self-attention mechanism model is as follows: the CPU processor is the AMD Ryzen 9 3950X 16-Core Processor, and the GPU is the NVIDIA GeForce RTX 3090; the software configuration: the computer operating system is Ubuntu 24.04, the CUDA version is 11.2, the neural network framework used is Pytorch, and the version is 2.1.0. The initial learning rate of the parameter setting is 10-4, and the optimizer strategy used is AdamW. The classification head is initialized with a Gaussian distribution initialization (0, 0.1). The number of training iterations is 40, and the Warmup strategy is used in the first 10 iterations. The window factor w in the training head dual-stream attention branch is set to 1.5 for scaling, and the parameters are adjusted appropriately.

[0059] The following presents the action recognition results on the THUMOS14 action recognition dataset. This dataset is collected based on YouTube videos and contains motion categories in daily life. There are 13,000 video clips in the dataset, and the number of action categories is 20. Here, the mean Average Precision (mAP) with different IoUs is selected to measure the action recognition accuracy for different categories, and the inference time is selected to measure the model running efficiency. Table 1 shows the different recognition results obtained by the above dataset in different algorithm recognition models, and the measurement metrics are the mean precision and the inference latency (ms).

[0060] As shown in Table 1, in this embodiment, a balance is achieved in terms of accuracy and speed in the action recognition task of the THUMOS14 dataset, which is superior to other recognition models, demonstrating the advantage of the method in this embodiment in terms of recognition efficiency.

[0061] Table 1 is a table of different recognition results obtained by the THUMOS14 dataset in different algorithm recognition models

[0062] Embodiment 2 Based on the same inventive concept, this embodiment discloses a human action recognition system based on a dual-stream self-attention mechanism, including: A data acquisition module for acquiring video data of human actions; A basic network module for extracting multiple modal features of video data through the basic network of the dual-stream self-attention mechanism model and fusing them with the time stream; A spatio-temporal memory network module for inputting the fused spatio-temporal features into the spatio-temporal memory network of the dual-stream self-attention mechanism model to perform spatio-temporal feature interaction at different scales and generate a preliminary action classification result; The action recognition module is used to estimate the action boundaries and the offset positions of the action centers of the preliminary action classification results through an action prediction head, and generate the final human action recognition results.

[0063] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0064] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more of the processes Figure 1 or a combination of multiple processes and / or blocks

[0065] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one or more of the processes Figure 1 or a combination of multiple processes and / or blocks

[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more of the processes Figure 1 or a combination of multiple processes and / or blocks

[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention. The above content is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or replacements, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A human action recognition method based on a dual-stream self-attention mechanism, characterized in that: The following steps are involved: Collect video data of human movements; Extracting multiple modal features of the video data through a basic network of a two-stream self-attention mechanism model and fusing them with the time stream; The fused spatiotemporal features are input into the spatiotemporal memory network of the two-stream self-attention mechanism model to interact with spatiotemporal features of different scales and generate preliminary action classification results. The action boundary and action center offset position of the preliminary action classification result are estimated through the action prediction head to generate the final human action recognition result.

2. The human action recognition method based on the dual-stream self-attention mechanism as claimed in claim 1, characterized in that: In the spatiotemporal memory network, the fused spatiotemporal features are subjected to local context feature enhancement through a number of convolution blocks, a part of the features output by the convolution blocks are encoded through a cross-channel interactive encoder, the encoded data are combined with the features output by another part of the convolution blocks, the combined features are subjected to feature dimensionality reduction, the reduced features are input into a number of spatiotemporal memory convolution blocks for feature encoding, and a preliminary action classification result is output after feature encoding.

3. The human action recognition method based on the dual-stream self-attention mechanism as claimed in claim 2, characterized in that: The feature dimensionality reduction is achieved through a two-stream Manhattan self-attention layer.

4. The human action recognition method based on the dual-stream self-attention mechanism as claimed in claim 2, characterized in that: The dual-stream Manhattan self-attention layer includes an instance-level feature extraction flow, a window-level feature extraction flow and an original path branch; the instance-level feature extraction flow distinguishes action and non-action features by expanding the feature distance of video features; the window-level feature extraction flow uses a receptive field larger than the original video data to understand the high-dimensional semantic information characteristics of the dual-stream self-attention mechanism model; the original path branch performs identity mapping on the original features and performs feature transfer of the deep network through the residual connection method.

5. The human action recognition method based on the dual-stream self-attention mechanism as claimed in claim 4, characterized in that: Part of the data in the instance-level feature extraction stream is input into the fully connected layer, and the other part of the data enters the Manhattan self-attention layer, and the data passing through the fully connected layer is multiplied with the data passing through the Manhattan self-attention layer to generate an output result; the window-level feature extraction stream is divided into three tributaries, and the first tributary is added to the second tributary after k groups of convolutions after convolution; the added data stream is multiplied with the third convolved tributary to generate an output result, and the output results of the instance-level feature extraction stream and the window-level feature extraction stream are added to the output results of the original path branch to generate the output result of the dual-stream Manhattan self-attention layer.

6. The human action recognition method based on the dual-stream self-attention mechanism as claimed in claim 1, characterized in that: The basic network uses RetNet to process video data. x , y The two-dimensional feature extraction performs two-dimensional local feature enhancement on the video data. The expression of the feature extracted from the video data is: in, is the Manhattan self-attention mechanism, X is the input feature matrix, Q is the query vector matrix, K is the key vector matrix, T is the matrix transpose, D is the exponential decay matrix, d is the matrix dimension, V is the value vector matrix, n and m are the two dimensions of the matrix, is an exponential decay factor.

7. The human action recognition method based on the dual-stream self-attention mechanism as claimed in claim 6, characterized in that: The basic network x , y The features of the two dimensions are sequence modeled, and the image data action sequence is temporally modeled through the Transformers layer with a retention mechanism. The calculation formula of the self-attention stage in the Transformers layer is: . in, is the query vector matrix decomposed in the H direction, is the key vector matrix decomposed in the H direction, B is the batch size, L is the length of the spatial decay matrix after flattening into a one-dimensional tensor, C is the number of matrix channels, W and H are the width and height of the exponential decay matrix, is the query vector matrix decomposed in the W direction, is the bond vector matrix decomposed in the W direction, is the spatial self-attention in the H direction, is the spatial self-attention in the W direction, is the exponential decay matrix decomposed in the H direction, is the exponential decay matrix decomposed inversely by W, is Manhattan self-attention, on which the basic topological structure of the feature extraction part is formed, and its expression is , It is local feature enhancement.

8. The human action recognition method based on the dual-stream self-attention mechanism as claimed in claim 1, characterized in that: The action prediction head includes an action start boundary branch, an action end boundary branch and an intermediate offset branch. The action start boundary branch is used to predict the response strength of the start boundary of each action branch; the action end boundary branch is used to predict the response strength of the end boundary of each action branch; the intermediate offset branch uses a certain instance as a reference standard, and takes the starting point or end point of the two adjacent local time sets before and after it as the response strength, calculates the response expectation value through a local window, and obtains the boundary prediction value of each instance.

9. The method for human action recognition based on a dual-stream self-attention mechanism as claimed in claim 1, characterized in that: The overall loss function of the two-stream self-attention mechanism model is as follows: in, is the overall loss function, is the number of positive samples, l is the length of the spatial feature pyramid, t is the length of time it takes to read the feature, is to find a norm; is a classification label; is the IoU distance in the time dimension between the predicted sample and GT, is the classification loss, is the regression loss, is the number of negative samples.

10. A human action recognition system based on a dual-stream self-attention mechanism, characterized in that: include: A data acquisition module, used to collect video data of human body movements; A basic network module, used for extracting multiple modal features of the video data through a basic network of a two-stream self-attention mechanism model, and fusing them with a time stream; The spatiotemporal memory network module is used to input the fused spatiotemporal features into the spatiotemporal memory network of the two-stream self-attention mechanism model, interact with spatiotemporal features of different scales, and generate preliminary action classification results; The action recognition module is used to estimate the action boundary and action center offset position of the preliminary action classification result through the action prediction head to generate the final human action recognition result.

Citation Information

Patent Citations

  • Real-time behavior identification method based on time attention mechanism and double-flow network

    CN113283298A

  • Human body interaction behavior recognition method based on space-time diagram convolution

    CN114694174A

  • Online action detection method and system based on non-autoregressive attention mechanism

    CN115984810A

  • Video semantic feature and extensible granularity sensing time sequence motion detection method and device

    CN117058595A

  • Image processing method and apparatus, device, and storage medium

    WO2021088556A1

Cited By

  • Three-dimensional hand posture estimation method fusing hand and object features

    CN121259920A