An Adaptive Frame Sampling Driven Gesture Recognition Method
Through the adaptive frame sampling driving method, inter-frame motion attention and self-attention algorithm are used to track the motion area and fuse important features, solving the problem that redundant information in the video frame sequence affects gesture recognition and improves recognition accuracy.
Patent Information
- Application Number
- CN202310038279.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-01-10
AI Technical Summary
The video frame sequence contains too much redundant information, which affects the recognition accuracy of the gesture recognition model, and the background information interferes with the model's identification of the motion area.
Adaptive frame sampling drive method is adopted to track the motion area through the inter-frame motion attention algorithm, the self-attention time downsampling algorithm fuses multi-frame features, and the self-attention space downsampling algorithm fuses local area features to eliminate redundant information.
It improves the accuracy of the gesture recognition model, focuses on motion feature extraction, eliminates redundant information, and improves recognition accuracy.
Smart Images

Figure CN116229567B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of gesture recognition, and specifically to a gesture recognition method driven by adaptive frame sampling. Background Art
[0002] In many scenarios, gestures are a basic means of communication. Gesture recognition enables a computer to understand the meaning of a target gesture, and it has a wide range of application scenarios in human-computer interaction. Dynamic gesture recognition based on computer vision refers to parsing the video captured by a camera through a specific algorithm and then classifying the gestures.
[0003] With the rise of deep learning, more and more neural networks for video classification have been created. However, since the frame sequence of a video contains too much redundant information, this redundant information affects the model's attention to effective features to a certain extent, resulting in the model losing classification accuracy. The same problem exists in dynamic gesture recognition, mainly in two aspects: First, due to reasons such as limited quality of the video capture device and too large frame rate settings, the data may contain some duplicate frames and blurred frames, and these redundant frames contain very little useful information. On the other hand, each frame image records both the human body and the complex background behind the person, while dynamic gesture recognition mainly focuses on the moving area, and the information in other positions usually has a negative impact on the recognition accuracy of the model. Summary of the Invention
[0004] The purpose of the present invention is to provide a gesture recognition method driven by adaptive frame sampling, which solves the problem that the frame sequence of a video contains too much redundant information, resulting in a negative impact on the recognition accuracy of the model.
[0005] The purpose of the present invention can be achieved through the following technical solutions:
[0006] A gesture recognition method driven by adaptive frame sampling, the gesture recognition method includes the following steps:
[0007] S1: Convert the frame sequence captured by the camera into a tensor, and use two convolutional layers as feature extraction layers to extract the features of each frame image in the tensor.
[0008] S2: Adopt an inter-frame motion attention algorithm to track the moving area in each frame in S1 according to the similarity of the local area patterns between two frames, and assign a greater attention weight to the moving area.
[0009] S3: Adopt a self-attention time downsampling algorithm to assign different weights to the same area of adjacent multiple frames according to importance, and then fuse the features of multiple frames into one frame by summing.
[0010] S4: Use the self-attention spatial downsampling algorithm to assign different weights to each point in the local area according to importance, and then fuse the features of different points in the local area into one point by summation.
[0011] S5: Input the features with low redundant information obtained in S4 into the existing gesture classification model to classify the gestures.
[0012] Furthermore, both the selection of multiple frames in S3 and the selection of the local area in S4 are performed through a sliding window with a step size of 2.
[0013] Furthermore, the process of using two convolutional layers as the feature extraction layer to extract the features of each frame image in the frame sequence is as follows:
[0014] Input each frame image of the frame sequence into the 1×1 convolutional layer and the 3×3 convolutional layer to extract spatial features. The parameters of the two convolutional layers are as follows:
[0015] In the 1×1 convolutional layer, the number of input channels is 3, the size of the convolutional kernel is 1×1, the number of convolutional kernels is 64, the step size is 1, and the padding is 0. In the 3×3 convolutional layer, the number of input channels is 64, the size of the convolutional kernel is 3×3, the number of convolutional kernels is 64, the step size is 1, and the padding is 1.
[0016] Furthermore, the calculation process of the inter-frame motion attention algorithm is as follows:
[0017] (1) Divide the output of the convolutional layer into equally sized blocks according to a window of size (2, 7, 7, 64). Suppose the dimension of the feature obtained after the convolutional layer is (D, H, W, 64), where D is the number of frames, H and W are the height and width of each frame image, and 64 is the number of channels. Then, after dividing by the window, we get blocks of size (2, 7, 7, 64).
[0018] (2) Further divide each block in the first dimension. Each block of size (2, 7, 7, 64) is divided into two small blocks of size (1, 7, 7, 64). Then, we get groups of blocks consisting of two small blocks.
[0019] Input the two small blocks in the same group into the fully connected layer to extract patterns. The expression is:
[0020]
[0021]
[0022] where Q i and K iThey are the outputs obtained by processing the $i$-th group of small blocks through two linear layers. $L1$ represents the first linear layer, and $L2$ represents the second linear layer. The number of input and output channels of both linear layers is 64. represents the first small block of the $i$-th group. represents the second small block of the $i$-th group.
[0023] Calculate the attention weights of the two small blocks respectively and apply them. First, calculate the similarity matrix, and the expression is:
[0024] Attn i = Q i @ T(K i )
[0025] where Attn i represents the similarity matrix of the two small blocks in the $i$-th group, @ represents matrix multiplication, and T() represents transposing the last two dimensions of the tensor.
[0026] Then, calculate the attention weights of the two small blocks respectively, and the expressions are:
[0027] AF i = R(Softmax(max(Attn i , -1)))
[0028] AL i = R(T(Softmax(max(Attn i , -2))))
[0029] where AF i and AL i are the attention weights of the first and second small blocks in the $i$-th group respectively. R represents replicating the data 64 times in the last dimension of the tensor, Softmax() represents the Softmax function, and max() represents finding the maximum value in a certain dimension.
[0030] (3) Apply the attention weights to the input features, and the expression is:
[0031] output i = RS(concate(AF i , AL i )) × input
[0032] where output i is the result of applying the inter-frame motion attention weights to the $i$-th block. RS() is to splice each block according to the original relative position, concate() represents splicing two groups of tensors in the first dimension, and input is the output of the 3×3 convolutional layer in S1.
[0033] Furthermore, the calculation process of the self-attention temporal downsampling algorithm is as follows:
[0034] (1) Divide the output of S2 into equally sized blocks using a sliding window of size (4, 1, 1, 64) with a stride of 2.
[0035] (2) Reconstruct each block into a block of dimension (4, 64), and then perform the self-attention downsampling operation. The expression is:
[0036] y j = T(Softmax(S(L3(x j )) @ T(L4(x j ))))) @ L5(x j )
[0037] where y j is the calculation result of the j-th block, S() is the sum of the data in the last dimension of the input tensor, L3, L4, and L5 are three fully connected layers with 64 input and output channels each, and x j represents the tensor corresponding to the j-th block;
[0038] (3) Use RS() to splice the calculation results of each block according to their original relative positions.
[0039] Furthermore, the calculation process of the self-attention spatial downsampling algorithm is as follows:
[0040] (1) Divide the output of S3 into equally sized blocks using a sliding window of size (3, 3, 3, 64) with a stride of 2.
[0041] (2) Merge the first three dimensions of each block into one dimension, so that the size of the block becomes (27, 64). The expression for calculating the attention matrix is:
[0042] a k = Softmax(S(L6(c k )) @ T(L7(c k ))))
[0043] where a k is the attention matrix of the k-th block, L6 and L7 are both fully connected layers with 64 input and output channels each, and c k represents the k-th tensor of size (27, 64).
[0044] (3) Calculate the attention application matrix. The expression is:
[0045] v k = L8(c k )
[0046] where vk Denote the attention application matrix corresponding to the k-th block. L8 is a fully connected layer with both the input and output channels being 64.
[0047] Adjust a k and v k into blocks of sizes (3, 9, 1) and (3, 9, 64) respectively, and then perform matrix multiplication to obtain the result of applying the self-attention spatial downsampling algorithm. The expression is:
[0048] sout k = T(a k ) @ v k
[0049] where Sout k is the result of spatial downsampling for the k-th block.
[0050] (4) Adjust Sout k into a block of size (3, 64), and then splice it according to the original relative position of v k .
[0051] Advantages of the present invention:
[0052] 1. The gesture recognition method of the present invention uses the proposed inter-frame motion attention algorithm to track the motion regions in each frame according to the similarity of local region patterns between two frames and assigns greater attention weights to the motion regions, enabling the existing gesture classification model to focus more on the extraction of motion features;
[0053] 2. The gesture recognition method of the present invention uses the proposed self-attention temporal downsampling algorithm to assign different weights to the same region of adjacent multiple frames according to importance, and then fuses the features of multiple frames into one frame by summation, which can assign greater weights to the more important frames of the same region, fuse the effective features of multiple frames, and largely eliminate redundant information;
[0054] 3. The gesture recognition method of the present invention uses the proposed self-attention spatial downsampling algorithm to assign different weights to each point in the local region according to importance, and then fuses the features of different points in the local region into one point by summation, which can fuse the spatial features in each frame according to the degree of importance, further refine the effective features, and eliminate redundant information. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The present invention will be further described below with reference to the accompanying drawings.
[0056] Figure 1 is the flowchart of the gesture recognition method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] An adaptive frame sampling-driven gesture recognition method, as Figure 1 shown, the gesture recognition method includes the following steps:
[0059] S1: Convert the frame sequence captured by the camera into a tensor, and use two convolutional layers as feature extraction layers to extract the features of each frame image in the tensor;
[0060] The process of using two convolutional layers as feature extraction layers to extract the features of each frame image in the frame sequence is as follows:
[0061] Input each frame image of the frame sequence into a 1×1 convolutional layer and a 3×3 convolutional layer to extract spatial features. The parameters of the two convolutional layers are as follows:
[0062] In the 1×1 convolutional layer, the number of input channels is 3, the size of the convolutional kernel is 1×1, the number of convolutional kernels is 64, the stride is 1, and the padding is 0; in the 3×3 convolutional layer, the number of input channels is 64, the size of the convolutional kernel is 3×3, the number of convolutional kernels is 64, the stride is 1, and the padding is 1.
[0063] S2: Adopt an inter-frame motion attention algorithm to track the motion area in each frame in S1 according to the similarity of local region patterns between two frames, and assign a greater attention weight to the motion area;
[0064] The calculation process of the inter-frame motion attention algorithm is as follows:
[0065] First, divide the output of the convolutional layer into blocks of the same size according to a window of size (2, 7, 7, 64). Suppose the dimension of the feature obtained after passing through the convolutional layer is (D, H, W, 64), where D is the number of frames, H and W are the height and width of each frame image, and 64 is the number of channels. Then, after dividing by the window, blocks of size (2, 7, 7, 64) are obtained;
[0066] Then, each block is further divided in the first dimension. Each block of size (2, 7, 7, 64) is divided into two small blocks of size (1, 7, 7, 64). Then, groups of blocks composed of two small blocks are obtained;
[0067] Input the two small blocks in the same group into the fully connected layer to extract patterns. The expression is:
[0068]
[0069]
[0070] Among them, Q i and K i are the outputs obtained by processing the i-th group of patches by two linear layers respectively. L1 represents the first linear layer, and L2 represents the second linear layer. The number of input and output channels of both linear layers is 64. represents the first patch of the i-th group, represents the second patch of the i-th group;
[0071] Calculate the attention weights of the two patches respectively and apply them. First, calculate the similarity matrix, and the expression is:
[0072] Attn i = Q i @ T(K i )
[0073] Among them, Attn i represents the similarity matrix of the two patches in the i-th group, @ represents matrix multiplication, and T() represents transposing the last two dimensions of the tensor;
[0074] Then calculate the attention weights of the two patches respectively, and the expression is:
[0075] AF i = R(Softmax(max(Attn i , -1)))
[0076] AL i = R(T(Softmax(max(Attn i , -2))))
[0077] Among them, AF i and AL i are the attention weights of the first and second patches in the i-th group respectively. R represents replicating the data 64 times in the last dimension of the tensor, Softmax() represents the Softmax function, and max() represents finding the maximum value of a certain dimension;
[0078] Finally, apply the attention weights to the input features, and the expression is:
[0079] output i = RS(concate(AF i , AL i )) × input
[0080] Among them, outputi is the result of applying the inter-frame motion attention weight to the i-th block. S() is to splice each block according to the original relative position, concate() means to splice two groups of tensors in the first dimension, and input is the output of the 3×3 convolutional layer in S1.
[0081] S3: The self-attention temporal downsampling algorithm assigns different weights to the same region of adjacent multiple frames according to importance, and then fuses the features of multiple frames into one frame by summation. Each time, a sliding window with a stride of 2 is used to select multiple frames, so that the number of frames can be halved;
[0082] The calculation process of the self-attention temporal downsampling algorithm is as follows:
[0083] First, divide the output of S2 into equally sized blocks according to a sliding window of size (4, 1, 1, 64) with a stride of 2;
[0084] Then, each block is reconstructed into a block of (4, 64) dimensions, and then the self-attention downsampling operation is performed. The expression is:
[0085] y j = T(Softmax(S(L3(x j ) @ T(L4(x j )))) @ L5(x j )
[0086] where, y j is the calculation result of the j-th block. S() is to sum the data of the last dimension of the input tensor. L3, L4, and L5 are three fully connected layers with 64 input and output channels each. x j represents the tensor corresponding to the j-th block;
[0087] Finally, use RS() to splice the calculation results of each block according to the original relative position.
[0088] S4: The self-attention spatial downsampling algorithm assigns different weights to each point in the local area according to importance, and then fuses the features of different points in the local area into one point by summation. Each time, a sliding window with a stride of 2 is used to select the local area, so that the spatial dimension size of the features can be halved, and features with low redundancy are obtained;
[0089] The calculation process of the self-attention spatial downsampling algorithm is as follows:
[0090] First, divide the output of S3 into equally sized blocks according to a sliding window of size (3, 3, 3, 64) with a stride of 2;
[0091] Then, merge the first three dimensions of each block into one dimension, and the size of the block becomes (27, 64). The expression for calculating the attention matrix is:
[0092] a k = Softmax(S(L6(c k ) @ T(L7(c k ))))
[0093] Among them, a k is the attention matrix of the k-th block. Both L6 and L7 are fully connected layers with 64 input and output channels, and c k represents the k-th tensor of size (27, 64);
[0094] Then, calculate the attention application matrix, and the expression is:
[0095] v k = L8(c k )
[0096] Among them, v k represents the attention application matrix corresponding to the k-th block, and L8 is a fully connected layer with 64 input and output channels;
[0097] Reshape a k and v k into blocks of sizes (3, 9, 1) and (3, 9, 64) respectively, and then perform matrix multiplication to obtain the result of applying the self-attention spatial downsampling algorithm. The expression is:
[0098] Sout k = T(a k ) @ v k
[0099] Among them, Sout k is the result of spatial downsampling of the k-th block;
[0100] Finally, reshape Sout k into a block of size (3, 64), and then splice it according to the original relative position of v k .
[0101] S5: Input the features with low redundant information obtained in S4 into the existing gesture classification model to classify the gestures.
[0102] By first sampling the frame sequence, excluding frames with a low proportion of valid information and highlighting the motion regions in the frames, the classification model can pay more attention to the valid features, thereby improving the gesture recognition accuracy; then using the self-attention mechanism to implement the frame sampling operation, adaptively sampling the frame sequence to eliminate redundant information, tracking the inter-frame motion, allocating a larger attention weight to the motion regions, and excluding inefficient frames, thereby improving the accuracy of the gesture recognition model.
[0103] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0104] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and the descriptions in the specification are only used to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all these changes and improvements fall within the scope of the present invention claimed.
Claims
1. An adaptive frame sampling-driven gesture recognition method, characterized in that The gesture recognition method includes the following steps: S1: Convert the frame sequence captured by the camera into a tensor, and use two convolutional layers as feature extraction layers to extract the features of each frame image in the tensor; S2: Adopt an inter-frame motion attention algorithm to track the motion area in each frame in S1 according to the similarity of local area patterns between two frames, and assign a greater attention weight to the motion area; S3: Adopt a self-attention temporal downsampling algorithm to assign different weights to the same area of adjacent multiple frames according to importance, and then fuse the features of multiple frames into one frame by summation; S4: Adopt a self-attention spatial downsampling algorithm to assign different weights to each point in the local area according to importance, and then fuse the features of different points in the local area into one point by summation; S5: Input the feature with low redundant information obtained in S4 into an existing gesture classification model to classify the gesture.
2. The gesture recognition method driven by adaptive frame sampling according to claim 1, wherein For the selection of multiple frames in S3 and the selection of the local area in S4, both are selected by a sliding window with a step size of 2.
3. The gesture recognition method driven by adaptive frame sampling according to claim 1, wherein The process of using two convolutional layers as feature extraction layers to extract the features of each frame image in the frame sequence is as follows: Input each frame image of the frame sequence into a 1×1 convolutional layer and a 3×3 convolutional layer to extract spatial features. The parameters of the two convolutional layers are as follows: In the 1×1 convolutional layer, the number of input channels is 3, the kernel size is 1×1, the number of kernels is 64, the step size is 1, and the padding is 0; in the 3×3 convolutional layer, the number of input channels is 64, the kernel size is 3×3, the number of kernels is 64, the step size is 1, and the padding is 1.
4. The gesture recognition method driven by adaptive frame sampling according to claim 1, wherein The calculation process of the inter-frame motion attention algorithm is as follows: (1)Divide the output of the convolutional layer into equally sized blocks according to a window of size (2, 7, 7, 64). Suppose the dimension of the feature obtained after the convolutional layer is (D, H, W, 64), where D is the number of frames, H and W are the height and width of each frame image, and 64 is the number of channels. Then, after dividing by the window, we get blocks of size (2, 7, 7, 64); (2) Further divide each block in the first dimension. Each block of size (2, 7, 7, 64) is divided into two small blocks of size (1, 7, 7, 64), and then a group of blocks consisting of two small blocks is obtained; Input two small blocks in the same group into the fully connected layer to extract patterns, and the expression is: Among them, Q i and K i are the outputs obtained by the two linear layers processing the i-th group of patches respectively. L1 represents the first linear layer, and L2 represents the second linear layer. The number of input and output channels of the two linear layers is 64, represents the first patch of the i-th group, represents the second patch of the i-th group; Obtain the respective attention weights of the two small blocks and apply them. First, calculate the similarity matrix, and the expression is: Attn i = Q i @T(K i ) Among them, Attn i represents the similarity matrix of two small blocks in the i-th group, @ represents matrix multiplication, and T() represents transposing the last two dimensions of the tensor; Then calculate the respective attention weights of the two small blocks respectively, and the expression is: AF i = R(Softmax(max(Attn i , -1))) AL i = R(T(Softmax(max(Attn i , -2)))) Among them, AF i and AL i are the attention weights of the first and second small blocks in the $i$-th group respectively. $R$ represents replicating the data 64 times in the last dimension of the tensor, Softmax() represents the Softmax function, and max() represents finding the maximum value in a certain dimension; (3) Apply the attention weight to the input feature, and the expression is: output i = RS(concate(AF i , AL i )) × input Among them, output i is the result of applying the inter-frame motion attention weight to the i-th block, RS() is to splice each block according to the original relative position, concate() means to splice two groups of tensors in the first dimension, and input is the output of the 3×3 convolutional layer in S1.
5. The gesture recognition method driven by adaptive frame sampling according to claim 2, wherein The calculation process of the self-attention temporal downsampling algorithm is as follows: (1) Divide the output of S2 into equally sized blocks according to a sliding window with a size of (4, 1, 1, 64) and a step size of 2: (2) Reconstruct each block into a block with a dimension of (4, 64), and then perform a self-attention downsampling operation, and the expression is: y j = T(Softmax(S(L3(x j ) @ T(L4(x j ))))) @ L5(x j ) where y j is the calculation result of the j-th block, S() is to sum the data of the last dimension of the input tensor, L3, L4, and L5 are three fully connected layers with 64 input and output channels each, and x j represents the tensor corresponding to the j-th block; (3) Use RS() to splice the calculation results of each block according to the original relative positions.
6. The gesture recognition method driven by adaptive frame sampling according to claim 2, wherein, The calculation process of the self-attention spatial downsampling algorithm is as follows: (1) Divide the output of S3 into equally sized blocks according to a sliding window with a size of (3, 3, 3, 64) and a step size of 2; (2) Merge the first three dimensions of each block into one dimension, then the size of the block becomes (27, 64), and the expression for calculating the attention matrix is: a k = Softmax(S(L6(c k )@T(L7(c k )))) where a k is the attention matrix of the k-th block, and both L6 and L7 are fully connected layers with 64 input and output channels, and c k represents the k-th tensor of size (27, 64); (3) Calculate the attention application matrix, and the expression is: v k = L8(c k ) Among them, v k represents the attention application matrix corresponding to the k-th block, and L8 is a fully connected layer with both input and output channel numbers of 64; Adjust a k and v k into blocks of sizes (3, 9, 1) and (3, 9, 64) respectively, and then perform matrix multiplication to obtain the result of applying the self-attention spatial downsampling algorithm. The expression is: Sout k = T(a k ) @ v k Among them, Sout k is the result of spatial downsampling of the k-th block; (4) Adjust Sout k to a block of size (3, 64), and then splice it according to the k original relative positions.
Citation Information
Patent Citations
Sign language recognition method based on space-time attention mechanism
CN111091045A
Action recognition method based on double-flow space-time attention mechanism
CN111627052A