Channel attention method and system in gesture recognition application

By adopting a channel attention module composed of asymmetric multi-scale convolution in gesture recognition applications, the problem of channel feature extraction imbalance and redundant information in the prior art is solved, and the generalization ability and robustness of the model are improved.

CN120148100APending Publication Date: 2025-06-13浪潮智能终端有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510097819.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, it is difficult for graph convolution networks to effectively select and weight the most useful channel features in gesture recognition applications, resulting in some channel features being overemphasized while some effective features being ignored. At the same time, square convolution has limitations, ignoring the differences between different scales and structures, and retaining redundant information.

Method used

A channel attention module based on asymmetric convolution is adopted, combined with vertical and horizontal convolution kernel skeletons, a channel attention module composed of asymmetric multi-scale convolution is proposed. Through multi-scale processing and information fusion, the model's dynamic perception ability of different channel information is improved.

Benefits of technology

This improves the model's attention to significance channels, reduces redundant information, and enhances the generalization and robustness of the model, especially when handling image flip and rotation operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148100A_ABST
    Figure CN120148100A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of gesture recognition, and particularly provides a channel attention method and system in gesture recognition application, based on an asymmetric convolution channel attention module, a vertical convolution kernel skeleton and a horizontal convolution kernel skeleton are adopted, and discriminative channel features of a feature map are enhanced; a multi-scale horizontal convolution kernel skeleton and a vertical convolution kernel skeleton are combined, a channel attention module formed by asymmetric multi-scale convolution is provided, refining processing is performed on input features on multiple scales, multi-scale features are reserved, information fusion is performed, the dynamic sensing ability of the model to different channel information is improved, and the accuracy of the model is improved. Attention on a significant channel is enhanced. Compared with the prior art, the method can solve the problem that part of channel features are excessively emphasized and part of effective channel features are ignored, and the problems that the attention of channels formed by square convolution ignores the difference between different scales and structures, redundant information is reserved, and channel feature extraction is single.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gesture recognition, and specifically provides a channel attention method and system in gesture recognition applications. Background Art

[0002] Gesture recognition technology is currently widely used in all aspects of production and life. In the field of smart home, users can achieve contactless control of smart devices at home through gestures, which is more convenient; in the medical field, doctors use gesture recognition for contactless operations, reducing the risk of infection and providing data support in rehabilitation training; in the gaming field, players directly control game characters through gestures, increasing the gaming experience and fun; in the security field, gesture recognition in surveillance is used to identify abnormal behaviors and automatically alarm, improving public safety levels; in the virtual reality field, gesture recognition enables users to interact with virtual objects in a more natural way, enhancing the immersive experience; in the social field, gesture recognition in video calls allows users to trigger special effects or stickers through specific gestures, increasing the interactive fun.

[0003] In the early days, gesture recognition was carried out by extracting manual features, such as feature extraction methods like HOG and SIFT. However, manual features require manual selection and targeted processing, which is time-consuming, laborious, and not generalizable. The birth of Convolutional Neural Networks (CNN) revolutionized the way of feature extraction, no longer relying on manual operations, thus promoting the development of fields such as image classification, face recognition, speech separation, and object detection. CNN is known for its ability to process large amounts of data, local perception, and weight sharing, and thus occupies an important position in many application scenarios. In order to capture the spatial graph connection relationship, experts began to apply skeleton point data to fields such as gesture and sign language recognition. However, since the topological graph constructed based on skeleton points does not support translational invariance, traditional CNN technologies are difficult to handle. In this case, Graph Convolutional Network (GCN) has shown its unique value, being able to effectively handle non-Euclidean data types such as topological graphs and capture complex spatio-temporal relationships in the spatial and temporal dimensions.

[0004] Convolutional networks were first applied to the field of action recognition and then improved based on this, promoting the development of the fields of gesture and sign language recognition. Yan et al. adopted a three-partition strategy and proposed a spatio-temporal graph convolution that fuses information in the spatial and temporal domains. Subsequent expert teams improved the spatial convolution module on this basis, using an adaptable graph convolution to solve the problem of insufficient flexibility in the skeletal point graph structure; introduced high-level semantic features to establish the connection between the spatial dimension and the temporal dimension; adopted decoupled graph convolution to increase the expressiveness of spatial fusion without increasing the time and memory costs; adopted shifted graph convolution to achieve feature exchange between channels through shift operations and pointwise convolutions, further improving the lightweight degree; and improved the neighborhood matrix in the spatial graph convolution to solve the problem of biased weights.

[0005] The application of GCN in the field of gesture recognition has demonstrated its significant advantages in processing hand joint movements and capturing the details of hand actions. With the continuous progress of deep learning technology, gesture recognition technology is developing towards a more intelligent, real-time, and accurate direction, providing strong support for achieving a safer and more convenient social environment.

[0006] In the existing technology's graph convolutional network, channel feature extraction is performed through graph convolution operations, and the feature vector of each node contains multiple channel features. In each layer of convolution, GCN aggregates the features of nodes through the adjacency matrix and updates the features of the target node according to the features of neighboring nodes. Specifically, the feature vector of a node is fused with the feature vectors of its adjacent nodes through weighted summation to capture the structural information of the graph. After each layer of convolution, a non-linear activation function (such as ReLU) is usually applied to enhance the expressive power of the model and further explore the potential of different channel features. As the number of layers increases, GCN gradually extracts higher-order channel features, providing rich and accurate node representations for downstream tasks.

[0007] The disadvantages are as follows: In the actual model, the contributions of different channel features to the model are different. In the graph convolutional network, all channels are usually treated equally, resulting in some channel features being overemphasized while some effective features may be ignored. This processing method cannot effectively select and weight the most useful channel features.

[0008] Most of the channel attention in the existing technology consists of square convolutions. Specifically, square convolutions usually refer to using a convolution kernel of size k×k, and the commonly used sizes are 3×3 or 5×5 convolution kernels, which operate on each channel. First, the input feature map generates an attention map for each channel through square convolution, and these maps reflect the importance of different channels in the local space. Then, the generated attention maps are used to weight the features of each channel, enhancing the important channel features and suppressing the unimportant channels by multiplying element-wise with the original feature map, effectively capturing the local relationships between channels.

[0009] The disadvantages are as follows: Square convolution has certain limitations. Especially when dealing with input data with complex structures, a large amount of redundant channel information will still be generated. Since square convolution usually uses a fixed convolution kernel size (such as 3×3 or 5×5 convolution), the extraction of channel features mainly depends on the local similarity and ignores the differences between different scales and structures. As a result, redundant information is retained in the channel features, reducing the model processing efficiency.

[0010] In summary, how to solve the problems of some channel features being overemphasized while some effective channel features are ignored, the channel attention composed of square convolution ignoring the differences between different scales and structures, retaining redundant information, and single-channel feature extraction is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0011] The present invention aims at the above-mentioned deficiencies of the prior art and provides a channel attention method with strong practicability in gesture recognition applications.

[0012] A further technical task of the present invention is to provide a channel attention system in gesture recognition applications with reasonable design, safety and applicability.

[0013] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0014] In the channel attention method for gesture recognition applications, a channel attention module based on asymmetric convolution is adopted, using a vertical convolution kernel skeleton and a horizontal convolution kernel skeleton to strengthen the discriminative channel features of the feature map.

[0015] By combining the horizontal convolution kernel skeletons and vertical convolution kernel skeletons of multiple scales, a channel attention module composed of asymmetric multi-scale convolution is proposed, which finely processes the input features at multiple scales, retains multi-scale features and performs information fusion, improves the model's dynamic perception ability of different channel information, and enhances the attention to significant channels.

[0016] Further, the asymmetric convolution refers to vertical convolution and horizontal convolution. The vertical convolution only retains the vertical convolution kernel skeleton in the convolution kernel. When the image is horizontally flipped, the vertical convolution will retain the correct features at symmetric positions.

[0017] Further, for the channel attention module AMC-CA based on asymmetric multi-branch convolution, the input feature map is F 0 , AMC-CA first compresses the spatial and temporal information, and the temporal information is denoted as SqueezeST(·), only retaining the channel information to obtain the feature Figure 1 , denoted as F C1 , which is expressed as:

[0018] FC1 = SqueezeST(F 0 ).

[0019] Furthermore, after passing through the fully connected layer FC, the number of channels is reduced from C to thereby reducing the computational amount, where r is the channel reduction coefficient; after activation by the non-linear activation function ReLU, the feature map F C2 :

[0020] F C2 = ReLU(FC(F C1 ))

[0021] Furthermore, 2m multi-scale asymmetric convolutions are used in parallel to extract features, with the convolutional kernels being k i ×1, 1×k i (k i = 3,..., 2m + 1, i = 1,..., m), denoted as β 1 , β 2 , L, β 2m There is a relational expression k i = 2×i + 1. Then, the output feature maps are concatenated by channel, denoted by ||, and the Sigmoid activation function is used to obtain the feature map F C3 :

[0022] F C3 = Sigmoid(β 1 (F C2 ) || β 2 (F C2 ) || L || β 2m (F C2 ));

[0023] Finally, the feature Figure 4 is restored to the initial dimension:

[0024] F C4 = scale(F C3 ).

[0025] In the channel attention system for gesture recognition applications, the channel attention module based on asymmetric convolution adopts vertical and horizontal convolutional kernel skeletons to strengthen the discriminative channel features of the feature map;

[0026] By combining horizontal and vertical convolutional kernel skeletons of multiple scales, a channel attention module composed of asymmetric multi-scale convolutions is proposed, which refines the input features at multiple scales, retains multi-scale features and performs information fusion, enhancing the model's dynamic perception ability of different channel information and increasing the attention to significant channels.

[0027] Furthermore, the asymmetric convolution refers to vertical convolution and horizontal convolution. In vertical convolution, only the vertical convolution kernel skeleton is retained in the convolution kernel. When the image is horizontally flipped, the vertical convolution will retain the correct features at symmetric positions.

[0028] Furthermore, for the channel attention module AMC-CA based on asymmetric multi-branch convolution, the input feature map is F 0 , and AMC-CA first compresses the spatial and temporal information. The temporal information is denoted as SqueezeST(·), and only the channel information is retained to obtain the feature Figure 1 , denoted as F C1 , which is expressed as:

[0029] F C1 = SqueezeST(F 0 ).

[0030] Furthermore, after passing through the fully connected layer FC, the number of channels is reduced from C to thereby reducing the computational amount. r is the channel reduction coefficient; after activation by the non-linear activation function ReLU, the feature map F C2 is obtained:

[0031] F C2 = ReLU(FC(F C1 )).

[0032] Furthermore, 2m multi-scale asymmetric convolutions are used in parallel to extract features. The convolution kernels are k i ×1, 1×k i (k i = 3,..., 2m + 1, i = 1,…, m), denoted as β 1 , β 2 , L, β 2m There is a relational expression k i = 2×i + 1. Then, the output feature maps are concatenated by channel, denoted by ||, and the Sigmoid activation function is used to obtain the feature map F C3 :

[0033] F C3 = Sigmoid(β 1 (F C2 ) || β 2 (F C2 ) || L || β 2m (F C2 ));

[0034] Finally, the feature Figure 4 is restored to the initial dimension:

[0035] F C4 = scale(F C3 ).

[0036] Compared with the prior art, the channel attention method and system in the gesture recognition application of the present invention have the following outstanding beneficial effects:

[0037] The present invention adopts a channel attention module to automatically assign weight values to each channel that match its importance. By weighting each feature, the model focuses more on effective channel features, improves the generalization ability of the model, and makes it perform more robustly when facing different data sets. By introducing the attention mechanism, not only the problem of feature imbalance is solved, but also the overall model performance is improved.

[0038] The present invention proposes a channel attention module composed of asymmetric convolutions. Using the convolutional kernel skeletons in the vertical and horizontal directions, it enhances the ability to extract discriminative channel features in the feature map, effectively reduces unnecessary redundant information at the same time, and enhances the discriminative channel features. It not only has advantages in terms of computational complexity and number of parameters, but also reduces the risk of misidentifying information after image transformation, and improves the adaptability to image flipping and rotation operations.

[0039] The present invention combines multi-scale horizontal and vertical convolutional kernel skeletons, and proposes a channel attention module based on asymmetric multi-scale convolutions. The module refines features at multiple scales, better adapts to diverse input data, thereby improving the overall performance and scope of application. Retaining multi-scale features and performing information fusion improves the model's dynamic perception ability of different channel information and suppresses the interference of redundant information. The AMC-CA module makes the model more efficient when processing diverse data and has stronger robustness to image rotation and flipping transformations. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Att Figure 1 is a schematic structural diagram of a channel attention module based on asymmetric multi-scale convolutions in the channel attention method in the gesture recognition application;

[0042] Att Figure 2 is the feature map extracted by 3×1 vertical convolution before and after the image is rotated 180 degrees in the channel attention method in the gesture recognition application;

[0043] Att Figure 3In the channel attention method for gesture recognition applications, before and after the image is rotated 180 degrees, a 3×3 square convolution is used to extract the feature map;

[0044] Appendix Figure 4 In the channel attention method for gesture recognition applications, before and after the image is rotated 90 degrees, a 5×1 vertical convolution is used to extract the feature map;

[0045] Appendix Figure 5 In the channel attention method for gesture recognition applications, before and after the image is rotated 90 degrees, a 5×5 square convolution is used to extract the feature map;

[0046] Appendix Figure 6 In the channel attention method for gesture recognition applications, before and after the image is flipped along the vertical axis, a 3×1 vertical convolution is used to extract the feature map;

[0047] Appendix Figure 7 In the channel attention method for gesture recognition applications, before and after the image is flipped along the vertical axis, a 3×3 square convolution is used to extract the feature map;

[0048] Appendix Figure 8 In the channel attention method for gesture recognition applications, before and after the image is flipped along the vertical axis, a 5×1 vertical convolution is used to extract the feature map;

[0049] Appendix Figure 9 In the channel attention method for gesture recognition applications, before and after the image is flipped along the vertical axis, a 5×5 vertical convolution is used to extract the feature map. Detailed implementation

[0050] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0051] The following gives a best embodiment:

[0052] In the channel attention method for gesture recognition applications in this embodiment, based on the channel attention module of asymmetric convolution, a vertical convolution kernel skeleton and a horizontal convolution kernel skeleton are used to strengthen the discriminative channel features of the feature map;

[0053] By combining the horizontal convolution kernel skeletons and vertical convolution kernel skeletons of multiple scales, a channel attention module composed of asymmetric multi-scale convolutions is proposed, which finely processes the input features at multiple scales, retains the multi-scale features and performs information fusion, improves the model's dynamic perception ability of different channel information, and enhances the attention to the significant channels.

[0054] Asymmetric convolution refers to vertical convolution and horizontal convolution. Vertical convolution is equivalent to only retaining the vertical convolution kernel skeleton in a square convolution kernel. When the image is horizontally flipped, vertical convolution will retain the correct features at symmetric positions, while square convolution may not necessarily do so.

[0055] As Figure 6 shown, when horizontally flipping along the vertical axis, the features corresponding to the 3×1 vertical convolution are equal to the original features at the symmetric positions (as shown in the part within the purple rectangular frame), so it retains the local correct features, which is of great significance for the correct recognition of gestures. As Figure 7 shown, the features corresponding to the 3×3 square convolution change (as shown in the purple rectangular frame part), resulting in meaningless or incorrect results. Generalizing to the K×1 vertical convolution can retain the original features at horizontally symmetric positions, enhancing the robustness of the image under horizontal flipping. The same applies to horizontal convolution, where the 1×K horizontal convolution enhances the robustness of the image under vertical flipping.

[0056] As Figure 8 shown, it is an example of the vertical convolution with K = 5. While it is highly likely that the K×K square convolution cannot retain the correct features. As Figure 9 shown, it is an example of the square convolution with K = 5. The same applies to horizontal convolution, where the 1×k horizontal convolution enhances the robustness of the image under vertical flipping. Therefore, horizontal convolution and vertical convolution reduce the risk of misidentifying incorrect features after image flipping transformation.

[0057] When the image rotates, asymmetric convolution has a greater possibility of retaining the original features. If the features within a local area are symmetric up and down, then the corresponding features can definitely be found after rotating 180 degrees. As Figure 2 shown, the features within the purple rectangular frame are symmetric up and down, and the same features still exist after rotating 180 degrees. After calculating through the 3×1 vertical convolution, the original information is retained. While the square convolution fails to capture the correct feature information, as shown in Figure 3 . If there are features symmetric about the diagonal within a local area, then the same features can still be found after rotating 90 degrees. As Figure 4 shown, the features within the purple frame still have the original features after rotating 90 degrees (the features within the yellow frame are obtained by rotation). However, as Figure 5 shown, the square convolution cannot capture the same information. Generalizing to horizontal convolution and the image rotating at a general angle, compared with square convolution, horizontal convolution and vertical convolution have a greater probability of retaining the correct local features.

[0058] As Figure 1 shown, the channel attention module AMC-CA based on asymmetric multi-branch convolution consists of a compression module, an asymmetric convolution module, an activation function, etc.

[0059] The input feature map is F 0, the AMC-CA module first compresses spatial and temporal information (denoted as SqueezeST(·)), only retaining channel information, to obtain a feature Figure 1 , denoted as F C1 , expressed as:

[0060] F C1 = SqueezeST(F 0 ) (1)

[0061] Then, through the fully connected layer FC, the number of channels is reduced from C to thereby reducing the computational amount (r is the channel reduction coefficient); after activation by the non-linear activation function ReLU, the feature map F C2 is obtained:

[0062] F C2 = ReLU(FC(F C1 )) (2)

[0063] 2m multi-scale asymmetric convolutions are used in parallel to extract features, with convolution kernels of k i ×1, 1×k i (k i = 3, …, 2m + 1, i = 1, ..., m), denoted as β 1 , β 2 , L, β 2m There is a relational expression k i = 2×i + 1. Then, the output feature maps are concatenated by channel (denoted by ||), and the Sigmoid activation function is used to obtain the feature map F C3 :

[0064] F C3 = Sigmoid(β 1 (F C2 ) || β 2 (F C2 ) || L || β 2m (F C2 )) (3)

[0065] Finally, the feature Figure 4 restores the initial dimension:

[0066] F C4 = scale(F C3 ) (4)

[0067] The AMC-CA module uses asymmetric convolutions, which is equivalent to a square convolution kernel k i ×k iOperations are performed on the horizontal or vertical skeletons in it, only strengthening the key channel features of the skeleton feature map. Multi-scale convolution is used for feature extraction, retaining multi-scale features and performing information fusion, enhancing the model's representation ability. The number of channels marked in the figure is the output channels of the current module, and the final feature map restores the dimension, so this module has the characteristics of plug-and-play.

[0068] Compared with square convolution, AMC-CA suppresses the interference of redundant information. It not only has advantages in terms of computational complexity and number of parameters, but also improves the model's perception ability of significant channel information and the robustness to vertical flipping and horizontal flipping of images. Ablation experiments show that when the convolution kernel sizes of 3×1, 1×3, 5×1, and 1×5 are selected, the model performance reaches the optimal.

[0069] The present invention applies the channel attention module based on asymmetric multi-scale convolution to three graph convolution-based backbone networks, namely Decouple-GCN, CTR-GCN, and SL-GCN. In the table, "1s" represents a single-stream network, and "4s" represents a four-stream network including joint stream, skeleton stream, joint motion stream, and skeleton motion stream. "+CA" represents the experimental effect after adding the AMC-CA module. The recognition accuracy adopts two indicators, Top-1 and Top-5. The former represents the accuracy rate that the category ranked first in the score in classification matches the actual result, and the latter represents the accuracy rate that the categories ranked top five in the score in classification match the actual result. The "accuracy" mentioned in the experimental results usually refers to the Top-1 accuracy. As can be seen from Table 1, the accuracy improvement of the AMC-CA module for the single-stream network is 1.28%-1.56%, and the accuracy improvement for the four-stream network is 0.65%-1.13%, with significant effects.

[0070] Table 1 Performance improvement of AMC-CA on the model in the WLASL2000 dataset

[0071]

[0072] Based on the above method, in the gesture recognition application in this embodiment, the channel attention system, the channel attention module based on asymmetric convolution, adopts vertical convolution kernel skeletons and horizontal convolution kernel skeletons to strengthen the discriminative channel features of the feature map;

[0073] By combining multi-scale horizontal convolution kernel skeletons and vertical convolution kernel skeletons, a channel attention module composed of asymmetric multi-scale convolution is proposed, which finely processes the input features at multiple scales, retains multi-scale features and performs information fusion, enhances the model's dynamic perception ability of different channel information, and enhances the attention to significant channels.

[0074] Among them, the asymmetric convolution refers to the vertical convolution and the horizontal convolution. The vertical convolution only retains the skeleton of the vertical convolution kernel. When the image is horizontally flipped, the vertical convolution will retain the correct features at the symmetric positions.

[0075] The channel attention module AMC-CA based on the asymmetric multi-branch convolution, the input feature map is F 0 , AMC-CA first compresses the spatial and temporal information, and the temporal information is denoted as SqueezeST(·), only retaining the channel information, to obtain the feature Figure 1 , denoted as F C1 , expressed as:

[0076] F C1 = SqueezeST(F 0 ).

[0077] Then, through the fully connected layer FC, the number of channels is reduced from C to thereby reducing the computational amount, where r is the channel reduction coefficient; after being activated by the non-linear activation function ReLU, the feature map F C2 is obtained:

[0078] F C2 = ReLU(FC(F C1 )).

[0079] 2m multi-scale asymmetric convolutions are used in parallel to extract features, and the convolution kernels are k i ×1, 1×k i (k i = 3,..., 2m + 1, i = 1,..., m), denoted as β 1 , β 2 , L, β 2m There is a relational expression k i = 2×i + 1. Then, the output feature maps are concatenated by channel, denoted by ||, and the Sigmoid activation function is used to obtain the feature map F C3 :

[0080] F C3 = Sigmoid(β 1 (F C2 ) || β 2 (F C2 ) || L || β 2m (F C2 ));

[0081] Finally, the feature Figure 4 restores the initial dimension:

[0082] F C4 = scale(F C3 ).

[0083] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any technical solution that conforms to the technical solutions described in the above specific embodiments of the present invention and any appropriate changes or substitutions made by any person of ordinary skill in the art shall fall within the patent protection scope of the present invention.

[0084] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A channel attention method in gesture recognition applications, characterized in that The channel attention module based on asymmetric convolution uses vertical and horizontal convolution kernel skeletons to enhance the discriminative channel features of feature maps; By combining the multi-scale horizontal convolution kernel skeleton and the vertical convolution kernel skeleton, a channel attention module composed of asymmetric multi-scale convolution is proposed. It refines the input features at multiple scales, retains multi-scale features and performs information fusion, improves the model's dynamic perception ability of different channel information, and enhances attention to significant channels.

2. The channel attention method in gesture recognition application according to claim 1, characterized in that: The asymmetric convolution refers to vertical convolution and horizontal convolution. The vertical convolution is to retain only the vertical convolution kernel skeleton in the convolution kernel. When the image is flipped horizontally, the vertical convolution will retain the correct features in a symmetrical position.

3. The channel attention method in gesture recognition application according to claim 2, characterized in that: AMC-CA is a channel attention module based on asymmetric multi-branch convolution. The input feature map is F0. AMC-CA first compresses the spatial and temporal information. The temporal information is recorded as S queezeST (·), only retaining the channel information, we get feature map 1, denoted as F C1 , expressed as: F C1 =SqueezeST(F0)。 4. The channel attention method in gesture recognition application according to claim 3, characterized in that: After the fully connected layer FC, the number of channels is reduced from C to Thereby reducing the amount of computation, r is the channel reduction coefficient; after activation by the nonlinear activation function ReLU, the feature map F is obtained C2 : F C2 =ReLU(FC(F C1 ))。 5. The channel attention method in gesture recognition application according to claim 3, characterized in that: 2m multi-scale asymmetric convolutions are used in parallel to extract features, and the convolution kernel is k i ×1, 1×k i (k i =3,...,2m+1, i=1,...,m), expressed as β1,β2,L,β 2m There exists a relation k i =2×i+1, and then concatenate the output feature maps by channel, represented by ||, and use the Sigmoid activation function to obtain the feature map F C3 : F C3 =Sigmoid(β1(F C2 )||β2(F C2 )||L||b 2m (F C2 )); Finally, feature map 4 restores the original dimension: F C4 =scale(F C3 )。 6. A channel attention system in gesture recognition applications, characterized in that The channel attention module based on asymmetric convolution uses vertical and horizontal convolution kernel skeletons to enhance the discriminative channel features of feature maps; By combining the multi-scale horizontal convolution kernel skeleton and the vertical convolution kernel skeleton, a channel attention module composed of asymmetric multi-scale convolution is proposed. It refines the input features at multiple scales, retains multi-scale features and performs information fusion, improves the model's dynamic perception ability of different channel information, and enhances attention to significant channels.

7. The channel attention system in gesture recognition application according to claim 6, characterized in that: The asymmetric convolution refers to vertical convolution and horizontal convolution. The vertical convolution is to retain only the vertical convolution kernel skeleton in the convolution kernel. When the image is flipped horizontally, the vertical convolution will retain the correct features in a symmetrical position.

8. The channel attention system in gesture recognition application according to claim 7, characterized in that: AMC-CA is a channel attention module based on asymmetric multi-branch convolution. The input feature map is F0. AMC-CA first compresses the spatial and temporal information. The temporal information is recorded as S queezeST (·), only retaining the channel information, we get feature map 1, denoted as F C1 , expressed as: F C1 =SqueezeST(F0)。 9. The channel attention system in gesture recognition application according to claim 8, characterized in that: After the fully connected layer FC, the number of channels is reduced from C to This reduces the amount of computation, r is the channel reduction coefficient; after activation by the nonlinear activation function ReLU, the feature map F is obtained C2 : F C2 =ReLU(FC(F C1 ))。 10. The channel attention system in gesture recognition application according to claim 9, characterized in that: 2m multi-scale asymmetric convolutions are used in parallel to extract features, and the convolution kernel is k i ×1, 1×k i (k i =3,...,2m+1, i=1,...,m), expressed as β1,β2,L,β 2m There exists a relation k i =2×i+1, and then concatenate the output feature maps by channel, represented by ||, and use the Sigmoid activation function to obtain the feature map F C3 : F C3 =Sigmoid(β1(F C2 )||β2(F C2 )||L||b 2m (F C2 )); Finally, feature map 4 restores the original dimension: F C4 =scale(F C3 )。

Citation Information

Cited By

  • Novel dual-path network architecture combining Mama and convolutional neural network

    CN120874937A