Dynamic gesture recognition method and device, equipment and storage medium

By adopting a target dual-path network in dynamic gesture recognition, combining cascaded small convolution kernels and 3D depth separation convolution layer to extract features, and using attention modules to fusion, the problem of large amount of model parameters and poor real-time recognition performance in the prior art is solved, and efficient and accurate gesture recognition is achieved.

CN119964246AActive Publication Date: 2025-05-09惠州市康冠汽车电子有限公司

Patent Information

Application Number
CN202510110379.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-09
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

In the prior art, the large amount of parameters of the dynamic gesture recognition model makes it difficult to deploy on resource-constrained edge devices, and information loss or delay may occur when processing rapidly changing gesture actions, affecting the accuracy and real-timeness of the recognition.

Method used

A target dual-path network is used to extract feature information through slow channels and fast channels. The slow channel uses a convolutional layer based on a cascading small convolution kernel, and the fast channel uses 3D depth to separate the convolutional layer. The attention module is used to allocate feature weights and fusion feature information of the two, and generate and recognize the fusion features.

Benefits of technology

It reduces computing costs, improves recognition speed and accuracy, and solves the problems of large model parameters, difficulty in deploying edge devices and poor real-time recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964246A_ABST
    Figure CN119964246A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic gesture recognition method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the steps of obtaining a to-be-recognized video containing a dynamic gesture, and performing key frame extraction on the to-be-recognized video to obtain a target key frame; inputting the target key frame into a target dual-path network to obtain first feature information output by a slow channel in the target dual-path network and second feature information output by a fast channel in the target dual-path network; wherein the convolutional layer of the slow channel is constructed based on a cascade small convolution kernel, and the convolutional layer of the fast channel is constructed based on 3D depth separable convolution; and performing feature weight distribution and feature fusion on the first feature information and the second feature information by using an attention module to obtain fused features, so that a classifier performs gesture type recognition according to the fused features. The calculation cost can be reduced, and the recognition speed and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a dynamic gesture recognition method, device, equipment and storage medium. Background Art

[0002] At present, human-computer interaction technology has evolved into one of the key factors driving the development of many fields. Dynamic gesture recognition, as a research direction with great potential, is gradually changing the way people interact with electronic devices, and its importance is becoming increasingly prominent. Traditional human-computer interaction methods, such as keyboards, mice, touch screens, etc., can no longer meet users' pursuit of immersive experience and operational convenience in some scenarios. Dynamic gesture recognition technology came into being, which allows users to communicate with devices through natural hand movements, greatly improving the intuitiveness and fluency of interaction; dynamic gesture recognition has a wide range of application scenarios, such as VR / AR, smart home systems, smart driving, medical fields, etc.

[0003] In traditional solutions, dynamic gesture recognition relies on traditional computer vision technologies, such as skin color-based segmentation, contour extraction, feature point tracking, etc. The algorithm is relatively simple, but the robustness of recognition in complex environments is poor, and it is easy to have false detection and missed detection, and the expressive ability of gestures is limited. To solve the above problems, the prior art proposes dynamic gesture recognition based on deep learning technology, such as convolutional neural network (CNN), recurrent neural network (RNN) and its variants. By learning a large amount of annotated gesture data, it can automatically extract high-level features of hand movements, thereby achieving more accurate and robust gesture recognition. However, the number of parameters in its model is large, making it impossible to deploy on resource-constrained edge devices (such as mobile devices and embedded devices), and it is difficult to achieve real-time and low power requirements; in addition, when processing rapidly changing gestures, information loss or delay may occur, affecting the accuracy and real-time performance of recognition. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a dynamic gesture recognition method, device, equipment and storage medium, which can reduce the computing cost and improve the recognition speed and accuracy. The specific scheme is as follows:

[0005] In a first aspect, the present application discloses a dynamic gesture recognition method, comprising:

[0006] Acquire a video to be identified that contains dynamic gestures, and extract key frames from the video to be identified to obtain target key frames;

[0007] Inputting the target key frame into the target dual-path network, obtaining first feature information output by a slow channel in the target dual-path network, and second feature information output by a fast channel in the target dual-path network; wherein the convolution layer of the slow channel is constructed based on cascaded small convolution kernels, and the convolution layer of the fast channel is constructed based on 3D depth-separable convolution;

[0008] The attention module is used to perform feature weight assignment and feature fusion on the first feature information and the second feature information to obtain fused features, so that the classifier can perform gesture type recognition based on the fused features.

[0009] Optionally, extracting key frames from the video to be identified to obtain target key frames includes:

[0010] Calculating the image entropy of each video frame in the video to be identified, and determining local extreme value points from the video to be identified according to all the image entropies; the local extreme value points include local maximum value points and local minimum value points;

[0011] Calculating the local density of each of the local extreme value points, and determining the minimum distance between each local extreme value point and the remaining local extreme value points based on the local density;

[0012] The cluster centers are determined according to the minimum distance and the first preset target quantity, and a target key frame is determined for each cluster center.

[0013] Optionally, inputting the target key frame into a target dual-path network to obtain first feature information output by a slow channel in the target dual-path network and second feature information output by a fast channel in the target dual-path network includes:

[0014] Inputting the target key frame into the slow channel; the slow channel is constructed in the order of data layer, convolution layer, pooling layer and residual layer;

[0015] Using the data layer of the slow channel, selecting a second preset target number of target key frames from the target key frames as inputs of the convolution layer of the slow channel, and obtaining first feature information of the target key frames according to outputs of the residual layer of the slow channel;

[0016] The target key frame is input into the fast channel, and the second feature information of the target key frame is obtained according to the output of the fast channel; the fast channel is constructed in the order of convolution layer, pooling layer and residual layer.

[0017] Optionally, the convolution layer of the slow channel is constructed based on multiple cascaded 3×3 convolution kernels.

[0018] Optionally, the 3D depth separable convolution includes a 3D depth convolution layer and a 3D point-by-point convolution layer;

[0019] The 3D deep convolution layer is used to perform independent convolution on each channel of the input feature map to obtain intermediate feature maps with the same number of channels;

[0020] The 3D point-by-point convolution layer is used to aggregate channel information of the intermediate feature map.

[0021] Optionally, the attention module is constructed by cascading a channel attention submodule and a spatial depth attention submodule.

[0022] Optionally, the using the attention module to perform feature weight assignment and feature fusion on the first feature information and the second feature information includes:

[0023] Using the channel attention submodule, assigning weights to the spatial features of the first feature information and the second feature information;

[0024] Using the spatial depth attention submodule, weights are assigned to the temporal features of the first feature information and the second feature information.

[0025] In a second aspect, the present application discloses a dynamic gesture recognition device, comprising:

[0026] A video acquisition module is used to acquire a video to be identified that contains dynamic gestures, and extract key frames from the video to be identified to obtain target key frames;

[0027] A feature extraction module, used to input the target key frame into a target dual-path network, obtain first feature information output by a slow channel in the target dual-path network, and second feature information output by a fast channel in the target dual-path network; wherein the convolution layer of the slow channel is constructed based on cascaded small convolution kernels, and the convolution layer of the fast channel is constructed based on 3D depth-separable convolution;

[0028] The feature fusion module is used to use the attention module to perform feature weight assignment and feature fusion on the first feature information and the second feature information to obtain fused features so that the classifier can use the fused features to perform gesture type recognition.

[0029] In a third aspect, the present application discloses an electronic device, including:

[0030] Memory, used to store computer programs;

[0031] The processor is used to execute the computer program to implement the above-mentioned dynamic gesture recognition method.

[0032] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein the computer program implements the aforementioned dynamic gesture recognition method when executed by a processor.

[0033] In the present application, a video to be identified containing dynamic gestures is obtained, and key frames are extracted from the video to be identified to obtain target key frames; the target key frames are input into a target dual-path network to obtain first feature information output by a slow channel in the target dual-path network, and second feature information output by a fast channel in the target dual-path network; wherein the convolution layer of the slow channel is constructed based on cascaded small convolution kernels, and the convolution layer of the fast channel is constructed based on 3D depth-separable convolution; an attention module is used to perform feature weight assignment and feature fusion on the first feature information and the second feature information to obtain fused features, so that the classifier can use the fused features to identify the gesture type. It can be seen that by extracting key frames and performing recognition based on key frames, useless and redundant data in model training can be reduced, and the generalization ability of the model and the recognition accuracy of the model can be improved; video-based recognition can avoid information loss or delay, efficiently process long sequences of gesture data, and improve recognition accuracy; by using cascaded small convolution kernels in the slow channel, the network's receptive field can be increased while reducing the number of model parameters, capturing data details and high-level features; by using 3D depth-separable convolution in the fast channel, the computing cost and the number of model parameters can be reduced, and the recognition speed can be improved; the attention module is used for feature weight allocation and feature fusion to enhance feature validity, enhance the adaptability of the model and the accuracy of reasoning recognition; and solve the problems of large number of model parameters, difficulty in deploying edge devices, and poor real-time recognition performance in existing algorithm models. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0035] Figure 1 A flow chart of a dynamic gesture recognition method provided by this application;

[0036] Figure 2 A specific key frame extraction schematic diagram provided for this application;

[0037] Figure 3 A specific 3D depth separable convolution method flow chart provided for this application;

[0038] Figure 4A schematic diagram of a specific feature fusion method provided in this application;

[0039] Figure 5 A schematic diagram of the structure of a dynamic gesture recognition system provided in this application;

[0040] Figure 6 A schematic diagram of the structure of a dynamic gesture recognition device provided in this application;

[0041] Figure 7 A structural diagram of an electronic device provided for this application. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] In the prior art, dynamic gesture recognition based on deep learning technology, such as convolutional neural networks, recurrent neural networks and their variants, can automatically extract high-level features of hand movements by learning a large amount of annotated gesture data, thereby achieving more accurate and robust gesture recognition. However, the large number of parameters in its model makes it impossible to deploy on resource-constrained edge devices, making it difficult to achieve real-time and low-power requirements; in addition, when processing rapidly changing gesture movements, information loss or delay may occur, affecting the accuracy and real-time performance of recognition. In order to overcome the above technical problems, the present application proposes a dynamic gesture recognition method that can reduce computing costs and improve recognition speed and accuracy.

[0044] The present application discloses a method for dynamic gesture recognition. Figure 1 As shown, the method may include the following steps:

[0045] Step S11: obtaining a video to be identified that contains dynamic gestures, and extracting key frames from the video to be identified to obtain target key frames.

[0046] In this embodiment, the acquired video to be recognized containing dynamic gestures is the original video and cannot be directly used as the input of the neural network. Each video needs to be decomposed into continuous video frames for processing. However, usually, the target action in the video only occupies a small part of the original video. At the same time, the moving target may also be blocked and interfered by a lot of background information, generating more invalid information or redundant information, which has a counterproductive effect on the training of the network model. Therefore, redundant information is removed by key frame extraction.

[0047] In some embodiments, the key frame extraction of the video to be identified to obtain the target key frame may include: calculating the image entropy of each video frame in the video to be identified, and determining local extreme points from the video to be identified based on all the image entropies; the local extreme points include local maximum points and local minimum points; calculating the local density of each of the local extreme points, and determining the minimum distance between each local extreme point and the remaining local extreme points based on the local density; determining the cluster center based on the minimum distance and a first preset target number, and determining a target key frame for each cluster center.

[0048] Specifically, the target key frame is extracted from the video using the image entropy and density clustering method. For example, in a video V, there are n consecutive video frames Fframes, and the target key frame Skeyframes can be expressed as: Skeyframes = Fframes (V); the key frame extraction steps are as follows:

[0049] S101: Calculate the image entropy of each frame in the video: where f i Represents a video frame, i is the index of the video frame, p fi (j) represents the image frame f i The probability density function of the image frame f i The grayscale pixel histogram is normalized to obtain the grayscale value, j represents the index of the grayscale value, that is, the different grayscale levels in the image, and the grayscale value range is 0-255. i ) represents the image frame f i The entropy value of .

[0050] S102: E(f i ) is mapped to a two-dimensional coordinate space through the formula:

[0051] The local maximum points and local minimum points in the two-dimensional coordinate space are calculated. The set of local maximum points and local minimum points of all video frames in the two-dimensional coordinate space is Pextreme; Pextreme = Pmax∪Pmin.

[0052] S103: Calculate the local density ρ of each extreme point in Pextreme, and calculate the minimum distance δ between the extreme point and other points based on ρ. Assume that the number of extreme points is N, where P k is an extreme point, P i (i=1,2,...,N-1) is the number of p k Other extreme points outside, dp k p i is the distance between two extreme points, and dc is the threshold distance. The calculation formulas for local density and minimum distance are as follows:

[0053]

[0054] That is, for each extreme point P i Calculate its difference from the extreme point P k Distance dp k p i By all P i Calculate and sum the local density to get the extreme point P k Density value Filter out Then calculate the points between these points and P k The distance is finally taken as the minimum value To some extent, it reflects the extreme point P k The proximity to a denser area is expressed as the minimum distance.

[0055] S104: Select N maximum values ​​δ as cluster centers and calculate key frames. For example, N=16 may be selected, i.e., 06 target key frames. Figure 2 The figure shows a video to be identified, and the framed part is the target key frame.

[0056] Step S12: Input the target key frame into the target dual-path network to obtain the first feature information output by the slow channel in the target dual-path network and the second feature information output by the fast channel in the target dual-path network; wherein the convolution layer of the slow channel is constructed based on cascaded small convolution kernels, and the convolution layer of the fast channel is constructed based on 3D depth-separable convolution.

[0057] In this embodiment, based on the dual-path network architecture as the main framework, targeted improvements are made on this basis to better adapt to dynamic gesture recognition tasks. The slow channel captures spatial semantic information in the video at a lower frame rate, focusing on the overall information and slowly changing information in the scene, using a larger number of channels, and learning high-level, long-term spatiotemporal information; the fast channel samples the video at a higher frame rate to capture details and rapidly changing information in the video, using a smaller number of channels, and focusing on local details of fast motion and short-term spatiotemporal information. In the existing dual-path network architecture, the slow channel uses a large convolution kernel, which increases the number of model parameters. This application uses cascaded small convolution kernels instead of large convolution kernels to reduce the number of model parameters while increasing the network's receptive field, capturing data details and advanced features; in the existing dual-path network architecture, the fast channel uses C3D convolution, resulting in a slow reasoning process, which cannot meet the needs of real-time recognition and cannot be applied to edge devices. This application uses 3D depth-separable convolution to improve the reasoning speed.

[0058] In this embodiment, the target key frame is input into the target dual-path network to obtain the first feature information output by the slow channel in the target dual-path network, including: inputting the target key frame into the slow channel; the slow channel is constructed in the order of data layer, convolution layer, pooling layer and residual layer; using the data layer of the slow channel, a second preset target number of target key frames are selected from the target key frame as the input of the convolution layer of the slow channel, and the first feature information of the target key frame is obtained according to the output of the residual layer of the slow channel; wherein the convolution layer of the slow channel is constructed based on multiple cascaded 3×3 convolution kernels. For example, the data layer of the slow channel (SlowPathway Network) uses a step size of 4×1×1, and the first dimension is set to 4. If 16 frames of key frames are input, 4 frames are selected as the actual input of the slow channel after passing through the data layer, and the last two dimensions of 1×1 are used to ensure that the size of each frame of the input image does not change. That is, it receives 4 video frames selected after grayscale conversion and key frame extraction as input. These frames are preprocessed to provide the network with a relatively concise and representative image sequence. The convolution layer is constructed by cascading small convolution kernels. Compared with the traditional large convolution kernels, the cascade of small convolution kernels can not only reduce the number of model parameters, but also increase the receptive field of the network, and more effectively capture the spatial semantic information of gestures and some slowly changing detail features; specifically, multiple 3×3 small convolution kernels can be cascaded to replace the larger 5×5 convolution kernel, which can reduce the parameters and learn the local features in the image more carefully. The slow channel uses a larger number of channels to fully learn high-level, long-term spatiotemporal information, and more channels allow the network to extract gesture features from different dimensions, such as different color channels, texture channels, etc., so as to better grasp the overall form and change trend of gestures.

[0059] In this embodiment, the target key frame is input into the target dual-path network to obtain the second feature information output by the fast channel in the target dual-path network, including: inputting the target key frame into the fast channel, and obtaining the second feature information of the target key frame according to the output of the fast channel; the fast channel is constructed in the order of convolution layer, pooling layer and residual layer. The 3D depth separable convolution includes a 3D depth convolution layer and a 3D point-by-point convolution layer; the 3D depth convolution layer is used to perform independent convolution on each channel on the input feature map to obtain an intermediate feature map of the same number as the number of channels; the 3D point-by-point convolution layer is used to aggregate channel information on the intermediate feature map. All the extracted target key frames are used as the input of the fast channel, so that the advantage of high frame rate can be used to capture the rapid changes in gestures and more subtle movement details in the video.

[0060] 3D depth separable convolution is applied in the convolution layer. 3D depth separable convolution is divided into 3D depth convolution and 3D point-by-point convolution, which greatly reduces the computational cost and the number of model parameters while effectively extracting spatiotemporal features. 3D depth convolution independently convolves each channel on the 3D feature map to obtain channel-independent intermediate feature maps, improving the real-time reasoning performance of the model without much loss in accuracy. The 3D depth convolution formula is as follows:

[0061]

[0062] Among them, W1 represents the weight of 3D depth convolution, V represents the input 3D feature map, i, j, u represent the position index, K, L, M represent the convolution kernel size, Represents the dot product of each element. The 3D point-by-point convolution will act on the channel-independent feature map obtained in the previous step Further aggregate channel information. The formula is defined as follows:

[0063]

[0064] Where W2 represents the weight of the 3D point-by-point convolution, and n represents the size of the convolution kernel. The 3D depthwise convolution and 3D point-by-point convolution are performed sequentially to form a complete 3D depthwise separable convolution process, for example Figure 3 As shown, the formula is as follows:

[0065] Conv sepConv (V) = Conv Point (Conv Depth (V));

[0066] The comparison of 3D separable convolution and traditional 3D convolution parameters is shown in Table 1:

[0067] Table 1. Comparison of parameters between 3D separable convolution and traditional 3D convolution

[0068] Convolution name Parameter quantity 3D Convolution <![CDATA[M*K 3 *C*C^]]> 3D Deep Convolution <![CDATA[M*K 3 *C]]> 3D point-by-point convolution M*C*C^ 3D Separable Convolution <![CDATA[M*K 3 *C+M*C*C^]]>

[0069] Among them, M = H*W*D, for example, H, W are set to 112, and D is 16. It can be seen from the above table that when K = 3, C^ = 32, the number of parameters of 3D separable convolution is only 1 / 14 of the number of parameters of 3D convolution, which significantly reduces the complexity of the model.

[0070] The Fast Pathway maintains a smaller number of channels because it focuses on the local details of fast motion and short-term spatiotemporal information. A smaller number of channels can reduce the amount of computation while ensuring the capture of key information, increase the network's operating speed, match the high frame rate input, and achieve effective capture of fast gestures.

[0071] The improved network structure and the feature map size output after each layer of the network are shown in Table 2. Where T represents the temporal depth, that is, the number of input video frames, S 2 Indicates the square size of the feature map.

[0072] Table 2. Improved slowfast network structure and its output

[0073]

[0074] The convolution layer of the slow channel violates the 1×7 2 It means that the convolution kernel size is 1×7×7, the number of output channels after convolution is 64, and the convolution step is 1×2×2. The maximum number of channels in the slow channel is 512, and the maximum number of channels in the fast channel is 256. Compared with the Input, the number of input video frames of the Data Layer layer of the fast channel will be reduced. Compared with the Input, the number of input video frames and image size of the DataLayer layer of the fast channel are unchanged. After the image passes through the improved slowfast network, the output data shape of each layer [BS, T, W, H, C] is shown in the following table:

[0075] Table 3. Shape of output feature map of each layer of the improved slowfast network

[0076] Stage Slow Pathway Network Fast Pathway Network Input.shape [32,16,112,112,1] [32,16,112,112,1] Data layer.shape [32,4,112,112.32] [32,16,112,112,1] Conv1.shape [32,4,56,56,64] [32,16,56,56,16] Pool1.shape [32,4,56,56,64] [32,16,56,56,16] Res2.shape [32,4,28,28,128] [32,16,28,28,64] Res3.shape [32,4,14,14,256] [32,16,14,14,128] Res4.shape [32,4,7,7,512] [32,16,7,7,256]

[0077] BS is BatchSize, which can be manually set according to the hardware resources of the training network. Take BS=32 as an example. T is the number of input video frames, W and H are the width and height of the input video frame / feature map, respectively, and C is the number of channels of the feature map. The number of channels of Input is 1, which means that the input is a grayscale image.

[0078] Step S13: using the attention module to perform feature weight assignment and feature fusion on the first feature information and the second feature information to obtain fused features, so that the classifier can perform gesture type recognition based on the fused features.

[0079] Finally, feature weight allocation and feature fusion are performed based on the first feature information and the second feature information to obtain fused features. In this embodiment, feature fusion is not just simple splicing, but also combined with an attention module, specifically a 3D spatiotemporal attention module, which can adaptively allocate weights according to the importance of different path features, fully aggregate spatial and temporal features, and enhance the effectiveness of features. In order to fully utilize and fuse the temporal and spatial features between continuous video frames, a 3D spatiotemporal attention module is proposed to improve the model recognition accuracy.

[0080] In some embodiments, the attention module is constructed by cascading a channel attention submodule and a spatial depth attention submodule. The above-mentioned use of the attention module to perform feature weight assignment and feature fusion on the first feature information and the second feature information includes: using the channel attention submodule to assign weights to the spatial features of the first feature information and the second feature information; using the spatial depth attention submodule to assign weights to the temporal features of the first feature information and the second feature information. That is, the channel attention submodule focuses on spatial features, such as background features, and the spatial depth attention submodule pays more attention to the association between different frames, that is, the features of the time dimension. For example, for some gesture parts that have obvious features in space but do not change obviously in time, the 3D spatiotemporal attention module will give higher weights to the corresponding features in the slow channel; and for the details of rapidly changing gesture movements, the weights of the features of the fast channel will be increased, thereby achieving more accurate feature fusion.

[0081] Specifically, the structure of the 3D spatiotemporal attention module is as follows Figure 4 As shown in the figure, the channel attention submodule and the spatial depth attention submodule are cascaded to obtain the 3D spatiotemporal attention module. The channel attention weight is calculated by the multi-layer perceptron MLP, and the spatial depth attention weight is calculated by acting on the multi-dimensional (spatial, depth) feature map through different convolution kernels. Finally, the complete 3D spatiotemporal attention module is formed by cascading the two modules. The weight calculation formula of the channel attention submodule is as follows:

[0082]

[0083] Among them, V represents the input feature map, Avg Pool () indicates average pooling, Max Pool () represents maximum pooling, σ() represents the activation function, which is used to introduce nonlinear characteristics into the model and enhance the expressive power of the model. The activation function is the Relu() function. is the weight calculated after channel attention. The intermediate feature map V′ is obtained. Furthermore, the weight calculation formula of the spatial depth attention submodule is as follows:

[0084]

[0085] Among them, f 1×7×7 represents the convolution operation, 1×7×7 is the convolution kernel, which represents the horizontal convolution of the feature map; f 7×1×1 represents a convolution operation with a convolution kernel of 7×1×1, which represents the convolution in the depth direction (i.e., time dimension) of the feature map; f 7×7×7 Represents the overall convolution operation. The final feature map output by the 3D spatiotemporal attention module is: The target dual-path network combines the feature information extracted by the two paths, which can more comprehensively and accurately classify various dynamic gestures. Finally, after the improved target dual-path network and the 3D spatiotemporal attention module, a representative feature vector is obtained, which is used as the input of the fully connected classification layer and sent to the classifier for gesture classification.

[0086] For example Figure 5 The figure shows a specific structure diagram of a dynamic gesture recognition system.

[0087] It consists of four parts: data acquisition module, key frame extraction module, network main framework and classifier. The specific recognition process is as follows: collect gesture data through the camera, and convert the acquired continuous video frames into grayscale; use key frame technology to extract key frames; select some key frames and input them into the slow channel of the target dual-path network to capture the spatial semantic information in the video at a lower frame rate; at the same time, through parallel processing, input all the extracted key frames into the fast channel, sample the video at a higher frame rate, and capture the details and rapidly changing information in the video; use the 3D spatiotemporal attention module to splice and fuse the feature information of the two paths to achieve a comprehensive understanding of the video, that is, consider the overall information of the video and the local rapidly changing information, and finally input the fused information into the classifier (Softmax) for gesture classification. The classification results include but are not limited to: no gesture, shaking hand, and thumbs up. Whether on the PC side or ported to the system-level chip (Soc), significant results have been achieved, and the recognition accuracy and real-time inference speed can meet the needs of multiple scenarios.

[0088] As can be seen from the above, in this embodiment, a video to be identified containing dynamic gestures is obtained, and key frames are extracted from the video to be identified to obtain target key frames; the target key frames are input into the target dual-path network to obtain the first feature information output by the slow channel in the target dual-path network, and the second feature information output by the fast channel in the target dual-path network; wherein the convolution layer of the slow channel is constructed based on cascaded small convolution kernels, and the convolution layer of the fast channel is constructed based on 3D depth-separable convolution; the attention module is used to assign feature weights and fuse the first feature information and the second feature information to obtain fused features, so that the classifier can use the fused features to identify the gesture type. It can be seen that by extracting key frames and performing recognition based on key frames, useless and redundant data in model training can be reduced, and the generalization ability of the model and the recognition accuracy of the model can be improved; video-based recognition can avoid information loss or delay, and efficiently process long sequences of gesture data; by using cascaded small convolution kernels in the slow channel, the receptive field of the network can be increased while reducing the number of model parameters, capturing data details and high-level features; by using 3D depth-separable convolution in the fast channel, the computing cost and the number of model parameters can be reduced; the attention module is used for feature weight allocation and feature fusion to enhance feature validity, enhance the adaptability of the model and the accuracy of reasoning and recognition; and solve the problems of large number of parameters in existing algorithm models, difficulty in deploying edge devices, and poor real-time recognition performance.

[0089] Accordingly, the present application also discloses a dynamic gesture recognition device, see Figure 6 As shown, the device comprises:

[0090] The video acquisition module 11 is used to acquire a video to be identified that contains dynamic gestures, and extract key frames from the video to be identified to obtain target key frames;

[0091] A feature extraction module 12 is used to input the target key frame into a target dual-path network to obtain first feature information output by a slow channel in the target dual-path network and second feature information output by a fast channel in the target dual-path network; wherein the convolution layer of the slow channel is constructed based on cascaded small convolution kernels, and the convolution layer of the fast channel is constructed based on 3D depth-separable convolution;

[0092] The feature fusion module 13 is used to use the attention module to perform feature weight assignment and feature fusion on the first feature information and the second feature information to obtain fused features so that the classifier can use the fused features to perform gesture type recognition.

[0093] As can be seen from the above, in this embodiment, a video to be identified containing dynamic gestures is obtained, and key frames are extracted from the video to be identified to obtain target key frames; the target key frames are input into the target dual-path network to obtain the first feature information output by the slow channel in the target dual-path network, and the second feature information output by the fast channel in the target dual-path network; wherein the convolution layer of the slow channel is constructed based on cascaded small convolution kernels, and the convolution layer of the fast channel is constructed based on 3D depth-separable convolution; the attention module is used to assign feature weights and fuse the first feature information and the second feature information to obtain fused features, so that the classifier can use the fused features to identify the gesture type. It can be seen that by extracting key frames and performing recognition based on key frames, useless and redundant data in model training can be reduced, and the generalization ability of the model and the recognition accuracy of the model can be improved; video-based recognition can avoid information loss or delay, and efficiently process long sequences of gesture data; by using cascaded small convolution kernels in the slow channel, the receptive field of the network can be increased while reducing the number of model parameters, capturing data details and high-level features; by using 3D depth-separable convolution in the fast channel, the computing cost and the number of model parameters can be reduced; the attention module is used for feature weight allocation and feature fusion to enhance feature validity, enhance the adaptability of the model and the accuracy of reasoning and recognition; and solve the problems of large number of parameters in existing algorithm models, difficulty in deploying edge devices, and poor real-time recognition performance.

[0094] In some specific embodiments, the video acquisition module 11 may specifically include:

[0095] An image entropy calculation unit, used to calculate the image entropy of each video frame in the video to be identified, and determine local extreme value points from the video to be identified according to all the image entropies; the local extreme value points include local maximum value points and local minimum value points;

[0096] a minimum distance determination unit, configured to calculate the local density of each of the local extreme value points, and determine the minimum distance between each local extreme value point and the remaining local extreme value points based on the local density;

[0097] The target key frame extraction unit is used to determine the cluster center according to the minimum distance and the first preset target number, and determine a target key frame for each cluster center.

[0098] In some specific embodiments, the feature extraction module 12 may specifically include:

[0099] An input unit, used for inputting the target key frame into the slow channel; the slow channel is constructed in the order of a data layer, a convolutional layer, a pooling layer and a residual layer;

[0100] A first feature information acquisition unit is used to select a second preset target number of target key frames from the target key frames using the data layer of the slow channel as inputs of the convolution layer of the slow channel, and obtain first feature information of the target key frames according to outputs of the residual layer of the slow channel;

[0101] The second feature information acquisition unit is used to input the target key frame into the fast channel, and obtain the second feature information of the target key frame according to the output of the fast channel; the fast channel is constructed in the order of convolution layer, pooling layer and residual layer.

[0102] In some specific embodiments, the convolution layer of the slow channel may be constructed based on multiple cascaded 3×3 convolution kernels.

[0103] In some specific embodiments, the 3D depth separable convolution may specifically include a 3D depth convolution layer and a 3D point-by-point convolution layer; the 3D depth convolution layer is used to perform independent convolution on each channel on the input feature map to obtain an intermediate feature map with a number equal to the number of channels; the 3D point-by-point convolution layer is used to aggregate channel information on the intermediate feature map.

[0104] In some specific embodiments, the attention module can be specifically constructed by cascading a channel attention submodule and a spatial depth attention submodule.

[0105] In some specific embodiments, the feature fusion module 13 may specifically include:

[0106] A first weight allocation unit, configured to allocate weights to spatial features of the first feature information and the second feature information using the channel attention submodule;

[0107] The second weight allocation unit is used to use the spatial depth attention sub-module to allocate weights to the temporal features of the first feature information and the second feature information.

[0108] Furthermore, the present application also discloses an electronic device, see Figure 7 As shown, the contents in the drawings should not be considered as any limitation on the scope of use of the present application.

[0109] Figure 7 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the dynamic gesture recognition method disclosed in any of the aforementioned embodiments.

[0110] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0111] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system 221, a computer program 222, and data 223 including a video to be identified, etc. The storage method can be temporary storage or permanent storage.

[0112] The operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, so as to realize the operation and processing of the massive data 223 in the memory 22 by the processor 21, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the dynamic gesture recognition method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.

[0113] Furthermore, an embodiment of the present application also discloses a computer storage medium, in which computer executable instructions are stored. When the computer executable instructions are loaded and executed by a processor, the steps of the dynamic gesture recognition method disclosed in any of the aforementioned embodiments are implemented.

[0114] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0115] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0116] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0117] The above is a detailed introduction to a dynamic gesture recognition method, device, equipment and storage medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A dynamic gesture recognition method, characterized in that: include: Acquire a video to be identified that contains dynamic gestures, and extract key frames from the video to be identified to obtain target key frames; Inputting the target key frame into the target dual-path network, obtaining first feature information output by a slow channel in the target dual-path network, and second feature information output by a fast channel in the target dual-path network; wherein the convolution layer of the slow channel is constructed based on cascaded small convolution kernels, and the convolution layer of the fast channel is constructed based on 3D depth-separable convolution; The attention module is used to perform feature weight assignment and feature fusion on the first feature information and the second feature information to obtain fused features, so that the classifier can perform gesture type recognition based on the fused features.

2. The dynamic gesture recognition method according to claim 1, characterized in that: The step of extracting key frames from the video to be identified to obtain target key frames includes: Calculating the image entropy of each video frame in the video to be identified, and determining local extreme value points from the video to be identified according to all the image entropies; the local extreme value points include local maximum value points and local minimum value points; Calculating the local density of each of the local extreme value points, and determining the minimum distance between each local extreme value point and the remaining local extreme value points based on the local density; The cluster centers are determined according to the minimum distance and the first preset target quantity, and a target key frame is determined for each cluster center.

3. The dynamic gesture recognition method according to claim 1, characterized in that: The step of inputting the target key frame into the target dual-path network to obtain first feature information output by a slow channel in the target dual-path network and second feature information output by a fast channel in the target dual-path network includes: Inputting the target key frame into the slow channel; the slow channel is constructed in the order of data layer, convolution layer, pooling layer and residual layer; Using the data layer of the slow channel, selecting a second preset target number of target key frames from the target key frames as inputs of the convolution layer of the slow channel, and obtaining first feature information of the target key frames according to outputs of the residual layer of the slow channel; The target key frame is input into the fast channel, and the second feature information of the target key frame is obtained according to the output of the fast channel; the fast channel is constructed in the order of convolution layer, pooling layer and residual layer.

4. The dynamic gesture recognition method according to claim 1, characterized in that: The convolution layer of the slow channel is constructed based on multiple cascaded 3×3 convolution kernels.

5. The dynamic gesture recognition method according to claim 1, characterized in that: The 3D depth separable convolution includes a 3D depth convolution layer and a 3D point-by-point convolution layer; The 3D deep convolution layer is used to perform independent convolution on each channel of the input feature map to obtain intermediate feature maps with the same number of channels; The 3D point-by-point convolution layer is used to aggregate channel information of the intermediate feature map.

6. The dynamic gesture recognition method according to any one of claims 1 to 5, characterized in that: The attention module is constructed by cascading a channel attention submodule and a spatial depth attention submodule.

7. The dynamic gesture recognition method according to claim 6, characterized in that: The using the attention module to perform feature weight allocation and feature fusion on the first feature information and the second feature information includes: Using the channel attention submodule, assigning weights to the spatial features of the first feature information and the second feature information; Using the spatial depth attention submodule, weights are assigned to the temporal features of the first feature information and the second feature information.

8. A dynamic gesture recognition device, characterized in that: include: A video acquisition module is used to acquire a video to be identified that contains dynamic gestures, and extract key frames from the video to be identified to obtain target key frames; A feature extraction module, used to input the target key frame into a target dual-path network, and obtain first feature information output by a slow channel in the target dual-path network, and second feature information output by a fast channel in the target dual-path network; wherein the convolution layer of the slow channel is constructed based on cascaded small convolution kernels, and the convolution layer of the fast channel is constructed based on 3D depth-separable convolution; The feature fusion module is used to use the attention module to perform feature weight assignment and feature fusion on the first feature information and the second feature information to obtain fused features so that the classifier can use the fused features to perform gesture type recognition.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the dynamic gesture recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein when the computer program is executed by a processor, the dynamic gesture recognition method as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Shark motion behavior analysis method and system based on key frame detection and semantic component segmentation

    CN112528823A

  • Taking identification method and device based on double-channel cross attention mechanism

    CN113936339A

  • Gesture recognition method and electronic equipment

    CN115661941A

  • Time bottleneck attention system structure for video action recognition

    CN116686017A

  • Leaf disease identification method and system based on double-path convolutional neural network

    CN117392552A

Cited By

  • Improved SlowFast network and application thereof in highway illegal behavior identification

    CN121259705A