Real-time gesture detection method and device

By adopting a gesture detection model based on a separable convolutional structure and a residual structure in mobile gesture detection, the problem that gesture recognition cannot meet the real-time requirements due to limited processing capabilities of mobile devices is solved, and efficient and real-time gesture recognition effect is achieved.

CN114612832BActive Publication Date: 2025-05-09BIGO TECH PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210249415.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2025-05-09
Estimated Expiration
2042-03-14

AI Technical Summary

Technical Problem

In the prior art, due to the large amount of convolutional neural network computing, the equipment processing capability is limited, which makes gesture recognition unable to meet the real-time requirements.

Method used

Using a gesture detection model based on separable convolutional structure and residual structure, the original feature maps of multiple different levels of the input image are acquired and multiple original feature maps are fused to reduce the calculation amount of feature extraction and enhance the detection ability of the target.

Benefits of technology

It effectively reduces the calculation amount of gesture detection, improves the real-timeness of gesture recognition, enhances the detection effect of small targets and fuzzy scenes, and meets the real-time computing efficiency and high-precision requirements of mobile terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612832B_ABST
    Figure CN114612832B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a real-time gesture detection method and device. The technical solution provided by the embodiment of the present application obtains an image to be identified and inputs the image to be identified into a gesture detection model for gesture recognition, and determines the gesture type and gesture position according to the gesture recognition result output by the gesture detection model. The gesture detection model extracts multiple layers of original feature maps of the input image based on a separable convolution structure and a residual structure, reduces the computational amount of feature extraction, reduces the computational amount of gesture detection, and fuses multiple original feature maps to obtain a fused feature map. The fused features are used to enhance the detection capability of the target to compensate for the performance loss caused by the reduction in the number of parameters, and at the same time enhance the detection effect for small targets and blurred scenes. Then, gesture recognition is performed according to the fused feature map and the gesture recognition result is output, which can effectively meet the real-time requirements of gesture recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology, and in particular to a real-time gesture detection method and device. Background Art

[0002] With the large-scale rise of live video and short video applications on mobile terminals, smart content applications on mobile terminals are becoming more and more widespread. As an important interactive method, gestures can be used for emotional expression, interactive entertainment, virtual games, and so on.

[0003] Gesture detection can directly obtain the position of the hand in the image and the type of gesture currently being made, which is of great significance for the interaction of live broadcast and short video applications. Existing gesture detection methods are mainly divided into two categories: gesture detection based on traditional features such as SIFT and gesture detection based on convolutional neural networks. The former calculates the position and category of gestures in the image by extracting some scale-invariant features in the image. However, such features are generally designed manually, and their ability to express the features contained in the image is very limited, which is prone to missed detection and false detection. The latter extracts image features through multi-layer convolutional neural networks, and then regresses the position and category of gestures in the image. However, the general convolutional neural network has a huge amount of calculation, and the computing power, memory, heat dissipation capacity, etc. of mobile devices are limited, and cannot be directly applied to scenes with high real-time requirements such as live broadcast. Summary of the invention

[0004] The embodiments of the present application provide a real-time gesture detection method and device to solve the technical problem in the prior art that gesture recognition on a mobile terminal cannot meet the real-time requirements due to the large amount of convolutional neural network calculations and limited device processing capabilities. The method reduces the amount of calculations for gesture detection and can effectively meet the real-time requirements of gesture recognition.

[0005] In a first aspect, an embodiment of the present application provides a real-time gesture detection method, comprising:

[0006] Obtain an image to be recognized;

[0007] Inputting the image to be recognized into a trained gesture detection model so that the gesture detection model outputs a gesture recognition result based on the image to be recognized, wherein the gesture detection model is configured to obtain a plurality of original feature maps of different levels of the input image based on a separable convolution structure and a residual structure, fuse the plurality of original feature maps to obtain a plurality of fused feature maps, perform gesture recognition based on the plurality of fused feature maps and output a gesture recognition result;

[0008] The gesture type and gesture position are determined based on the gesture recognition result output by the gesture detection model.

[0009] In a second aspect, an embodiment of the present application provides a real-time gesture detection device, including an image acquisition module, a gesture recognition module and a gesture determination module, wherein:

[0010] The image acquisition module is configured to acquire an image to be identified;

[0011] The gesture recognition module is configured to input the image to be recognized into a trained gesture detection model so that the gesture detection model outputs a gesture recognition result based on the image to be recognized, and the gesture detection model is configured to obtain a plurality of original feature maps of different levels of the input image based on a separable convolution structure and a residual structure, fuse the plurality of the original feature maps to obtain a plurality of fused feature maps, perform gesture recognition based on the plurality of the fused feature maps and output a gesture recognition result;

[0012] The gesture determination module is configured to determine a gesture type and a gesture position based on a gesture recognition result output by the gesture detection model.

[0013] In a third aspect, an embodiment of the present application provides a real-time gesture detection device, including: a memory and one or more processors;

[0014] The memory is used to store one or more programs;

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the real-time gesture detection method as described in the first aspect.

[0016] In a fourth aspect, an embodiment of the present application provides a storage medium storing computer executable instructions, which, when executed by a computer processor, are used to perform the real-time gesture detection method as described in the first aspect.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor of a device reads and executes the computer program from the computer-readable storage medium, so that the device performs the real-time gesture detection method as described in the first aspect.

[0018] The embodiment of the present application obtains an image to be recognized and inputs the image to be recognized into a gesture detection model for gesture recognition, and determines the gesture type and gesture position according to the gesture recognition result output by the gesture detection model. The gesture detection model extracts original feature maps of multiple levels of the input image based on a separable convolution structure and a residual structure, reduces the computational amount of feature extraction, reduces the computational amount of gesture detection, and fuses multiple original feature maps to obtain a fused feature map. The fused features are used to enhance the detection capability of the target to compensate for the performance loss caused by the reduction in the number of parameters, and at the same time enhance the detection effect for small targets and blurred scenes. Gesture recognition is then performed according to the fused feature map and the gesture recognition result is output, which can effectively meet the real-time requirements of gesture recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of a real-time gesture detection method provided by an embodiment of the present application;

[0020] Figure 2 It is a schematic diagram of a process of extracting features from an input image provided by an embodiment of the present application;

[0021] Figure 3 This is a schematic diagram of a basic feature extraction network structure provided in an embodiment of the present application;

[0022] Figure 4 It is a schematic diagram of a fusion process of an original feature map provided in an embodiment of the present application;

[0023] Figure 5 This is a schematic diagram of a feature fusion network structure provided in an embodiment of the present application;

[0024] Figure 6 It is a schematic diagram of a flow chart of performing gesture recognition on a fused feature map provided in an embodiment of the present application;

[0025] Figure 7 This is a schematic diagram of a separate detection head network structure provided in an embodiment of the present application;

[0026] Figure 8 It is a schematic diagram of the relationship between a fusion feature map and a priori frame provided in an embodiment of the present application;

[0027] Fig. 9 is a structural diagram of a real-time gesture detection device provided in an embodiment of the present application;

[0028] Fig.10 It is a structural diagram of a real-time gesture detection device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical scheme and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for the convenience of description, only the part related to the present application but not all the contents are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow chart describes each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of each operation can be rearranged. The above process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The above process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc.

[0030] Figure 1 A flow chart of a real-time gesture detection method provided in an embodiment of the present application is given. The real-time gesture detection method provided in an embodiment of the present application can be executed by a real-time gesture detection device, which can be implemented by hardware and / or software and integrated in a real-time gesture detection device.

[0031] The following description is made by taking a real-time gesture detection device executing a real-time gesture detection method as an example. Figure 1 , the real-time gesture detection method comprises:

[0032] S101: Acquire an image to be recognized.

[0033] The image to be identified can be obtained through a video or image acquired through the Internet or a local gallery, or can be obtained through real-time shooting by a camera module mounted on the real-time gesture detection device. For example, a video application (such as a live video software) is installed on the real-time gesture detection device (such as a mobile terminal), and each frame of the image is used as the image to be identified while shooting the video. After determining the gesture type and gesture position on the image to be identified, the next step of processing can be performed based on the gesture type and gesture position.

[0034] Taking live video streaming software as an example, a gesture detection model is configured in the live video streaming software. When it is necessary to render relevant special effects based on the host's gestures, the captured video frame is obtained as the image to be recognized and submitted to the gesture detection model for gesture recognition. The gesture type and gesture position are determined based on the gesture recognition result output by the gesture detection model, and the special effect type is determined based on the gesture type and the rendering position of the special effect is determined based on the gesture position.

[0035] S102: Input the image to be recognized into the trained gesture detection model so that the gesture detection model outputs a gesture recognition result based on the image to be recognized. The gesture detection model is configured to obtain original feature maps of multiple different levels of the input image based on a separable convolution structure and a residual structure, fuse the multiple original feature maps to obtain multiple fused feature maps, perform gesture recognition based on the multiple fused feature maps and output the gesture recognition result.

[0036] Exemplarily, a trained gesture detection model is configured in the real-time gesture detection device. After obtaining the images to be recognized, the images to be recognized are input into the gesture detection model in sequence. The gesture detection model performs gesture recognition based on the received images to be recognized and outputs corresponding gesture recognition results.

[0037] Among them, the gesture detection model is built based on the separable convolution structure and the residual structure. When the gesture detection model performs gesture recognition based on the received image to be recognized, the original feature maps of multiple different levels of the input image (i.e., the image to be recognized) are obtained based on the separable convolution structure and the residual structure, and the multiple original feature maps are fused to obtain multiple fused feature maps. Gesture recognition is performed based on the multiple fused feature maps and the gesture recognition results are output. The gesture detection model provided by this solution extracts the original feature maps of multiple levels of the input image based on the separable convolution structure and the residual structure, effectively reducing the computational amount of feature extraction and gesture detection, and fuses multiple original feature maps to obtain fused feature maps. The fused features are used to enhance the detection capability of the target to compensate for the performance loss caused by the reduction in the number of parameters, while enhancing the detection effect for small targets and blurred scenes, which can effectively meet the real-time requirements of gesture recognition.

[0038] In one possible embodiment, the gesture detection model provided by the present application includes a hierarchical feature extraction network, a feature fusion network, and a separate detection head network connected in sequence. The hierarchical feature extraction network is configured to obtain multiple original feature maps of different levels of the input image based on a separable convolution structure and a residual structure, the feature fusion network is configured to fuse multiple original feature maps output by the hierarchical feature extraction network to obtain multiple fused feature maps, and the separate detection head network is configured to perform gesture recognition based on multiple fused feature maps and output gesture recognition results. In one embodiment, the gesture recognition results output by the separate detection head network include predicted gesture categories, gesture confidence, and predicted gesture positions.

[0039] In one embodiment, the hierarchical feature extraction network provided by the present application includes a plurality of serial basic feature extraction networks, and the basic feature extraction network of each level is configured to extract features from the input image to obtain the original feature map of the corresponding level. And the size of the original feature map output by the basic feature extraction network of each level is halved relative to the size of the input image, and the number of channels of the original feature map (the number of channels of the convolution structure) is doubled relative to the number of channels of the input image.

[0040] Among them, the output image of the basic feature extraction network of one level is used as the input image of the basic feature extraction network of the next level. For example, the basic feature grid of the first level uses the acquired image to be recognized as the input image, halves the input corresponding to the image to be recognized (halves the size) and doubles the channels (doubles the number of channels of the input convolution structure), extracts features from the image to be recognized and outputs the original feature map of the first level. Further, the original feature map of the first level is used as the input image of the basic feature extraction network of the second level, halves the input corresponding to the original feature map of the first level (halves the size) and doubles the channels (doubles the number of channels of the input convolution structure), extracts features from the original feature map of the first level and outputs the original feature map of the second level, and so on to obtain the original feature map of each level.

[0041] In one possible embodiment, the hierarchical feature extraction network of the present application includes 5 serial basic feature extraction networks, that is, the hierarchical feature extraction network is composed of 5 layers of serial basic feature extraction networks, and the original feature map obtained by each layer of basic feature extraction network is reduced by half relative to the size (length and width) of the input image. Correspondingly, the downsampling step size of the entire hierarchical feature extraction network is 32. After feature extraction, the input image obtains an original feature map with a length and width reduced by 32 times (the downsampling step size is 32). These original feature maps are characterized by being highly abstract and having rich high-level visual features.

[0042] In one embodiment, Figure 2 As shown in the flowchart of feature extraction of an input image, the basic feature extraction network provided in the present application specifically includes steps S1021-S1023 when extracting features of an input image:

[0043] S1021: Perform a convolution structure channel halving operation on the input image through the basic convolution module, and perform feature extraction on the input image after the convolution structure channel halving through the separable convolution module to obtain a feature extraction result.

[0044] Exemplarily, the basic feature extraction network is built based on a basic convolution module and a separable convolution module, wherein the basic convolution module can be used to change the number of channels of the input convolution structure, and the separable convolution module can be used for main feature extraction.

[0045] After receiving the input image (the input image of the first-level basic feature extraction network is the image to be recognized, and the input image of the subsequent-level basic feature extraction network is the original feature map output by the previous-level basic feature extraction network), the basic feature extraction network uses the basic convolution module to halve the convolution structure channels of the input image to reduce the amount of feature extraction calculations, and sends the input image after the convolution structure channels are halved to the separable convolution module for feature extraction to obtain the feature extraction result.

[0046] S1022: performing element-by-element addition on the input image and the feature extraction result after the convolution structure channel is halved to obtain an element-by-element addition result, and performing a confusion operation on the element-by-element addition result through a basic convolution module to obtain an element-by-element confusion result.

[0047] Exemplarily, after obtaining the feature extraction result obtained by the separable convolution module through feature extraction of the input image after the convolution structure channel is halved, the input image on which the convolution structure channel is halved by the previous basic convolution module is added element by element with the feature extraction result output by the separable convolution module (for example, the input image after the channel is halved and the corresponding pixels of the feature extraction result are added) to obtain the element-by-element addition result, and the basic convolution module is used to perform a confusion operation on the element-by-element addition result to obtain the element-by-element confusion result.

[0048] S1023: Perform string concatenation on the element-wise addition confusion result and the input image after the convolution structure channel is halved to obtain a concatenated result, and downsample the concatenated result to obtain the original feature map.

[0049] Exemplarily, after performing a confusion operation on the element-by-element addition result to obtain an element-by-element addition confusion result, the element-by-element addition confusion result and the input image with halved channels of the previous basic convolution module are further concatenated to obtain a connection result, and the connection result is further downsampled (assuming the downsampling step is 2) to obtain the original feature map of the basic feature extraction network at the current level with the input image input halved (the size is halved) and the channels doubled (the number of channels in the convolution structure is doubled).

[0050] In one embodiment, an efficient separable convolution (DwConv, depthwise separableconvolution) and residual structure can be used to construct a basic feature extraction network. Figure 3As shown in the schematic diagram of a basic feature extraction network structure provided, the basic feature extraction network (Layer in the figure) provided by this solution is constructed based on the basic convolution module (CBL in the figure) and the separable convolution module (DwUnit in the figure). Among them, the basic convolution module includes a 1*1 convolution kernel (1x1 Conv in the figure), a BatchNorm normalization unit (BatchNorm in the figure) and a LeakyReLU activation function unit (LeakyReLU in the figure) connected in sequence. Among them, the nonlinear activation function used by the LeakyReLU activation function unit is optimized by the ReLU activation function. Compared with other activation functions, it has the advantages of high computational efficiency and fast convergence speed, and reduces the sparsity of the ReLU activation function.

[0051] Among them, the separable convolution module includes a first basic convolution module (the CBL of the previous level of DwCBL in the figure), a feature extraction module (DwCBL in the figure) and a second basic convolution module (the CBL of the next level of DwCBL in the figure) connected in sequence. Among them, the feature extraction module includes a 3*3 depth-separable convolution kernel (3x3 DwConv in the figure), a BatchNorm normalization unit (BatchNorm in the figure) and a LeakyReLU activation function unit (LeakyReLU in the figure) connected in sequence. Among them, the feature extraction module DwConv is different from the traditional convolution. Each channel of the DwConv convolution kernel only performs convolution calculation with part of the channels of the input feature (the number of channels involved in the calculation can be pre-set), which greatly reduces the amount of calculation, but the feature extraction ability of the feature extraction module DwCBL is also weakened. Therefore, before using the feature extraction module DwConv, the basic convolution module CBL is used to increase the number of channels, and the basic convolution module CBL is used to reduce the number of channels after the feature extraction module DwConv.

[0052] After building the basic convolution module CBL and the separable convolution module DwUnit, the basic feature extraction network Layer is built based on the basic convolution module CBL and the separable convolution module DwUnit. In the figure, input is an image receiving module for receiving the input image. After the image receiving module input, a basic convolution module CBL is used to perform a convolution structure channel halving operation on the input image to halve the size of the input image. The left side of the basic feature extraction network Layer in the figure is a residual structure mainly using the separable convolution module DwUnit, and the right side does not perform other operations after the input is halved. In the residual structure on the left, after the image receiving module input is connected to the basic convolution module CBL, the separable convolution module DwUnit, the element addition module Add, the basic convolution module CBL, the channel connection module concat and the separable convolution module DwUnit with a stride of 2 (Strident=2) are connected in sequence. After the image receiving module input is connected to the basic convolution module CBL, the right side is connected to the channel connection module concat, forming a basic feature extraction network in the hierarchical feature extraction network.

[0053] Based on the above basic feature extraction network Layer, after the image receiving module input receives the input image, it performs channel halving through the basic convolution modules CBL on both sides. The input image after channel halving is subjected to feature extraction on the input image after channel halving of the convolution structure on the left side through the separable convolution module DwUnit block to obtain the feature extraction result. Then, the input image after channel halving of the convolution structure and the feature extraction result are added element by element in the element addition module Add to obtain the element by element addition result. The basic convolution module CBL after the element addition module Add performs a confusion operation on the element by element addition result to obtain the element addition confusion result. Further, in the channel connection module concat, the element addition confusion result output by the basic convolution module CBL after the element addition module Add and the input image after channel halving of the convolution structure output by the basic convolution module CBL on the right are string concatenated (i.e., the outputs of the basic convolution modules CBL on the left and right are concat-connected in the channel dimension) to obtain the connection result. Finally, the connection result is downsampled by the separable convolution module DwUnit with a step size of 2 to obtain the original feature map with the input halved relative to the input image and the channels doubled. The basic feature extraction network provided by this solution only performs convolution operations on the data of the left half of the channels, which reduces the amount of calculation by half. At the same time, the residual structure can well maintain the data transmission of the deep network. In one embodiment, five of the above basic feature extraction networks are used to form a hierarchical feature extraction network. The length and width of the original feature map obtained in each layer are reduced by half, and the downsampling step length of the entire network is 32.

[0054] It is understandable that in the above-mentioned hierarchical feature extraction network, due to multi-layer downsampling and scale (size) limitations, the final original feature map will lose some basic features and some targets. In order to ensure gesture detection in blurred scenes or small targets, the original feature maps of different levels can be fused, and feature fusion confusion can be used to enhance the ability of gesture recognition.

[0055] In related technologies, traditional feature fusion is similar to the FPN (Feature Pyramid Networks multi-level feature fusion, a top-down feature fusion method), which has many processing steps and complex calculations, making it difficult to achieve good real-time performance on mobile terminals. This solution aims to address the shortcomings of existing gesture detection methods in detection accuracy and computational efficiency by proposing a lightweight feature pyramid network structure to fuse the multi-layer original feature maps output by the hierarchical feature extraction network, efficiently fuse low-level pixel features and high-level abstract information, and complement each other to enhance the detection effect of small targets and occluded targets, which can meet the real-time computing efficiency and high-precision requirements of mobile terminals.

[0056] In one embodiment, when the feature fusion network fuses multiple original feature maps output by the hierarchical feature extraction network to obtain multiple fused feature maps, the feature fusion network specifically fuses the last three layers of original feature maps output by the hierarchical feature extraction network to obtain three fused feature maps. Exemplarily, the fusion method for fusing the original feature maps can adopt an element-wise (feature multiplication and addition) fusion method.

[0057] In one possible embodiment, Figure 4 As shown in the schematic diagram of the fusion process of the original feature map provided, when the feature fusion network fuses the last three layers of the original feature maps output by the hierarchical feature extraction network to obtain multiple fused feature maps, it includes steps S1024-S1026:

[0058] S1024: Perform downsampling step-size halving and channel-halving operations on the last layer original feature map output by the hierarchical feature extraction network to obtain a first intermediate feature map, and perform element-by-element addition of the first intermediate feature map and the penultimate layer original feature map output by the hierarchical feature extraction network to obtain a second fused feature map.

[0059] Exemplarily, the last three layers of original feature maps output by the hierarchical feature extraction network are used as the basis for fusion. Since the sizes of the original feature maps at different levels are different, taking the hierarchical feature extraction network with a 5-layer basic feature extraction network as an example, the downsampling steps of the original feature maps in the last three stages are x8, x16 and x32, respectively, and the corresponding numbers of channels are 128, 256 and 512, respectively. Before fusing the original feature maps, the downsampling step and the number of channels need to be processed so that the two original feature maps used for fusion are at the desired downsampling step and number of channels.

[0060] It can be understood that the downsampling step size and the number of channels of the last layer of the original feature map are twice the downsampling step size and the number of channels of the penultimate layer of the original feature map. Based on this, for the fusion of the last layer of the original feature map output by the hierarchical feature extraction network and the penultimate layer of the original feature map, this solution performs a downsampling step size halving and a channel halving operation on the last layer of the original feature map to obtain a first intermediate feature map, and performs element-by-element addition of the first intermediate feature map and the penultimate layer of the original feature map output by the hierarchical feature extraction network (for example, the corresponding pixels of the first intermediate feature map and the penultimate layer of the original feature map are added) to obtain a second fused feature map. In one embodiment, after obtaining the second fused feature map, the second fused feature map can be further subjected to feature confusion processing to further enhance the feature expression capability of the second fused feature map.

[0061] S1025: Perform downsampling step size halving and channel halving operations on the second fused feature map to obtain a second intermediate feature map, and perform element-by-element addition of the second intermediate feature map and the third-to-last original feature map output by the hierarchical feature extraction network to obtain a third fused feature map.

[0062] In a possible embodiment, the fusion processing of the second-to-last layer original feature map and the third-to-last layer original feature map output by the hierarchical feature extraction network can be performed according to the above-mentioned fusion of the last layer original feature map and the second-to-last layer original feature map.

[0063] Considering that the second fused feature map fuses the features of the last layer of original feature maps and the second-to-last layer of original feature maps, its feature expression capability is stronger. Based on this, the second fused feature map can be used to replace the second-to-last layer of original feature maps at this stage, that is, the fusion processing of the second fused feature map and the third-to-last layer of original feature maps is used. That is, the second fused feature map is downsampled by half and the channel is halved to obtain the second intermediate feature map, and the second intermediate feature map and the third-to-last layer of original feature map output by the hierarchical feature extraction network are element-by-element added (for example, the corresponding pixels of the second intermediate feature map and the third-to-last layer of original feature map are added) to obtain the third fused feature map. In one embodiment, after obtaining the third fused feature map, the third fused feature map can be further subjected to feature confusion processing to further enhance the feature expression capability of the third fused feature map.

[0064] S1026: Perform a downsampling step-size doubling operation on the second fused feature map to obtain a third intermediate feature map, and perform element-by-element addition of the third intermediate feature map and the last layer of the original feature map output by the hierarchical feature extraction network to obtain a first fused feature map.

[0065] For the fusion processing of the last layer of original feature map output by the hierarchical feature extraction network and the second fused feature map, the second fused feature map is downsampled and the step size is doubled to obtain a third intermediate feature map, and the third intermediate feature map and the last layer of original feature map output by the hierarchical feature extraction network are element-by-element added (for example, the third intermediate feature map and the corresponding pixels of the last layer of original feature map are added) to obtain an enhanced high-level feature map, i.e., the first fused feature map, which can be used to detect large targets in the image to be identified.

[0066] like Figure 5 As shown in the schematic diagram of a feature fusion network structure provided, it is assumed that F5, F4 and F3 in the figure are the original feature maps of the last layer, the second to last layer and the third to last layer of the hierarchical feature extraction network output, respectively. The downsampling steps of the original feature maps F5, F4 and F3 are x32, x16 and x8, respectively, and the number of channels is 512, 256 and 128, respectively. For the original feature map F5, the x2 upsampling module (UpSample) and the basic convolution module (1x1 CBL) are used to halve the downsampling step size (reduce the downsampling step size to x16) and halve the channels (reduce the number of channels to 256) of the original feature map F5 to obtain the first intermediate feature map P5, and the first intermediate feature map P5 and the original feature map F4 are fused in an element-by-element addition manner, and further use the 3x3 conv with stride=1 (3x3 DwCBL in the figure) to confuse the fused feature map to obtain the second fused feature map FF2.

[0067] Furthermore, the upsampling module (UpSample) and the basic convolution module (1x1 CBL) are used to respectively reduce the downsampling step size by half (reducing the downsampling step size to x8) and the channel by half (reducing the number of channels to 128) on the second fused feature map FF2 to obtain the second intermediate feature map P4, and the second intermediate feature map P4 and the original feature map F3 are fused by element-by-element addition, and further the 3x3 conv with stride=1 (3x3 DwCBL in the figure) is used to perform feature confusion on the fused feature map to obtain the third fused feature map FF3.

[0068] Furthermore, the 3x3 DwCBL with stride=2 is used to double the downsampling step size of the second fused feature map (3x3 conv is used to downsample once, and the downsampling step size is increased to x32), and then the second fused feature map FF2 and the original feature map F5 are fused by element-by-element addition to obtain the first fused feature map FF1. Among them, the second fused feature map FF2 and the third fused feature map FF3 both adopt the forward feature fusion method, especially the third fused feature map FF3 combines the perceptual features of the original feature maps F3, F4, F5, etc., and has a larger visual receptive field, which can better detect small targets and process blurred scenes, while the enhanced first fused feature map FF1 is mainly used to detect large targets, and the second fused feature map FF2 takes both into account. The three fused feature maps complement each other and effectively improve the performance of gesture detection.

[0069] In the related art, the existing end-to-end object detection network generally uses a fully connected layer or 1x1 conv to directly regress the target category and position information on the feature map. However, considering that some category features of gesture targets are similar, this method has defects in gesture detection. For example, extending two fingers and three fingers, especially in blurred scenes, will lead to a high false detection rate. Based on this, the separate detection head network of this solution performs gesture detection processing on multiple fused feature maps respectively. Figure 6 As shown in the flowchart of performing gesture recognition on a fused feature map, the separate detection head network provided in this solution includes steps S1027-S1028 when performing gesture recognition based on multiple fused feature maps and outputting gesture recognition results:

[0070] S1027: For each fused feature map, separate the fused feature map through a basic convolution module to obtain a first separated feature map, a second separated feature map, and a third separated feature map.

[0071] S1028: Determine a predicted gesture category according to the first separated feature map, determine a gesture confidence according to the second separated feature map, and determine a predicted gesture position according to the third separated feature map.

[0072] Exemplarily, for each fused feature map (including the first fused feature map FF1, the second fused feature map FF2 and the third fused feature map FF3 provided above), three basic convolution modules (1x1 CBL) are used to separate the fused feature map to obtain three branches, which are the first separated feature map, the second separated feature map and the third separated feature map. These three branches can be used to predict gesture category, gesture confidence and hand position respectively, and finally the three branches are merged as the final output.

[0073] Furthermore, a 1x1 conv convolution kernel can be used to determine the predicted gesture category from the first separated feature map, a 1x1 conv convolution kernel can be used to determine the gesture confidence from the second separated feature map, and a 1x1 conv convolution kernel can be used to determine the predicted gesture position from the third separated feature map. Finally, the outputs corresponding to the three branches are connected to the concat connection to output the gesture recognition results including the predicted gesture category, gesture confidence and predicted gesture position.

[0074] In one embodiment, before separating and fusing the feature map, the basic convolution module (1x1 CBL) can be used to reduce the number of channels corresponding to the separated feature map and reduce the amount of calculation. After obtaining the predicted gesture category, the softmax normalization module can be used to normalize the predicted gesture category. After obtaining the gesture confidence, the sigmoid normalization module can be used to normalize the gesture confidence to between 0 and 1, that is, if the normalized gesture confidence corresponding value is greater than 0.5, it means that the prior frame contains a valid target, and if it is less than 0.5, it means that the prior frame does not contain a valid target.

[0075] like Figure 7As shown in the schematic diagram of the structure of a separate detection head network, after obtaining the first fusion feature map FF1, the second fusion feature map FF2 and the third fusion feature map FF3, for each fusion feature map (FF in the figure), first use 1x1 CBL to reduce the number of channels and reduce the amount of calculation, and then use 3 1x1 CBL separations to obtain 3 branches, namely the first separation feature map, the second separation feature map and the third separation feature map. For the first separation feature map, use 1x1 conv normalization to obtain an output with the same number of preset categories (the preset number of categories is the number marked in the fusion feature map, for example, if 10 gestures need to be recognized, the different probabilities of the 10 gestures are output respectively, and the one with the largest probability is considered to correspond to the predicted gesture category), and then use softmax to normalize the probability of the category, and determine the category with the largest probability as the predicted gesture category. For the second separated feature map, 1x1 conv is used for normalization to obtain the gesture confidence, and the sigmoid function is used to normalize the gesture confidence to between 0 and 1. If the output is greater than 0.5, it means that the prior frame contains a valid target; if it is less than 0.5, it means that the prior frame does not contain a valid target. For the second separated feature map, 1x1 conv is used for normalization to obtain the predicted gesture position. Finally, the three branches are connected through concat, and the gesture recognition result including the predicted gesture category, gesture confidence and predicted gesture position is output.

[0076] In one embodiment, in order to obtain a more accurate predicted gesture position, this solution can use a grid position encoding based on a priori box (anchor) to represent the position information. The priori box is a target, and what needs to be predicted is the target position (the position of the target box, that is, the predicted box containing the target), but the direct prediction position range is too large. This solution sets the priori box, and the predicted target position is the priori box + offset (encoding). The predicted gesture position is represented based on the grid position encoding of the target box. The grid position encoding is used to represent the encoded coordinates of the target box in the feature grid, and the feature grid is obtained by dividing the fused feature map according to the set unit length.

[0077] The predicted gesture position is determined based on the decoded coordinates, decoded size, and downsampling step size of the target frame on the fused feature map. That is, the global absolute coordinates of the target frame on the corresponding fused feature map are determined based on the predicted decoded coordinates and decoded size of the target frame, and then the global absolute coordinates are multiplied by the downsampling step size of the fused feature map to obtain the global absolute coordinates of the target frame on the image to be identified. The global absolute coordinates of the target frame on the image to be identified are the predicted gesture position.

[0078] like Figure 8A schematic diagram of the relationship between a fused feature map and a priori frame is provided. In this embodiment, a priori frame (dashed frame) and a fused feature map are shown in the figure. Assuming that the length and width of the fused feature map are both N, that is, the size of the fused feature map is NxN, the fused feature map is divided into NxN feature grids (cells), and the length and width of each feature grid are 1. Three priori frames of different sizes are set in each feature grid (in this scheme, one image to be identified corresponds to three fused feature maps, and correspondingly, there are 9 priori frames of different sizes). Considering that direct prediction of the gesture position will cause serious drift, slow training convergence speed and large error, this scheme uses relative offset coordinate encoding. During the training process, the encoding result is predicted. During use, the predicted result is decoded to obtain the global absolute coordinates of the target (preset gesture) in the image to be identified.

[0079] In one embodiment, the coordinates of the upper left corner of the current feature grid are marked as (c x , c y ), the center coordinate of the prior box (t x , t y ) represents the offset from the upper left corner of the current feature grid, and the sigmoid function is used to set it between 0 and 1 (because the scale of each feature grid is recorded as 1). Based on this, the decoding coordinates provided by this solution can be determined based on the following formula:

[0080] b x =σ(t x )+c x

[0081] b y =σ(t y )+c y

[0082] Among them, (b x , b y ) is the decoded coordinates of the center coordinates of the target box on the fusion feature map, (c x , c y ) is the coordinate of the upper left corner of the current feature grid, σ(t x ) and σ(t y ) is the offset of the prior frame from the upper left corner of the current feature grid, (t x , t y ) is the encoded coordinates of the center coordinates of the prior box on the fused feature map.

[0083] The decoding size provided by this solution can be determined based on the following formula:

[0084]

[0085]

[0086] Among them, b h , and b w is the length and width of the decoded size of the target box, p h and p w is the length and width of the encoding size of the prior box, t h and t w is the exponential coefficient obtained by training the gesture detection model. x 、b y 、b h , and b w They are the global absolute coordinates of the target frame on the corresponding fused feature map. The global absolute coordinates are multiplied by the downsampling step size of the fused feature map to obtain the global absolute coordinates of the target frame on the image to be recognized. The global absolute coordinates of the target frame on the image to be recognized are the predicted gesture positions.

[0087] In one embodiment, for the training of the gesture detection model, different types of gesture pictures are collected, and the gesture targets in the pictures are manually annotated (including gesture type and gesture position), and then a training set and a validation set are constructed. Through back propagation and gradient descent methods, the parameters of the gesture detection model are iteratively trained and continuously updated based on the loss function. After the gesture detection model converges on the validation set, the parameters of the gesture detection model are saved and the model file of the gesture detection model is output. On a real-time gesture detection device such as a mobile application product, the saved gesture detection model file is loaded through a neural network inference framework, and the image to be identified is used as input to perform forward calculation of the gesture detection model, so that the gesture category and position contained in the image to be identified can be obtained. These results (gesture category and position) can be used as input signals for other technical chains such as special effects rendering to meet various mobile application requirements.

[0088] This solution adopts an end-to-end network structure. Correspondingly, the gesture detection model is trained in an end-to-end supervised training method, and the stochastic gradient descent method can be used for optimization and solution. The detection network used in this solution has three prediction branches that already have prior frames, so an optimized joint training method can be used to train the gesture detection model. Based on this, the gesture detection model is trained based on a joint loss function, where the joint loss function is determined based on whether the prior frame contains the target, the coordinate error between the prior frames, and the loss value of the predicted target and the matched prior frame.

[0089] In a possible embodiment, the joint loss function provided by this solution can be determined based on the following formula:

[0090]

[0091] in:

[0092]

[0093]

[0094]

[0095] Among them, W is the width of the fused feature map, H is the length of the fused feature map, A is the number of prior boxes for each point on the fused feature map, maxiou is the maximum overlap ratio of each prior box and all real targets, thresh is the set overlap ratio screening threshold, and λ noobj is the weight of the negative sample loss function set, is the coordinate of the kth priori box at the point with width i and length j on the current fusion feature map, o is the target score corresponding to the priori box, t is the number of training times, and λ prior is the weight of the warmup loss function, represents the coordinates of the kth prior frame, r represents the preset coordinates, is the coordinate of the prior box, Indicates that this part only calculates the loss value of the box that matches a real target, λ coord is the loss function weight of the coordinate, truth r is the coordinate value of the labeled target in the training sample, λ obj is the loss function weight of whether to include the target, is the IOU score between the prior box and the labeled target, λ class is the loss function weight for category prediction, truth c is the predicted target category, is the category of the prior box.

[0096] Among them, the first loss function loss1 is used to determine whether the predicted box contains the target. First, it is necessary to calculate the intersection-over-union (IoU) of each predicted box and all the marked groundtruths, and take the maximum value maxiou. If the value is less than the preset threshold (preset hyperparameter, such as 0.65), then the predicted box is marked as the background category, so it is necessary to calculate the confidence error of noobj (negative sample). Among them, the real target is the gesture marked on the sample image.

[0097] The second loss function, loss2, is used to calculate the coordinate error between the prior box and the predicted box, but only the first 12,800 iterations are calculated (this process is called the warmup process. The warmup method is used to enhance the shape convergence effect of the predicted box and effectively speed up the overall training speed). The second loss function is mainly set to enable the gesture detection model to quickly learn the length and width of the prior box and speed up the convergence of the overall training.

[0098] The third loss function loss3 is used to calculate various loss values ​​of the predicted target and a matching ground-truth target. Because each feature grid on the fusion feature map predicts 3 target boxes, and the number of real targets on a map is very small, and each real target corresponds to only one prediction box to be predicted, that is, a positive sample, and the remaining prediction boxes are negative samples. In order to distinguish whether the prediction box is a positive sample or a negative sample, this scheme can use a matching method to distinguish positive and negative samples: for a real target, first determine which feature grid its center point will fall in, and then calculate the IoU value of the three prior boxes of this feature grid and the real target (since the coordinates are not considered when calculating the IoU value, only the shape is considered, their upper left corner can be offset to the zero position before calculation), and select the prior box with the largest IoU as the match. Correspondingly, the prediction box corresponding to this prior box is the positive sample for subsequent calculations. All prediction boxes that are not matched by the true target are negative samples, so the number of negative samples is particularly large. In order to balance the positive and negative samples, this solution only selects the prediction boxes whose maxiou is less than the threshold for calculation according to the setting of the first loss function, and the rest of the prediction boxes are discarded. For the loss part of the positive sample, it is also divided into three parts for calculation, corresponding to the three branches of prediction (the first separation feature map, the second separation feature map and the third separation feature map): the first item is to calculate the coordinate loss of the prediction box and the true target, using the square difference loss function; the second item is the confidence loss, the smaller the IoU, the larger the loss function value; the third item is the classification loss, the category target corresponding to the true target is 1, and the other category targets are 0, and the cross entropy loss function is used to calculate the output result of softmax. This solution uses a joint optimized joint loss function to directly train the entire gesture detection model at one time, and sets a corresponding matching mechanism for positive and negative samples to reduce the situation where the training effect is poor due to the imbalance of the number of positive and negative samples.

[0099] S103: Determine the gesture type and gesture position based on the gesture recognition result output by the gesture detection model.

[0100] Exemplarily, after receiving the image to be recognized, the gesture detection model performs gesture recognition on the image to be recognized and outputs the corresponding gesture recognition result. It can determine whether a set type of gesture is recognized on the image to be recognized, and when a set type of gesture is recognized, the recognized gesture type and the gesture position corresponding to each gesture type according to the gesture recognition result output by the gesture detection model. For example, the predicted gesture category, gesture confidence and predicted gesture position in the gesture recognition result are determined, the gesture category and predicted gesture position whose corresponding gesture confidence reaches the set confidence threshold are determined, and the corresponding gesture category and predicted gesture position are determined as the gesture type and gesture position. When the gesture confidence corresponding to each gesture category and predicted gesture position is less than the set confidence threshold, it is determined that the target gesture is not recognized in the image to be recognized.

[0101] In a possible embodiment, after determining the gesture type and gesture position, the gesture response mode and gesture response position can be determined based on the gesture type. Determining the gesture response mode can be determining the special effect type for special effect rendering, and correspondingly, the gesture response position can be the rendering position of the corresponding special effect. In one embodiment, multiple different types of special effect information can be configured in the real-time gesture detection device, and special effect rendering can be performed according to the special effect information and the corresponding special effect can be displayed on the interactive interface.

[0102] Taking the live video broadcast software installed on the real-time gesture detection device as an example, a gesture detection model and special effect information of "heartbeat" type are configured in the live video broadcast software. When the host starts the live video broadcast software for live broadcast, he makes a "heart" gesture in the live broadcast screen. At this time, the live video broadcast software submits the real-time collected video frames to the gesture detection model, and the gesture detection model outputs a gesture recognition result indicating that the "heart" gesture type is detected at a certain position. The live video broadcast software determines that the "heart" gesture is detected based on the gesture recognition result, and can determine that the corresponding special effect type is "heartbeat". After determining the gesture type and gesture position, the "heartbeat" special effect is rendered and displayed at the position of the "heart" gesture according to the corresponding special effect information, thereby enriching the interactive experience between the host and the audience and realizing various mobile application needs.

[0103] In the above, by obtaining an image to be recognized and inputting the image to be recognized into a gesture detection model for gesture recognition, and determining the gesture type and gesture position according to the gesture recognition result output by the gesture detection model, the gesture detection model extracts original feature maps of multiple levels of the input image based on a separable convolution structure and a residual structure, reduces the computational amount of feature extraction, reduces the computational amount of gesture detection, and fuses multiple original feature maps to obtain a fused feature map, uses fused features to enhance the detection capability of the target to compensate for the performance loss caused by the reduction in the parameter amount, and at the same time enhances the detection effect for small targets and blurred scenes, and then performs gesture recognition according to the fused feature map and outputs the gesture recognition result, which can effectively meet the real-time requirements of gesture recognition. At the same time, by reducing the model parameter amount through separable convolution, and using the channel-level residual structure to reduce the number of input channels of the convolution calculation, the gesture detection model is lightweight, the model parameter amount and the computational amount are effectively reduced, and the training objectives such as warmup, position, and category are unified into a joint loss function for joint optimization, which accelerates the model convergence and operation efficiency. The residual structure and feature fusion are used to compensate for the performance loss caused by the reduction of parameters, while enhancing the detection effect for small targets and blurred scenes, effectively solving the problem of poor performance of end-to-end detection for small targets and blurred backgrounds. The prediction of gesture position is represented by encoding to reduce the prediction error caused by the difference in coordinate extreme values, while accelerating the convergence of training. This solution does not use connection layers such as full connection and pooling. The features of the image to be identified are extracted through a convolutional neural network. The positions and categories of all gestures in the image to be identified are regressed and output based on the features. In order to solve the problem that the traditional convolutional neural network has too much computational complexity, the deep separable convolution and feature pyramid structure are used to take into account both computational efficiency and the granularity of feature extraction, which can effectively reduce the computational scale of the network while ensuring the accuracy of the neural network, and can achieve better gesture recognition effects on mobile applications.

[0104] Fig. 9 is a structural diagram of a real-time gesture detection device provided by an embodiment of the present application. Fig. 9 The real-time gesture detection device includes an image acquisition module 21, a gesture recognition module 22 and a gesture determination module 23.

[0105] Among them, the image acquisition module 21 is configured to acquire the image to be recognized; the gesture recognition module 22 is configured to input the image to be recognized into the trained gesture detection model so that the gesture detection model outputs a gesture recognition result based on the image to be recognized, and the gesture detection model is configured to acquire multiple original feature maps of different levels of the input image based on the separable convolution structure and the residual structure, fuse the multiple original feature maps to obtain multiple fused feature maps, and perform gesture recognition based on the multiple fused feature maps and output the gesture recognition result; the gesture determination module 23 is configured to determine the gesture type and gesture position based on the gesture recognition result output by the gesture detection model.

[0106] In the above, by obtaining an image to be recognized and inputting the image to be recognized into a gesture detection model for gesture recognition, and determining the gesture type and gesture position according to the gesture recognition result output by the gesture detection model, the gesture detection model extracts multiple levels of original feature maps of the input image based on a separable convolution structure and a residual structure, reduces the computational amount of feature extraction, reduces the computational amount of gesture detection, and fuses multiple original feature maps to obtain a fused feature map. The fused features are used to enhance the detection capability of the target to compensate for the performance loss caused by the reduction in the number of parameters, and at the same time enhance the detection effect for small targets and blurred scenes. Gesture recognition is then performed according to the fused feature map and the gesture recognition result is output, which can effectively meet the real-time requirements of gesture recognition.

[0107] In one possible embodiment, the gesture detection model includes a hierarchical feature extraction network, a feature fusion network, and a separate detection head network, wherein:

[0108] A hierarchical feature extraction network configured to obtain original feature maps of multiple different levels of an input image based on a separable convolutional structure and a residual structure;

[0109] A feature fusion network is configured to fuse multiple original feature maps output by the hierarchical feature extraction network to obtain multiple fused feature maps;

[0110] The separate detection head network is configured to perform gesture recognition based on multiple fused feature maps and output gesture recognition results, wherein the gesture recognition results include predicted gesture category, gesture confidence, and predicted gesture position.

[0111] In one possible embodiment, the hierarchical feature extraction network includes multiple serial basic feature extraction networks, and the basic feature extraction network of each level is configured to perform feature extraction on the input image to obtain an original feature map of the corresponding level, wherein the size of the original feature map is halved relative to the size of the input image, and the number of channels of the original feature map is doubled relative to the number of channels of the input image.

[0112] In a possible embodiment, the basic feature extraction network includes a feature extraction module, an element addition confusion module and a data connection module, wherein:

[0113] A feature extraction module is configured to perform a convolution structure channel halving operation on an input image through a basic convolution module, and perform feature extraction on the input image after the convolution structure channel halving operation through a separable convolution module to obtain a feature extraction result;

[0114] An element addition obfuscation module is configured to perform element-by-element addition on the input image and the feature extraction result after the convolution structure channel is halved to obtain an element-by-element addition result, and perform an obfuscation operation on the element-by-element addition result through a basic convolution module to obtain an element-by-element obfuscation result;

[0115] The data connection module is configured to perform string connection on the element addition confusion result and the input image after the convolution structure channel is halved to obtain a connection result, and downsample the connection result to obtain the original feature map.

[0116] In a possible embodiment, the basic convolution module includes a 1*1 convolution kernel, a BatchNorm normalization unit and a LeakyReLU activation function unit connected in sequence, the separable convolution module includes a first basic convolution module, a feature extraction module and a second basic convolution module connected in sequence, and the feature extraction module includes a 3*3 depth-separable convolution kernel, a BatchNorm normalization unit and a LeakyReLU activation function unit connected in sequence.

[0117] In a possible embodiment, the hierarchical feature extraction network includes five serial basic feature extraction networks.

[0118] In a possible embodiment, the feature fusion network is configured to fuse the last three layers of original feature maps output by the hierarchical feature extraction network to obtain three fused feature maps.

[0119] In a possible embodiment, the feature fusion network includes a first fusion module, a second fusion module and a third fusion module, wherein:

[0120] The second fusion module is configured to perform downsampling step-halving and channel-halving operations on the last layer original feature map output by the hierarchical feature extraction network to obtain a first intermediate feature map, and perform element-by-element addition of the first intermediate feature map and the penultimate layer original feature map output by the hierarchical feature extraction network to obtain a second fused feature map;

[0121] A third fusion module is configured to perform downsampling step-size halving and channel-halving operations on the second fused feature map to obtain a second intermediate feature map, and perform element-by-element addition of the second intermediate feature map and the third-to-last original feature map output by the hierarchical feature extraction network to obtain a third fused feature map;

[0122] The first fusion module is configured to perform a downsampling step-size doubling operation on the second fused feature map to obtain a third intermediate feature map, and perform element-by-element addition of the third intermediate feature map and the last layer original feature map output by the hierarchical feature extraction network to obtain a first fused feature map.

[0123] In one possible embodiment, the separate detection head network includes a feature map separation module and a gesture prediction module, wherein:

[0124] A feature map separation module is configured to separate the fused feature map through the basic convolution module for each fused feature map to obtain a first separated feature map, a second separated feature map and a third separated feature map;

[0125] The gesture prediction module is configured to determine the predicted gesture category according to the first separated feature map, determine the gesture confidence according to the second separated feature map, and determine the predicted gesture position according to the third separated feature map.

[0126] In a possible embodiment, the predicted gesture position is represented based on a grid position code of a target box, where the grid position code is configured to represent the encoding coordinates of the target box in a feature grid, where the feature grid is obtained by dividing a fused feature map according to a set unit length.

[0127] In a possible embodiment, the predicted gesture position is determined based on the decoded coordinates and decoded size of the target box on the fused feature map and the downsampling step size of the fused feature map.

[0128] In one possible embodiment, the decoded coordinates are determined based on the following formula:

[0129] b x =σ(t x )+c x

[0130] b y =σ(t y )+c y

[0131] Among them, (b x , b y ) is the decoded coordinates of the center coordinates of the target box on the fusion feature map, (c x , c y ) is the coordinate of the upper left corner of the current feature grid, σ(t x ) and σ(t y ) is the offset of the prior frame from the upper left corner of the current feature grid, (t x , t y ) is the encoding coordinate of the center coordinate of the prior frame on the fusion feature map;

[0132] The decoded size is determined based on the following formula:

[0133]

[0134]

[0135] Among them, b h , and b w is the length and width of the decoded size of the target box, p h and p w is the length and width of the encoding size of the prior box, t h and t w Exponential coefficient obtained by training the gesture detection model.

[0136] In one possible embodiment, the gesture detection model is trained based on a joint loss function, which is determined based on whether the prior frame contains the target, the coordinate error between the prior frame and the prior frame, and the loss value of the predicted target and the matched prior frame.

[0137] In one possible embodiment, the joint loss function is determined based on the following formula:

[0138]

[0139] in:

[0140]

[0141]

[0142]

[0143] Among them, W is the width of the fused feature map, H is the length of the fused feature map, A is the number of prior boxes for each point on the fused feature map, maxiou is the maximum overlap ratio of each prior box and all real targets, thresh is the set overlap ratio screening threshold, and λ noobj is the weight of the negative sample loss function set, is the coordinate of the kth priori box at the point with width i and length j on the current fusion feature map, o is the target score corresponding to the priori box, t is the number of training times, and λ prior is the weight of the warmup loss function, represents the coordinates of the kth prior frame, r represents the preset coordinates, is the coordinate of the prior box, Indicates that this part only calculates the loss value of the box that matches a real target, λ coord is the loss function weight of the coordinate, truth r is the coordinate value of the labeled target in the training sample, λobj is the loss function weight of whether to include the target, is the IOU score between the prior box and the labeled target, λ class is the loss function weight for category prediction, truth c is the predicted target category, is the category of the prior box.

[0144] It is worth noting that in the embodiment of the above-mentioned real-time gesture detection device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present invention.

[0145] The embodiment of the present application further provides a real-time gesture detection device, which can be integrated with the real-time gesture detection apparatus provided in the embodiment of the present application. Fig.10 is a structural diagram of a real-time gesture detection device provided by an embodiment of the present application. Fig.10 The real-time gesture detection device includes: an input device 33, an output device 34, a memory 32 and one or more processors 31; the memory 32 is used to store one or more programs; when one or more programs are executed by one or more processors 31, the one or more processors 31 implement the real-time gesture detection method provided in the above embodiment. The real-time gesture detection device, equipment and computer provided above can be used to execute the real-time gesture detection method provided in any of the above embodiments, and have corresponding functions and beneficial effects.

[0146] The embodiments of the present application also provide a storage medium storing computer executable instructions, which are used to execute the real-time gesture detection method provided in the above embodiments when executed by a computer processor. Of course, the computer executable instructions of a storage medium storing computer executable instructions provided in the embodiments of the present application are not limited to the real-time gesture detection method provided above, and can also execute related operations in the real-time gesture detection method provided in any embodiment of the present application. The real-time gesture detection apparatus, device and storage medium provided in the above embodiments can execute the real-time gesture detection method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, please refer to the real-time gesture detection method provided in any embodiment of the present application.

[0147] In some possible implementations, various aspects of the method provided by the present disclosure may also be implemented in the form of a program product, which includes a program code. When the above-mentioned program product is run on a computer device, the above-mentioned program code is used to enable the above-mentioned computer device to execute the steps of the method according to various exemplary embodiments of the present disclosure described above in this specification. For example, the above-mentioned computer device can execute the real-time gesture detection method recorded in the embodiments of the present disclosure.

Claims

1. A real-time gesture detection method, characterized in that: include: Obtain an image to be recognized; Inputting the image to be recognized into a trained gesture detection model so that the gesture detection model outputs a gesture recognition result based on the image to be recognized, wherein the gesture detection model is configured to obtain a plurality of original feature maps of different levels of the input image based on a separable convolution structure and a residual structure, fuse the plurality of original feature maps to obtain a plurality of fused feature maps, perform gesture recognition based on the plurality of fused feature maps and output a gesture recognition result; Determine the gesture type and gesture position based on the gesture recognition result output by the gesture detection model; The gesture detection model includes a hierarchical feature extraction network, a feature fusion network and a separate detection head network, wherein: The hierarchical feature extraction network is configured to obtain original feature maps of multiple different levels of the input image based on a separable convolution structure and a residual structure; the feature fusion network is configured to perform a downsampling step size halving and a channel halving operation on the original feature map of the last layer output by the hierarchical feature extraction network to obtain a first intermediate feature map, and perform element-by-element addition of the first intermediate feature map and the penultimate original feature map output by the hierarchical feature extraction network to obtain a second fused feature map; perform a downsampling step size halving and a channel halving operation on the second fused feature map to obtain a second intermediate feature map, and The second intermediate feature map and the third-to-last original feature map output by the hierarchical feature extraction network are added element by element to obtain a third fused feature map; the second fused feature map is downsampled and doubled in step size to obtain a third intermediate feature map, and the third intermediate feature map and the last layer of the original feature map output by the hierarchical feature extraction network are added element by element to obtain a first fused feature map; the separate detection head network is configured to perform gesture recognition based on a plurality of the fused feature maps and output a gesture recognition result, wherein the gesture recognition result includes a predicted gesture category, a gesture confidence, and a predicted gesture position; When the separate detection head network performs gesture recognition based on the multiple fused feature maps and outputs the gesture recognition result, it includes: for each of the fused feature maps, separating the fused feature map through a basic convolution module to obtain a first separated feature map, a second separated feature map and a third separated feature map; using a 1x1 conv convolution kernel to determine the predicted gesture category from the first separated feature map, and using a softmax normalization module to normalize the predicted gesture category; using a 1x1 conv convolution kernel to determine the gesture confidence from the second separated feature map, and using a sigmoid normalization module to normalize the gesture confidence to between 0 and 1; using a 1x1 conv convolution kernel to determine the predicted gesture position from the third separated feature map; and outputting the gesture recognition result including the predicted gesture category, the gesture confidence and the predicted gesture position.

2. The real-time gesture detection method according to claim 1, characterized in that: The hierarchical feature extraction network includes multiple serial basic feature extraction networks, and the basic feature extraction network at each level is configured to perform feature extraction on an input image to obtain an original feature map of the corresponding level, wherein the size of the original feature map is halved relative to the size of the input image, and the number of channels of the original feature map is doubled relative to the number of channels of the input image.

3. The real-time gesture detection method according to claim 2, characterized in that: When extracting features from an input image, the basic feature extraction network includes: Performing a convolution structure channel halving operation on the input image through the basic convolution module, and performing feature extraction on the input image after the convolution structure channel halving through the separable convolution module to obtain a feature extraction result; Performing element-by-element addition on the input image after the convolution structure channel is halved and the feature extraction result to obtain an element-by-element addition result, and performing a confusion operation on the element-by-element addition result through a basic convolution module to obtain an element-by-element confusion result; The element-addition confusion result and the input image after the convolution structure channel is halved are connected by string to obtain a connection result, and the connection result is downsampled to obtain the original feature map.

4. The real-time gesture detection method according to claim 3, characterized in that: The basic convolution module includes a 1*1 convolution kernel, a BatchNorm normalization unit and a LeakyReLU activation function unit connected in sequence, the separable convolution module includes a first basic convolution module, a feature extraction module and a second basic convolution module connected in sequence, and the feature extraction module includes a 3*3 depth-separable convolution kernel, a BatchNorm normalization unit and a LeakyReLU activation function unit connected in sequence.

5. The real-time gesture detection method according to claim 2, characterized in that: The hierarchical feature extraction network includes five serial basic feature extraction networks.

6. The real-time gesture detection method according to claim 1, characterized in that: The predicted gesture position is represented based on a grid position code of a target frame, wherein the grid position code is configured to represent the encoding coordinates of the target frame in a feature grid, wherein the feature grid is obtained by dividing a fused feature map according to a set unit length.

7. The real-time gesture detection method according to claim 6, characterized in that: The predicted gesture position is determined based on the decoded coordinates and decoded size of the target frame on the fused feature map and the downsampling step size of the fused feature map.

8. The real-time gesture detection method according to claim 7, characterized in that: The decoded coordinates are determined based on the following formula: b x =σ(t x )+c x b y =σ(t y )+c y Among them, (b x , b y ) is the decoded coordinates of the center coordinates of the target box on the fusion feature map, (c x , c y ) is the coordinate of the upper left corner of the current feature grid, σ(t x ) and σ(t y ) is the offset of the prior frame from the upper left corner of the current feature grid, (t x , t y ) is the encoding coordinate of the center coordinate of the prior frame on the fusion feature map; The decoded size is determined based on the following formula: Among them, b h , and b w is the length and width of the decoded size of the target box, p h and p w is the length and width of the encoding size of the prior box, t h and t w Exponential coefficient obtained by training the gesture detection model.

9. The real-time gesture detection method according to claim 6, characterized in that: The gesture detection model is trained based on a joint loss function, which is determined based on whether the prior frame contains a target, the coordinate error between the prior frames, and the loss value of the predicted target and the matched prior frame.

10. The real-time gesture detection method according to claim 9, characterized in that: The joint loss function is determined based on the following formula: in: Among them, W is the width of the fused feature map, H is the length of the fused feature map, A is the number of prior boxes for each point on the fused feature map, maxiou is the maximum overlap ratio of each prior box and all real targets, thresh is the set overlap ratio screening threshold, and λ noobj is the weight of the negative sample loss function set, is the coordinate of the kth priori box at the point with width i and length j on the current fusion feature map, o is the target score corresponding to the priori box, t is the number of training times, and λ prior is the weight of the warmup loss function, represents the coordinates of the kth prior frame, r represents the preset coordinates, is the coordinate of the prior box, Indicates that this part only calculates the loss value of the box that matches a real target, λ coord is the loss function weight of the coordinate, truth r is the coordinate value of the labeled target in the training sample, λ obj is the loss function weight of whether to include the target, is the IOU score between the prior box and the labeled target, λ class is the loss function weight for category prediction, truth c is the predicted target category, is the category of the prior box.

11. A real-time gesture detection device, characterized in that: It includes an image acquisition module, a gesture recognition module and a gesture determination module, wherein: The image acquisition module is configured to acquire an image to be identified; The gesture recognition module is configured to input the image to be recognized into a trained gesture detection model so that the gesture detection model outputs a gesture recognition result based on the image to be recognized, and the gesture detection model is configured to obtain a plurality of original feature maps of different levels of the input image based on a separable convolution structure and a residual structure, fuse the plurality of the original feature maps to obtain a plurality of fused feature maps, perform gesture recognition based on the plurality of the fused feature maps and output a gesture recognition result; The gesture determination module is configured to determine a gesture type and a gesture position based on a gesture recognition result output by the gesture detection model; The gesture detection model includes a hierarchical feature extraction network, a feature fusion network and a separate detection head network, wherein: The hierarchical feature extraction network is configured to obtain original feature maps of multiple different levels of the input image based on a separable convolution structure and a residual structure; the feature fusion network is configured to perform a downsampling step size halving and a channel halving operation on the original feature map of the last layer output by the hierarchical feature extraction network to obtain a first intermediate feature map, and perform element-by-element addition of the first intermediate feature map and the penultimate original feature map output by the hierarchical feature extraction network to obtain a second fused feature map; perform a downsampling step size halving and a channel halving operation on the second fused feature map to obtain a second intermediate feature map, and The second intermediate feature map and the third-to-last original feature map output by the hierarchical feature extraction network are added element by element to obtain a third fused feature map; the second fused feature map is downsampled and doubled in step size to obtain a third intermediate feature map, and the third intermediate feature map and the last layer of the original feature map output by the hierarchical feature extraction network are added element by element to obtain a first fused feature map; the separate detection head network is configured to perform gesture recognition based on a plurality of the fused feature maps and output a gesture recognition result, wherein the gesture recognition result includes a predicted gesture category, a gesture confidence, and a predicted gesture position; When the separate detection head network performs gesture recognition based on the multiple fused feature maps and outputs the gesture recognition result, it includes: for each of the fused feature maps, separating the fused feature map through a basic convolution module to obtain a first separated feature map, a second separated feature map and a third separated feature map; using a 1x1 conv convolution kernel to determine the predicted gesture category from the first separated feature map, and using a softmax normalization module to normalize the predicted gesture category; using a 1x1 conv convolution kernel to determine the gesture confidence from the second separated feature map, and using a sigmoid normalization module to normalize the gesture confidence to between 0 and 1; using a 1x1 conv convolution kernel to determine the predicted gesture position from the third separated feature map; and outputting the gesture recognition result including the predicted gesture category, the gesture confidence and the predicted gesture position.

12. A real-time gesture detection device, characterized in that: include: memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the real-time gesture detection method as described in any one of claims 1-10.

13. A storage medium storing computer executable instructions, characterized in that: When the computer executable instructions are executed by a computer processor, they are used to perform the real-time gesture detection method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the real-time gesture detection method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Image processing method and device and processing equipment

    CN109740534A

  • Downhole pipeline abnormal target identification method and system based on densely connected Yolov3 network

    CN113298181A