A gesture recognition method based on a multi-core dynamic attention mechanism

By introducing a multi-core dynamic attention mechanism and YOLOv5 network into the gesture recognition method, combining RGB and depth images for feature fusion, the shortcomings of existing gesture recognition methods in lighting and complex gesture recognition are solved, and higher accuracy and robustness are achieved.

CN117152838BActive Publication Date: 2025-06-13HEBEI UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311098247.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-29
Publication Date
2025-06-13
Estimated Expiration
2043-08-29

AI Technical Summary

Technical Problem

Existing gesture recognition methods are easily affected by factors such as lighting and occlusion, and have poor recognition of complex gestures, and are insufficient in real-time and accuracy.

Method used

Using a gesture recognition method based on the multi-core dynamic attention mechanism, gesture features are extracted through a parallel dual-branch YOLOv5 network, combined with RGB images and depth images, and feature fusion and adjustment are used for multi-core dynamic attention mechanism to improve the accuracy and robustness of recognition.

Benefits of technology

It improves the accuracy and robustness of gesture recognition, can better adapt to lighting changes and environmental changes, enhances the recognition ability of complex gestures, and provides a more intuitive and convenient operation method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117152838B_ABST
    Figure CN117152838B_ABST
Patent Text Reader

Abstract

The present invention relates to a gesture recognition method based on a multi-core dynamic attention mechanism, comprising the following steps: S1, constructing a gesture recognition model; S2, acquiring the RGB image and depth image of a gesture; S3, extracting gesture features; S4, detecting the position of the gesture; S5, recognizing the gesture; S6, sending the recognition result of the gesture to a robot terminal in the format of a message message. The method of the present invention uses a multi-core dynamic attention mechanism for multi-modal feature extraction, which can better extract and fuse the gesture features of the RGB image and the gesture features of the depth image, and obtain a method with better gesture recognition effect. Applying this method, static gestures captured by a RealSense depth camera can be recognized in real time. At the same time, the present invention enables people to control the operation of a mobile operation robot in real time through gestures, having the technical effects of improving the gesture recognition effect and enhancing the interaction experience, and having a positive significance for promoting the development of fields such as human-computer interaction, virtual reality, and smart home.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a target detection method, specifically a gesture recognition method based on a multi-core dynamic attention mechanism. Background Art

[0002] Gesture recognition technology has a wide range of applications in the fields of human-computer interaction, virtual reality, smart home, etc. Through gesture recognition, users can use natural gesture actions to interact with devices to provide a more intuitive and convenient operation control method. Gesture recognition involves technical fields such as computer vision, pattern recognition, and machine learning. With the continuous increase in the application of mobile robots, gesture recognition, as a natural and intuitive interaction method, has been widely applied to the control field of mobile robots. Gesture recognition technology can convert human action instructions into control commands for robots by analyzing and understanding the shape and motion information of human gestures. Traditional gesture recognition methods are mainly based on image processing and machine learning algorithms, but limited by the diversity and complexity of gestures, these methods have certain deficiencies in real-time performance and accuracy.

[0003] Traditional monocular gesture recognition methods usually include the following steps: obtaining an RGB image containing a gesture, detecting and segmenting the hand in the RGB image, extracting gesture features in the image, and gesture recognition. Such gesture recognition methods have certain limitations. One is that since the RGB image can only provide two-dimensional image information and lacks the perception of hand depth information, it is easily affected by factors such as lighting and occlusion; the other is that the recognition effect for complex gestures is not high, such as hand rotation and subtle movements.

[0004] In multi-modal learning, traditional modal fusion often uses a static fusion method. However, the contribution degree of each modality may change with the change of the scene and content. Static fusion methods usually use fixed weights and thus cannot adapt to these changes, which will lead to a decline in the accuracy performance of the gesture recognition model, a reduction in robustness, and information loss. There may be complex interrelationships and dependencies between different modalities. Traditional static fusion methods may only capture a part of these relationships, which limits the representation ability of the gesture recognition model. In addition, multi-modal data is also affected by various factors such as lighting, perspective, and object size. Traditional fusion methods may not be able to adapt to these complex and variable influencing factors, resulting in performance fluctuations of the gesture recognition model in different environments and scenarios. Summary of the Invention

[0005] The purpose of the present invention is to provide a gesture recognition method based on a multi-core dynamic attention mechanism to solve the problems that the existing recognition methods are easily affected and have poor recognition effects on complex gestures.

[0006] The object of the present invention is achieved as follows:

[0007] A gesture recognition method based on a multi-core dynamic attention mechanism, comprising the following steps:

[0008] S1. Construct a gesture recognition model: Use the parallel dual-branch YOLOv5 network as the gesture recognition model; the YOLOv5 network includes a backbone network for extracting gesture features and a detection head for predicting the position and category of the gesture.

[0009] S2. Obtain the RGB image and depth image of the gesture: Use a depth camera to obtain the RGB image and depth image of the operation control gesture made by the controller.

[0010] S3. Extract gesture features: Utilize the multi-core dynamic attention mechanism, and input the RGB image and depth image of the gesture into the gesture recognition model through forward propagation, respectively extract the image features therein, and then fuse the image features extracted from the RGB image and the image features extracted from the depth image; the extracted image features contain the semantic information of the gesture.

[0011] S4. Detect the position of the gesture: Predict the gesture in the extracted image features through the detection head in the gesture recognition model to obtain the position of the gesture box, the category of the gesture box, and the confidence of the gesture box.

[0012] S5. Recognize the gesture: According to the obtained position and confidence of the gesture box, screen the gesture box, and use the non-maximum suppression algorithm (NMS) to eliminate the overlapping gesture boxes. The recognition result including the position of the gesture, the category of the gesture, and the confidence of the gesture is included in the finally output gesture box.

[0013] S6. Send the recognition result of the gesture to the robot terminal in the format of a message message.

[0014] Further, the working mode of the multi-core dynamic attention mechanism of the gesture recognition model in step S3 is as follows: The multi-core dynamic attention convolution dynamically adjusts the weight and bias parameters of the convolution by using the attention weight to perform weighted averaging on the convolution parameters, calculates the attention weights of different modalities through the SE module, and then uses these weights to perform weighted summation on the features of different modalities to obtain the fused features.

[0015] In step 3, the multi-core dynamic attention mechanism is used to extract the image features of the gesture, which can accurately obtain the RGB image and depth image containing the gesture, and automatically adjust the weights during the process of gesture recognition using the multi-core dynamic attention mechanism, so that the feature information of each gesture can be better mined and utilized. By extracting and fusing the features in the RGB image and depth image, the gesture can be described more comprehensively and accurately, thus realizing accurate gesture recognition.

[0016] Furthermore, the specific working mode of the multi-core dynamic attention mechanism of the gesture recognition model is as follows:

[0017] S3-1-1 Generate K different weights, each weight corresponding to a convolutional kernel, and the convolutional calculation of each convolutional kernel on the input is expressed as:

[0018] output[k] = Conv2d(x, W[k]) + b[k]

[0019] where x is the input feature map, W[k] is the k-th convolutional kernel, b[k] is the k-th bias, and output[k] is the output feature map obtained by adding the bias b[k] after the k-th convolutional kernel W[k] performs a convolutional operation on the input feature map x.

[0020] S3-1-2 Multiply the weight of each convolutional kernel by the input feature map.

[0021] S3-1-3 All the outputs are accumulated together to form the final output as:

[0022]

[0023] where π[k] is the k-th weight obtained through the attention mechanism.

[0024] Furthermore, the working mode of the feature fusion of the gesture recognition model in step S3 is as follows:

[0025] S3-2-1 Obtain the RGB image of the gesture through the gesture recognition model as: RGB in ∈R C×H×W ; obtain the depth image of the gesture through the gesture recognition model as: Depth in ∈R C×H×W , and after the RGB image and the depth image each pass through a gated convolutional layer, we get:

[0026] RGB′ in = σ(Conv 1x1 (RGB in ))·RGB in

[0027] Depth′ in = σ(Conv1x1 (Depth in ))·Depth in

[0028] Among them, σ represents the activation function, and Conv 1x1 (x) represents a 1×1 convolution operation on the input x.

[0029] After the gating operation, adaptive average pooling (AAP) is performed to generate a cross-modal spatial descriptor: X = (X 1 ,..., X k ,..., X 2C ):

[0030] X = AAP(RGB′ in ||Depth′ in )

[0031] Among them, || represents concatenating RGB in and Depth in .

[0032] The cross-modal attention vector of the S3-2-2 depth input is learned by the following formula:

[0033] W rgb = σ(Conv 1x1 (ReLU(DM(X))))

[0034] W depth = σ(Conv 1x1 (ReLU(DM(X))))

[0035] Among them, W rgb is the weight representing the RGB gesture feature, W depth is the weight representing the Depth gesture feature, DM(x) represents the proposed multi-core dynamic attention module, and Conv 1x1 (x) represents a 1×1 convolution operation.

[0036] The gesture feature maps RGB′ in and Depth′ in after the gating operation are respectively multiplied by the channel weights W rgb and W depth to adjust or enhance the gesture features:

[0037] RGBf = Wrgb·RGB′ in ;

[0038] Depth f = W depth ·Depth′ in .

[0039] Convolve RGBf and Depth in S3-2-4 f to obtain the RGB-D feature map RGBD f :

[0040] RGBD f = Conv 1x1 ([RGB f ; Depth f )

[0041] In S3-2-5, for the RGB-D feature map RGBD f , obtain the attention weights a of RGB and depth features through multi-core dynamic attention convolution rgb and a depth :

[0042] a rgb = DM rgb (RGBD f )

[0043] a depth = DM depth (RGBD f )

[0044] where DM rgb () represents performing multi-core dynamic attention convolution on RGB features; DM depth () represents performing multi-core dynamic attention convolution on depth features; a rgb is the weight assigned to each position in the RGB gesture feature map, and a depth is the weight assigned to each position in the depth gesture feature map

[0045] In S3-2-6, perform softmax normalization on the spatial attention weights of each modality to obtain the final spatial attention weights:

[0046]

[0047]

[0048] where and

[0049] In S3-2-7, the gesture feature maps RGB' in and Depth' in after gated operation are respectively multiplied by the corresponding attention weights, and then added together to obtain the fused gesture feature RGBD out :

[0050] RGBD out = A rgb·RGB′ in +A depth ·Depth′ in 。

[0051] The gesture recognition method of the present invention can adaptively adjust the weights of the RGB modality and the depth modality by using a multi-core dynamic attention mechanism, so as to better mine and utilize the feature information of each gesture. Compared with the traditional static fusion method, the method of dynamically aggregating multiple convolutional kernels can more flexibly capture and model the complex interactions between the RGB modality and the depth modality, improving the accuracy and robustness of the gesture features after multi-modal fusion. By adaptively adjusting the convolutional kernels and using the attention mechanism, this method can automatically perceive and adapt to changes in the environment and scene such as illumination changes and object size differences, so it has strong robustness and is less affected by illumination conditions. To facilitate deployment on mobile robots, the present invention selects the lightweight network YOLOv5 as the backbone network for gesture recognition. Such a choice can reduce the consumption of computing resources and make the method run more efficiently on mobile robots. Therefore, the gesture recognition method of the present invention not only has low computing resource requirements, but also has stronger adaptability and is very suitable for application on mobile robots.

[0052] The present invention well solves the problem of how to accurately extract and fuse the gesture features in the RGB image and the depth image by using a multi-core dynamic attention mechanism. By introducing the multi-core dynamic attention mechanism, the weights of the RGB modality and the depth modality can be flexibly adjusted, and the feature information of each gesture can be more effectively mined and utilized, thereby improving the accuracy and robustness of gesture recognition. The gesture recognition method of the present invention can adaptively adjust the convolutional kernels and introduce the attention mechanism to adapt to changes in the environment and scene such as illumination changes and object size differences, so it has strong robustness and is less affected by illumination conditions.

[0053] The robot gesture interaction system constructed according to the present invention provides a convenient and reliable way to control the robot system. By establishing a robot gesture interaction database, constructing the mapping relationship between the databases, establishing a static gesture data set and annotations, and receiving the message of the gesture recognition result, intelligent gesture interaction can be realized. Such a way to control the robot system enables users to interact with the robot naturally and intuitively through gestures, providing a convenient and reliable control method and further enhancing the user experience.

[0054] The present invention uses the method of simultaneously acquiring RGB images and depth images containing gestures for gesture recognition. At the same time, feature extraction is a key link in gesture recognition, and the effect of feature extraction affects the result of gesture recognition. The present invention uses the method of simultaneously acquiring RGB images and depth images containing gestures for gesture recognition, which means that it is necessary to extract the gesture features of RGB images and depth images simultaneously. The present invention uses multimodal feature extraction to extract and fuse the features of the two types of images.

[0055] The present invention uses the method of a multi-core dynamic attention mechanism for multimodal feature extraction, which can better extract and fuse the gesture features of RGB images and depth images, and obtain a method with better gesture recognition effect. Applying this method, static gestures captured by a RealSense depth camera can be recognized in real time. At the same time, the present invention enables people to control the operation of a mobile operation robot in real time through gestures.

[0056] The present invention adopts a multi-core dynamic attention mechanism. By simultaneously acquiring RGB images and depth images containing gestures, multimodal feature extraction and fusion are carried out, so that gesture recognition has better accuracy and real-time performance. It has the technical effects of improving the gesture recognition effect and enhancing the interaction experience, as well as the advantages of providing a more intuitive and convenient operation method, and has a positive significance for promoting the development of fields such as human-computer interaction, virtual reality, and smart home. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is the overall architecture diagram of the gesture recognition method of the present invention.

[0058] Figure 2 is the mapping relationship diagram of static gestures with motion patterns, tool / operation patterns, and user intentions.

[0059] Figure 3 is the structural diagram of the gesture recognition model.

[0060] Figure 4 are the structural diagrams of each module; among them, (a) is the Focus module, (b) is the CBS module, (c) is the C3 module, and (d) is the structural diagram of the SPP module.

[0061] Figure 5 is the structural diagram of the feature fusion module.

[0062] Figure 6 is the structural diagram of the multi-core dynamic attention mechanism.

[0063] Figure 7 is the flowchart of static gesture recognition. DETAILED DESCRIPTION OF THE INVENTION

[0064] The application object of the present invention is a mobile operation robot, especially a robot for performing explosive disposal tasks. During the handling of crisis events involving explosives in public places such as subways, shopping malls, and airports, in order to avoid triggering an explosion, it is usually necessary to implement wireless signal shielding measures at the scene. In this scenario, compared with traditional wired or wireless handle control schemes, a better way is to use gesture remote control of the robot. This is because gesture control of the robot has a more natural interaction method and a more convenient usage method, and can better adapt to the working environment where network signals are shielded.

[0065] As Figure 1 shown, the present invention mainly involves the following three aspects: 1. The mapping relationship between the operation instructions of the mobile robot and static gestures; 2. A static gesture recognition method based on a multi-core attention mechanism; 3. The human-computer interaction between the gesture executor and the mobile operation robot. When the gesture executor makes a pre-set static gesture, the meaning of the gesture can be recognized through the gesture recognition method of the present invention and sent to the robot side in the form of a message. After receiving the message, the robot side compares the gesture information in the message and searches for the operation instruction corresponding to the gesture in the mapping relationship, and finally executes the corresponding operation.

[0066] The basis for the implementation of the present invention is the robot gesture interaction system. Building this robot gesture interaction system mainly includes the following aspects:

[0067] I. Establish a robot gesture interaction database.

[0068] The robot gesture interaction database includes four parts: a motion mode database, a tool / operation mode database, a user intention database, and a static gesture database. The purpose of establishing these databases is to record information such as robot motion, operation mode, user intention, and static gestures.

[0069] The motion mode database records various motion modes of the robot, including moving forward left, moving forward, moving forward right, moving backward left, moving backward, moving backward right, the middle arm of the robotic arm moving up, the middle arm of the robotic arm moving down, the front arm of the robotic arm moving up, the front arm of the robotic arm moving down, the chassis of the robotic arm moving left, and the chassis of the robotic arm moving right.

[0070] The tool / operation mode database records the operation modes of the end effector of the robotic arm. For a wheeled mobile operation robot, the end of the robotic arm is a clamp, so the database also includes the operation modes of the opening and closing of the clamp.

[0071] The user intention database includes three ways of expressing user intentions: "yes", "no", and "uncertain, please repeat". These intention expression methods are used for information exchange and understanding between the user and the mobile operation robot.

[0072] The static gesture database records 17 different static gestures through which the user can interact with the robot. Figure 2 The types of gestures used in the specific operations are given.

[0073] II. Construct the mapping relationships between the databases.

[0074] The mapping relationships between the databases include the mapping relationship between the static gesture and the user intention, the mapping relationship between the static gesture and the tool / operation mode, and the mapping relationship between the static gesture and the robot motion mode.

[0075] In the robot gesture interaction system, there are mapping relationships between the databases, and the purpose is to realize the conversion and matching between the static gesture and the robot motion mode, the static gesture and the tool / operation mode, and the static gesture and the user intention.

[0076] First, establish the mapping relationship between the static gesture and the user intention. Through the recognition of the static gesture, the robot gesture interaction system can judge whether the user's intention is "yes", "no", or "uncertain, please repeat". Second, establish the mapping relationship between the static gesture and the tool / operation mode. Through the recognition of the static gesture, the robot gesture interaction system can determine the specific operation that the user hopes to perform, such as opening the clip or closing the clip, etc. Finally, establish the mapping relationship between the static gesture and the robot motion mode. Through the recognition of the static gesture, the robot gesture interaction system can determine the motion mode that the robot should adopt, such as moving forward, moving backward, or rotating, etc. The establishment of the above mapping relationships enables the robot gesture interaction system to accurately perform corresponding operations and interactions according to the user's gesture commands. Figure 2 The legend of all the mapping relationships is given.

[0077] III. Establish a static gesture dataset.

[0078] Use the functions in the pyrealsense2 package to call the camera to capture various preset static gestures, and at the same time collect the RGB images and depth images of these static gestures to generate a static gesture dataset.

[0079] To train the gesture recognition model, use the Intel RealSense D455 depth camera to capture the static gesture dataset. Through the calibration method provided by the RealSense SDK, calibrate the internal and external parameters of the camera to ensure the accuracy of the data.

[0080] Gesture images are collected at distances of 0.4 m, 0.8 m, and 1.2 m, with no less than 300, 300, and 200 groups respectively. Each group of images includes RGB images and depth images with a resolution of 640×480 pixels. To ensure data diversity and robustness, each gesture is recorded by multiple different participants in a laboratory environment. At the same time, gestures of the left hand and the right hand are regarded as the same gesture category. Finally, 6232 groups of gesture images are selected from the static gesture dataset as the training set, 5284 groups of gesture images are selected as the validation set, and 5074 groups of gesture images are selected as the test set. These datasets form the static gesture data, which will be used to train and evaluate the performance of the gesture recognition model.

[0081] IV. Annotation of the static gesture dataset.

[0082] Annotate the RGB images and depth images of each gesture image in the static gesture dataset, including the position where the gesture appears and the category information of the gesture, so as to create a computer file containing the annotation information.

[0083] For the static gesture dataset, annotation is required to train the gesture recognition model. The annotation process includes annotating the position where the gesture appears and the category of the gesture. When annotating the position where the gesture appears, use an annotation tool (such as labelme) to select the area where the gesture appears in the image and generate the corresponding annotation file. At the same time, use a rectangular box to mark the position of the gesture for subsequent processing. For the annotation of the gesture category, corresponding labels are created for each gesture category and associated with the position where the gesture appears to obtain a dataset file containing the annotation information.

[0084] V. Receive the message of the gesture recognition result and perform corresponding operations according to the mapping relationship.

[0085] On the terminal of the wheeled mobile operation robot, after the robot gesture interaction system receives the recognition result sent by the gesture recognition model, it will compare the gesture information in the recognition result and search for the corresponding operation instruction of the gesture in the mapping relationship. According to the mapping relationship, the robot gesture interaction system will execute the operations corresponding to the gesture, such as the movement of the robot, the switching of tools / operation modes, and the understanding of user intentions.

[0086] The robot gesture interaction system constructed through the above steps can realize the natural and intuitive operation control of the user on the mobile operation robot.

[0087] As Figure 7 shown, the gesture recognition method based on the multi-core dynamic attention mechanism of the present invention includes the following steps:

[0088] S1. Construct a gesture recognition model.

[0089] The YOLOv5 network with a parallel dual-branch structure is used as the gesture recognition model. As Figure 3 shown, the YOLOv5 network includes a backbone network and a detection head. The backbone network is mainly responsible for extracting gesture features, while the detection head is used to predict the position and category of the gesture. The key modules of the YOLOv5 network include the Focus module, CBS module, C3 module, and SPP (Spatial Pyramid Pooling) module, and the corresponding structures are as Figure 4 shown.

[0090] Among them, the Focus module amplifies the number of channels to 4 times the original by slicing the input image and obtains the downsampled feature map through a single convolution operation. This downsampling operation not only reduces the number of model parameters but also improves the inference speed. The CBS module is the basic component of the YOLOv5 network, which combines two-dimensional convolution, batch normalization, and the SiLU activation function. The C3 module is a key part of constructing the backbone network of the YOLOv5 network. It consists of multiple CBS modules and forms a residual connection. The SPP module performs max-pooling operations on the feature map using kernels of different sizes to solve the problems of inconsistent input image sizes and object size variations.

[0091] The backbone network is used to extract image features at different scales, including low-level features (such as texture, edges, etc.) extracted from the shallow layer and high-level semantic features extracted from the deep layer. As Figure 5 shown, the present invention improves the backbone network and introduces a feature fusion module. Its structure is in the backbone part of YOLOv5, and a feature fusion module is added after each C3 module, so that RGB gesture features and depth gesture features can be better extracted and fused.

[0092] The feature fusion module is used to combine RGB gesture features and depth gesture features to obtain a richer and more accurate feature representation. After each C3 module, the RGB and depth gesture features are concatenated and further fused through a series of convolution operations. This enables the gesture recognition model to better utilize the complementarity between RGB and depth gesture information and improve the detection and recognition performance of gestures.

[0093] S2. Obtain the RGB image and depth image of the gesture.

[0094] To obtain data containing gesture-related information, the present invention uses a RealSense depth camera to simultaneously obtain the RGB image and depth image of the gesture. In this way, rich gesture visual information and depth information can be obtained to assist the gesture recognition process.

[0095] S3. Extract gesture features.

[0096] Using a multi-core dynamic attention mechanism, the RGB image and depth gesture image of the gesture are input into the gesture recognition model through forward propagation to extract the image features therein respectively, and then the image features extracted from the RGB image and the image features extracted from the depth image are fused to better describe and represent the gesture, and the semantic information of the gesture is included in these extracted image features.

[0097] S3-1 Construction of the multi-core dynamic attention mechanism.

[0098] The multi-core dynamic attention convolution uses attention weights to perform weighted averaging on the convolution parameters, dynamically adjusts the weights and bias parameters of the convolution, calculates the attention weights of different modalities through the SE module, and then uses these weights to perform weighted summation on the features of different modalities to obtain the fused features.

[0099] Figure 6 In it, the part in the dotted box is the improved SE attention mechanism for obtaining the weight π of each convolution kernel. Its operation method is: First, multiply the K convolution kernels element-wise with the corresponding weights to obtain the weighted convolution kernels; then, connect the weighted convolution kernels, and then process the connected result through batch normalization and the ReLU activation function; finally, output the improved feature map. Its specific operation method is:

[0100] S3-1-1 Generate K different weights, each weight corresponds to a convolution kernel, and the convolution calculation of each convolution kernel on the input can be expressed as:

[0101] output[k] = Conv2d(x, W[k]) + b[k]

[0102] Where, x is the input feature map, W[k] is the k-th convolution kernel, b[k] is the k-th bias, and output[k] is the output feature map obtained by adding the bias b[k] after the k-th convolution kernel W[k] performs convolution operation on the input feature map x.

[0103] S3-1-2 Multiply the weight of each convolution kernel by the input feature map.

[0104] S3-1-3 All the outputs are accumulated together to form the final output:

[0105]

[0106] Where, α[k] is the k-th weight obtained through the attention mechanism. Compared with the standard convolutional layer, this dynamic convolution enables the gesture recognition model to adaptively adjust its parameters according to the characteristics of each input sample, thereby improving the performance of the gesture recognition model.

[0107] Construction of the S3-2 Feature Fusion Module.

[0108] As Figure 6 shown, the fusion module first performs spatial dimension fusion and refinement on RGB and depth features through the feature fusion stage, enabling the gesture recognition model to capture and utilize the long-range spatial dependencies between these two modalities. Through the fusion and refinement process, the gesture recognition model can better understand the distribution and morphological changes of gestures in the image. Next, through the attention-guided feature fusion stage, feature fusion is performed in the channel dimension. This step enables the gesture recognition model to capture and utilize the cross-channel context dependencies between the RGB and depth modalities. The feature fusion module will extract important channel information through the attention mechanism, thereby enhancing the gesture recognition model's perception ability of the correlation between different modalities. Its working method is as follows:

[0109] S3-2-1 Obtain the RGB image of the gesture through the gesture recognition model as: RGB in ∈R C×H×W , and obtain the depth image of the gesture as: Depth in ∈R C×H×W ; After the RGB image and the depth image each pass through a separate gated convolutional layer, we get:

[0110] RGB′ in =σ(Conv 1x1 (RGB in ))·RGB in

[0111] Depth′ in =σ(Conv 1x1 (Depth in ))·Depth in

[0112] where σ represents the activation function, and Conv 1x1 (x) represents a 1×1 convolution operation on the input x.

[0113] After the gating operation, adaptive average pooling (AAP) is performed to generate a cross-modal spatial descriptor: X = (X 1 ,..., X k ,..., X 2C ):

[0114] X = AAP(RGB′ in ||Depth′ in )

[0115] where || represents concatenating RGB in and Depth in .

[0116] The cross-modal attention vector of the S3-2-2 depth input is learned by the following formula:

[0117] W rgb = σ(Conv 1x1 (ReLU(DM(X))))

[0118] W depth = σ(Conv 1x1 (ReLU(DM(X))))

[0119] where W rgb represents the weight of the RGB gesture feature, W depth represents the weight of the Depth gesture feature, DM(x) represents the proposed multi-core dynamic attention module, and Conv 1x1 (x) represents a 1×1 convolution operation.

[0120] The gesture feature maps RGB′ in and Depth′ in after the gating operation are respectively multiplied by the channel weights W rgb and W depth to adjust or enhance the gesture features:

[0121] RGBf = Wrgb·RGB′ in

[0122] Depth f = W depth ·Depth′ in

[0123] This can ensure that important features can be obtained from both inputs and can be adjusted according to the needs of the gesture recognition model.

[0124] By multiplying the gesture feature maps (RGB and depth maps) element-wise with the corresponding channel weights, weighted processing of different channels can be achieved. The purpose of this weighted operation is to highlight or weaken the information of specific channels, thereby adjusting or enhancing the gesture features. The purpose of multiplying the weights is to introduce the correlation or weight relationship between channels in the gesture feature map. By appropriately selecting and adjusting the weight values, the representation of the gesture feature map can be changed, making the features of some channels more obvious while the features of other channels are weakened or ignored. Multiplying the channel weights can adjust the expression of the gesture features according to the needs, so that the influence degree of the information of different channels in the feature representation is adjusted, thereby performing weighted processing on the gesture features. This operation helps to extract more discriminative and important features, providing more accurate and effective input for subsequent gesture recognition or other tasks.

[0125] Convolve RGBf and Depth in S3-2-4 f to fuse them and obtain the RGB-D feature map RGBD f :

[0126] RGBD f = Conv 1x1 ([RGB f ; Depth f )

[0127] Thus, a unified gesture feature descriptor RGBD is generated using the gesture RGB image and depth image f .

[0128] For the RGB-D feature map RGBD in S3-2-5 f , obtain the attention weights a of the RGB and depth features through multi-core dynamic attention convolution rgb and a depth :

[0129] a rgb = DM rgb (RGBD f )

[0130] a depth = DM depth (RGBD f )

[0131] where DM rgb () represents performing multi-core dynamic attention convolution on RGB features; DM depth () represents performing multi-core dynamic attention convolution on depth features; a rgb is the weight assigned to each position in the RGB gesture feature map, and a depth is the weight assigned to each position in the depth gesture feature map

[0132] Normalize the spatial attention weights of each modality by softmax to obtain the final spatial attention weights

[0133]

[0134]

[0135] where and

[0136] The gesture feature maps RGB' in and Depth' inMultiply by the corresponding attention weights respectively and then sum them up to obtain the fused gesture feature RGBD out :

[0137] RGBD out = A rgb ·RGB′ in + A depth ·Depth′ in

[0138] Input the fused gesture feature RGBD out into the subsequent gesture recognition model.

[0139] Through the design and improvement of the above fusion module, the present invention can effectively fuse RGB and depth gesture features, and make full use of the spatial and channel correlations between them. In this way, the network can more comprehensively understand the gesture data and improve the performance and accuracy of the gesture recognition task. The introduction of the feature fusion module provides the network with a more abundant information expression ability and further optimizes the performance of the gesture recognition model. Therefore, while improving gesture recognition, this module also provides an effective solution for the field of multi-modal data fusion and analysis.

[0140] S4. Detect the position of the gesture.

[0141] Through the detection head in the gesture recognition model, accurate prediction is made on the extracted gesture features, so as to obtain the accurate position of the gesture box and the corresponding gesture category. By carefully analyzing and decoding the detection head, the robot gesture interaction system can obtain the accurate position of the gesture in the image and identify the specific category to which the gesture belongs.

[0142] S5. Recognize the gesture.

[0143] According to the predicted gesture box position and the corresponding confidence score, further processing and screening are carried out:

[0144] S5-1. The system will sort all the detected gesture boxes according to the confidence score to retain the boxes with higher confidence.

[0145] S5-2. The system uses non-maximum suppression (NMS) technology to effectively eliminate the overlapping boxes and only retain the most representative and accurate gesture boxes.

[0146] The output gesture box will contain detailed information such as the accurate position of the gesture, the corresponding category label, and the confidence.

[0147] S6. Send the recognition result of the gesture to the robot terminal in the format of a message message.

[0148] For the convenience of subsequent applications and system integration, the system will send the results of gesture recognition in the form of a message. In this way, the system can integrate the recognized gesture results and interactively transfer them to other systems or applications.

Claims

1. A gesture recognition method based on a multi-core dynamic attention mechanism, characterized in that, it includes the following steps: S1. Build a gesture recognition model: Use the parallel dual-branch YOLOv5 network as the gesture recognition model; the YOLOv5 network includes a backbone network for extracting gesture features and a detection head for predicting the position and category of gestures; S2. Obtain the RGB image and depth image of the gesture: Use a depth camera to obtain the RGB image and depth image of the operation control gesture made by the controller; S3. Extract gesture features: Utilize the multi-core dynamic attention mechanism, and through forward propagation, input the RGB image and depth image of the gesture into the gesture recognition model, respectively extract the image features therein, and then fuse the image features extracted from the RGB image and the image features extracted from the depth image; the extracted image features contain the semantic information of the gesture; S4. Detect the position of the gesture: Predict the gesture in the extracted image features through the detection head in the gesture recognition model to obtain the position of the gesture box, the category of the gesture box, and the confidence of the gesture box; S5. Recognize the gesture: According to the obtained position and confidence of the gesture box, screen the gesture box, and use the non-maximum suppression algorithm to eliminate the overlapping gesture boxes. The final output gesture box contains the recognition result including the position of the gesture, the category of the gesture, and the confidence of the gesture; S6. Send the recognition result of the gesture to the robot terminal in the format of a message message; The working mode of the feature fusion of the gesture recognition model in step S3 includes: The RGB image of the gesture obtained by the gesture recognition model in S3-2-1 is: RGB in ∈R C×H×W ; The depth image of the gesture obtained by the gesture recognition model is: Depth in ∈R C×H×W , After the RGB image and the depth image each pass through a gated convolutional layer, we get: RGB′ in = σ(Conv 1×1 (RGB in ))·RGB in Depth′ in = σ(Conv 1×1 (Depth in ))·Depth in Among them, σ represents the activation function, and Conv 1×1 (x) represents performing a 1×1 convolution operation on the input x; After the gating operation, perform adaptive average pooling to generate a cross-modal spatial descriptor: X = AAP(RGB′ in || Depth′ in ) Among them, || indicates concatenating RGB in and Depth in for splicing; The cross-modal attention vector of the S3-2-2 depth input is learned by the following formula: W rgb = σ(Conv 1×1 (ReLU(DM(X)))) W depth = σ(Conv 1×1 (ReLU(DM(X)))) Among them, W rgb represents the weight of RGB gesture features, and W depth represents the weight of Depth gesture features. DM(x) represents the proposed multi-core dynamic attention module, and Conv 1×1 (x) represents a 1×1 convolution operation; The gesture feature maps RGB' in and Depth' in are multiplied by their respective channel weights W rgb and W depth to adjust or enhance the gesture features: RGBf = Wrgb·RGB′ in ; Depth f = W depth ·Depth' in ; Use convolution to fuse RGBf and Depth f to obtain the RGB-D feature map RGBD f : RGBD f = Conv 1×1 ([RGB f ; Depth f ) S3-2-5 For the RGB-D feature map RGBD f , the attention weights a rgb and a depth : a rgb = DM rgb (RGBD f ) a depth = DM depth (RGBD f ) Among them, DM rgb () represents performing multi-core dynamic attention convolution on RGB features; DM depth () represents performing multi-core dynamic attention convolution on depth features; a rgb is the weight assigned to each position in the RGB gesture feature map, a depth is the weight assigned to each position in the depth gesture feature map.

2. The gesture recognition method according to claim 1, characterized in that, The working mode of the multi-core dynamic attention mechanism of the gesture recognition model in step S3 is: The multi-core dynamic attention convolution dynamically adjusts the weight and bias parameters of the convolution by using the attention weight to perform weighted averaging on the convolution parameters, calculates the attention weights of different modalities through the SE module, and then uses these weights to perform weighted summation on the features of different modalities to obtain the fused features.

3. The gesture recognition method according to claim 2, characterized in that, The specific working mode of the multi-core dynamic attention mechanism of the gesture recognition model is: S3-1-1 Generate K different weights, each weight corresponds to a convolution kernel, and the convolution calculation of each convolution kernel on the input is expressed as: output[k] = Conv2d(x, W[k]) + b[k] where x is the input feature map, W[k] is the kth convolution kernel, b[k] is the kth bias, and output[k] is the output feature map obtained by adding the bias b[k] after the kth convolution kernel W[k] performs convolution operation on the input feature map x; S3-1-2 Multiply the weight of each convolution kernel by the input feature map; S3-1-3 All the outputs are accumulated together to form the final output as: where π[k] is the kth weight obtained through the attention mechanism.

4. The gesture recognition method according to claim 1, characterized in that, the working mode of feature fusion of the gesture recognition model in step S3 further includes: S3-2-6 performing softmax normalization on the spatial attention weights of each modality to obtain the final spatial attention weights: Among them, The gesture feature maps RGB′ in and Depth′ in after gating operation are respectively multiplied by the corresponding attention weights and then added together to obtain the fused gesture feature RGBD out : RGBD out = A rgb · RGB′ in + A depth · Depth′ in 。

Citation Information

Patent Citations

  • Gesture recognition method based on color information and depth information fusion

    CN110502981A

  • MPE-YOLOv5-based gesture recognition method

    CN116052266A