Gesture recognition method, gesture control method and device

By acquiring gesture image data for fingertip detection and feature extraction, combined with the decision-making layer to fuse gesture recognition feature vectors, the problem of low gesture recognition accuracy is solved, and higher recognition accuracy and reliability are achieved.

CN120375413APending Publication Date: 2025-07-25ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510238306.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing gesture recognition methods are low in accuracy, which affects the efficiency of gesture control.

Method used

By obtaining the object gesture image data, fingertip detection is performed, the fingertip recognition feature vector is obtained, and feature extraction is performed based on the gesture classification model, combining the fingertip recognition feature vector and gesture recognition feature vector at the decision-making level to perform analysis and processing to improve the recognition accuracy.

Benefits of technology

It improves the accuracy and reliability of gesture recognition, eliminates ambiguity, enhances the ability to understand gesture structure, and avoids misidentification behavior during algorithm generalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375413A_ABST
    Figure CN120375413A_ABST
Patent Text Reader

Abstract

The invention discloses a gesture recognition method and device and a gesture control method and device. The gesture recognition method comprises the following steps: acquiring object gesture image data; performing fingertip detection on the gesture image data to obtain a fingertip recognition feature vector; performing feature extraction on the gesture image data based on a gesture classification model to obtain a gesture recognition feature vector; and fusing the fingertip recognition feature vector and the gesture recognition feature vector in a decision-making layer, and performing analysis processing to obtain a gesture recognition result. The gesture recognition accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of device control, and particularly to a gesture recognition method, a gesture control method, and a device thereof. Background Art

[0002] Gesture Control technology realizes interactive operations on devices or systems by recognizing and interpreting users' gestures, and has the advantages of intuitiveness and naturalness. It is widely used in multiple fields, greatly improving the user experience and interaction efficiency.

[0003] In the long-term research and development process, the inventors of this application found that the gesture recognition accuracy in current gesture control methods is relatively low, thus affecting the gesture control efficiency. Summary of the Invention

[0004] This application provides a gesture recognition method and a device thereof, which can improve the gesture recognition accuracy.

[0005] To achieve the above object, this application provides a gesture recognition method, which includes:

[0006] Obtain object gesture image data;

[0007] Perform fingertip detection on the gesture image data to obtain a fingertip recognition feature vector;

[0008] Extract features from the gesture image data based on a gesture classification model to obtain a gesture recognition feature vector;

[0009] Analyze and process the fingertip recognition feature vector and the gesture recognition feature vector after fusing them at the decision-making layer to obtain a gesture recognition result.

[0010] To achieve the above object, this application provides a gesture control method, which includes:

[0011] Determine the gesture category of an object based on the above gesture recognition method;

[0012] Control the controlled device based on the gesture category.

[0013] To achieve the above object, this application further provides an electronic device, which includes a processor; the processor is used to execute instructions to implement the steps of the above method.

[0014] To achieve the above object, this application further provides a computer-readable storage medium, which is used to store instructions / program data, and the instructions / program data can be executed to implement the above method.

[0015] The gesture recognition method of this application obtains the gesture image data of an object; performs fingertip detection on the gesture image data to obtain a fingertip recognition feature vector; extracts features from the gesture image data based on a gesture classification model to obtain a gesture recognition feature vector; fuses the fingertip recognition feature vector and the gesture recognition feature vector at the decision-making level to obtain a gesture recognition result. Such a feature fusion method that combines a gesture recognition algorithm and a fingertip detection algorithm is used to avoid misrecognition behaviors that may occur during the generalization process of the algorithm. By introducing the fingertip vector feature, the key points of the hand are accurately located to eliminate ambiguity, improve the reliability of overall gesture recognition, and further enhance the algorithm's ability to understand the gesture structure, thus improving the accuracy of gesture recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of this application, form a part of this application, and the schematic embodiments and descriptions thereof are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0017] Figure 1 is a flowchart of an embodiment of the gesture recognition method of this application;

[0018] Figure 2 is a flowchart of an embodiment of the gesture recognition method of this application;

[0019] Figure 3 is a flowchart of the fingertip detection method in the gesture recognition method of this application;

[0020] Figure 4 is a flowchart of another embodiment of the gesture recognition method of this application;

[0021] Figure 5 is a schematic structural diagram of an embodiment of an electronic device of this application;

[0022] Figure 6 is a schematic structural diagram of an embodiment of a computer-readable storage medium of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application. Additionally, unless otherwise specified (e.g., "or alternatively" or "or in an alternative"), the term "or" as used herein refers to a non-exclusive "or" (i.e., "and / or"). Moreover, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments.

[0024] The present application proposes a gesture recognition method. This gesture recognition method acquires object gesture image data; performs fingertip detection on the gesture image data to obtain a fingertip recognition feature vector; extracts features from the gesture image data based on a gesture classification model to obtain a gesture recognition feature vector; and fuses the fingertip recognition feature vector and the gesture recognition feature vector at the decision layer to obtain a gesture recognition result. Such a feature fusion method that combines a gesture recognition algorithm and a fingertip detection algorithm is used to avoid misrecognition behaviors that may occur during the generalization process of the algorithm. By introducing the fingertip vector feature, the key points of the hand are accurately located to eliminate ambiguity, improve the reliability of overall gesture recognition, and further enhance the algorithm's understanding ability of the gesture structure, thereby improving the accuracy of gesture recognition.

[0025] Specifically, as Figure 1 shown, a gesture recognition method according to an implementation manner proposed by the present application specifically includes the following steps. It should be noted that the following step numbers are only used for simplified description and are not intended to limit the execution order of the steps. The steps of this implementation manner can be arbitrarily changed in the execution order without violating the technical idea of the present application.

[0026] S101: Acquire object gesture image data.

[0027] Object gesture image data can be acquired first to facilitate subsequent recognition of the gestures of the object in the object gesture image data.

[0028] In one implementation manner, the position of the object's hand can be determined, and then image data acquisition can be performed towards the position of the object's hand to obtain object gesture image data.

[0029] In this implementation manner, the position of the object's hand can be determined based on the image data acquired by the imaging device, and then the shooting angle of the imaging device can be controlled so that the imaging field of the shooting device is aligned with the object's hand, thereby enabling the imaging device to perform image data acquisition towards the position of the object's hand to obtain object gesture image data.

[0030] In another implementation, gesture image data of an object can be intercepted from the image data of the object.

[0031] Optionally, hand detection can be performed based on the object image data to determine the position of the object's hand; then, the gesture image data of the object can be determined using the position of the object's hand. For example, image data acquisition is performed towards the position of the object's hand to obtain the object gesture image data, or for example, the gesture image data of the object is intercepted from the image data of the object according to the position of the object's hand.

[0032] Among them, there are various methods for determining the position of the object's hand based on the object image data, which are not limited here.

[0033] In one embodiment, hand detection can be directly performed on the image data to determine the position of the object's hand.

[0034] Optionally, computer vision algorithms or deep learning models can be used to detect the hand position.

[0035] In one implementation, a rule-based method can be used to detect the position of the object's hand. For example, traditional image processing techniques such as skin color segmentation and edge detection are used to identify the hand region to determine the position of the object's hand.

[0036] In another implementation, a deep learning model can be used to detect the position of the object's hand. Among them, a pre-trained convolutional neural network (CNN) model, such as MediaPipe, OpenPose, YOLO, etc., can be used to detect the key points or bounding boxes of the hand to determine the position of the object's hand.

[0037] In another embodiment, object detection can be first performed on the image data, and then hand detection of the object can be performed in the detected object body region to determine the position of the object's hand. In combination with the characteristics of the actual environment, the detection of the object's hand is enlarged to object detection. Through the detection method of first global and then local, the detection of the object's hand is upgraded to object body detection, and the detection speed and accuracy are improved in the way of feature amplification. In this way, by combining the object detection method and the object hand detection method, the object body is first located within a wide range to exclude interference factors from non-object body regions (such as outdoor interference factors), thereby improving the detection accuracy of the object hand position.

[0038] Among them, object detection technologies such as image recognition technology can be used to identify the object in front of the imaging device.

[0039] Among them, during the object detection process, a Gaussian heatmap of the probability of the object's hand appearance can also be determined; when detecting the object's hand, the object's hand is detected from the region with a high probability to the region with a low probability in the Gaussian heatmap. In this way, by accurately positioning the local ROI (region of interest), the hand is detected within a small range to more accurately detect the object's hand. This process ensures that the system can be efficiently positioned to improve the recognition efficiency and accuracy of the hand.

[0040] In one example, the step of determining the position of the object's hand based on the object image data may include: whether the object is recognized; if so, based on the Gaussian heatmap of the probability of the object's hand appearance, each region in the Gaussian heatmap is sequentially used as the region of interest in the order from high to low probability, and it is detected whether there is an object's hand in the region of interest until an object's hand is detected in the region of interest.

[0041] In addition, as Figure 2 shown, considering that generally the gesture of the object can be distinguished clearly only when the object is facing forward, when the object is recognized, it can also be confirmed whether the object in the image data is facing forward (also called the front side), that is, it is determined whether the object image data is collected from the front view by the imaging device.

[0042] In this way, in another example, the step of determining the position of the object's hand based on the object image data may include: whether the object is recognized; if the object is recognized, it is determined whether the object's front side is recognized; if the object's front side is recognized, based on the Gaussian heatmap of the probability of the object's hand appearance, each region in the Gaussian heatmap is sequentially used as the region of interest in the order from high to low probability, and it is detected whether there is an object's hand in the region of interest until an object's hand is detected in the region of interest; if the object or the object's front side is not recognized, continue to perform object recognition on the real-time collected image.

[0043] Among them, the gesture image data of the object can be gesture static image data including one frame of image, or gesture dynamic image data including multiple frames of images.

[0044] Optionally, once the hand is detected, a tracking algorithm (such as KCF, CSRT, DeepSort, etc.) can be used to continuously track the movement of the hand to avoid re-detection each time, so as to obtain the gesture dynamic image data of the object.

[0045] S102: Perform fingertip detection on the gesture image data to obtain a fingertip recognition feature vector.

[0046] After obtaining the gesture image data of the object, fingertip detection can be performed on the gesture image data to obtain a fingertip recognition feature vector.

[0047] Optionally, the fingertip recognition feature vector may include an object finger fingertip position feature vector.

[0048] Among them, there are various detection methods for the feature vector of the fingertip position of the target finger, which are not limited herein.

[0049] For example, the fingertip detection of the gesture image data can be performed by a fingertip detection method based on color segmentation, a fingertip detection method based on depth information, a fingertip detection method based on a deep learning model, a fingertip detection method based on geometric features, a fingertip detection method based on optical flow, a fingertip detection method based on a skeleton model, a fingertip detection method based on multimodal fitting, a fingertip detection method based on thermal imaging, a fingertip detection method based on contour analysis, or a fingertip detection method based on template matching to obtain the feature vector of the fingertip position of the target finger.

[0050] In the process of obtaining the feature vector of the fingertip position of the target finger through fingertip detection, the gesture image data can be processed by a hand feature extraction unit (such as the fingertip feature extraction unit, finger feature extraction unit, or overall hand feature extraction unit in the fingertip detection methods such as the above-mentioned fingertip detection method based on a deep learning model) to obtain a feature matrix; then the position of the finger fingertip is determined based on the element values in the feature matrix. When the overall hand feature extraction unit extracts features, the element values corresponding to the pixels in the hand area in the feature matrix are the first value, and the element values corresponding to the pixels outside the hand area in the feature matrix are the second value. When the finger feature extraction unit extracts features, the element values corresponding to the pixels in the finger area in the feature matrix are the first value, and the element values corresponding to the pixels outside the finger area in the feature matrix are the second value. When the fingertip feature extraction unit extracts features, the element values corresponding to the pixels in the fingertip area in the feature matrix are the first value, and the element values corresponding to the pixels outside the fingertip area in the feature matrix are the second value.

[0051] Optionally, determining the position of the finger tip based on the element values in the feature matrix may include: finding the position where the first value is the first value in a preset row of the feature matrix and recording the position in a linked list, then finding the position where the value is the first value in the sequence (the second value, the first value) and recording it in the linked list to obtain the first linked list data, where the preset row is located in the middle region of the feature matrix; finding the position with the smallest column number in the first linked list data and determining whether it is at the boundary of the feature matrix; if it is, using the first linked list data as the second linked list data; if not, traversing from bottom to top in a preset column of the feature matrix and determining the position where the first value is the first value and the position where the value is the first value in the sequence (the first value, the second value) to obtain at least one intermediate position, where the preset column is located in the left edge region of the feature matrix; when the number of the intermediate positions is greater than or equal to 2, determining whether the distance between the current intermediate position and its next intermediate position is less than or equal to a preset value, and the current intermediate position is initially the first intermediate position; if it is less than or equal to, adding the next intermediate position of the current intermediate position to the first linked list data, then updating the next intermediate position of the current intermediate position to the current intermediate position, and returning to the step of determining whether the distance between the current intermediate position and its next intermediate position is less than or equal to the preset value; until all the intermediate positions are traversed, or the distance between the current intermediate position and its next intermediate position is greater than the preset value, to obtain the second linked list data; performing an operation of finding the contour vertex upward for each position in the second linked list data and changing the corresponding linked list position to the position of the vertex to obtain the third linked list data, where the positions in the third linked list data are the positions of the finger tips of the object.

[0052] Exemplarily, as Figure 3As shown, first, through an image processing method, a 25×25 feature matrix is obtained, which represents the features of the gesture image. Locate the starting point of the finger: In the 8th row of the feature matrix, find the first position with a value of 1 and record this position. Then, find the position with a value of 1 in the sequence 0, 1 and record it as well. Store this position information in a linked list. Check the positions in the linked list: Check the linked list storing the positions. Find the position with the smallest column number and determine whether it is on the boundary. If it is on the boundary, keep the linked list unchanged. If it is not on the boundary, vertically search for 1 from bottom to top at column number 2 and the position with a value of 1 in the sequence 1, 0. Calculate the distance between the two 1s. If the distance is greater than 4, keep the linked list unchanged. Otherwise, add the position of the second found 1 to the linked list. By detecting the position of 1 in the sequence 1, 0, it helps to identify the interval between adjacent fingers. By calculating the distance between the two 1s, if the distance is less than or equal to 4, to ensure that the interval between adjacent fingers is within a reasonable range and avoid misidentifying them as one finger. Traverse the positions in the linked list: Traverse the position information stored in the linked list. For each position, perform the operation of searching for the contour vertex upwards and change the corresponding linked list position to the coordinates of the vertex.

[0053] Optionally, the fingertip recognition feature vector may include the feature vector of the center of gravity of the object's hand and / or the data related to the center of gravity of the object's hand.

[0054] Among them, the feature vector of the center of gravity of the object's hand may refer to the feature vector of the position of the center of gravity of the object's hand. Optionally, the position of the center of gravity of the object's hand can be calculated based on the positions of the pixel points of the object's hand contour, and thus the feature vector of the position of the center of gravity of the object's hand can be obtained. Exemplarily, in the case where the feature matrix is obtained by the hand overall feature extraction unit extracting features from the object gesture image data, the position information with the first value can be averaged to obtain the feature vector of the position of the center of gravity of the object's hand.

[0055] The feature vector of the data related to the center of gravity of the object's hand may include the feature vector of the orientation of the finger tips. Exemplarily, the angle parameter between the object finger and the preset direction can be calculated and used as the feature vector of the orientation of the finger tips. Among them, the preset direction can be the vertical direction or the horizontal direction, etc. In a specific example, the first vector can be calculated based on the position of the object finger tip and the position of the center of gravity of the hand, and the dot product operation is performed between the first vector and the vertical direction vector b(1, 0) of the gesture image data to calculate the cosine value of the angle between the finger and the vertical direction.

[0056] S103: Extract features from the gesture image data based on the gesture classification model to obtain the gesture recognition feature vector.

[0057] After obtaining the gesture image data of the object, the features of the gesture image data can be extracted based on the gesture classification model to obtain the gesture recognition feature vector.

[0058] Among them, there are various methods for the gesture classification model, which are not limited here.

[0059] As Figure 4 shown, for example, it can be a C3D model. Among them, the C3D network model is optimized based on the 3D-CNN network model and can have a good recognition effect. However, during the generalization process, the accuracy of the neural network for detection decreases, and the performance will gradually improve. And during the neural network recognition process, when the confidence level is lower than 0.6, more than 50% of its classification results are incorrect. Some of them are misclassifications, where the gesture category is misidentified, and the other part is the interference of non-standard gestures. Based on this, the present application fuses the finger tip feature vector and the gesture recognition feature vector, and analyzes and processes the fused feature vector. In this way, the gesture recognition result is obtained after fusing the fingertip recognition feature vector and the gesture recognition feature vector at the decision level. Compared with the data-level fusion and feature-level fusion, which have problems such as mutual interference between different dimensions and model complexity, the decision-level fusion has good robustness to external influencing factors and fully considers the differences in different dimensions. Therefore, using the decision-level fusion method for fusion helps the model understand the object behavior and gesture decision, and thus this fusion strategy can improve the running speed and the overall recognition accuracy of the model.

[0060] For another example, the gesture classification model can be a CNN model or an RNN model, etc.

[0061] Among them, the gesture classification model can include a feature extraction unit and an output unit. Among them, the feature extraction unit is used to extract features from the gesture image data to obtain a gesture recognition feature vector; the output unit is used to process based on the gesture recognition feature vector to obtain the probability distribution of the gesture category, so as to obtain the gesture category recognition result.

[0062] Optionally, the feature extraction unit can include a convolutional layer, a pooling layer, and a fully connected layer. For example, the C3D model contains multiple 3D convolutional layers and can perform convolution in three dimensions: time, height, and width. The number of convolutional kernels (i.e., the number of filters) increases layer by layer to extract higher-level features; Pooling layer: After each 3D convolutional layer, a 3D max pooling layer is usually added to reduce the dimension of the feature map, reduce the amount of calculation, and prevent overfitting; Fully connected layer: After multiple convolutional and pooling operations, the C3D model flattens the feature map and further extracts high-level features through several fully connected layers.

[0063] In a specific example, the gesture image data may include color gesture image data (such as RGB image data) and / or depth gesture image data. Thus, the color gesture image data and the depth gesture image data can be used as inputs. The gesture classification model is a C3D model, which includes 8 convolutional layers, 5 pooling layers, 2 fully connected layers, and a Softmax output layer. The convolutional kernels of the convolutional layers are all 3×3×3, and the stride is 1×1×1. Each convolutional layer uses the Mish activation function. Except that the kernel and stride of the first pooling layer are set to 2×2×1, the kernels and strides of the remaining pooling layers are set to 2×2×2. The output unit of the fully connected layer is 4096. In addition, dropout can be used after the fully connected layer to prevent overfitting.

[0064] S104: Analyze and process the fingertip recognition feature vector and the gesture recognition feature vector after fusing them at the decision layer to obtain the gesture recognition result.

[0065] Optionally, the finger fingertip feature vector can be fused with the gesture recognition feature vector, and the fused feature vector can be analyzed and processed. That is, the finger fingertip feature vector is fused into the decision layer of the gesture classification model as a dimension of the image algorithm for analysis and gesture result analysis output. When classifying gestures, by introducing the fingertip vector feature, the key points of the hand can be accurately located to eliminate ambiguity, improve the accuracy and reliability of the hand features, enhance the reliability of the overall gesture recognition, and further improve the algorithm's ability to understand the gesture structure. Thus, a feature fusion method combining the gesture classification algorithm and the fingertip detection algorithm is adopted to avoid misrecognition behaviors that may occur during the generalization process of the algorithm. Therefore, by introducing the fingertip vector feature, the reliability of the overall gesture recognition can be fundamentally improved, making the algorithm more adaptable to different scenarios and complex gestures.

[0066] The output unit (i.e., the output layer) of the gesture classification model can be used to process the fused feature vector for analysis and processing to obtain the gesture category recognition result.

[0067] In another implementation, other processing modules can be used to process the fused feature vector for analysis and processing to obtain the gesture category recognition result.

[0068] This application can also provide a gesture control method, which includes: performing gesture recognition based on the above gesture recognition method to obtain a gesture recognition result; controlling the controlled device based on the gesture recognition result.

[0069] For example, if the controlled device is an outdoor advertising screen, the object gesture image data can be obtained by first using a camera to capture the image in front of the advertising screen. Specifically, after recognizing the gesture of the object in front of the screen, the first focus point is aligned with the object, that is, the position where the object's hand appears on the screen. Then, the gesture recognition method described above is used to perform gesture recognition on the object gesture image data to obtain a gesture recognition result. Subsequently, the advertising screen is controlled using the gesture recognition result.

[0070] For another example, if the controlled device is an air conditioner, the air conditioner can be controlled based on the gesture recognition result of the image in front of the air conditioner.

[0071] Please refer to Figure 5 , Figure 5 , which is a schematic structural diagram of an embodiment of the electronic device of the present application. The electronic device 20 includes a processor 22, and the processor 22 is used to execute instructions to implement the above method. For the specific implementation process, please refer to the description of the above embodiment, which will not be elaborated here.

[0072] The processor 22 can also be referred to as a CPU (Central Processing Unit). The processor 22 may be an integrated circuit chip with signal processing capabilities. The processor 22 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 22 can also be any conventional processor, etc.

[0073] The electronic device 20 may further include a memory 21 for storing instructions and data required for the operation of the processor 22.

[0074] The processor 22 is used to execute instructions to implement the method provided by any embodiment of the above method of the present application and any non-conflicting combination.

[0075] Among them, the electronic device of the present application can be an outdoor advertising screen or a household appliance (such as an air conditioner), etc.

[0076] Please refer to Figure 6 , Figure 6This is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present application. The computer-readable storage medium 30 in the embodiment of the present application stores instruction / program data 31, and when the instruction / program data 31 is executed, it implements the methods provided by any one embodiment of the gesture recognition method and the gesture control method of the present application and any non-conflicting combination. In one embodiment, the instruction / program data 31 can form a program file and be stored in the above storage medium 30 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium 30 includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0077] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0078] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0079] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.

[0080] The above are only the implementation manners of this application, and do not thus limit the patent scope of this application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of this application.

Claims

1. A gesture recognition method, characterized in that, The method includes: Obtaining object gesture image data; Performing fingertip detection on the gesture image data to obtain a fingertip recognition feature vector; Performing feature extraction on the gesture image data based on a gesture classification model to obtain a gesture recognition feature vector; Fusing the fingertip recognition feature vector and the gesture recognition feature vector at the decision-making layer and then performing analysis and processing to obtain a gesture recognition result.

2. The method according to claim 1, characterized in that, The fingertip recognition feature vector includes the fingertip position feature vector and / or the fingertip orientation feature vector of each finger of the object's hand.

3. The method according to claim 2, wherein The performing fingertip detection on the gesture image data to obtain a fingertip recognition feature vector includes: Performing feature extraction on the gesture image data to obtain a feature matrix, where the element values corresponding to the pixels in the hand region of the feature matrix are the first value, and the element values corresponding to the pixels in the non-hand region of the feature matrix are the second value; Determining the fingertip positions of the object's fingers according to the feature matrix; Calculating the average value of the position information with the first value to obtain the center-of-gravity position of the object's hand; Calculating a first vector for each finger according to the fingertip position of each finger and the center-of-gravity position of the hand; Calculating the angle between the first vector of each finger and a preset direction vector, and the angle corresponding to each finger is the fingertip orientation feature vector of each finger.

4. The method according to claim 3, wherein The determining the fingertip positions of the object's fingers according to the feature matrix includes: Finding the first position with the first value and the positions with the first value in the sequence of the second value and the first value in a preset row in the feature matrix and recording them in a linked list to obtain first linked list data, where the preset row is located in the middle region of the feature matrix; Finding the position with the smallest column number in the first linked list data and determining whether it is at the boundary of the feature matrix; If it is at the boundary of the feature matrix, using the first linked list data as the second linked list data; If it is not, traversing from bottom to top in a preset column of the feature matrix and determining the first position with the first value and the subsequent positions with the first value in the sequence of the first value and the second value to obtain two intermediate positions, where the preset column is located in the left edge region of the feature matrix; Determining whether the distance between the two intermediate positions is less than or equal to a preset value; If it is less than or equal, adding the subsequent positions with the first value in the sequence of the first value and the second value to the first linked list data to obtain the second linked list data, otherwise the second linked list data is the first linked list data; Performing an operation of searching for the contour vertex upward on each position in the second linked list data and changing the corresponding linked list position to the position of the vertex to obtain third linked list data, where the positions in the third linked list data are the fingertip positions of the object's fingers.

5. The method according to claim 1, wherein The obtaining object gesture image data includes: Performing object detection on the image data; Performing object hand detection in the detected object body region to determine the position of the object's hand; Obtaining object gesture image data based on the position of the object's hand.

6. The method according to claim 5, wherein The performing object detection on the image data includes: Determining a Gaussian heat map of the probability of the object's hand appearance; Performing object hand detection on the body area of the detected object to determine the position of the object's hand includes: Performing object hand detection from the area with high probability to the area with low probability in the Gaussian heat map to obtain the position of the object's hand.

7. The method according to claim 5, wherein After performing object detection on the image data, it includes: If an object is recognized, determining whether the front of the object is recognized; If the front of the object is recognized, performing the step of performing object hand detection on the body area of the detected object to determine the position of the object's hand.

8. A gesture control method, characterized in that The method includes: Determining the gesture category of the object based on the method according to any one of claims 1-7; Controlling the controlled device based on the gesture category.

9. An electronic device, characterized in that, The electronic device includes a processor; the processor is configured to execute instructions to implement the steps of the method according to any one of claims 1-8.

10. A computer-readable storage medium having a program and / or instructions stored thereon, characterized in that, When the program and / or instructions are executed, the steps of the method according to any one of claims 1-8 are implemented.