Gesture recognition method and device, driving equipment and readable storage medium

By extracting the coordinates of hand key points to obtain gesture information and matching, the problem that the gesture recognition model cannot be flexibly updated in the prior art is solved, and the need for rapid updates and custom gestures is achieved.

CN120260108APending Publication Date: 2025-07-04APTIV ELECTRONICS (SUZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410004162.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the gesture recognition model cannot flexibly match the newly added or changed gesture types, resulting in a long update cycle and cannot meet the user's needs for customized gestures.

Method used

By extracting the coordinates of the key points of the hand, and matching them with the preset gesture information, the recognition of gesture types is achieved, avoiding the complex model training process.

Benefits of technology

It realizes fast updates and flexible matching of gestures, shortens the development cycle, and meets the needs of user-defined gestures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260108A_ABST
    Figure CN120260108A_ABST
Patent Text Reader

Abstract

The invention discloses a gesture recognition method and device, driving equipment and a readable storage medium. The method comprises the following steps: acquiring a to-be-recognized image containing a first gesture of a user; extracting first coordinates of a plurality of hand key points from the to-be-recognized image; acquiring first gesture information based on the first coordinates of the plurality of hand key points; matching the first gesture information with each piece of preset second gesture information, and obtaining target gesture information meeting a preset matching condition from each piece of second gesture information; and obtaining the type of the first gesture based on a second gesture corresponding to the target gesture information. According to the method and the device, the gesture can be newly added or changed by modifying the multiple pieces of second gesture information, a complex model training process is not needed, the updating period is short, the method and the device are easy to implement, various newly added and changed gesture types can be flexibly matched, and the requirements of custom gestures can be well met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of intelligent driving, and in particular, to a gesture recognition method, apparatus, driving device, and readable storage medium. Background Art

[0002] The development of driving devices, such as automobiles, has entered the era of networking and intelligence, and multimodal human-computer interaction has gradually become a necessity. Gesture is a natural and intuitive human behavior language. Compared with other parts of the body, gestures can often convey more behavioral information. Therefore, gesture recognition is also the most powerful and effective way in human-computer interaction. Users can interact with the driving device through different gestures so that the driving device can obtain corresponding operation instructions.

[0003] Driving devices usually use computer vision to perform gesture recognition, that is, collect the user's hand images, and analyze and infer the hand images through a gesture recognition model to obtain the recognition result of the gesture type. Among them, the gesture recognition model is trained by using these labeled hand images after collecting hand images in advance and performing manual classification and marking during the development stage.

[0004] However, the trained gesture recognition model can only recognize fixed gesture types. When a new gesture type needs to be added, it is usually necessary to collect new gesture images, mark and classify them, and then merge them with the original images to retrain the gesture recognition model so that the gesture recognition model can recognize the new gesture type. This process not only has a large workload but also a long update cycle because it requires repeated execution of development work, and thus cannot meet the flexible needs of user-defined gestures. Summary of the Invention

[0005] Embodiments of the present application provide a gesture recognition method, apparatus, driving device, and readable storage medium to shorten the cycle of adding or changing gestures and flexibly meet the needs of user-defined gestures.

[0006] To solve the above technical problems, embodiments of the present application disclose the following technical solutions:

[0007] In a first aspect, a gesture recognition method is provided, which is applied to a driving device. The method includes:

[0008] Obtain a to-be-recognized image including a user's first gesture;

[0009] Extract first coordinates of multiple hand key points from the to-be-recognized image;

[0010] Obtain first gesture information based on the first coordinates of the multiple hand key points;

[0011] Match the first gesture information with each of the pre-set second gesture information, and obtain target gesture information that meets the preset matching conditions from each of the second gesture information;

[0012] Based on the second gesture corresponding to the target gesture information, obtain the type of the first gesture.

[0013] In a second aspect, a gesture recognition device is provided, which is configured in a driving device. The device includes:

[0014] An acquisition unit, configured to acquire a to-be-recognized image including a user's first gesture;

[0015] An inference unit, configured to extract first coordinates of multiple hand key points from the to-be-recognized image;

[0016] A processing unit, configured to obtain first gesture information based on the first coordinates of the multiple hand key points;

[0017] A matching unit, configured to match the first gesture information with each of the pre-set second gesture information, and obtain target gesture information that meets the preset matching conditions from each of the second gesture information;

[0018] An identification unit, configured to obtain the type of the first gesture based on the second gesture corresponding to the target gesture information.

[0019] In a third aspect, a gesture recognition device for a driving device is provided, including:

[0020] A memory, configured to store a program;

[0021] A processor, configured to execute the program stored in the memory;

[0022] When the program stored in the memory is executed, the processor executes the gesture recognition method according to any one of the first aspect.

[0023] In a fourth aspect, a driving device is provided, including the gesture recognition device according to the second aspect.

[0024] In a fifth aspect, a computer-readable storage medium is provided. The computer-readable medium stores instructions for a computing device to execute. When the computing device executes the instructions, the method according to any one of the first aspect is implemented.

[0025] One of the above technical solutions has the following advantages or beneficial effects:

[0026] Compared with the prior art, a gesture recognition method of the present application includes: obtaining a to-be-recognized image containing a user's first gesture; extracting first coordinates of multiple hand key points from the to-be-recognized image; obtaining first gesture information based on the first coordinates of the multiple hand key points; matching the first gesture information with each pre-set second gesture information, and obtaining target gesture information that meets a preset matching condition from each second gesture information; and obtaining the type of the first gesture based on the second gesture corresponding to the target gesture information. The gesture recognition method provided by the present application obtains first gesture information through the first coordinates of multiple extracted hand key points, and matches the first gesture information with multiple pre-set second gesture information to identify the type of the first gesture. Thus, the addition or change of gestures can be achieved by modifying the multiple second gesture information, without the need for a complex model training process, with a shorter update cycle, easy to implement, capable of flexibly matching various newly added and changed gesture types, and can better meet the needs of custom gestures.

[0027] The gesture recognition device provided by the present application does not require a complex model training process, has a shorter update cycle, is easy to implement, can flexibly match various newly added and changed gesture types, and can better meet the needs of custom gestures.

[0028] The gesture recognition device of the driving device provided by the present application does not require a complex model training process, has a shorter update cycle, is easy to implement, can flexibly match various newly added and changed gesture types, and can better meet the needs of custom gestures.

[0029] The driving device provided by the present application can flexibly match various newly added and changed gesture types, both the product development cycle and the update cycle are shorter, can effectively recognize gestures, and can better meet the custom needs of users.

[0030] The computer-readable storage medium provided by the present application can flexibly match various newly added and changed gesture types, both the product development cycle and the update cycle are shorter, can effectively recognize gestures, and can better meet the custom needs of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0032] Figure 1 It is a schematic diagram of the model development process of a traditional computer vision-based gesture recognition method;

[0033] Figure 2Schematic diagram of the model inference process of traditional computer vision-based gesture recognition methods;

[0034] Figure 3 Schematic diagram of the internal application scenario of the driving device according to an embodiment of the present application;

[0035] Figure 4 For Figure 3 Enlarged structural schematic diagram of area A in

[0036] Figure 5 Schematic diagram of the overall process of the gesture recognition method according to an embodiment of the present application;

[0037] Figure 6 Schematic diagram of an example of hand key points provided by an embodiment of the present application;

[0038] Figure 7 Schematic diagram of the method for obtaining the first pose parameter in an embodiment of the present application;

[0039] Figure 8 Schematic diagram of the palm coordinate system in an embodiment of the present application;

[0040] Figure 9 Schematic diagram of the method for obtaining the first feature parameter in an embodiment of the present application;

[0041] Figure 10 Schematic diagram of an example of the first gesture in an embodiment of the present application;

[0042] Figure 11 Schematic diagrams of various examples of the second gesture in an embodiment of the present application;

[0043] Figure 12 Schematic diagram of the process of matching the first gesture information and the second gesture information in an embodiment of the present application;

[0044] Figure 13 Schematic diagram of the process from step 503 to step 505 in an embodiment of the present application;

[0045] Figure 14 Schematic diagram of the separation of the fixed model and the variable application in the gesture recognition method according to an embodiment of the present application;

[0046] Figure 15 Schematic diagram of the structure of the gesture recognition device according to an embodiment of the present application;

[0047] Figure 16 Schematic diagram of the hardware structure of the gesture recognition device of the driving device according to an embodiment of the present application;

[0048] Reference numerals:

[0049] 100 - Gesture recognition device; 110 - Acquisition unit; 120 - Inference unit; 130 - Processing unit; 140 - Matching unit; 150 - Recognition unit; 200 - Camera; 2201 - Memory; 2202 - Processor. Detailed implementation manners

[0050] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0051] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present application. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present application, "a plurality" means two or more, and "at least one" means one, two or more, unless otherwise specifically defined.

[0052] A driving device, such as a car, can usually integrate gesture recognition with various interaction methods such as speech recognition and face recognition to completely understand the user's gesture language and judge the action that the user wants to perform according to the user's natural gesture, reducing the cost for the user to learn the operation. Among them, gesture recognition is the most powerful and effective way in human-computer interaction. Gestures can be divided into two categories: static gesture recognition and dynamic gesture recognition. Static gesture recognition distinguishes different gestures by different hand shapes, while dynamic gesture recognition extracts the behavioral meaning based on the movement trajectory and movement time of the palm and the arm and then recognizes the corresponding action. Static gesture recognition has a lower computational cost compared to dynamic gesture recognition, and only a single-frame image needs to be processed for each recognition, so static gesture recognition is widely used in various human-computer interaction fields.

[0053] The recognition methods of static gesture recognition can be mainly divided into two categories: sensor-based recognition and computer vision-based recognition. Sensor-based gesture recognition mainly obtains various data of the hand through a microcontroller or specific sensors, such as data gloves, depth camera Kinect (peripheral device), Leap Motion (controlled by the jumping motion of gestures and actions), etc. Sensor-based gesture recognition methods can obtain higher recognition accuracy, but these methods require dedicated devices, and these devices are often relatively expensive. The computer vision-based gesture recognition method only needs an ordinary RGB (Red, Green, Blue) camera to capture an image containing the hand, and then performs gesture segmentation, binarization, feature extraction, and gesture classification on the image to obtain the specific static gesture recognition result.

[0054] Please refer to Figure 1 , Figure 1 which illustrates the model development process of the traditional computer vision-based gesture recognition method. The traditional computer vision-based gesture recognition method often pre-collects and preprocesses the original image, then manually classifies and marks the gestures on the original image, and then uses these marked hand images for training to obtain and store the gesture recognition model. Therefore, if a new gesture needs to be added, the sampled images of the new gesture need to be merged with the original images and retrained and classified again. Since this process requires repeated execution of development work, the workload is large and the cycle is long, and it is very inflexible to add new gestures.

[0055] Please refer to Figure 2 , Figure 2 which illustrates the model inference process of the traditional computer vision-based gesture recognition method. Since the traditional gesture recognition method defines and trains the gesture recognition model in the development stage, the trained gesture recognition model can only recognize fixed gesture types and cannot match new or changed gesture types. If new or changed gestures are to be added, it involves various links such as gesture image sampling, image marking and classification, and model training. The workload is large and time-consuming, and it cannot flexibly meet the needs of custom gestures, affecting the user experience.

[0056] In view of this, the embodiments of the present application provide a driving device that obtains first gesture information by extracting the first coordinates of multiple hand key points, and matches the first gesture information with a preset plurality of second gesture information to identify the type of the first gesture. Furthermore, the addition or change of gestures can be realized by modifying the plurality of second gesture information, without the need for a complex model training process, shortening the gesture update cycle, and flexibly matching various newly added and changed gesture types, thereby solving at least part of the above technical problems.

[0057] Please refer to Figure 3 andFigure 4 , Figure 3 illustrates the internal application scenario of the driving device according to an embodiment of the present application. Figure 4 illustrates Figure 3 the enlarged structure of area A in. The driving device according to an embodiment of the present application may be equipped with an intelligent cockpit system B, and the intelligent cockpit system B can integrate various intelligent technologies and functions to provide a more comfortable and convenient driving experience for users. The driving device according to an embodiment of the present application may include a gesture recognition device 100 and a camera 200. The gesture recognition device 100 can be deployed to run in operating systems such as Linux, QNX, and Android, and is used to implement gesture recognition (Gesture Recognition) and gesture definition (Gesture Definition) using the gesture recognition method according to an embodiment of the present application. In some examples, the gesture recognition device 100 can be deployed in the intelligent cockpit system. The number of cameras 200 can be multiple, and each camera 200 is configured to collect the to-be-recognized images of the corresponding shooting area inside the driving device. Exemplarily, corresponding cameras 200 can be respectively arranged at the driver's seat, the passenger seat, and the rear row position to meet the needs of different passengers. The present application embodiment does not specifically limit the number and position of the cameras 200.

[0058] The following introduces the gesture recognition method according to an embodiment of the present application with reference to the accompanying drawings.

[0059] Please refer to Figure 5 , Figure 5 illustrates the overall process of the gesture recognition method according to an embodiment of the present application. This gesture recognition method is applied to the driving device and is specifically used for the gesture recognition device 100, and specifically includes the following steps:

[0060] Step 501: Obtain the to-be-recognized image containing the user's first gesture.

[0061] Specifically, the user can make the first gesture in any shooting area of the camera 200, and the camera 200 collects the to-be-recognized image, and the to-be-recognized image includes the first gesture. It can be understood that the first gesture is used to control the corresponding functions of the driving device.

[0062] Specifically, the user can make the first gesture with the left hand or the right hand, and no specific limitation is made.

[0063] In some examples, after performing step 501 and before performing step 502, related technologies can be used to preprocess the to-be-recognized image, and the preprocessing can include color conversion, size conversion, etc., so that the to-be-recognized image meets the subsequent processing format requirements.

[0064] Step 502: Extract the first coordinates of multiple hand key points from the to-be-recognized image.

[0065] Please refer to Figure 6 , Figure 6 which illustrates an example of hand key points provided by an embodiment of the present application. In some examples, multiple hand key points may at least include a wrist (WRIST) key point, a carpometacarpal joint (CMC) key point, metacarpophalangeal joint (MCP) key points of each finger, distal interphalangeal joint (DIP) key points of each finger, and tip (TIP) key points of each finger. Exemplarily, the wrist key point may be the midpoint of the wrist, denoted by point 0; the carpometacarpal joint key point (THUMB_CMC, also known as the base joint of the thumb) is denoted by point 1; the metacarpophalangeal joint key points of each finger include the metacarpophalangeal joint key point of the thumb (THUMB_MCP, denoted by point 2), the metacarpophalangeal joint key point of the index finger (INDEX_FINGER_MCP, denoted by point 5), the metacarpophalangeal joint key point of the middle finger (MIDDLE_FINGER_MCP, denoted by point 9), the metacarpophalangeal joint key point of the ring finger (RING_FINGER_MCP, denoted by point 13), and the metacarpophalangeal joint key point of the little finger (PINKY_MCP, denoted by point 17); the distal interphalangeal joint key points of each finger include the distal interphalangeal joint key point of the thumb (THUMB_IP, denoted by point 3), the distal interphalangeal joint key point of the index finger (INDEX_FINGER_DIP, denoted by point 7), the distal interphalangeal joint key point of the middle finger (MIDDLE_FINGER_DIP, denoted by point 11), the distal interphalangeal joint key point of the ring finger (RING_FINGER_DIP, denoted by point 15), and the distal interphalangeal joint key point of the little finger (PINKY_DIP, denoted by point 19); the tip key points of each finger include the tip key point of the thumb (THUMB_TIP, denoted by point 4), the tip key point of the index finger (INDEX_FINGER_TIP, denoted by point 8), the tip key point of the middle finger (MIDDLE_FINGER_TIP, denoted by point 12), the tip key point of the ring finger (RING_FINGER_TIP, denoted by point 16), and the tip key point of the little finger (PINKY_TIP, denoted by point 20). It should be noted that since the structure of the thumb is slightly different from that of other fingers, the interphalangeal joint (IP) of the thumb is used as the key point of the proximal interphalangeal joint of the thumb in the embodiment of the present application. In this way, by extracting at least 17 hand key points, the characteristics of various gestures can be better reflected, which is beneficial to ensuring the accuracy of subsequent gesture analysis.

[0066] In other examples, the multiple hand key points may further include the key points of the distal interphalangeal joints (Proximal Interphalangeal Joint, PIP) of each finger. For example, they include: the key point of the distal interphalangeal joint of the index finger (INDEX_FINGER_PIP, represented by dot 6), the key point of the distal interphalangeal joint of the middle finger (MIDDLE_FINGER_PIP, represented by dot 10), the key point of the distal interphalangeal joint of the ring finger (RING_FINGER_PIP, represented by dot 14), and the key point of the distal interphalangeal joint of the little finger (PINKY_PIP, represented by dot 18). In this way, 21 hand key points can be extracted, which can fully represent the characteristics of various gestures, thus making the analysis more accurate.

[0067] In some embodiments, the first coordinates of each hand key point can be extracted in the following manner:

[0068] The hand key point recognition model processes the image to be recognized to obtain the first coordinates of multiple hand key points. Among them, the hand key point recognition model is pre-trained through multiple hand images with hand key point annotations.

[0069] It can be understood that the first coordinates are used to represent the three-dimensional coordinates in the original space coordinate system. Exemplarily, the original space coordinate system can be a right-handed coordinate system with the upper left corner of the image to be recognized as the origin.

[0070] In some examples, the hand key point recognition model can adopt the Media-pipe hand tracking model. This model can extract the three-dimensional coordinates of 21 hand key points from the image to be recognized.

[0071] In addition, the hand key point recognition model can also output the left / right hand type corresponding to the first gesture, that is, output whether the first gesture corresponds to the left hand or the right hand. This can facilitate the subsequent calculation of the palm orientation. It can be understood that the left hand and the right hand are mirror images of each other.

[0072] It can be understood that the category of the hand making the first gesture can be the same as the category of the hand of the predefined second gesture and the category of the hand used to train the hand key point recognition model, which helps to further improve the recognition accuracy.

[0073] Through the above method, the hand key point recognition model only outputs the coordinates of 21 hand key points and the left / right hand type, without complex model definitions, and the model training and computing resource occupation are small. Thus, the processor resource occupation can be reduced, which is beneficial to the application on driving devices.

[0074] Step 503: Obtain the first gesture information based on the first coordinates of multiple hand key points.

[0075] In some embodiments, the first gesture information may include first posture parameters, and the first posture parameters may include at least one parameter among the palm axis direction, the palm cutting plane direction, and the palm center orientation.

[0076] Please refer to Figure 7 , Figure 7 which illustrates the manner of obtaining the first posture parameters in the embodiments of the present application. In some examples, the first posture parameters can be obtained through the following steps:

[0077] Step 1: Based on the first coordinates of the wrist key point and the first coordinates of the middle finger root key point, obtain the palm axis direction.

[0078] Exemplarily, Figure 7 shows a view from the back of the right hand. The vector P09 from point 0 to point 9 can be used to represent the palm axis direction.

[0079] Step 2: Based on the first coordinates of the index finger root key point and the first coordinates of the little finger root key point, obtain the palm cutting plane direction.

[0080] Exemplarily, the vector P175 from point 17 to point 5 can be used to represent the palm cutting plane direction.

[0081] Step 3: Based on the first coordinates of the index finger root key point, the first coordinates of the little finger root key point, the first coordinates of the wrist key point, and the first coordinates of the middle finger root key point, obtain the palm center orientation.

[0082] Exemplarily, with the vector P175 as the X direction and the vector P09 as the Y direction, a right - hand coordinate system is used to obtain the normal vector P0 of the plane formed by the vectors P175 and P09. If the first gesture corresponds to the right hand, then the palm center orientation is the same as the direction of the normal vector P0; if the first gesture corresponds to the left hand, then the palm center orientation is opposite to the direction of the normal vector P0. By calculating the angles between the vectors P09, P175, P0 and the original space coordinate system X(1,0,0), Y(01,0), Z(0,0,1), the corresponding specific spatial directions can be obtained.

[0083] In this way, through the above - mentioned manner, the front (FORWARD), back (BACK), up (UP), down (DOWN), left (LEFT), and right (RIGHT) can be distinguished by the palm axis direction, the palm cutting plane direction, and the palm center orientation, and further, the gesture postures with pointing information can be identified.

[0084] In some embodiments, the first gesture information may further include first feature parameters, and the first feature parameters may include at least one parameter among the positions of each fingertip, the distance between any two fingers, the slope of each finger, and the curvature of each finger.

[0085] In some examples, the first feature parameter can be obtained through the following steps:

[0086] Step 1: Construct a palm coordinate system.

[0087] Please refer to Figure 8 , Figure 8 which illustrates the palm coordinate system of the embodiments of the present application. In some examples, the palm coordinate system can use the wrist key point, i.e., point 0, as the origin O, and use the direction from the wrist key point to the base key point of the index finger, i.e., the vector P05 from point 0 to point 5, as the X-axis direction. The unit vector of the X-axis is denoted as M0. Use the direction from the wrist key point to the base key point of the little finger, i.e., the vector P017 from point 0 to point 17, as the Y-axis direction. The unit vector of the Y-axis is denoted as M1. The unit vectors M0 and M1 form the XY plane of the palm coordinate system. Based on the right-hand rule, obtain the normal vector perpendicular to the vector P05 and the vector P017 as the Z-axis direction. The unit vector of the Z-axis is denoted as M2. The palm coordinate system formed by the unit vectors M0, M1, and M2 is a right-hand coordinate system.

[0088] Step 2: Convert the first coordinates of each hand key point into second coordinates in the palm coordinate system.

[0089] In some examples, first, obtain the transformation matrix M, where M = [M0, M1, M2] -1 .

[0090] Specifically, the X-axis vector and Y-axis vector of the palm coordinate system can be represented by the following formula (1):

[0091]

[0092] The Z-axis vector can be calculated by cross product, and is specifically represented by the following formula (2):

[0093]

[0094] Construct a matrix R with the X-axis vector, Y-axis vector, and Z-axis vector as column vectors, as shown in the following formula (3):

[0095]

[0096] Then the inverse matrix of R is the transformation matrix M, i.e., M = R -1 .

[0097] Then, convert the first coordinates of each hand key point into second coordinates in the palm coordinate system through the transformation matrix M, specifically as shown in the following formula (4):

[0098] A′ = M · A (4);

[0099] In Formula (4), A is the first coordinate in the original space coordinate system, and A' is the second coordinate in the palm coordinate system.

[0100] By the above method, coordinate conversion based on the palm coordinate system can avoid the problem that the recognition accuracy is affected by individual differences such as hand size when directly calculating in the original space coordinate system.

[0101] Step 3: Obtain the first feature parameter based on multiple second coordinates.

[0102] Please refer to Figure 9 , Figure 9 which illustrates the method for obtaining the first feature parameter in the embodiment of the present application. In some examples, the first feature parameter can be obtained based on multiple second coordinates through the following steps:

[0103] The first step: Based on the second coordinates of the key points of the fingertips of each finger, the second coordinates of the key point of the wrist, and the second coordinates of the key points of the roots of each finger, obtain the fingertip positions of each finger.

[0104] In some examples, the first distance from the key point of the fingertip of each finger to the key point of the wrist and the second distance from the key point of the root of each finger to the key point of the wrist can be obtained, and then the ratio of the first distance to the second distance is determined as the fingertip position of the corresponding finger. This fingertip position can reflect the distance feature of the corresponding finger. Taking the index finger in the first gesture made by the left hand currently as an example, in this first gesture, the first distance from the key point of the fingertip of the index finger (point 8) to the key point of the wrist (point 0) is L08, and the second distance from the key point of the root of the index finger (point 5) to the key point of the wrist (point 0) is L05. The fingertip position L1 of the index finger = L08 / L05.

[0105] Exemplarily, according to the difference in the fingertip position, the fingertip position can be divided into far, moderate, and near. Specifically, a first threshold and a second threshold can be defined first. When the fingertip position is less than the first threshold, the fingertip position can be determined to be near. When the fingertip position is greater than the second threshold, the fingertip position can be determined to be far. When the fingertip position is between the first threshold and the second threshold, the fingertip position can be determined to be moderate. The first threshold can be set to 1.0, and the second threshold can be set to 1.7.

[0106] The second step: Based on the second coordinates of the key points of the fingertips of each finger, obtain the distance between any two fingers.

[0107] In some examples, based on the second coordinates of the key points of the fingertips of each finger, the first vector pointing from one finger to another finger can be obtained, and the length of the first vector is determined as the distance between the two fingers. Taking the distance between the ring finger and the little finger in the first gesture made by the left hand currently as an example, the length of the vector P1620 is determined as the distance D16-20 between the ring finger and the little finger.

[0108] Exemplarily, according to the difference in the distance between any two fingers, the relationship between any two fingers can be divided into far and near. To determine the relationship between two fingers, a third threshold can be defined first. When the distance between the two fingers is greater than the third threshold, it is determined that the two fingers are separated, and the relationship between the fingers can be determined as far; when the distance between the two fingers is less than the third threshold, it is determined that the two fingers are closed together, and the relationship between the fingers can be determined as near. The typical value of the third threshold can be set according to requirements. For example, when calculating the distance between the thumb and other fingers, the third threshold can be the length of the first joint of the thumb; when calculating the distance between the index finger and other fingers, the third threshold can be the length of the first joint of the index finger.

[0109] In the third step, based on the second coordinates of the key points of each finger tip, the second coordinates of the key points of the proximal interphalangeal joints of each finger, the second coordinates of the wrist key points, and the second coordinates of the key points of the finger roots of each finger, the bending degree of each finger and the slope of each finger are obtained.

[0110] In some embodiments, for any one of the index finger, middle finger, ring finger, and little finger, a slope feature can be constructed through the angle between the straight line formed by the key points of the finger tip and the key points of the finger root and the palm surface. Specifically, based on the second coordinates of the key points of the finger tip and the key points of the finger root, the second vector from the finger root to the finger tip is obtained. Then, based on the second coordinates of the wrist key points and the second coordinates of the key points of the middle finger root, the third vector from the wrist key point to the middle finger root is obtained, and the angle between the second vector and the third vector is determined as the finger slope θ. Taking the middle finger in the first gesture made by the left hand currently as an example, the angle ∠12 between the second vector P912 and the third vector P09 is determined as the finger slope θ. For the thumb, a slope feature can be constructed through the angle between the straight line formed by the key points of the proximal interphalangeal joint (point 3) and the key points of the finger root (point 2), and the straight line formed by the key points of the finger root (point 2) and the key points of the index finger root (point 5). Specifically, based on the second coordinates of the key points of the proximal interphalangeal joint of the thumb (point 3) and the second coordinates of the key points of the thumb root (point 2), the fourth vector P32 is obtained. Then, based on the second coordinates of the key points of the thumb root (point 2) and the second coordinates of the key points of the index finger root (point 5), the fifth vector P25 is obtained, and the angle ∠325 between the fourth vector P32 and the fifth vector P25 is determined as the finger slope θ.

[0111] Specifically, the angle α between two three-dimensional vectors A and B can be determined by the following formula (5):

[0112]

[0113] In formula (5), the numerator represents the dot product (inner product) of vector A and vector B, and the denominator represents the norms (modulus and length) of vector A and vector B.

[0114] Exemplarily, according to the different finger slopes θ, the finger slopes can be divided into parallel, perpendicular, and inclined. For fingers other than the thumb, a fourth threshold and a fifth threshold can be defined first. When the finger slope θ is less than the fourth threshold, it is determined that the finger is parallel to the palm surface; when the finger slope θ is greater than the fifth threshold, it is determined that the finger is inclined to the palm surface; when the finger slope θ is between the fourth threshold and the fifth threshold, it is determined that the finger is perpendicular to the palm surface. The typical values of the fourth threshold and the fifth threshold can be set according to requirements. For example, the fourth threshold can be set to 20 degrees, and the fifth threshold can be set to 110 degrees. For the thumb, the fourth threshold can be set to 40 degrees, and the fifth threshold can be set to 80 degrees. That is to say, if the finger slope θ of the thumb is less than 40 degrees, it can be considered that the thumb is close to the palm surface, that is, parallel to the palm surface; if it is greater than 80 degrees, it can be considered that the thumb is perpendicular to the palm surface; if it is greater than or equal to 40 degrees and less than or equal to 80 degrees, it can be considered that the thumb is inclined to the palm surface.

[0115] In some embodiments, for any one of the index finger, middle finger, ring finger, and little finger, the bending degree feature can be constructed by the angle between the straight line formed by the finger tip key point and the proximal interphalangeal joint key point and the straight line formed by the wrist key point and the finger root key point. Specifically, based on the second coordinates of the finger tip key point and the second coordinates of the proximal interphalangeal joint key point, the sixth vector of the first phalanx can be obtained, and then based on the second coordinates of the wrist key point and the second coordinates of the finger root key point, the seventh vector from the wrist key point to the finger root can be obtained. The angle between the sixth vector and the seventh vector is determined as the finger bending degree. Taking the index finger in the first gesture made by the left hand currently as an example, the angle ∠0578 between the sixth vector P78 and the seventh vector P05 is determined as the finger bending degree. For the thumb, based on the second coordinates of the thumb tip key point (point 4) and the second coordinates of the thumb proximal interphalangeal joint key point (point 3), the sixth vector of the first phalanx can be obtained, and then based on the second coordinates of the thumb proximal interphalangeal joint key point (point 3) and the second coordinates of the thumb finger root key point (point 2), the seventh vector can be obtained. The angle ∠432 between the sixth vector and the seventh vector is determined as the finger bending degree. Specifically, the angle between the two vectors can be obtained according to formula (5), which will not be elaborated here.

[0116] Exemplarily, according to different finger bending degrees, the finger bending degrees can be divided into bending, slightly bending, and straightening. For fingers other than the thumb, a sixth threshold and a seventh threshold can be defined first. When the finger bending degree is less than or equal to the sixth threshold, the finger is determined to be in a straightened state; when the finger bending degree is greater than or equal to the seventh threshold, the finger is determined to be in a bent state; when the finger bending degree is between the sixth threshold and the seventh threshold, the finger is determined to be in a slightly bent state. The typical values of the sixth threshold and the seventh threshold can be set according to requirements. For example, the sixth threshold can be set to 30 degrees, and the seventh threshold can be set to 180 degrees. For the thumb, the sixth threshold can be set to 100 degrees, and the seventh threshold can be set to 150 degrees.

[0117] It can be understood that for the above-mentioned first characteristic parameters and first posture parameters, except for the set states, if there is no need to calculate, they can be set to an ignored state. For example, for a fist gesture, if there is no need to calculate the palm orientation, the palm orientation can be in an ignored state. It should be determined flexibly according to the specific situation.

[0118] Please refer to Figure 10 , Figure 10 which illustrates an example of the first gesture in an embodiment of the present application. Taking the first gesture made by the left hand shown in the figure as an example, based on the first posture parameters such as the palm axis direction, the palm section direction, and the palm orientation obtained in step 503, as well as the first characteristic parameters such as the positions of each fingertip, the distance between any two fingers, the slope of each finger, and the bending degree of each finger, the first gesture information of the first gesture obtained is shown in Table 1 below.

[0119] Table 1: Example of First Gesture Information

[0120]

[0121] In Table 1, Near means near, Far means far, Bias means tilt, Parallel means parallel, Bend means bend, Straight means straighten, Ignore means not considered, and / means meaningless.

[0122] Through the above method, based on the hand key point coordinates, the parameters of the first gesture are calculated from multiple dimensions. The amount of calculation data is small and the calculation is simple, which can save a large amount of processor resources.

[0123] Step 504: Match the first gesture information with each pre-set second gesture information, and obtain the target gesture information that meets the preset matching conditions from each second gesture information.

[0124] Specifically, the second gesture information may include second posture parameters and second feature parameters. The second posture parameters include the palm axis direction, the palm cutting plane direction, and the palm center orientation. The second feature parameters include the positions of each fingertip, the distance between any two fingers, the slope of each finger, and the curvature of each finger. That is to say, the gesture definition dimensions of the second gesture information and the first gesture information are the same.

[0125] In some examples, each second gesture information and the corresponding second gesture are stored in a preset gesture configuration library. The gesture configuration library may be configured with a custom interface, and the custom interface is configured to obtain third gesture information and update each second gesture information based on the third gesture information.

[0126] Among them, the third gesture information may be new gesture information or update gesture information used to update a certain second gesture information.

[0127] It can be understood that each second gesture information and the corresponding second gesture in the gesture configuration library can be set in advance, and specifically, the right hand can be used for definition. The custom interface can be opened to users, and users can input the corresponding third gesture information according to their own habits to update the gesture configuration library, so that gesture recognition is more in line with user habits, the recognition accuracy is higher, and the user experience and usage flexibility can be greatly improved.

[0128] Please refer to Figure 11 , Figure 11 , which shows various examples of the second gesture in the embodiments of the present application. The examples of each second gesture and the corresponding application scenarios can be specifically seen in Table 2 below.

[0129] Table 2: Examples of the second gesture, application scenarios, and remarks

[0130]

[0131]

[0132] The examples of each second gesture in Table 2 and the corresponding second gesture information can be specifically seen in Table 3 below.

[0133] Table 3: Examples of the second gesture and the second gesture information

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142] In Table 3, Near means near, Far means far, Bias means tilt, Parallel means parallel, Vertical means vertical, Bend means bend, Microbend means microbend, Straight means straight, Ignore means not considered, and / means meaningless.

[0143] As shown in the above table, the desired second gesture can be defined using the following dimensions: the distance between the fingertip and the center of the wrist, the distance between fingertips, the fingertip slope, the finger bend, the orientation of the palm center, the palm axis, and the direction of the palm section.

[0144] In some embodiments, the preset matching conditions may include fuzzy matching conditions and exact matching conditions.

[0145] Please refer to Figure 12 , Figure 12 which illustrates the process of matching the first gesture information and the second gesture information in the embodiments of the present application. Specifically, the target gesture information can be obtained from each second gesture information through the following steps:

[0146] First step, perform fuzzy matching between the first gesture information and each second gesture information respectively, and determine the second gesture information that meets the fuzzy matching conditions as the candidate gesture information.

[0147] Specifically, it can be detected whether there is un-matched second gesture information in the gesture configuration library. If so, take out a second gesture information and compare it with the first gesture information parameter by parameter. Among them, if the parameters are exactly the same, it is determined as a match. If the comparison results of all parameters in this second gesture information are all matches or ignored (i.e., Ignore), then this second gesture information is determined as the candidate gesture information and added to the candidate list. Otherwise, it is not added to the candidate list, and the comparison of the next second gesture information is restarted until all second gesture information in the gesture configuration library has been traversed.

[0148] Second step, perform exact matching between the first gesture information and each candidate gesture information respectively, and determine the candidate gesture information that meets the exact matching conditions as the target gesture information.

[0149] Specifically, the candidate gesture information can be compared with the first gesture information parameter by parameter again. If the comparison results of all parameters in this candidate gesture information are all matches, then this candidate gesture information is determined as the target gesture information.

[0150] In other embodiments, other matching strategies may also be adopted, which are not specifically limited.

[0151] In this way, through the above method, the gesture configuration in the embodiments of the present application is separated from the hand key point recognition model. The definition of the gesture is only reflected in the configuration file and can be updated later. To change the gesture, only the configuration file needs to be updated, and there is no need to retrain the hand key point recognition model repeatedly. Therefore, the development cycle is shorter, and a gesture definition interface can also be developed for users to customize gestures by themselves, which has extremely strong flexibility.

[0152] Step 505: Based on the second gesture corresponding to the target gesture information, obtain the type of the first gesture.

[0153] Specifically, determine the second gesture corresponding to the target gesture information as the type of the first gesture. For example, if after matching, the second gesture corresponding to the target gesture information is a fist gesture, then the type of the first gesture is also determined as a fist.

[0154] In addition, in some embodiments, after executing step 505, the driving device may execute an operation corresponding to the first gesture in the current scenario in response to the type of the first gesture. For example, when the current scenario is multimedia playback and the user needs to switch to the next playback item, the user can show a first gesture to the right at the camera 200. When the driving device recognizes that the type of the first gesture is to the right, the current playback item is switched to the next playback item. It can be understood that the same gesture may trigger different operations in different scenarios, and different gestures may also trigger the same operation in different scenarios, which can be specifically set according to actual requirements.

[0155] To more clearly illustrate the method of the embodiments of the present application, please refer to Figure 13 , Figure 13 which shows the flow of steps 503 to 505 in the embodiments of the present application. After converting the three-dimensional coordinates output by the hand key point recognition model into palm coordinates, hand pose operations and hand feature operations are respectively performed, and then they are matched with each second gesture information in the configuration library of the gesture definition configuration to obtain the final recognition type.

[0156] Please refer to Figure 14 , Figure 14 which shows the separation schematic of the fixed model and the variable application in the gesture recognition method of the embodiments of the present application. After the hand key point recognition model is pre-trained using multiple hand images with hand key point annotations, it is set on the driving device and does not need to be retrained and updated repeatedly. Each second gesture in the gesture configuration library can be continuously modified and added according to actual requirements, user habits, or application iteration updates.

[0157] It can be understood that the gesture recognition method according to the embodiments of the present application obtains the three-dimensional coordinates of 21 hand key points from the gesture image through the hand key point recognition model, defines the gesture using spatial geometry in several dimensions, realizes the separation of the fixed model and the variable application, and matches the geometric features of the obtained hand key points with the defined gesture features, so as to obtain the custom gesture type, which can solve the problem of retraining required when changing the gesture definition or adding new gestures, shorten the development cycle. In addition, by separating the model data and the gesture application, after the product is released, when the data of the trained hand key point recognition model remains unchanged, only the definition description of the gesture needs to be added or updated to implement the addition or change of gestures. Even a user interface can be opened for users to customize gestures, enabling users to flexibly define gesture features according to their own needs and obtain the corresponding gesture categories during gesture matching.

[0158] Correspondingly, please refer to Figure 15 , Figure 15 which schematically shows the structure of the gesture recognition device according to the embodiments of the present application. The gesture recognition device 100 provided by the embodiments of the present application is configured in the driving device. The gesture recognition device 100 includes an acquisition unit 110, an inference unit 120, a processing unit 130, a matching unit 140, and an identification unit 150.

[0159] The acquisition unit 110 is configured to acquire an image to be recognized including the user's first gesture;

[0160] The inference unit 120 is configured to extract the first coordinates of multiple hand key points from the image to be recognized;

[0161] The processing unit 130 is configured to obtain first gesture information based on the first coordinates of the multiple hand key points;

[0162] The matching unit 140 is configured to match the first gesture information with each pre-set second gesture information, and obtain target gesture information that meets the preset matching conditions from each of the second gesture information;

[0163] The identification unit 150 is configured to obtain the type of the first gesture based on the second gesture corresponding to the target gesture information.

[0164] In some embodiments, the first gesture information includes first attitude parameters, and the first attitude parameters include at least one parameter among the palm axis direction, the palm section direction, and the palm center orientation.

[0165] In some embodiments, the multiple hand key points include a wrist key point and the root key points of each finger;

[0166] The processing unit 130 is specifically configured to:

[0167] Based on the first coordinates of the wrist key points and the first coordinates of the midpoint of the middle finger root key points, obtain the palm axis direction;

[0168] Based on the first coordinates of the root key points of the index finger and the root key points of the little finger, obtain the palm section direction;

[0169] Based on the first coordinates of the root key points of the index finger, the root key points of the little finger, the wrist key points, and the root key points of the middle finger, obtain the palm orientation.

[0170] In some embodiments, the first gesture information further includes a first feature parameter, and the first feature parameter includes at least one parameter among the positions of each fingertip, the distance between any two fingers, the slope of each finger, and the curvature of each finger.

[0171] In some embodiments, the processing unit 130 is further specifically configured to:

[0172] Construct a palm coordinate system;

[0173] Convert the first coordinates of each hand key point into second coordinates in the palm coordinate system;

[0174] Based on multiple second coordinates, obtain the first feature parameter.

[0175] In some embodiments, the multiple hand key points further include the key points of each finger tip and the key points of each proximal interphalangeal joint of each finger;

[0176] The processing unit 130 is specifically configured to:

[0177] Based on the second coordinates of the key points of each finger tip, the second coordinates of the wrist key points, and the second coordinates of the root key points of each finger, obtain the positions of the fingertips of each finger;

[0178] Based on the second coordinates of the key points of each finger tip, obtain the distance between any two fingers;

[0179] Based on the second coordinates of the key points of each finger tip, the second coordinates of the key points of each proximal interphalangeal joint of each finger, the second coordinates of the wrist key points, and the second coordinates of the root key points of each finger, obtain the curvature and the slope of each finger.

[0180] In some embodiments, the processing unit 130 is specifically configured to:

[0181] Taking the wrist key point as the origin, the direction from the wrist key point to the root key point of the index finger as the X-axis direction, and the direction from the wrist key point to the root key point of the little finger as the Y-axis direction, obtain a palm coordinate system, and the palm coordinate system is a right-hand coordinate system.

[0182] In some embodiments, the inference unit 120 is specifically configured to:

[0183] The hand key point recognition model is used to process the image to be recognized, and the first coordinates of multiple hand key points are obtained. The hand key point recognition model is pre-trained through multiple hand images with hand key point annotations.

[0184] In some embodiments, the preset matching conditions include a fuzzy matching condition and an exact matching condition;

[0185] The matching unit 140 is specifically configured to:

[0186] Perform fuzzy matching on the first gesture information and each second gesture information respectively, and determine the second gesture information that meets the fuzzy matching condition as the candidate gesture information;

[0187] Perform exact matching on the first gesture information and each candidate gesture information respectively, and determine the candidate gesture information that meets the exact matching condition as the target gesture information.

[0188] In some embodiments, each second gesture information and the corresponding second gesture are stored in a preset gesture configuration library;

[0189] The gesture configuration library is configured with a custom interface, and the custom interface is configured to obtain third gesture information and update each second gesture information based on the third gesture information.

[0190] In some embodiments, the device further includes:

[0191] An operation unit, configured to execute an operation corresponding to the first gesture in the current scenario in response to the type of the first gesture.

[0192] It can be understood that the gesture recognition device 100 of the embodiment of the present application can obtain the first gesture information through the first coordinates of multiple extracted hand key points, and match the first gesture information with a preset plurality of second gesture information to identify the type of the first gesture. Therefore, the addition or change of gestures can be realized by modifying the plurality of second gesture information, without a complex model training process, with a short update cycle, easy to implement, capable of flexibly matching various newly added and changed gesture types, and can better meet the needs of custom gestures.

[0193] Correspondingly, please refer to Figure 16 , Figure 16 which illustrates the hardware structure of the gesture recognition device of the driving device according to the embodiment of the present application. The embodiment of the present application further provides a gesture recognition device for a driving device, including a memory 2201 and a processor 2202. The memory 2201 is used to store programs. The processor 2202 is used to execute the programs stored in the memory 2201. When the programs stored in the memory 2201 are executed, the processor 2202 executes the gesture recognition method of the foregoing embodiments of the present application.

[0194] Correspondingly, the driving device provided by the embodiment of the present application can flexibly match various newly added and changed gesture types, and both the product development cycle and the update cycle are relatively short. It can effectively recognize gestures and can also better meet the user's customization requirements.

[0195] Correspondingly, the embodiment of the present application further provides a computer-readable storage medium. The computer-readable medium stores instructions for a computing device to execute. When the computing device executes the instructions, the gesture recognition method as described in the foregoing embodiments of the present application is implemented.

[0196] The foregoing has introduced in detail a gesture recognition method, device, driving device, and readable storage medium provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the technical solution and its core idea of the present application; those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A gesture recognition method, characterized in that, Applied to a driving device, the method includes: Obtain a to-be-recognized image containing a user's first gesture; Extract the first coordinates of multiple hand key points from the to-be-recognized image; Based on the first coordinates of the multiple hand key points, obtain first gesture information; Match the first gesture information with each pre-set second gesture information, and obtain target gesture information that meets the preset matching conditions from each of the second gesture information; Based on the second gesture corresponding to the target gesture information, obtain the type of the first gesture.

2. The gesture recognition method according to claim 1, wherein The first gesture information includes first posture parameters, and the first posture parameters include at least one parameter among the palm axis direction, the palm cutting plane direction, and the palm center orientation.

3. A gesture recognition method according to claim 2, characterized in that, The multiple hand key points include a wrist key point and the root key points of each finger; The obtaining of the first gesture information based on the first coordinates of the multiple hand key points includes: Based on the first coordinates of the wrist key point and the root key point of the middle finger, obtain the palm axis direction; Based on the first coordinates of the root key point of the index finger and the root key point of the little finger, obtain the palm cutting plane direction; Based on the first coordinates of the root key point of the index finger, the first coordinates of the root key point of the little finger, the first coordinates of the wrist key point, and the first coordinates of the root key point of the middle finger, obtain the palm center orientation.

4. A gesture recognition method according to claim 3, characterized in that, The first gesture information further includes first feature parameters, and the first feature parameters include at least one parameter among the positions of each fingertip, the distance between any two fingers, the slope of each finger, and the curvature of each finger.

5. A gesture recognition method according to claim 4, characterized in that, The obtaining of the first gesture information based on the first coordinates of the multiple hand key points further includes: Construct a palm coordinate system; Convert the first coordinates of each of the hand key points into second coordinates in the palm coordinate system; Based on the multiple second coordinates, obtain the first feature parameters.

6. A gesture recognition method according to claim 5, characterized in that The multiple hand key points further include the key points of each finger tip and the key points of each proximal interphalangeal joint of each finger; The obtaining of the first feature parameters based on the multiple second coordinates includes: Based on the second coordinates of each finger tip key point, the second coordinates of the wrist key point, and the second coordinates of each finger root key point, obtain the positions of the fingertips of each finger; Based on the second coordinates of each finger tip key point, obtain the distance between any two fingers; Based on the second coordinates of each finger tip key point, the second coordinates of each proximal interphalangeal joint key point of each finger, the second coordinates of the wrist key point, and the second coordinates of each finger root key point, obtain the curvature and the slope of each finger.

7. A gesture recognition method according to claim 5, characterized in that, The constructing of the palm coordinate system includes: Taking the wrist key point as the origin, taking the direction from the wrist key point towards the root key point of the index finger as the X-axis direction, and taking the direction from the wrist key point towards the root key point of the little finger as the Y-axis direction, obtain the palm coordinate system, and the palm coordinate system is a right-hand coordinate system.

8. A gesture recognition method according to claim 1, characterized in that, The extracting of the first coordinates of multiple hand key points from the to-be-recognized image includes: Process the image to be recognized through a hand key point recognition model to obtain the first coordinates of multiple hand key points, where the hand key point recognition model is pre-trained through multiple hand images with hand key point annotations.

9. A gesture recognition method according to claim 1, wherein The preset matching conditions include a fuzzy matching condition and an exact matching condition; The step of obtaining the target gesture information from each of the second gesture information includes: Perform fuzzy matching between the first gesture information and each of the second gesture information respectively, and determine the second gesture information that meets the fuzzy matching condition as candidate gesture information; Perform exact matching between the first gesture information and each of the candidate gesture information respectively, and determine the candidate gesture information that meets the exact matching condition as the target gesture information.

10. A gesture recognition method according to claim 1, characterized in that, Each of the second gesture information and the second gestures corresponding to the second gesture information are stored in a preset gesture configuration library; The gesture configuration library is configured with a custom interface, and the custom interface is configured to obtain third gesture information and update each of the second gesture information based on the third gesture information.

11. A gesture recognition method according to claim 1, characterized in that, The method further includes: In response to the type of the first gesture, perform an operation corresponding to the first gesture in the current scenario.

12. A gesture recognition device, characterized in that, Configured in a driving device, the device includes: An acquisition unit for acquiring an image to be recognized including a user's first gesture; An inference unit for extracting the first coordinates of multiple hand key points from the image to be recognized; A processing unit for obtaining first gesture information based on the first coordinates of multiple hand key points; A matching unit for matching the first gesture information with each of the pre-set second gesture information, and obtaining target gesture information that meets the preset matching conditions from each of the second gesture information; An identification unit for obtaining the type of the first gesture based on the second gesture corresponding to the target gesture information.

13. A gesture recognition device for a driving device, characterized in that, Includes: A memory for storing programs; A processor for executing the programs stored in the memory; When the program stored in the memory is executed, the processor executes the gesture recognition method according to any one of claims 1-11.

14. A driving device, characterized in that, Includes the gesture recognition device according to claim 12.

15. A computer-readable storage medium, characterized in that, The computer-readable medium stores instructions for a computing device to execute, and when the computing device executes the instructions, the method according to any one of claims 1-11 is implemented.