Gesture classification method and related device
By combining information from two-dimensional hand images and three-dimensional key points for gesture classification, the problem of insufficient accuracy in non-planar gesture classification is solved, and the accuracy of gesture classification is improved without increasing hardware costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2023-10-28
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies, without increasing hardware costs, have poor accuracy in classifying non-planar gestures, making it difficult to improve them effectively.
The first gesture classification information is determined based on the two-dimensional image of the hand, and the second gesture classification information is determined by combining the three-dimensional key points of the hand. Finally, the third gesture classification information of the two-dimensional image of the hand is determined based on both, and the gesture classification is performed by using the texture information of the two-dimensional image and the spatial information of the three-dimensional key points.
Without increasing hardware costs, it significantly improves the accuracy of gesture classification and can better represent the differences between planar and three-dimensional gestures.
Smart Images

Figure CN119942629B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent control, and more particularly to a gesture classification method and related equipment. Background Technology
[0002] Air gesture control on a mobile phone relies on the phone's front-facing optical sensor to acquire hand information. Based on this information, it calculates hand gestures and movement trajectories to remotely control the phone. Static gesture classification is a crucial sub-function of air gesture control; different gestures and their variations serve as input to the air gesture control decision-making logic.
[0003] Static gesture classification models rely on two-dimensional (2D) images as input and perform well in classifying planar gestures such as the palm and back of the hand. However, for non-planar gestures like clenching a fist or grasping, 2D images lose depth information in three-dimensional space and cannot accurately represent these gestures. Therefore, 2D image-based gesture classification models have poor accuracy in classifying non-planar gestures. Some mobile phones with front-facing depth sensors can use depth images as input to train gesture classification models. While depth images retain depth information compared to 2D images, deploying depth sensors requires increased hardware costs and technological investment.
[0004] Therefore, improving the accuracy of non-planar gesture classification without increasing hardware costs is an urgent problem to be solved. Summary of the Invention
[0005] This application provides a gesture classification method and related equipment, which can improve the accuracy of gesture classification without increasing hardware costs.
[0006] Firstly, a gesture classification method is provided, which can be executed by a gesture classification device or by a chip in the gesture classification device.
[0007] The above gesture classification method includes the following steps: determining first gesture classification information based on a two-dimensional image of the hand; determining second gesture classification information based on three-dimensional key points of the hand, wherein the three-dimensional key points are obtained from the two-dimensional image of the hand; and determining third gesture classification information based on the first and second gesture classification information of the two-dimensional image of the hand.
[0008] In this solution, three-dimensional key points of the hand are obtained using two-dimensional hand images, which can reduce the cost of acquiring hand depth information without increasing hardware costs. Then, a first gesture classification information is determined based on the two-dimensional hand images, and a second gesture classification information is determined based on the three-dimensional hand key points. Finally, a third gesture classification information is determined based on the first and second gesture classification information. By using the rich texture information of the two-dimensional images and the spatial information of the three-dimensional key points, it is compatible with both planar and three-dimensional gestures, which can effectively improve the accuracy of gesture classification.
[0009] In one possible implementation of the first aspect, determining the second gesture classification information based on three-dimensional keypoints of the hand includes: establishing a distance matrix based on the distance between each pair of three-dimensional keypoints of the hand; normalizing the coordinates of the three-dimensional keypoints of the hand and a first keypoint to obtain a relative coordinate vector, wherein the first keypoint is any one of the three-dimensional keypoints of the hand; and determining the second gesture classification information based on the distance matrix and the relative coordinate vector.
[0010] In this scheme, a distance matrix is established based on the distance between each pair of three-dimensional key points of the hand, and the three-dimensional key points of the hand are normalized based on the coordinates of the first key point, that is, transformed into the coordinate system of the first key point to obtain the relative coordinate vector. Then, the second gesture classification information is determined based on the distance matrix and the relative coordinate vector. This application uses the distance information and relative position information between the three-dimensional key points of the hand, which can better represent the differences in geometric space between different gestures and effectively improve the accuracy of the second gesture classification information.
[0011] In one possible implementation of the first aspect, the first gesture classification information includes a first gesture category and a corresponding first confidence level. The second gesture classification information includes a second gesture category and a corresponding second confidence level; the third gesture classification information includes a third gesture category and a corresponding third confidence level. Determining the third gesture classification information of a two-dimensional hand image based on the first and second gesture classification information includes: when the first and second gesture categories are the same, determining either the first or second gesture category as the third gesture category; and determining the third confidence level as the linear weighted value of the first confidence level and the first weight, and the second confidence level and the second weight.
[0012] In this application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are the same, the first gesture category or the second gesture category is determined as the third gesture category, and a linear weighted operation is performed based on the first confidence level and the first weight, the second confidence level and the second weight to determine the third confidence level.
[0013] In one possible implementation of the first aspect, the determination of the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: when the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a first preset gesture category combination, the first gesture category is determined as the third gesture category, and the difference between the product of the first confidence level, the second confidence level, and the third weight is determined as the third confidence level. The first preset gesture category combination can be set according to the situation. A preset gesture category combination refers to one first gesture category and one second gesture category. For example, if the first gesture category is palm and the second gesture category is fist, then palm-fist is a first preset gesture category combination. The number of first preset gesture category combinations can be more than one.
[0014] In this application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a combination of the first preset gesture categories, the first gesture category is determined as the third gesture category, and the first confidence is "penalized" to obtain the third confidence. That is, the difference obtained by subtracting (the product of the second confidence and the third weight) from the first confidence is determined as the third confidence, making the third confidence more reasonable and credible.
[0015] In one possible implementation of the first aspect, the determination of the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: when the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of second preset gesture categories, the second gesture category is determined as the third gesture category, and the difference between the product of the second confidence level and the first confidence level and the fourth weight is determined as the third confidence level. The second preset gesture category combination can be set according to the situation, and the number of second preset gesture category combinations can be one or more.
[0016] In this application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a combination of the second preset gesture categories, the second gesture category is determined as the third gesture category, and the second confidence is "penalized" to obtain the third confidence. That is, the difference obtained by subtracting (the product of the first confidence and the fourth weight) from the second confidence is determined as the third confidence, making the third confidence more reasonable and credible.
[0017] In one possible implementation of the first aspect, the determination of the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: when the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the product of the higher confidence level and the lower confidence level among the first confidence level and the second confidence level plus the fifth weight is determined as the third confidence level. The third preset gesture category combination can be set according to the situation, and the number of second preset gesture category combinations can be one or more.
[0018] In this application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a combination of the third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the product of the higher confidence level and the lower confidence level among the first confidence level and the fifth weight is determined as the third confidence level.
[0019] In one possible implementation of the first aspect, the gesture classification method further includes: outputting third gesture classification information.
[0020] In this application, output may refer to outputting the third gesture classification information to other devices or modules for further control or processing based on the third gesture classification information; output may also refer to outputting the third gesture classification information to the user so that the user is aware of the third gesture classification information of the two-dimensional image of the hand.
[0021] Secondly, this application also provides a gesture classification device, the device including a unit or module for performing the gesture classification method described in the first aspect.
[0022] Thirdly, this application also provides a terminal device, including a processor and a memory, wherein the processor and the memory are connected, wherein the memory is used to store program code, and the processor is used to call the program code to perform the gesture classification method described in the first aspect.
[0023] Fourthly, this application also provides a computer-readable storage medium storing a computer program that is executed by a processor to implement the gesture classification method described in the first aspect.
[0024] Fifthly, this application also provides a computer program product containing instructions that, when the computer program product is run on a computer, cause the computer to execute the gesture classification method described in the first aspect.
[0025] In a sixth aspect, this application also provides a chip, the chip including a processor and a data interface, the processor reading instructions stored in a memory through the data interface to execute the gesture classification method described in the first aspect.
[0026] Optionally, as one implementation, the chip may further include a memory storing instructions, and the processor is configured to execute the instructions stored in the memory. When the instructions are executed, the processor is configured to execute the gesture classification method described in the first aspect. Attached Figure Description
[0027] The accompanying drawings used in the embodiments of this application are described below.
[0028] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0029] Figure 2 This is a schematic diagram illustrating the application of a gesture classification method provided in an embodiment of this application;
[0030] Figure 3 This is a flowchart illustrating a gesture classification method provided in an embodiment of this application;
[0031] Figure 4 This is a schematic diagram illustrating the specific process of a gesture classification method provided in an embodiment of this application;
[0032] Figure 5 This is a schematic diagram of a relative coordinate vector provided in an embodiment of this application;
[0033] Figure 6 This is a schematic diagram of a gesture provided in an embodiment of this application;
[0034] Figure 7 This is a schematic diagram of the structure of a gesture classification device provided in an embodiment of this application;
[0035] Figure 8 This is a schematic diagram of another gesture classification device provided in an embodiment of this application. Detailed Implementation
[0036] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0037] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0038] In this application, "at least one" in the embodiments refers to one or more items, and "more than one" refers to two or more items. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, (a and b), (a and c), (b and c), or (a and b and c), where a, b, and c can be single or multiple. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. The step numbers in the embodiments of this application (such as step S1, step S21, etc.) are only for distinguishing different steps and do not limit the order of execution between steps.
[0039] Furthermore, unless otherwise stated, the use of ordinal numbers such as "first" and "second" in the embodiments of this application is for distinguishing multiple objects and is not for limiting the order, sequence, priority, or importance of multiple objects. For example, "first device" and "second device" are only for ease of description and do not indicate that the first device and the second device are different in structure, importance, etc. In some embodiments, the first device and the second device may also be the same device.
[0040] In the above embodiments, the term "when..." can be interpreted, depending on the context, as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". The above descriptions are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the concept and principles of this application should be included within the protection scope of this application.
[0041] To facilitate understanding, the relevant terms and concepts involved in the embodiments of this application will be introduced below.
[0042] Gesture recognition is a topic in computer science and language technology, aiming to recognize human gestures using mathematical algorithms. Gestures can originate from any body movement or state, but are typically derived from the face or hands. Users can use simple gestures to control or interact with devices without touching them. Gesture recognition can be viewed as a way for computers to understand human language, thus building a richer bridge between machines and humans than traditional text-based user interfaces or even graphical user interfaces (GUIs).
[0043] Gesture recognition enables people to communicate with machines and interact naturally without any mechanical devices. Using the concept of gesture recognition, pointing a finger at a computer screen will cause the cursor to move accordingly. Whether the gesture is static or dynamic, the recognition process first involves acquiring an image, detecting and segmenting the hand, analyzing the gesture, and then performing static or dynamic gesture recognition.
[0044] Among these, gesture classification models based on 2D images do not include depth information in their input. While their classification performance is adequate for general planar gestures, it falls short for complex gestures. Gesture classification models based on depth images include both planar and depth information in their input, but their reliance on depth information acquired by 3D sensors makes them significantly constrained by hardware limitations. Therefore, improving the accuracy of non-planar gesture classification without increasing hardware costs is a pressing issue.
[0045] To address the aforementioned technical issues, this application proposes a gesture classification method that can improve the accuracy of gesture classification without increasing hardware costs. In this embodiment, only a two-dimensional image of the hand is needed to achieve high-precision hand gesture classification.
[0046] The gesture classification method in this application embodiment can be executed by a gesture classification device or by a chip in the gesture classification device.
[0047] The aforementioned gesture classification device can be a mobile phone, desktop computer, tablet computer, wearable device (such as a smart bracelet, smartwatch), television, AR / VR, robot, robotic arm, smart home device, monitoring equipment, vehicle terminal, vehicle autonomous driving system, drone, and other electronic devices. The embodiments of this application do not limit the specific technology or device form used in the gesture classification device. The gesture classification device in the embodiments of this application can be any electronic device with gesture recognition function.
[0048] The gesture classification method of this application can be applied to any scenario requiring gesture recognition, such as the air gesture function of a mobile phone, or any other human-computer interaction scenario that requires the recognition of static gestures. For example, in smart products that require gesture control, such as large-screen TVs and head-mounted devices.
[0049] The exemplary electronic device provided in the embodiments of this application will be introduced first below.
[0050] Figure 1 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application.
[0051] The following uses an electronic device as an example to illustrate the embodiments in detail. It should be understood that the electronic device may have more than Figure 1The more or fewer components shown can be combined into two or more components, or they can have different component configurations. Figure 1 The various components shown can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0052] The electronic device may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, and a multispectral sensor (not shown), etc.
[0053] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0054] The controller can serve as the nerve center and command center of an electronic device. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.
[0055] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0056] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0057] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL).
[0058] The I2S interface can be used for audio communication.
[0059] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals.
[0060] The UART interface is a general-purpose serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication.
[0061] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193. The MIPI interface includes the camera serial interface (CSI) and the display serial interface (DSI).
[0062] The GPIO interface can be configured via software. The GPIO interface can be configured as either control signals or data signals.
[0063] The SIM interface can be used to communicate with the SIM card interface 195 to transmit data to or read data from the SIM card.
[0064] USB interface 130 is an interface that conforms to the USB standard specification, specifically it can be a Mini USB interface, Micro USB interface, USB Type C interface, etc.
[0065] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a limitation on the structure of the electronic device. In other embodiments of this application, the electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0066] The charging management module 140 is used to receive charging input from the charger.
[0067] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110 to provide power to external memory, display 194, camera 193 and wireless communication module 160, etc.
[0068] The wireless communication function of electronic devices can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0069] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in an electronic device can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.
[0070] The mobile communication module 150 can provide solutions for wireless communication applications in electronic devices, including 2G / 3G / 4G / 5G. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1.
[0071] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0072] The wireless communication module 160 can provide solutions for wireless communication applications in electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), infrared (IR) technology, and other wireless communication solutions.
[0073] In some embodiments, antenna 1 of the electronic device is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling the electronic device to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), etc.
[0074] Electronic devices implement display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0075] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include one or N displays 194, where N is a positive integer greater than 1.
[0076] Electronic devices can achieve shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0077] The ISP (Image Signal Processor) processes data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light signals are transmitted through the lens to the camera's photosensitive element. The light signals are converted into electrical signals, and the camera's photosensitive element transmits these electrical signals to the ISP for processing, transforming them into an image visible to the naked eye. The ISP can also perform algorithmic optimizations on image noise, brightness, etc. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be located within the camera 193. This photosensitive element can also be referred to as an image sensor.
[0078] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1.
[0079] A digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals.
[0080] Video codecs are used to compress or decompress digital video. Electronic devices can support one or more video codecs. This allows the electronic device to play or record video in various encoded formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0081] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as gesture recognition, image recognition, facial recognition, speech recognition, and text understanding.
[0082] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device.
[0083] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area.
[0084] The electronic device can implement audio functions through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor. Examples include music playback and recording. In this embodiment, the electronic device may include n microphones 170C, where n is a positive integer greater than or equal to 2.
[0085] The audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal.
[0086] The 180L ambient light sensor is used to detect ambient light levels. Electronic devices can adaptively adjust the brightness of the display screen 194 based on the detected ambient light level. The 180L ambient light sensor can also be used to automatically adjust the white balance when taking photos.
[0087] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can be corresponding to touch operations applied to different applications (such as taking photos, playing audio, etc.).
[0088] In this embodiment of the application, the processor 110 can call computer instructions stored in the internal memory 121 to cause the electronic device to execute the gesture classification method in this embodiment of the application.
[0089] The following uses a mobile phone as an example to describe the application process of the gesture classification method in this application:
[0090] For example, refer to Figure 2 , Figure 2 This is a schematic diagram illustrating the application of a gesture classification method provided in this application embodiment; when user Y is using a mobile phone and the air unlock function is enabled, assuming the current screen of the phone is as follows... Figure 2 As shown in (B), user Y makes a hand gesture in front of the phone's camera so that the camera can capture a two-dimensional image of the hand. The phone processes the two-dimensional hand image according to the gesture classification method of this application embodiment to obtain gesture classification information of the two-dimensional hand image. This gesture classification information includes the gesture category and the corresponding confidence level. The phone compares the gesture category with the preset screen unlock gesture category. When the two are the same, the phone unlocks the screen and enters the unlock page. Otherwise, it remains on the waiting page. For example, assuming the gesture category of the current user Y's gesture is a fist, such as... Figure 2 As shown in (C), and the preset screen gesture for unlocking is also a fist, then the phone is controlled to enter the unlock page, as shown in (C). Figure 2 As shown in (D) in the diagram.
[0091] For example, suppose a user has enabled the air screenshot function on their phone. The user makes a gesture in front of the phone's camera so that the camera can capture a two-dimensional image of the hand. The phone processes the two-dimensional hand image according to the gesture classification method of this application embodiment to obtain gesture classification information for the two-dimensional hand image. This gesture classification information includes a gesture category and a corresponding confidence level. The phone compares this gesture category with a preset air screenshot gesture category. If they match, the phone captures the current screen page. Otherwise, no response is made. For example, suppose the gesture category of the user's gesture in front of the camera is "palm," and the preset air screenshot gesture category is also "palm," then the phone captures the current screen page to obtain a screenshot.
[0092] The gesture classification method of this application embodiment will be described in detail below.
[0093] In this embodiment of the application, the gesture classification method is implemented by a gesture classification device.
[0094] refer to Figure 3 , Figure 3This is a flowchart illustrating a gesture classification method provided in an embodiment of this application; the gesture classification method 300 includes the following steps:
[0095] 301. The gesture classification device determines the first gesture classification information based on a two-dimensional image of the hand.
[0096] Specifically, the first gesture classification information indicates at least the gesture category of the two-dimensional image of the hand and the confidence level corresponding to the gesture category.
[0097] 302. The gesture classification device determines the second gesture classification information based on the three-dimensional key points of the hand.
[0098] refer to Figure 4 , Figure 4 This is a schematic flowchart illustrating a gesture classification method provided in this application embodiment; the aforementioned three-dimensional key points of the hand are obtained based on a two-dimensional image of the hand. For example, a three-dimensional key point extraction model can be used to process the two-dimensional image of the hand to predict the aforementioned three-dimensional key points. For instance, the three-dimensional key point extraction model outputs the three-dimensional key points of the hand in sequence, such as the index of the key point at the base of the palm being 0, while the index of the key point at the top of the thumb is 20.
[0099] Specifically, the second gesture classification information indicates at least the gesture category of the two-dimensional image of the hand and the confidence level corresponding to the gesture category.
[0100] 303. The gesture classification device determines the third gesture classification information of the two-dimensional image of the hand based on the first gesture classification information and the second gesture classification information.
[0101] In this embodiment, three-dimensional key points of the hand are obtained using a two-dimensional image of the hand, which can reduce the cost of acquiring hand depth information without increasing hardware costs. Then, a first gesture classification information is determined based on the two-dimensional image of the hand, and a second gesture classification information is determined based on the three-dimensional key points of the hand. Finally, a third gesture classification information is determined based on the first and second gesture classification information. By using the rich texture information of the two-dimensional image and the spatial information of the three-dimensional key points, it is compatible with planar gestures and three-dimensional gestures, which can effectively improve the accuracy of gesture classification.
[0102] In one possible implementation, the gesture classification method 300 further includes:
[0103] The gesture classification device outputs third gesture classification information.
[0104] In this embodiment of the application, the output may refer to outputting the third gesture classification information to other devices or modules so as to perform other control or processing based on the third gesture classification information; taking a mobile phone as an example, such as controlling the unlocking of the mobile phone or controlling the screenshot of the mobile phone.
[0105] Output can also refer to providing the user with third-party gesture classification information, so that the user is aware of the third-party gesture classification information of the two-dimensional image of the hand. The specific methods of providing this output to the user can include display output or voice output, etc.
[0106] In one possible implementation, refer to Figure 4 In step 301 above, the gesture classification device can first preprocess the two-dimensional hand image to obtain a two-dimensional hand image that meets uniform requirements (which can be set according to actual conditions). Assuming the two-dimensional hand image is a grayscale image, the preprocessing includes resizing and padding to adjust the grayscale image to a uniform size. Alternatively, if the two-dimensional hand image is a color image, the preprocessing, in addition to the aforementioned resizing and padding, also includes converting the color image to a grayscale image.
[0107] Further, refer to Figure 4 In step 301 above, the gesture classification device uses a first gesture recognition model to process the preprocessed two-dimensional hand image to obtain first gesture classification information. The first gesture recognition model uses the texture features of the two-dimensional hand image to recognize hand gestures.
[0108] For example, the first gesture recognition model could be a MobileNet, ShuffleNet, or similar model. MobileNet is a lightweight convolutional neural network whose main goal is to minimize model size and computational complexity while maintaining model accuracy. MobileNet's design philosophy is to use depthwise separable convolutional layers instead of traditional convolutional layers to reduce computation and model size. MobileNet's depthwise separable convolutional layers consist of depthwise convolutional layers and pointwise convolutional layers. Depthwise convolutional layers consider only the spatial relationships within each channel, while pointwise convolutional layers consider only the channel relationships at each location. This separation allows MobileNet to learn spatial and channel features with fewer parameters and less computation, thereby reducing model size and computational complexity.
[0109] The main idea behind ShuffleNet is to reduce computation and model size by using channel rearrangement and group convolution. Channel rearrangement divides the input channels into different groups, performs convolution on each group, and then merges the results. This method reduces computation while increasing the interaction between channels, helping to improve model accuracy. ShuffleNet also employs a special group convolution operation called "channel-wise group convolution," which maintains the relationship between each channel while reducing computation. Channel-wise group convolution splits each input channel into multiple subgroups, performs convolution on each subgroup, and then merges the subgroups to obtain the output channel.
[0110] In one possible implementation, refer to Figure 4 Step 302 above includes:
[0111] 321. Gesture classification devices establish a distance matrix based on the distance between the three-dimensional key points of each pair of hands.
[0112] For example, for the three-dimensional key point p of the hand a and p b p a and p b The Euclidean distance between the two is expressed as:
[0113]
[0114] The distance matrix established based on the pairwise Euclidean distances between all 3D keypoints (assuming there are N keypoints) in the 2D image of the hand is as follows:
[0115]
[0116] 322. The gesture classification device normalizes the coordinates of the three-dimensional key points of the hand and the first key point to obtain a relative coordinate vector. The first key point is any one of the three-dimensional key points of the hand.
[0117] For example, assuming the first keypoint is the keypoint with index 0 in the 3D keypoints of the hand, the other keypoints are normalized using the first keypoint as the origin to obtain a relative coordinate vector. (Reference) Figure 5 , Figure 5 This is a schematic diagram of a relative coordinate vector provided in an embodiment of this application; there are 21 three-dimensional key points of the hand in the two-dimensional image of the hand, and the indices of the key points are 0, 1, 2, 3, ..., 20. Taking the key point with index 0 as the origin, the key points with indices 1 to 20 are transformed into the coordinate system of the first key point to obtain the relative coordinate vector.
[0118] 323. The gesture classification device determines the second gesture classification information based on the distance matrix and the relative coordinate vector.
[0119] Specifically, the gesture classification device uses a second gesture recognition model to process the distance matrix and relative coordinate vector to obtain second gesture classification information. The second gesture recognition model uses the spatial information of three-dimensional key points to perform gesture recognition.
[0120] For example, the second gesture recognition model can be a convolutional neural network (CNN) model, a multilayer perceptron, etc.
[0121] In this embodiment, the distance and relative position information between three-dimensional key points of the hand are used, which can better characterize the differences in geometric space between different gestures and effectively improve the accuracy of the second gesture classification information.
[0122] refer to Figure 4 The gesture classification device determines the final third gesture classification information of the two-dimensional hand image based on the first and second gesture classification information to complete the gesture recognition and classification.
[0123] In one possible implementation, the first gesture classification information includes a first gesture category and a corresponding first confidence level. The first gesture category is the gesture category with the highest confidence level output by the first gesture recognition model. The second gesture classification information includes a second gesture category and a corresponding second confidence level. The second gesture category is the gesture category with the highest confidence level output by the second gesture recognition model. The third gesture classification information includes a third gesture category and a corresponding third confidence level.
[0124] For example, refer to Figure 6 , Figure 6 This is a schematic diagram of a gesture provided in an embodiment of this application; the first gesture category or the second gesture category includes the palm (e.g., hand). Figure 6 (as shown in (A)), back of hand, fist (as shown in (A)). Figure 2 (as shown in (C)), pinch open (as shown in the middle) Figure 6 (as shown in (B)), pinch closed, OK, thumb up (as shown in (B)). Figure 6 (as shown in (C)) and the hook (as shown in (C)). Figure 6 (as shown in (D) in the text), and others.
[0125] Accordingly, step 303 above includes:
[0126] When the first gesture category and the second gesture category are the same, the gesture classification device determines the first gesture category or the second gesture category as the third gesture category, and determines the linear weighted value of the first confidence level and the first weight, the second confidence level and the second weight as the third confidence level.
[0127] In this embodiment, when the gesture categories obtained based on the two-dimensional image and the three-dimensional keypoints are the same, the first gesture category or the second gesture category is determined as the third gesture category. A linear weighted calculation is then performed based on the first confidence level and the first weight, and the second confidence level and the second weight to determine the third confidence level, thus obtaining a result with a higher confidence level. The specific values of the first weight and the second weight can be set according to the situation, such that the third confidence level is less than or equal to one, without any particular limitation. For example, both the first weight and the second weight are less than or equal to 0.5. Another example is that the first weight is 1, and the second weight is less than or equal to 0.5. The second weight can be set to different values depending on the first gesture category or the second gesture category. For example, when the first gesture category or the second gesture category is palm, back of hand, fist, or others, the second weight is 0.1. When the first gesture category or the second gesture category is pinch open or pinch close, the second weight is 0.2.
[0128] In one possible implementation, step 303 above further includes:
[0129] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the first preset gesture categories, the gesture classification device determines the first gesture category as the third gesture category, and determines the difference between the product of the first confidence level, the second confidence level and the third weight as the third confidence level.
[0130] The specific value of the aforementioned third weight can be set according to the actual situation. For example, the third weight can be less than or equal to 0.5. Alternatively, the third weight can be set to different values depending on the type of the second gesture. For instance, if the second gesture type is palm, back of hand, fist, or others, the third weight is 0.3. If the second gesture type is pinch open or pinch close, the third weight is 0.5.
[0131] The first preset gesture category combination can be set as needed. A preset gesture category combination refers to a first gesture category and a second gesture category. For example, if the first gesture category is palm and the second gesture category is fist, then palm-fist is a first preset gesture category combination. There can be more than one first preset gesture category combination.
[0132] In this embodiment, when the gesture categories obtained based on the two-dimensional image and the three-dimensional keypoints are different, and the first gesture category and the second gesture category are a combination of the first preset gesture categories, the first gesture category is determined as the third gesture category, and the first confidence level is "penalized" to obtain the third confidence level. That is, the difference between the first confidence level and (the product of the second confidence level and the third weight) is determined as the third confidence level, making the third confidence level more reasonable and credible. For example, if the original first confidence level is 0.9, after the penalty, the third confidence level is 0.8, making the confidence level more reasonable and credible.
[0133] In one possible implementation, step 303 above further includes:
[0134] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the second preset gesture categories, the gesture classification device determines the second gesture category as the third gesture category, and determines the difference between the product of the second confidence level, the first confidence level, and the fourth weight as the third confidence level.
[0135] The specific value of the fourth weight can be set according to the actual situation. For example, the fourth weight can be less than or equal to 0.5. Alternatively, the fourth weight can be set to different values depending on the first gesture category. For instance, if the first gesture category is "pinch open" or "pinch close," the fourth weight is 0.3. If the first gesture category is "palm," "back of hand," "fist," or others, the fourth weight is 0.4.
[0136] The second preset gesture category combination can be set as needed, and there can be more than one second preset gesture category combination. The first preset gesture category combination and the second preset gesture category combination are different.
[0137] In this embodiment of the application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a combination of the second preset gesture categories, the second gesture category is determined as the third gesture category, and the second confidence is "penalized" to obtain the third confidence. That is, the difference obtained by subtracting (the product of the first confidence and the fourth weight) from the second confidence is determined as the third confidence, so that the third confidence is more reasonable and credible.
[0138] In this embodiment of the application, when the first gesture category and the second gesture category are different, one of the first gesture category and the second gesture category is selected as the final category result, and the confidence value of the selected category is subtracted from the confidence value of the discarded category and multiplied by a weight to achieve the purpose of punishment.
[0139] In one possible implementation, step 303 above further includes:
[0140] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the third preset gesture category, the gesture classification device determines the gesture category with the higher confidence level among the first confidence level and the second confidence level as the third gesture category, and determines the difference between the product of the higher confidence level and the lower confidence level among the first confidence level and the fifth weight as the third confidence level.
[0141] The third preset gesture category combination can be set according to the situation, and the number of second preset gesture category combinations can be more than one. The third preset gesture category combination is different from the first and second preset gesture category combinations. The specific value of the fifth weight can be set according to the actual situation. For example, the fifth weight is less than or equal to 0.5. Alternatively, the fifth weight can be set to different values depending on the gesture category with the lowest confidence level. For example, if the gesture category with the lowest confidence level is pinch open or pinch close, the fifth weight is 0.2. If the gesture category with the lowest confidence level is palm, back of hand, fist, or others, the fifth weight is 0.4.
[0142] In this embodiment of the application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a combination of the third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the product of the higher confidence level and the lower confidence level among the first confidence level and the fifth weight is determined as the third confidence level.
[0143] In summary, when deciding on the final third gesture category, this embodiment of the application can use customized decision logic to select a more reliable classification result based on the shape characteristics of different gesture categories, thereby improving the accuracy of gesture classification.
[0144] For example, refer to Table 1 below. Table 1 is a decision logic table, and its contents can be adjusted according to the actual situation. In Table 1, "0" indicates that the classification results of the first gesture recognition model and the second gesture recognition model are the same, that is, the first gesture category and the second gesture category are the same. In this case, the first gesture category or the second gesture category is taken as the third gesture category. The first confidence, the first weight, the second confidence, and the second weight are linearly weighted to obtain the final third confidence. See formula (2) below. For example, in formula (2), the first weight is 1, and α... i This is the second weight, and the second weight is set to different values depending on the category of the first gesture.
[0145] In Table 1, “1” indicates that the classification result based on the two-dimensional image of the hand is used as the output, that is, the first gesture category is used as the third gesture category, and the third confidence is calculated according to the following formula (2).
[0146] In Table 1, “2” indicates that the classification results based on the three-dimensional key points of the hand are used as the output, that is, the second gesture category is used as the third gesture category, and the third confidence is calculated according to the following formula (2).
[0147] In Table 1, “3” indicates that the higher confidence value of the first gesture recognition model and the second gesture recognition model is used as the output, and the third confidence value is calculated according to the following formula (2).
[0148] Table 1 Decision Logic Table
[0149]
[0150] For example, the contents of Table 1 can be represented by formula (1) as follows:
[0151] Third gesture category:
[0152] classid = δ i,j argmax(out img )+(1-δ i,j argmax(out keypoint (1)
[0153] Here, argmax() represents the index of the gesture category with the highest confidence level in the model output. img The output of the first gesture recognition model represents the gesture category, out. keypoint This represents the gesture category output by the second gesture recognition model. `i` represents the index number of the gesture category output by the first gesture recognition model. For example, the gesture categories output by the first gesture recognition model include palm (e.g., represented by 000), back of hand (e.g., represented by 001), fist (e.g., represented by 010), open / closed (e.g., represented by 011), closed / closed (e.g., represented by 100), and others (e.g., represented by 101). `j` represents the index number of the gesture category output by the second gesture recognition model. For example, the gesture categories output by the second gesture recognition model include palm (e.g., represented by 000), fist (e.g., represented by 010), open / closed (e.g., represented by 011), closed / closed (e.g., represented by 100), and others (e.g., represented by 101). If the third gesture category is the same as the first gesture category, then δ... i,j The value is 1, otherwise δ i,j The value is 0. δ i,j This can be adjusted according to the actual situation. The index number of the gesture category can be set according to the actual situation, without any special restrictions.
[0154] Third confidence level:
[0155] confidence = λ i,j (max(out img )+α i max(out keypoint ))+(1-λ i,j )(confidence choosed -βconfidence discard (2)
[0156] in, β is the third or fourth weight.
[0157] confidence choosed =δ i,j max(out img )+(1-δ i,j max(out) keypoint )
[0158] confidence discard =(1-δ i,j max(out) img )+δ i,j max(out keypoint )
[0159] For example, if the first gesture category is palm and the second gesture category is palm, referring to Table 1, we know that if the third gesture category is the first gesture category, then δ i,j The value is 1, and the third gesture category classid = argmax(out img The third confidence level is max(out). img )+α i max(out keypoint ), max(out img ) is the first confidence level.
[0160] max(out keypoint () represents the second confidence level.
[0161] For example, if the first gesture category is palm and the second gesture category is pinching open, referring to Table 1, we can determine that the final third gesture category is palm, then δ i,j The value is 1, and the third gesture category classid = argmax(out img The third confidence level is calculated as the confidence level of the palm - β * the confidence level of the hand being clenched and opened, where β is the third weight.
[0162] For example, if the first gesture category is fist and the second gesture category is clenched fist, referring to Table 1, we can determine that the final third gesture category is clenched fist. Therefore, δ i,j The value is 0, and the third gesture category classid = argmax(out keypoint The third confidence level is the confidence level of the pinched closure - β * the confidence level of the fist, where β is the fourth weight.
[0163] Compared to existing technologies, this invention addresses the issue from a different perspective, improving the classification accuracy of static gestures by relying solely on 2D hand images without increasing hardware costs. First, 3D hand keypoints are obtained from the 2D hand image through a deep learning model regression. Considering that the distance between each 3D keypoint varies for different gestures, and that the physical coordinate offsets of other points relative to the first keypoint also differ, a distance matrix vector and a relative coordinate vector are constructed based on the 3D hand keypoints and used as input to a second gesture recognition model. The 2D hand image serves as input to the first gesture recognition model. The outputs of both models are ultimately processed through a decision logic to obtain the final result. In summary, this invention relies solely on the 2D hand image while also utilizing the 3D hand keypoint information estimated by the 3D keypoint extraction model, combining both to improve the classification accuracy of non-planar gestures.
[0164] The methods of the embodiments of this application have been described in detail above. The apparatus provided by the embodiments of this application is described below.
[0165] Figure 7 A schematic diagram of a possible device provided in an embodiment of this application. Wherein, Figure 7 The gesture classification device shown can be used to implement the functions of the gesture classification method embodiments described above, and therefore can also achieve the beneficial effects of the gesture classification method embodiments described above. In the embodiments of this application, the gesture classification device can be an electronic device, or it can be a module (such as a chip) applied in an electronic device.
[0166] like Figure 7 As shown, the gesture classification device 700 includes a determining module 701 and a classification module 702. The gesture classification device 700 is used to implement the above... Figure 3 The gesture classification method embodiment shown herein functions as described. Alternatively, the gesture classification device 700 may include components for implementing the above. Figure 3 Any module of a function or operation in the gesture classification method embodiment shown can be implemented in whole or in part by software, hardware, firmware or any combination thereof.
[0167] When gesture classification device 700 is used to implement Figure 3In the illustrated method embodiment, the determining module 701 is used to determine first gesture classification information based on a two-dimensional image of the hand. The determining module 701 is also used to determine second gesture classification information based on three-dimensional key points of the hand, which are obtained from the two-dimensional image of the hand. The classification module 702 is used to determine third gesture classification information of the two-dimensional image of the hand based on the first and second gesture classification information.
[0168] In this embodiment, three-dimensional key points of the hand are obtained using a two-dimensional image of the hand, which can reduce the cost of acquiring hand depth information without increasing hardware costs. Then, a first gesture classification information is determined based on the two-dimensional image of the hand, and a second gesture classification information is determined based on the three-dimensional key points of the hand. Finally, a third gesture classification information is determined based on the first and second gesture classification information. By using the rich texture information of the two-dimensional image and the spatial information of the three-dimensional key points, it is compatible with planar gestures and three-dimensional gestures, which can effectively improve the accuracy of gesture classification.
[0169] In one possible implementation, refer to Figure 7 The gesture classification device 700 also includes an output module 703 for outputting third gesture classification information.
[0170] In one possible implementation, the determining module 701 is specifically used for determining the second gesture classification information based on three-dimensional key points of the hand as follows:
[0171] A distance matrix is established based on the distance between the three-dimensional key points of each pair of hands.
[0172] The coordinates of the three-dimensional keypoints of the hand and the first keypoint are normalized to obtain a relative coordinate vector. The first keypoint is any one of the three-dimensional keypoints of the hand.
[0173] The second gesture classification information is determined based on the distance matrix and the relative coordinate vector.
[0174] In one possible implementation, the first gesture classification information includes a first gesture category and a corresponding first confidence level. The second gesture classification information includes a second gesture category and a corresponding second confidence level; the third gesture classification information includes a third gesture category and a corresponding third confidence level.
[0175] The aforementioned classification module 702 is specifically used for:
[0176] When the first gesture category and the second gesture category are the same, the first gesture category or the second gesture category is determined as the third gesture category, and the linear weighted value of the first confidence level and the first weight, the second confidence level and the second weight is determined as the third confidence level.
[0177] In one possible implementation, the classification module 702 is further specifically used for:
[0178] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the first preset gesture categories, the first gesture category is determined as the third gesture category, and the difference between the product of the first confidence level, the second confidence level, and the third weight is determined as the third confidence level.
[0179] The first preset gesture category combination can be set as needed. A preset gesture category combination refers to a first gesture category and a second gesture category. For example, if the first gesture category is palm and the second gesture category is fist, then palm-fist is a first preset gesture category combination. There can be more than one first preset gesture category combination.
[0180] In one possible implementation, the classification module 702 is further specifically used for:
[0181] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the second preset gesture categories, the second gesture category is determined as the third gesture category, and the difference between the product of the second confidence level, the first confidence level, and the fourth weight is determined as the third confidence level.
[0182] The second preset gesture category combination can be set according to the situation, and the number of the second preset gesture category combinations can be more than one.
[0183] In one possible implementation, the classification module 702 is further specifically used for:
[0184] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the product of the higher confidence level and the lower confidence level among the first confidence level and the fifth weight is determined as the third confidence level.
[0185] The third preset gesture category combination can be set according to the situation, and the number of the second preset gesture category combinations can be more than one.
[0186] For a description of each of the above modules, please refer to the description in the foregoing embodiments, which will not be repeated here.
[0187] refer to Figure 8 , Figure 8This is a schematic diagram of another gesture classification device provided in an embodiment of this application. The gesture classification device 800 includes a memory 801, a processor 802, a communication interface 804, and a bus 803. The memory 801, processor 802, and communication interface 804 are interconnected via the bus 803.
[0188] Optionally, the gesture classification device 800 further includes a display screen (not shown), which is communicatively connected to the memory 801, processor 802, and communication interface 804 via a bus 803. The display screen is used to output information and interact with the user. Optionally, the gesture classification device further includes an output module (not shown), which is communicatively connected to the memory 801, processor 802, and communication interface 804 via a bus 803. The output module is used to output audio. This output module can be a speaker. For example, after the gesture classification device 800 determines the third gesture classification information, it can output the third gesture classification information through the output module, such as voice output or display output.
[0189] The memory 801 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 801 may store a program, and when the program stored in the memory 801 is executed by the processor 802, the processor 802 and the communication interface 804 are used to execute the various steps of the gesture classification method of any embodiment of this application.
[0190] The processor 802 may be a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), graphics processing unit (GPU), or one or more integrated circuits, for executing related programs to implement the functions required by the units in the gesture classification device of any embodiment of this application, or to execute the gesture classification method of any embodiment of this application.
[0191] The processor 802 can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of the gesture classification method in any embodiment of this application can be completed by the integrated logic circuitry in the hardware of the processor 802 or by instructions in software form. The processor 802 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the gesture classification method in any embodiment of this application can be directly implemented by the hardware processor, or implemented by a combination of hardware and software modules in the processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 801. Processor 802 reads the information in memory 801 and, in conjunction with its hardware, performs the functions required by the units included in the gesture classification device of any embodiment of this application, or executes the gesture classification method of any embodiment of this application.
[0192] The communication interface 804 uses a transceiver device, such as, but not limited to, a transceiver, to enable communication between the gesture classification device 800 and other devices or communication networks. For example, third gesture classification information can be sent to other devices through the communication interface 804.
[0193] Bus 803 may include a pathway for transmitting information between various components of gesture classification device 800 (e.g., memory 801, processor 802, communication interface 804). In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interface, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0194] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0195] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0196] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media can be read-only memory (ROM), random access memory (RAM), or magnetic media, such as floppy disks, hard disks, magnetic tapes, magnetic disks, or optical media, such as digital versatile discs (DVDs), or semiconductor media, such as solid state disks (SSDs).
[0197] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A gesture classification method, characterized in that, include: The first gesture classification information is determined based on the two-dimensional image of the hand; The first gesture classification information includes a first gesture category and a corresponding first confidence level; The second gesture classification information is determined based on the three-dimensional key points of the hand, which are obtained from the two-dimensional image of the hand; the second gesture classification information includes a second gesture category and a corresponding second confidence level. A third gesture classification information is determined based on the first gesture classification information and the second gesture classification information of the hand two-dimensional image; The third gesture classification information includes the third gesture category and the corresponding third confidence level; The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information includes: The first gesture category and the second gesture category are different. One of the first gesture category and the second gesture category is selected as the third gesture category. The confidence level corresponding to the selected category is subtracted from the first value to determine the third confidence level. The first value is the confidence level corresponding to the discarded category multiplied by a weight.
2. The method according to claim 1, characterized in that, The determination of the second gesture classification information based on three-dimensional key points of the hand includes: A distance matrix is established based on the distances between each pair of the aforementioned three-dimensional key points of the hand; The coordinates of the three-dimensional key points of the hand and the first key point are normalized to obtain a relative coordinate vector, where the first key point is any one of the three-dimensional key points of the hand. The second gesture classification information is determined based on the distance matrix and the relative coordinate vector.
3. The method according to claim 1 or 2, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information includes: When the first gesture category and the second gesture category are the same, the first gesture category or the second gesture category is determined as the third gesture category, and the linear weighted value of the first confidence level and the first weight, the second confidence level and the second weight is determined as the third confidence level.
4. The method according to claim 3, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the first preset gesture categories, the first gesture category is determined as the third gesture category, and the difference between the product of the first confidence level, the second confidence level, and the third weight is determined as the third confidence level.
5. The method according to claim 1 or 2, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the second preset gesture categories, the second gesture category is determined as the third gesture category, and the difference between the second confidence level and the product of the first confidence level and the fourth weight is determined as the third confidence level.
6. The method according to claim 3, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the second preset gesture categories, the second gesture category is determined as the third gesture category, and the difference between the second confidence level and the product of the first confidence level and the fourth weight is determined as the third confidence level.
7. The method according to claim 4, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the second preset gesture categories, the second gesture category is determined as the third gesture category, and the difference between the second confidence level and the product of the first confidence level and the fourth weight is determined as the third confidence level.
8. The method according to claim 1 or 2, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the higher confidence level and the product of the lower confidence level among the first confidence level and the second confidence level and the fifth weight is determined as the third confidence level.
9. The method according to claim 3, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the higher confidence level and the product of the lower confidence level among the first confidence level and the second confidence level and the fifth weight is determined as the third confidence level.
10. The method according to claim 4, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the higher confidence level and the product of the lower confidence level among the first confidence level and the second confidence level and the fifth weight is determined as the third confidence level.
11. The method according to claim 5, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the higher confidence level and the product of the lower confidence level among the first confidence level and the second confidence level and the fifth weight is determined as the third confidence level.
12. The method according to claim 6, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the higher confidence level and the product of the lower confidence level among the first confidence level and the second confidence level and the fifth weight is determined as the third confidence level.
13. The method according to claim 7, characterized in that, The step of determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a combination of the third preset gesture categories, the gesture category with the higher confidence level among the first confidence level and the second confidence level is determined as the third gesture category, and the difference between the higher confidence level and the product of the lower confidence level among the first confidence level and the second confidence level and the fifth weight is determined as the third confidence level.
14. The method according to claim 1 or 2, characterized in that, The method further includes: Output the third gesture classification information.
15. The method according to claim 3, characterized in that, The method further includes: Output the third gesture classification information.
16. The method according to claim 4, characterized in that, The method further includes: Output the third gesture classification information.
17. The method according to claim 5, characterized in that, The method further includes: Output the third gesture classification information.
18. The method according to claim 6, characterized in that, The method further includes: Output the third gesture classification information.
19. The method according to claim 7, characterized in that, The method further includes: Output the third gesture classification information.
20. The method according to claim 8, characterized in that, The method further includes: Output the third gesture classification information.
21. The method according to claim 9, characterized in that, The method further includes: Output the third gesture classification information.
22. The method according to claim 10, characterized in that, The method further includes: Output the third gesture classification information.
23. The method according to claim 11, characterized in that, The method further includes: Output the third gesture classification information.
24. The method according to claim 12, characterized in that, The method further includes: Output the third gesture classification information.
25. The method according to claim 13, characterized in that, The method further includes: Output the third gesture classification information.
26. A gesture classification device, characterized in that, The device includes a unit or module for performing the gesture classification method according to any one of claims 1 to 25.
27. A gesture classification device, characterized in that, The device includes a processor and a memory, wherein the processor and the memory are connected together, wherein the memory is used to store program code, and the processor is used to call the program code to execute the gesture classification method as described in any one of claims 1 to 25.
28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement the gesture classification method as described in any one of claims 1 to 25.