Gesture classification method and related equipment

By combining the two-dimensional hand image and three-dimensional key points information, the problem of low non-planar gesture classification accuracy in the prior art is solved, and the effect of improving gesture classification accuracy without increasing hardware costs is achieved.

CN119942629AActive Publication Date: 2025-05-06HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311433623.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-28
Publication Date
2025-05-06
Estimated Expiration
2043-10-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the classification accuracy of non-planar gestures without increasing hardware costs, especially because the 2D image lacks depth information, resulting in poor classification effect of non-planar gestures.

Method used

By using the two-dimensional hand image to determine the first gesture classification information, and determining the second gesture classification information based on the three-dimensional key points of the hand, combining the texture information of the two-dimensional image and the spatial information of the three-dimensional key points, determining the third gesture classification information of the two-dimensional image.

Benefits of technology

Without increasing hardware costs, the accuracy of gesture classification is improved, and it can better be compatible with plane and three-dimensional gestures, and the classification accuracy of non-planar gestures is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942629A_ABST
    Figure CN119942629A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a gesture classification method and related equipment, and relates to the field of intelligent control. According to the method, the three-dimensional key points of the hand are obtained by using the two-dimensional image of the hand, so that the cost of obtaining the depth information of the hand can be reduced on the premise of not increasing the hardware cost; determining first gesture classification information based on the hand two-dimensional image, and determining second gesture classification information based on the hand three-dimensional key point; third gesture classification information of the hand two-dimensional image is determined based on the first gesture classification information and the second gesture classification information, rich texture information of the two-dimensional image and space information of three-dimensional key points are used, plane gestures and three-dimensional gestures can be compatible, and the gesture classification precision can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of intelligent control, and in particular to a gesture classification method and related equipment. Background Art

[0002] The air gesture operation function of the mobile phone relies on the optical sensor on the front of the mobile phone to obtain hand information, and then calculates the hand gestures and movement trajectories based on the hand information, so as to achieve the purpose of remote control of the mobile phone. The static gesture classification function is an important sub-item of the air gesture operation function. Different gestures and gesture change information are the input information of the air gesture operation decision logic.

[0003] Among them, the static gesture classification model algorithm relies on the input of two-dimensional (2D) images, and has a good classification effect on planar gestures such as palms and backs of hands. For non-planar gestures such as fisting and grabbing, since 2D images lose the depth information in three-dimensional (3D) space, they cannot represent non-planar gestures well, so the gesture classification model based on 2D images has poor classification accuracy for non-planar gestures. Some mobile phones with front-mounted depth sensors can obtain depth images as input to train gesture classification models. Compared with 2D images, although depth images retain depth information, at the hardware level, the deployment of depth sensors requires more hardware costs and technical investment.

[0004] Therefore, improving the accuracy of non-planar gesture classification without increasing hardware costs is an urgent problem to be solved. Summary of the invention

[0005] The present application provides a gesture classification method and related devices, which can improve the accuracy of gesture classification without increasing hardware costs.

[0006] In a first aspect, a gesture classification method is provided. The method can be executed by a gesture classification device, or by a chip in the gesture classification device.

[0007] The gesture classification method comprises the following steps: determining first gesture classification information based on a two-dimensional hand image; determining second gesture classification information based on three-dimensional key points of the hand, wherein the three-dimensional key points of the hand are obtained based on the two-dimensional hand image; and determining third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information.

[0008] In this solution, the three-dimensional key points of the hand are obtained by using the two-dimensional image of the hand, which can reduce the cost of obtaining the depth information of the hand without increasing the hardware cost; then the first gesture classification information is determined based on the two-dimensional image of the hand, and the second gesture classification information is determined based on the three-dimensional key points of the hand; then the third gesture classification information of the two-dimensional image of the hand is determined based on the first gesture classification information and the second gesture classification information. The rich texture information of the two-dimensional image and the spatial information of the three-dimensional key points are used to be compatible with planar gestures and three-dimensional gestures, which can effectively improve the accuracy of gesture classification.

[0009] In a possible implementation of the first aspect, the determining of the second gesture classification information based on the three-dimensional key points of the hand includes: establishing a distance matrix based on the distances between the three-dimensional key points of the hand. Performing normalization processing based on the coordinates of the three-dimensional key points of the hand and the first key points to obtain a relative coordinate vector, wherein the first key point is any one of the three-dimensional key points of the hand. Determining the second gesture classification information based on the distance matrix and the relative coordinate vector.

[0010] In this scheme, a distance matrix is ​​established based on the distance between two three-dimensional key points of the hand, and the three-dimensional key points of the hand are normalized based on the coordinates of the first key point, that is, they are converted into the coordinate system of the first key point to obtain a relative coordinate vector, and then the second gesture classification information is determined based on the distance matrix and the relative coordinate vector. This application uses the distance information and relative position information between the three-dimensional key points of the hand, which can better characterize the differences in geometric space between different gestures, and can effectively improve the accuracy of the second gesture classification information.

[0011] In a possible implementation of the first aspect, the first gesture classification information includes a first gesture category and a corresponding first confidence. The second gesture classification information includes a second gesture category and a corresponding second confidence; the third gesture classification information includes a third gesture category and a corresponding third confidence. The third gesture classification information of the two-dimensional hand image is determined based on the first gesture classification information and the second gesture classification information, including: when the first gesture category and the second gesture category are the same, the first gesture category or the second gesture category is determined as the third gesture category, and the linear weighted operation value of the first confidence and the first weight, the second confidence and the second weight is determined as the third confidence.

[0012] In the present application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are the same, the first gesture category or the second gesture category is determined as the third gesture category, and a linear weighted operation is performed based on the first confidence and the first weight, the second confidence and the second weight to determine the third confidence.

[0013] In a possible implementation of the first aspect, the third gesture classification information of the two-dimensional hand image is determined based on the first gesture classification information and the second gesture classification information, further comprising: when the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a first preset gesture category combination, the first gesture category is determined as the third gesture category, and the difference between the first confidence and the product of the second confidence and the third weight is determined as the third confidence. The first preset gesture category combination can be set according to the situation, and the preset gesture category combination refers to a first gesture category and a second gesture category. For example, the first gesture category is palm, and the second gesture category is fist, then palm-fist is a first preset gesture category combination. The number of first preset gesture category combinations can be more than one.

[0014] In the present application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a combination of the first preset gesture category, the first gesture category is determined as the third gesture category, and the first confidence level is "punished" to obtain the third confidence level, that is, the difference obtained by subtracting the first confidence level from (the product of the second confidence level and the third weight) is determined as the third confidence level, so that the third confidence level is more reasonable and credible.

[0015] In a possible implementation of the first aspect, the third gesture classification information of the two-dimensional hand image is determined based on the first gesture classification information and the second gesture classification information, further comprising: when the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a second preset gesture category combination, the second gesture category is determined as the third gesture category, and the difference between the second confidence and the product of the first confidence and the fourth weight is determined as the third confidence. The second preset gesture category combination can be set according to the situation, and the number of the second preset gesture category combinations can be more than one.

[0016] In the present application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a second preset gesture category combination, the second gesture category is determined as the third gesture category, and the second confidence level is "punished" to obtain the third confidence level, that is, the difference obtained by subtracting the second confidence level from (the product of the first confidence level and the fourth weight) is determined as the third confidence level, so that the third confidence level is more reasonable and credible.

[0017] In a possible implementation of the first aspect, the third gesture classification information of the two-dimensional hand image is determined based on the first gesture classification information and the second gesture classification information, further comprising: when the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a third preset gesture category combination, the gesture category of the one with a higher confidence in the first confidence and the second confidence is determined as the third gesture category, and the difference between the one with a higher confidence and the product of the one with a lower confidence in the first confidence and the second confidence and the fifth weight is determined as the third confidence. The third preset gesture category combination can be set according to the situation, and the number of the second preset gesture category combinations can be more than one.

[0018] In the present application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a third preset gesture category combination, the gesture category of the one with higher confidence between the first confidence and the second confidence is determined as the third gesture category, and the difference between the one with higher confidence and the product of the first confidence and the second confidence and the fifth weight is determined as the third confidence.

[0019] In a possible implementation of the first aspect, the gesture classification method further includes: outputting third gesture classification information.

[0020] In the present application, output may refer to outputting the third gesture classification information to other devices or modules so as to perform other control or processing based on the third gesture classification information; output may also refer to outputting the third gesture classification information to the user so that the user knows the third gesture classification information of the two-dimensional image of the hand.

[0021] In a second aspect, the present application further provides a gesture classification device, which includes a unit or module for executing the gesture classification method described in the first aspect.

[0022] In a third aspect, the present application also provides a terminal device, including a processor and a memory, wherein the processor and the memory are connected, wherein the memory is used to store program code, and the processor is used to call the program code according to the gesture classification method described in the first aspect.

[0023] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the gesture classification method described in the first aspect.

[0024] In a fifth aspect, the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the gesture classification method described in the first aspect.

[0025] In a sixth aspect, the present application further provides a chip, comprising a processor and a data interface, wherein the processor reads instructions stored in a memory through the data interface to execute the gesture classification method described in the first aspect.

[0026] Optionally, as an implementation method, the chip may further include a memory, in which instructions are stored, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to execute the gesture classification method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The following is an introduction to the drawings used in the embodiments of the present application.

[0028] Figure 1 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application;

[0029] Figure 2 is an application diagram of a gesture classification method provided in an embodiment of the present application;

[0030] Figure 3 is a flowchart of a gesture classification method provided in an embodiment of the present application;

[0031] Figure 4 It is a specific flow chart of a gesture classification method provided in an embodiment of the present application;

[0032] Figure 5 is a schematic diagram of a relative coordinate vector provided in an embodiment of the present application;

[0033] Figure 6 is a schematic diagram of a gesture provided in an embodiment of the present application;

[0034] Figure 7 is a structural schematic diagram of a gesture classification device provided in an embodiment of the present application;

[0035] Figure 8 It is a structural schematic diagram of another gesture classification device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The technical solution in this application will be described below in conjunction with the accompanying drawings.

[0037] In the embodiments of the present application, the words "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0038] The "at least one" mentioned in the embodiments of the present application refers to one or more, and "multiple" refers to two or more. "At least one of the following" or its similar expression refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, (a and b), (a and c), (b and c), or (a and b and c), where a, b, c can be single or multiple. "And / or" describes the association relationship of the associated objects, indicating that three relationships can exist. For example, A and / or B can be represented by: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are an "or" relationship. The serial numbers of the steps of the embodiments of the present application (such as step S1, step S21, etc.) are only for distinguishing different steps, and do not limit the order of execution between the steps.

[0039] Furthermore, unless otherwise specified, the ordinal numbers such as "first" and "second" used in the embodiments of the present application are used to distinguish multiple objects, and are not used to limit the order, timing, priority or importance of multiple objects. For example, the first device and the second device are only for the convenience of description, and do not represent the difference in structure, importance, etc. between the first device and the second device. In some embodiments, the first device and the second device can also be the same device.

[0040] In the above embodiments, the term "when..." can be interpreted as meaning "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modification, equivalent substitution, improvement, etc. made within the concept and principle of the present application shall be included in the protection scope of the present application.

[0041] To facilitate understanding, relevant terms and other related concepts involved in the embodiments of the present application are first introduced below.

[0042] Gesture recognition is a topic in computer science and language technology that aims to recognize human gestures through mathematical algorithms. Gestures can originate from any body movement or state, but usually originate from the face or hands. Users can use simple gestures to control or interact with devices without touching them. Gesture recognition can be seen as a way for computers to understand human body language, thus building a richer bridge between machines and humans than primitive text user interfaces or even graphical user interfaces (GUIs).

[0043] Gesture recognition enables people to communicate with machines and interact naturally without any mechanical devices. Using the concept of gesture recognition, one can point a finger at the computer screen so that the cursor will move accordingly. Whether the gesture is static or dynamic, the recognition sequence first requires image acquisition, hand detection and segmentation, gesture analysis, and then static or dynamic gesture recognition.

[0044] Among them, the input information of the gesture classification model based on 2D images does not contain depth information. For general planar gestures, the classification effect can meet the needs of the scene, but for complex gestures, the classification effect is poor. The input information of the gesture classification model based on depth images contains planar information and depth information, but because it relies on the depth information collected by 3D sensors, it is greatly restricted by hardware. Therefore, without increasing the hardware cost, improving the accuracy of non-planar gesture classification is an urgent problem to be solved.

[0045] Based on the above technical problems, the present application proposes a gesture classification method, which can improve the accuracy of gesture classification without increasing hardware costs. In the embodiment of the present application, only a two-dimensional image of the hand is needed to achieve high-precision hand gesture classification.

[0046] The gesture classification method of the embodiment of the present application may be executed by a gesture classification device, or by a chip in the gesture classification device.

[0047] The above-mentioned gesture classification device can be a mobile phone, desktop computer, tablet computer, wearable device (such as smart bracelet, smart watch), TV, AR / VR, robot, mechanical arm, smart home device, monitoring equipment, vehicle terminal, vehicle automatic driving system, drone and other electronic devices. The embodiments of this application do not limit the specific technology and specific device form adopted by the gesture classification device. The gesture classification device of the embodiment of this application is any electronic device with gesture recognition function.

[0048] The gesture classification method of the embodiment of the present application can be applied to any scenario where gesture recognition is required, such as in the air operation function of a mobile phone, and can also be applied to any other human-computer interaction scenario where static gestures need to be recognized. For example, in smart products such as large-screen TVs and head-mounted devices that require gesture control.

[0049] The following first introduces an exemplary electronic device provided by an embodiment of the present application.

[0050] Figure 1 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0051] The following is a detailed description of the embodiment using an electronic device as an example. It should be understood that the electronic device may have more than Figure 1More or fewer components may be shown, two or more components may be combined, or there may be a different configuration of components. Figure 1 The various components shown in the EMBODIMENTS 2000 may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.

[0052] The electronic device may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194 and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, a multispectral sensor (not shown), etc.

[0053] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0054] The controller can be the nerve center and command center of the electronic device. The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0055] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or cyclically used. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0056] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0057] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL).

[0058] The I2S interface can be used for audio communication.

[0059] The PCM interface can also be used for audio communication to sample, quantize and encode analog signals.

[0060] The UART interface is a universal serial data bus used for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication.

[0061] The MIPI interface can be used to connect the processor 110 with peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), and the like.

[0062] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal.

[0063] The SIM interface can be used to communicate with the SIM card interface 195 to implement the function of transmitting data to the SIM card or reading data in the SIM card.

[0064] The USB interface 130 is an interface that complies with USB standard specifications, and may specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc.

[0065] It is understandable that the interface connection relationship between the modules illustrated in the embodiment of the present invention is only a schematic illustration and does not constitute a structural limitation of the electronic device. In other embodiments of the present application, the electronic device may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0066] The charging management module 140 is used to receive charging input from a charger.

[0067] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110 to provide power for the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc.

[0068] The wireless communication function of the electronic device can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor.

[0069] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of the antennas.

[0070] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc., applied to electronic devices. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1.

[0071] The modem processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be sent into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After the low-frequency baseband signal is processed by the baseband processor, it is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker 170A, a receiver 170B, etc.), or displays an image or video through a display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.

[0072] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), infrared technology (IR), etc. for application in electronic devices.

[0073] In some embodiments, the antenna 1 of the electronic device is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include the global system for mobile communications (GSM), general packet radio service (GPRS), etc.

[0074] The electronic device implements the display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.

[0075] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode or an active-matrix organic light emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0076] The electronic device can realize the shooting function through ISP, camera 193, video codec, GPU, display screen 194 and application processor.

[0077] ISP is used to process the data fed back by camera 193. For example, when taking a photo, the shutter is opened, and the light signal is transmitted to the camera photosensitive element through the lens, and the light signal is converted into an electrical signal. The camera photosensitive element transmits the electrical signal to ISP for processing and converts it into an image visible to the naked eye. ISP can also perform algorithm optimization on the noise and brightness of the image. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, ISP can be set in camera 193. The photosensitive element can also be called an image sensor.

[0078] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0079] Digital signal processors are used to process digital signals. In addition to processing digital image signals, they can also process other digital signals.

[0080] Video codecs are used to compress or decompress digital videos. Electronic devices can support one or more video codecs. In this way, electronic devices can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0081] NPU is a neural network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input information and continuously self-learn. NPU can realize applications such as intelligent cognition of electronic devices, such as gesture recognition, image recognition, face recognition, voice recognition, text understanding, etc.

[0082] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device.

[0083] The internal memory 121 may be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area.

[0084] The electronic device can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. For example, music playing, recording, etc. In this embodiment, the electronic device may include n microphones 170C, where n is a positive integer greater than or equal to 2.

[0085] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals.

[0086] The ambient light sensor 180L is used to sense the brightness of the ambient light. The electronic device can adaptively adjust the brightness of the display screen 194 according to the perceived brightness of the ambient light. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures.

[0087] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects.

[0088] In the embodiment of the present application, the processor 110 can call the computer instructions stored in the internal memory 121 to enable the electronic device to execute the gesture classification method in the embodiment of the present application.

[0089] The following uses a mobile phone as an example to describe the application process of the gesture classification method of the embodiment of the present application:

[0090] For example, reference Figure 2 , Figure 2 : is an application diagram of a gesture classification method provided in an embodiment of the present application; when user Y is using a mobile phone, when the air unlocking function of the mobile phone is turned on, assuming that the current page of the mobile phone is as follows Figure 2 As shown in (B), user Y makes a gesture in front of the camera of the mobile phone so that the camera of the mobile phone can capture a two-dimensional image of the hand. The mobile phone processes the two-dimensional image of the hand according to the gesture classification method of the embodiment of the present application to obtain gesture classification information of the two-dimensional image of the hand, and the gesture classification information includes the gesture category and the corresponding confidence. The mobile phone compares the gesture category with the preset screen air unlock gesture category. When the two are the same, the screen of the mobile phone is controlled to unlock and enter the unlock page. Otherwise, it stays on the page to be unlocked. For example, assuming that the gesture category of the gesture currently made by user Y is a fist, such as Figure 2 As shown in (C) in the figure, and the preset screen air unlock gesture category is also a fist, the phone is controlled to enter the unlock page, such as Figure 2 As shown in (D) in .

[0091] As another example, suppose that the user has turned on the air screenshot function of the mobile phone, and the user makes a gesture in front of the camera of the mobile phone so that the camera of the mobile phone can capture a two-dimensional image of the hand. The mobile phone processes the two-dimensional image of the hand according to the gesture classification method of the embodiment of the present application to obtain gesture classification information of the two-dimensional image of the hand, and the gesture classification information includes the gesture category and the corresponding confidence. The mobile phone compares the gesture category with the preset air screenshot gesture category. When the comparison shows that the two are the same, the mobile phone is controlled to capture the current screen page. Otherwise, no response is made. For example, assuming that the gesture category of the gesture made by the user in front of the camera is palm, and the preset air screenshot gesture category is also palm, the mobile phone is controlled to capture the current screen page to obtain a mobile phone screenshot.

[0092] The gesture classification method of the embodiment of the present application is described in detail below.

[0093] In the embodiment of the present application, the execution subject of the gesture classification method takes a gesture classification device as an example.

[0094] refer to Figure 3 , Figure 33 is a flow chart of a gesture classification method provided in an embodiment of the present application; the gesture classification method 300 includes the following steps:

[0095] 301. A gesture classification device determines first gesture classification information based on a two-dimensional hand image.

[0096] Specifically, the first gesture classification information at least indicates the gesture category of the two-dimensional hand image and the confidence level corresponding to the gesture category.

[0097] 302. The gesture classification device determines second gesture classification information based on the three-dimensional key points of the hand.

[0098] refer to Figure 4 , Figure 4 : is a specific flow chart of a gesture classification method provided in an embodiment of the present application; the above three-dimensional key points of the hand are obtained based on the two-dimensional image of the hand. Exemplarily, the two-dimensional image of the hand can be processed by a three-dimensional key point extraction model to predict the above three-dimensional key points of the hand. For example, the three-dimensional key point extraction model outputs the three-dimensional key points of the hand in sequence, such as the key point at the base of the palm has an index of 0, and the key point at the top of the thumb has an index of 20.

[0099] Specifically, the second gesture classification information at least indicates the gesture category of the two-dimensional hand image and the confidence level corresponding to the gesture category.

[0100] 303. The gesture classification device determines third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information.

[0101] In an embodiment of the present application, the three-dimensional key points of the hand are obtained by using a two-dimensional image of the hand, so that the cost of obtaining the depth information of the hand can be reduced without increasing the hardware cost; then, the first gesture classification information is determined based on the two-dimensional image of the hand, and the second gesture classification information is determined based on the three-dimensional key points of the hand; then, the third gesture classification information of the two-dimensional image of the hand is determined based on the first gesture classification information and the second gesture classification information. The rich texture information of the two-dimensional image and the spatial information of the three-dimensional key points are used to be compatible with planar gestures and three-dimensional gestures, and the accuracy of gesture classification can be effectively improved.

[0102] In a possible implementation, the gesture classification method 300 further includes:

[0103] The gesture classification device outputs third gesture classification information.

[0104] In an embodiment of the present application, output may refer to outputting the third gesture classification information to other devices or modules to perform other control or processing according to the third gesture classification information; taking a mobile phone as an example, for example, controlling the unlocking of the mobile phone or controlling the screenshot of the mobile phone.

[0105] Outputting may also refer to outputting the third gesture classification information to the user so that the user knows the third gesture classification information of the two-dimensional hand image. The specific manner of outputting to the user may be display output or voice output.

[0106] In one possible implementation, reference Figure 4 In the above step 301, the gesture classification device may first pre-process the two-dimensional hand image to obtain a two-dimensional hand image that meets uniform requirements (which may be set according to actual conditions). Assuming that the two-dimensional hand image is a grayscale image, the pre-processing includes resizing and padding to adjust the grayscale image to an image of uniform size. For another example, when the two-dimensional hand image is a color image, the pre-processing includes converting the color image into a grayscale image in addition to the above-mentioned resizing and padding.

[0107] Further, refer to Figure 4 In the above step 301, the gesture classification device processes the preprocessed two-dimensional hand image using the first gesture recognition model to obtain first gesture classification information. The first gesture recognition model recognizes hand gestures using texture features of the two-dimensional hand image.

[0108] Exemplarily, the first gesture recognition model can be a model such as Mobilenet, ShuffleNet, etc. MobileNet is a lightweight convolutional neural network, and its main goal is to reduce the size and computational complexity of the model as much as possible while maintaining the accuracy of the model. The design idea of ​​MobileNet is to use a depth-separable convolutional layer to replace the traditional convolutional layer to reduce the amount of calculation and model size. The depth-separable convolutional layer of MobileNet is composed of a depth-wise convolutional layer and a point-by-point convolutional layer. The depth-wise convolutional layer only considers the spatial relationship within each channel, while the point-by-point convolutional layer only considers the channel relationship at each position. This separation method allows MobileNet to learn spatial and channel features with fewer parameters and computational complexity, thereby reducing the size and computational complexity of the model.

[0109] The main idea of ​​ShuffleNet is to use channel rearrangement and group convolution to reduce the amount of computation and model size. The purpose of channel rearrangement is to divide the input channels into different groups, perform convolution on each group of channels, and then merge the results. This method can increase the interaction between channels while reducing the amount of computation, which helps to improve the accuracy of the model. ShuffleNet also uses a special group convolution operation called "channel-by-channel group convolution", which can maintain the relationship between each channel while reducing the amount of computation. Channel-by-channel group convolution is achieved by splitting each input channel into multiple subgroups, performing convolution operations on each subgroup, and then merging the subgroups to obtain the output channel.

[0110] In one possible implementation, reference Figure 4 , the above step 302 includes:

[0111] 321. The gesture classification device establishes a distance matrix based on the distance between two three-dimensional key points of the hands.

[0112] For example, for the 3D key point p of the hand a and p b , p a and p b The Euclidean distance between the two is expressed as:

[0113]

[0114] The distance matrix established based on the Euclidean distance between all the three-dimensional key points of the hand in the two-dimensional hand image (assuming there are N key points) is as follows:

[0115]

[0116] 322. The gesture classification device performs normalization processing based on the coordinates of the three-dimensional key points of the hand and the first key point to obtain a relative coordinate vector, where the first key point is any key point among the three-dimensional key points of the hand.

[0117] For example, assuming that the first key point is the key point with index 0 in the three-dimensional key points of the hand, the first key point is taken as the origin, and the other key points are normalized to obtain the relative coordinate vector. Figure 5 , Figure 5 : is a schematic diagram of a relative coordinate vector provided by an embodiment of the present application; there are 21 three-dimensional key points of the hand in the two-dimensional hand image, and the indexes of the key points are 0, 1, 2, 3, ..., 20. With the key point with index 0 as the origin, the key points with indexes 1 to 20 are transformed into the coordinate system of the first key point to obtain the relative coordinate vector.

[0118] 323. The gesture classification device determines second gesture classification information based on the distance matrix and the relative coordinate vector.

[0119] Specifically, the gesture classification device processes the distance matrix and the relative coordinate vector using the second gesture recognition model to obtain second gesture classification information, and the second gesture recognition model performs gesture recognition using the spatial information of the three-dimensional key points.

[0120] Exemplarily, the second gesture recognition model may be a convolutional neural network (CNN) model, a multi-layer perceptron, or the like.

[0121] In the embodiment of the present application, the distance information and relative position information between the three-dimensional key points of the hand are used to better characterize the differences in geometric space between different gestures, and can effectively improve the accuracy of the second gesture classification information.

[0122] refer to Figure 4 The gesture classification device determines the final third gesture classification information of the hand two-dimensional image based on the first gesture classification information and the second gesture classification information to complete the gesture recognition classification.

[0123] In a possible implementation, the first gesture classification information includes a first gesture category and a corresponding first confidence. The first gesture category is a gesture category with the highest confidence output by the first gesture recognition model. The second gesture classification information includes a second gesture category and a corresponding second confidence. The second gesture category is a gesture category with the highest confidence output by the second gesture recognition model. The third gesture classification information includes a third gesture category and a corresponding third confidence.

[0124] For example, reference Figure 6 , Figure 6 is a schematic diagram of a gesture provided by an embodiment of the present application; the first gesture category or the second gesture category includes a palm (such as Figure 6 (A) in the figure), back of hand, fist (as shown in Figure 2 (C) in the figure), kneading and opening (as shown in Figure 6 (B) in the figure), pinch to close, OK, thumbs up (as shown in Figure 6 (C) in the figure), the hook finger (as shown in Figure 6 (as shown in (D)), others, etc.

[0125] Accordingly, the above step 303 includes:

[0126] When the first gesture category and the second gesture category are the same, the gesture classification device determines the first gesture category or the second gesture category as the third gesture category, and determines the linear weighted operation value of the first confidence and the first weight, the second confidence and the second weight as the third confidence.

[0127] In an embodiment of the present application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are the same, the first gesture category or the second gesture category is determined as the third gesture category, and a linear weighted operation is performed based on the first confidence and the first weight, the second confidence and the second weight to determine the third confidence, so as to obtain a result with a higher confidence. The specific values ​​of the first weight and the second weight can be set according to the situation, so that the third confidence is less than or equal to one, without special limitation. For example, the first weight and the second weight are both less than or equal to 0.5. For another example, the first weight is 1 and the second weight is less than or equal to 0.5. The second weight can be set to different values ​​according to the first gesture category or the second gesture category. For example, when the first gesture category or the second gesture category is palm, back of hand, fist or other, the second weight is 0.1. When the first gesture category or the second gesture category is pinch open or pinch closed, the second weight is 0.2.

[0128] In a possible implementation manner, the above step 303 further includes:

[0129] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a first preset gesture category combination, the gesture classification device determines the first gesture category as a third gesture category, and determines the difference between the first confidence and the product of the second confidence and the third weight as the third confidence.

[0130] The specific value of the third weight can be set according to actual conditions, for example, the third weight is less than or equal to 0.5. For another example, the third weight can be set to different values ​​according to different second gesture categories. For example, when the second gesture category is palm, back of hand, fist or others, the third weight is 0.3. When the second gesture category is pinch open or pinch close, the third weight is 0.5.

[0131] The first preset gesture category combination can be set according to the situation. The preset gesture category combination refers to a first gesture category and a second gesture category. For example, if the first gesture category is palm and the second gesture category is fist, then palm-fist is a first preset gesture category combination. The number of first preset gesture category combinations can be more than one.

[0132] In the embodiment of the present application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a combination of the first preset gesture categories, the first gesture category is determined as the third gesture category, and the first confidence is "penalized" to obtain the third confidence, that is, the difference obtained by subtracting (the product of the second confidence and the third weight) from the first confidence is determined as the third confidence, so that the third confidence is more reasonable and credible. For example, the original first confidence is 0.9, and after the penalty, the third confidence is 0.8, which makes the confidence more reasonable and credible.

[0133] In a possible implementation manner, the above step 303 further includes:

[0134] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a second preset gesture category combination, the gesture classification device determines the second gesture category as a third gesture category, and determines the difference between the second confidence and the product of the first confidence and the fourth weight as the third confidence.

[0135] The specific value of the fourth weight can be set according to actual conditions, for example, the fourth weight is less than or equal to 0.5. For another example, the fourth weight can be set to different values ​​according to different first gesture categories. For example, when the first gesture category is pinch open or pinch close, the fourth weight is 0.3. When the first gesture category is palm, back of hand, fist or others, the fourth weight is 0.4.

[0136] The second preset gesture category combination may be set according to the situation, and the number of the second preset gesture category combination may be more than one. The first preset gesture category combination and the second preset gesture category combination are different.

[0137] In an embodiment of the present application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a second preset gesture category combination, the second gesture category is determined as a third gesture category, and the second confidence level is "punished" to obtain a third confidence level, that is, the difference obtained by subtracting the second confidence level from (the product of the first confidence level and the fourth weight) is determined as the third confidence level, so that the third confidence level is more reasonable and credible.

[0138] In an embodiment of the present application, when the first gesture category and the second gesture category are different, one of the first gesture category and the second gesture category is selected as the final category result, and the confidence value of the selected category minus the confidence value of the discarded category is multiplied by a weight to achieve the purpose of punishment.

[0139] In a possible implementation manner, the above step 303 further includes:

[0140] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a third preset gesture category combination, the gesture classification device determines the gesture category of the one with higher confidence between the first confidence and the second confidence as the third gesture category, and determines the difference between the one with higher confidence and the product of the first confidence and the second confidence and the fifth weight as the third confidence.

[0141] Among them, the third preset gesture category combination can be set according to the situation, and the number of the second preset gesture category combination can be more than one. The third preset gesture category combination is different from the first preset gesture category combination and the second preset gesture category combination. The specific value of the fifth weight can be set according to the actual situation. For example, the fifth weight is less than or equal to 0.5. For another example, the fifth weight is set to different values ​​according to a gesture category with lower confidence. For example, when the gesture category with lower confidence is pinch open or pinch closed, the fifth weight is 0.2. When the gesture category with lower confidence is palm, back of hand, fist or other, the fifth weight is 0.4.

[0142] In an embodiment of the present application, when the gesture categories obtained based on the two-dimensional image and the three-dimensional key points are different, and the first gesture category and the second gesture category are a third preset gesture category combination, the gesture category of the one with higher confidence in the first confidence and the second confidence is determined as the third gesture category, and the difference between the one with higher confidence and the product of the first confidence and the second confidence and the fifth weight is determined as the third confidence.

[0143] From the above, it can be seen that when deciding the final third gesture category, in the embodiment of the present application, customized decision logic can be used to select a more reliable classification result according to the shape characteristics of different gesture categories to improve the accuracy of gesture classification.

[0144] For example, refer to Table 1 below, which is a decision logic table. The content of Table 1 can be adjusted according to actual conditions. Wherein, "0" in Table 1 indicates that the classification results of the first gesture recognition model and the second gesture recognition model are the same, that is, the first gesture category and the second gesture category are the same. At this time, the first gesture category or the second gesture category is used as the third gesture category; the first confidence, the first weight, the second confidence, and the second weight are linearly weighted to obtain the final third confidence. Referring to the following formula (2), for example, in formula (2), the first weight is 1, α i is the second weight, and the second weight is set to different values ​​according to the different first gesture categories.

[0145] "1" in Table 1 indicates that the classification result based on the two-dimensional hand image is used as the output, that is, the first gesture category is taken as the third gesture category, and the third confidence is calculated according to the following formula (2).

[0146] The "2" in Table 1 indicates that the classification result based on the three-dimensional key points of the hand is used as the output, that is, the second gesture category is taken as the third gesture category, and the third confidence is calculated according to the following formula (2).

[0147] "3" in Table 1 indicates that the higher confidence value of the first gesture recognition model and the second gesture recognition model is used as the output, and the third confidence value is calculated according to the following formula (2).

[0148] Table 1 Decision logic table

[0149]

[0150] For example, the content in Table 1 can be expressed by formula (1) as follows:

[0151] Third gesture category:

[0152] classid=δ i,j argmax(out img )+(1-δ i,j )argmax(out keypoint ) (1)

[0153] Among them, argmax() represents the index corresponding to the gesture category with the maximum confidence output by the model. out img Indicates the gesture category output by the first gesture recognition model, out keypoint represents the gesture category output by the second gesture recognition model. And i represents the index number represented by the gesture category output by the first gesture recognition model. For example, the gesture categories output by the first gesture recognition model include palm (for example, represented by 000), back of hand (for example, represented by 001), fist (for example, represented by 010), pinch open (for example, represented by 011), pinch closed (for example, represented by 100), and others (for example, represented by 101). j represents the index number represented by the gesture category output by the second gesture recognition model. For example, the gesture categories output by the second gesture recognition model include palm (for example, represented by 000), fist (for example, represented by 010), pinch open (for example, represented by 011), pinch closed (for example, represented by 100), and others (for example, represented by 101). If the third gesture category is the first gesture category, then δ i,j The value is 1, otherwise δ i,j The value is 0. i,j The index number of the gesture category can be set according to the actual situation without any special limitation.

[0154] Third confidence level:

[0155] confidence=λ i,j (max(out img )+α i max(out keypoint ))+(1-λ i,j )(confidence choosed -βconfidence discard ) (2)

[0156] in, β is the third weight or the fourth weight.

[0157] confidence choosed =δ i,j max(out img )+(1-δ i,j )max(out keypoint )

[0158] confidence discard =(1-δ i,j )max(out img )+δ i,j max(out keypoint )

[0159] For example, when the first gesture category is palm and the second gesture category is palm, it can be seen from Table 1 that if the third gesture category is the first gesture category, then δ i,j The value is 1, the third gesture category classid = argmax (out img ). The third confidence level is max(out img )+α i max(out keypoint ), max(out img ) is the first confidence level,

[0160] max(out keypoint ) is the second confidence level.

[0161] For another example, when the first gesture category is palm and the second gesture category is pinch open, it can be seen from Table 1 that the third gesture category is finally determined to be palm, then δ i,j The value is 1, the third gesture category classid = argmax (out img ). The third confidence is the confidence of the palm - β* the confidence of pinching open, where β is the third weight.

[0162] For another example, when the first gesture category is fist and the second gesture category is pinch to close, it can be seen from Table 1 that the third gesture category is finally determined to be pinch to close, then δ i,j The value is 0, the third gesture category classid = argmax (out keypoint ). The third confidence is the confidence of pinch closure - β*confidence of fist, β is the fourth weight.

[0163] Compared with the prior art, the embodiment of the present invention starts from a different angle. Without increasing the hardware cost, the classification accuracy of static gestures can be improved by relying only on the two-dimensional image of the hand. First, the three-dimensional key points of the hand are regressed by a deep learning model through the two-dimensional image of the hand. Considering that for different gestures, the distance between each three-dimensional key point is different, and at the same time, the physical coordinate offsets of other points relative to the first key point are different. Therefore, a distance matrix vector and a relative coordinate vector are constructed according to the three-dimensional key points of the hand as the input of the second gesture recognition model. The two-dimensional image of the hand is used as the input of the first gesture recognition model. The outputs of the first gesture recognition model and the second gesture recognition model are finally obtained through a decision logic to obtain the final result. In summary, the embodiment of the present application only relies on the two-dimensional image of the hand, and at the same time utilizes the three-dimensional key point information of the hand estimated by the three-dimensional key point extraction model, and combines the information of the two to improve the classification accuracy of non-planar gestures.

[0164] The method of the embodiment of the present application is described in detail above, and the device provided by the embodiment of the present application is introduced below.

[0165] Figure 7 A schematic diagram of a possible device structure provided in an embodiment of the present application. Figure 7 The gesture classification device shown can be used to implement the functions of the above-mentioned gesture classification method embodiment, and thus can also achieve the beneficial effects of the above-mentioned gesture classification method embodiment. In the embodiment of the present application, the gesture classification device can be an electronic device, or a module (such as a chip) applied to an electronic device.

[0166] like Figure 7 As shown, the gesture classification device 700 includes a determination module 701 and a classification module 702. The gesture classification device 700 is used to implement the above Figure 3 Alternatively, the gesture classification device 700 may include a device for implementing the above-mentioned Figure 3 A module for any function or operation in the gesture classification method embodiment shown in the figure may be implemented in whole or in part by software, hardware, firmware or any combination thereof.

[0167] When the gesture classification device 700 is used to implement Figure 3In the method embodiment shown, the determining module 701 is used to determine the first gesture classification information based on the two-dimensional hand image. The determining module 701 is also used to determine the second gesture classification information based on the three-dimensional key points of the hand, and the three-dimensional key points of the hand are obtained based on the two-dimensional hand image. The classification module 702 is used to determine the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information.

[0168] In an embodiment of the present application, the three-dimensional key points of the hand are obtained by using a two-dimensional image of the hand, so that the cost of obtaining the depth information of the hand can be reduced without increasing the hardware cost; then, the first gesture classification information is determined based on the two-dimensional image of the hand, and the second gesture classification information is determined based on the three-dimensional key points of the hand; then, the third gesture classification information of the two-dimensional image of the hand is determined based on the first gesture classification information and the second gesture classification information. The rich texture information of the two-dimensional image and the spatial information of the three-dimensional key points are used to be compatible with planar gestures and three-dimensional gestures, and the accuracy of gesture classification can be effectively improved.

[0169] In one possible implementation, reference Figure 7 The gesture classification device 700 further includes: an output module 703, configured to output third gesture classification information.

[0170] In a possible implementation manner, the determination module 701 is specifically used to determine the second gesture classification information based on the three-dimensional key points of the hand:

[0171] A distance matrix is ​​established based on the distance between two 3D key points of the hand.

[0172] Normalization processing is performed based on the coordinates of the three-dimensional key points of the hand and the first key point to obtain a relative coordinate vector, and the first key point is any key point among the three-dimensional key points of the hand.

[0173] Second gesture classification information is determined based on the distance matrix and the relative coordinate vector.

[0174] In a possible implementation, the first gesture classification information includes a first gesture category and a corresponding first confidence level. The second gesture classification information includes a second gesture category and a corresponding second confidence level. The third gesture classification information includes a third gesture category and a corresponding third confidence level.

[0175] The classification module 702 is specifically used for:

[0176] When the first gesture category and the second gesture category are the same, the first gesture category or the second gesture category is determined as the third gesture category, and the linear weighted operation value of the first confidence and the first weight, the second confidence and the second weight is determined as the third confidence.

[0177] In a possible implementation, the classification module 702 is further specifically configured to:

[0178] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a first preset gesture category combination, the first gesture category is determined as the third gesture category, and the difference between the first confidence and the product of the second confidence and the third weight is determined as the third confidence.

[0179] The first preset gesture category combination can be set according to the situation. The preset gesture category combination refers to a first gesture category and a second gesture category. For example, if the first gesture category is palm and the second gesture category is fist, then palm-fist is a first preset gesture category combination. The number of first preset gesture category combinations can be more than one.

[0180] In a possible implementation, the classification module 702 is further specifically configured to:

[0181] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a second preset gesture category combination, the second gesture category is determined as the third gesture category, and the difference between the second confidence and the product of the first confidence and the fourth weight is determined as the third confidence.

[0182] The second preset gesture category combination may be set according to circumstances, and the number of the second preset gesture category combination may be more than one.

[0183] In a possible implementation, the classification module 702 is further specifically configured to:

[0184] When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a third preset gesture category combination, the gesture category of the one with higher confidence between the first confidence and the second confidence is determined as the third gesture category, and the difference between the one with higher confidence and the product of the first confidence and the second confidence and the fifth weight is determined as the third confidence.

[0185] The third preset gesture category combination may be set according to circumstances, and the number of the second preset gesture category combination may be more than one.

[0186] For the introduction of the above modules, please refer to the description of the above embodiments, which will not be repeated here.

[0187] refer to Figure 8 , Figure 8800 is a schematic diagram of another gesture classification device provided in an embodiment of the present application. The gesture classification device 800 includes a memory 801, a processor 802, a communication interface 804 and a bus 803. The memory 801, the processor 802 and the communication interface 804 are connected to each other through the bus 803.

[0188] Optionally, the gesture classification device 800 further includes a display screen (not shown), which is connected to the memory 801, the processor 802, and the communication interface 804 via the bus 803. The display screen is used to output information and interact with the user. Optionally, the gesture classification device further includes an output module (not shown), which is connected to the memory 801, the processor 802, and the communication interface 804 via the bus 803, and the output module is used to output audio. The output module can be a speaker. Exemplarily, after the gesture classification device 800 determines the third gesture classification information, the third gesture classification information can be output through the output module, such as voice output or display output.

[0189] The memory 801 may be a read-only memory (ROM), a static storage device, a dynamic storage device or a random access memory (RAM). The memory 801 may store a program. When the program stored in the memory 801 is executed by the processor 802, the processor 802 and the communication interface 804 are used to execute the various steps of the gesture classification method of any embodiment of the present application.

[0190] The processor 802 can adopt a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits to execute relevant programs to implement the functions required to be performed by the units in the gesture classification device of any embodiment of the present application, or to execute the gesture classification method of any embodiment of the present application.

[0191] The processor 802 may also be an integrated circuit chip with signal processing capability. In the implementation process, each step of the gesture classification method of any embodiment of the present application may be completed by the hardware integrated logic circuit or software instructions in the processor 802. The above-mentioned processor 802 may also be a general-purpose processor, a digital signal processor (Digital Signal Processing, DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the gesture classification method in combination with any embodiment of the present application may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 801, and the processor 802 reads the information in the memory 801, and combines its hardware to complete the functions required to be performed by the units included in the gesture classification device of any embodiment of the present application, or executes the gesture classification method of any embodiment of the present application.

[0192] The communication interface 804 uses a transceiver such as but not limited to a transceiver to implement communication between the gesture classification device 800 and other devices or a communication network. For example, the third gesture classification information can be sent to other devices through the communication interface 804.

[0193] The bus 803 may include a path for transmitting information between the various components of the gesture classification device 800 (e.g., the memory 801, the processor 802, and the communication interface 804). In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0194] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0196] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that contains one or more available media integrated. The available medium can be a read-only memory (ROM), or a random access memory (RAM), or a magnetic medium, such as a floppy disk, a hard disk, a tape, a disk, or an optical medium, such as a digital versatile disc (DVD), or a semiconductor medium, such as a solid state drive (SSD), etc.

[0197] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A gesture classification method, characterized in that: include: determining first gesture classification information based on the two-dimensional hand image; Determining second gesture classification information based on the three-dimensional key points of the hand, where the three-dimensional key points of the hand are obtained based on the two-dimensional image of the hand; Third gesture classification information of the two-dimensional hand image is determined based on the first gesture classification information and the second gesture classification information.

2. The method according to claim 1, characterized in that The determining of the second gesture classification information based on the three-dimensional key points of the hand includes: Establishing a distance matrix based on the distances between the three-dimensional key points of the hand; Performing normalization processing based on the coordinates of the three-dimensional key points of the hand and the first key point to obtain a relative coordinate vector, wherein the first key point is any one of the three-dimensional key points of the hand; The second gesture classification information is determined based on the distance matrix and the relative coordinate vector.

3. The method according to claim 1 or 2, characterized in that: The first gesture classification information includes a first gesture category and a corresponding first confidence level; the second gesture classification information includes a second gesture category and a corresponding second confidence level; the third gesture classification information includes a third gesture category and a corresponding third confidence level; The determining the third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information includes: When the first gesture category and the second gesture category are the same, the first gesture category or the second gesture category is determined as the third gesture category, and the linear weighted operation value of the first confidence and first weight, the second confidence and second weight is determined as the third confidence.

4. The method according to claim 3, characterized in that The determining of third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a first preset gesture category combination, the first gesture category is determined as the third gesture category, and the difference between the first confidence and the product of the second confidence and the third weight is determined as the third confidence.

5. The method according to claim 3 or 4, characterized in that: The determining of third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a second preset gesture category combination, the second gesture category is determined as the third gesture category, and the difference between the second confidence and the product of the first confidence and a fourth weight is determined as the third confidence.

6. The method according to any one of claims 3 to 5, characterized in that: The determining of third gesture classification information of the two-dimensional hand image based on the first gesture classification information and the second gesture classification information further includes: When the first gesture category and the second gesture category are different, and the first gesture category and the second gesture category are a third preset gesture category combination, the gesture category of the higher confidence one of the first confidence and the second confidence is determined as the third gesture category, and the difference between the higher confidence one and the product of the first confidence and the lower confidence one of the second confidence and the fifth weight is determined as the third confidence.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: The third gesture classification information is output.

8. A gesture classification device, characterized in that: The device comprises a unit or a module for executing the gesture classification method according to any one of claims 1 to 7.

9. A gesture classification device, characterized in that: It comprises a processor and a memory, wherein the processor and the memory are connected, wherein the memory is used to store program codes, and the processor is used to call the program codes to execute the gesture classification method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the gesture classification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • dynamic gesture recognition method and system based on a multi-mode 3D convolutional neural network

    CN109871781A

  • Gesture recognition method and device, intelligent glasses and storage medium

    CN116416676A

  • Multi-view gesture recognition method and device, computer equipment and storage medium

    CN116665245A

  • Path recognition method, path recognition device, path recognition program, and path recognition program recording medium

    US20220004263A1