Over-the-air interaction gesture recognition method and terminal equipment

By first identifying the starting gesture of air-interval interaction at a low frame rate on the terminal device, and then identifying specific gestures at a high frame rate, the high power consumption and computing overhead problems caused by air-interval gesture recognition in the prior art are solved, and a more efficient recognition process is achieved.

CN119942628AActive Publication Date: 2025-05-06HONOR DEVICE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202311426343.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-05-06
Estimated Expiration
2043-10-27

AI Technical Summary

Technical Problem

The existing terminal devices that use air-to-air gesture interaction need to identify each frame of images after turning on the air-to-air interaction function, resulting in large power consumption and calculation overhead.

Method used

By first obtaining the first picture at a lower frame rate when the terminal device is in the screen-lit state, it is possible to identify whether the starting gesture exists. If there is, the second picture is then acquired at a higher frame rate to recognize the inter-air interactive gesture.

Benefits of technology

It reduces the duration of terminal devices identifying inter-air interactive gestures, reduces power consumption and computing overhead, and improves the accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942628A_ABST
    Figure CN119942628A_ABST
Patent Text Reader

Abstract

The invention provides an air interaction gesture recognition method and terminal equipment, and relates to the technical field of terminals, and the method is applied to the terminal equipment. And when the terminal device is in the screen-on state, the terminal device obtains the first picture at the first frame rate. In response to the fact that the first picture comprises the first gesture, the terminal equipment obtains a second picture at a second frame rate, the second frame rate is higher than the first frame rate, and then the air interaction gesture is recognized based on the gesture classification and the hand key point of each frame of second picture in the multiple frames of second pictures. The higher the frame rate is, the larger the number of pictures acquired in unit time is, and the higher the resource memory calculation power consumption required to be calculated by the terminal equipment is. Therefore, in the process, whether the terminal equipment executes the process of acquiring the second picture and identifying the air interaction gesture based on the second picture is determined based on whether the first picture comprises the first gesture, so that the process of identifying the air interaction gesture by the terminal equipment can be reduced, and the power consumption and the overhead of computing resources are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to a method and terminal device for air-to-air interactive gesture recognition. Background Art

[0002] For scenarios where it is inconvenient for users to perform contact operations on terminal devices, such as mobile phones and tablets, some terminal devices provide a way to interact using air gestures. The terminal device can capture images through a camera and recognize air gestures included in the image. When a specific air gesture is recognized, the terminal device can perform an operation corresponding to the specific air gesture. In this way, users can achieve contactless interaction with the terminal device through air gestures.

[0003] In the existing solutions using air gesture interaction, after the terminal device turns on the air interaction function, it must perform air interaction gesture recognition on each frame of the collected image, which will generate high power consumption. In addition, the process of air interaction gesture recognition will also occupy a large computing overhead. Summary of the invention

[0004] In view of this, the present application provides a method and terminal device for recognizing air-interaction gestures, which can reduce the time taken by the terminal device to recognize air-interaction gestures, thereby reducing the power consumption and computing overhead of the terminal device to recognize air-interaction gestures.

[0005] In the first aspect, the present application provides a method for identifying an air interaction gesture, which is applied to a terminal device. When the terminal device is in a screen-on state, the method includes: acquiring a first picture at a first frame rate, wherein the first picture may be captured by a camera of the terminal device. In response to the first picture including a first gesture, acquiring a second picture at a second frame rate, the second frame rate being higher than the first frame rate; and identifying an air interaction gesture based on the gesture classification and hand key points of each second frame in multiple frames of the second picture.

[0006] Among them, the starting action of the air interaction gesture is the first gesture. That is to say, when the terminal device recognizes the first gesture, it indicates that the user has the intention to interact with the terminal device through the air. Therefore, after the terminal device recognizes that the first gesture is included in the first picture, it can obtain the second picture and recognize the air interaction gesture in the second picture. In this way, the terminal device can reduce the time required to execute the recognition of the air interaction gesture in the second picture, thereby reducing the power consumption and computing overhead of the terminal device in recognizing the air interaction gesture.

[0007] In some examples, the first gesture may be a gesture unrelated to the air interaction gesture. For example, the air interaction gesture is an upward slide gesture, and the first gesture is a clenched fist. In some examples, the air interaction gesture is a dynamic gesture, and the first gesture may be the starting action for making the air interaction gesture. For example, the air interaction gesture is a downward slide gesture, and the user's initial action for making the downward slide gesture is to point the palm with the fingers facing up toward the camera, and accordingly, the first gesture is the palm with the fingers facing up. In this way, in the case of a downward slide gesture made by the user, the first picture obtained by the terminal device includes the first gesture, and the second picture obtained subsequently includes the air interaction gesture, avoiding the need for the user to make two gestures and achieving more convenient and quick air interaction.

[0008] Furthermore, the higher the frame rate of the image acquisition, the faster the processing speed of the terminal device is required to ensure that the image can be processed by the technology. Then in the above implementation, the frame rate at which the terminal device acquires the second image is higher than the frame rate at which the first image is acquired, which can ensure that the terminal device processes the first image at a lower speed and processes the second image at a higher speed. In the case where the above process can reduce the time it takes for the terminal device to process the second image, the power consumption of the terminal device can be further reduced.

[0009] After recognizing the air interaction gesture, the terminal device can perform operations corresponding to the air interaction gesture to realize air interaction between the user and the terminal device.

[0010] In a possible implementation of the first aspect, the size of the second picture is larger than the size of the first picture. The larger the size of the picture, the higher the resolution of the picture. In other words, the second picture has a higher definition than the first picture. The terminal device can identify whether the first gesture exists based on the first picture with a lower resolution, and can accurately identify the air interaction gesture based on the second picture with a higher resolution. In this way, the accuracy of the recognition result can be guaranteed.

[0011] In another possible implementation of the first aspect, based on the gesture classification and hand key points of each second frame in multiple frames of second pictures, identifying air interaction gestures, including: in response to the first second frame in the multiple frames of second pictures including a hand, based on the gesture classification and hand key points of each second frame in the multiple frames of second pictures, identifying air gestures.

[0012] In the above implementation, gesture classification can represent the state of the gesture, and the hand key points can represent the three-dimensional position of the hand in space. Therefore, the terminal device can determine the changing trend of the gesture based on each frame of the second picture in multiple frames, and accurately determine the air interaction gesture.

[0013] In a possible implementation of the first aspect, in response to the first frame of the second picture in the multiple frames of the second picture not including the hand, the step of acquiring the first picture at the first frame rate is returned. If the first frame of the second picture in the multiple frames of the second picture does not include the hand, it means that there is currently a misrecognition situation. Therefore, the terminal device can restart acquiring the first picture to avoid misrecognition causing the terminal device to increase the duration of recognizing the air interaction gesture based on the second picture, thereby reducing power consumption.

[0014] In a possible implementation of the first aspect, the method further includes: cropping a third image from each second image in the plurality of second images, the third image including a hand region, and then the terminal device can obtain a gesture classification and a hand key point based on the hand region in the third image. The size of the third image cropped from the second image is smaller, and the terminal device can perform faster recognition based on the third size, thereby reducing the recognition time, and thus, the effect of reducing power consumption can also be achieved.

[0015] In a possible implementation of the first aspect, after obtaining the hand key points of the k-th second picture in the multiple frames of second pictures, the method further includes: detecting whether the hand in the k-th second picture is complete based on the hand key points; k is a positive integer. In response to the hand in the k-th second picture being complete, obtaining a palm frame based on the hand key points in the k-th second picture.

[0016] After obtaining a palm frame based on the hand key points in the second picture of the kth frame, the terminal device can crop the second picture of the k+ith frame based on the palm frame to obtain a third picture, where i is a positive integer. The third picture is obtained by cropping the second picture of the k+ith frame based on the palm frame.

[0017] In the above implementation process, since the speed of hand movement is slower than the speed of second picture acquisition, that is, the moving distance between the hands in multiple consecutive second picture frames is small. In other words, the terminal device can consider that the hand in the k+i-th second picture frame is in the same or similar position as the hand in the k-th second picture frame. Then, when the hand in the k-th second picture frame acquired by the terminal device is complete, the terminal device will also consider that the hand in the k+i-th second picture frame is complete, and the position of the hand is the same or similar to the position of the hand in the k+i-th second picture frame. Therefore, the terminal device can directly crop the k+i-th second picture frame based on the palm frame, thereby improving the picture processing speed.

[0018] In a possible implementation manner of the first aspect, in response to an incomplete hand in the second picture of the kth frame, the terminal device may crop a third picture corresponding to the second picture of the k+ith frame based on the second picture of the k+ith frame.

[0019] In the above implementation process, when the hand in the second picture of the kth frame is incomplete, the terminal device will think that the hand in the second picture of the k+ith frame may be incomplete. Since the position of the hand in the second picture of the kth frame is the same as or similar to the position of the hand in the second picture of the k+ith frame, the terminal device can directly crop the third picture corresponding to the second picture of the k+ith frame based on the second picture of the k+ith frame, thereby improving the accuracy of image processing.

[0020] In another possible implementation manner of the first aspect, cropping the third picture from each second picture in the plurality of second pictures includes: the terminal device crops the third picture corresponding to the kth second picture based on the kth second picture. The terminal device may directly crop the third picture corresponding to the kth second picture based on the kth second picture to improve the accuracy of the third picture.

[0021] In a possible implementation manner of the first aspect, the method further includes: in response to detecting a moving action, the terminal device acquires a first picture at a first frame rate.

[0022] In a possible implementation manner of the first aspect, when the first picture has multiple frames, determining that the first picture includes the first gesture includes: determining that N consecutive first pictures in the multiple first pictures all include the first gesture, where N is a positive integer greater than 2. In this way, the terminal device can identify errors and improve robustness in actual applications by determining whether N consecutive first pictures all include the first gesture.

[0023] In a possible implementation of the first aspect, before the terminal device acquires the first picture at the first frame rate, the method further includes: when the terminal device is in a screen-off state, acquiring a fourth picture at a third frame rate, and in response to the fourth picture including a second gesture, the terminal device switches to a screen-on state, wherein the second gesture is a gesture for waking up the terminal device. In this way, when the terminal device is in a screen-off state, it can interact with the user to light up the screen, so that the user can interact with the terminal device using an air-interaction gesture, thereby improving the convenience of interaction.

[0024] In some designs, the third frame rate at which the terminal device obtains the fourth image is lower than the first frame rate. When the terminal device is in the screen-off state, the user has less need to interact with the terminal device remotely, so the terminal device can obtain the fourth image and recognize the second gesture therein at a slower speed, thereby effectively reducing the power consumption of the terminal device.

[0025] In a possible implementation manner of the first aspect, determining that the fourth picture includes the second gesture includes: determining that M consecutive fourth pictures in the plurality of fourth pictures all include the second gesture, where M is a positive integer greater than 2. In this way, the terminal device can identify errors and improve robustness in actual applications by determining whether M consecutive fourth pictures all include the second gesture.

[0026] In some implementations, when the terminal device is in a state where the screen is off, in response to detecting a movement action, the terminal device can obtain a fourth image at a third frequency. Through the above approach, the power consumption of the terminal device can be reduced.

[0027] In some implementations, the size of the fourth image is smaller than the size of the first image. The larger the size of the image, the higher the resolution of the image and the more memory it occupies. In this way, the terminal device can process the fourth image that occupies less memory faster, and only process the first image when the second gesture exists in the fourth image. In this way, the recognition speed of the terminal device can be improved, thereby reducing power consumption.

[0028] In some implementations, the air interaction gestures include a left swipe, a right swipe, a palm flip, and a two-finger pinch.

[0029] In a second aspect, the present application provides a terminal device, comprising a display screen, a memory and one or more processors; the display screen, the memory and the processor are coupled; the display screen is used to display an image generated by the processor, and the memory is used to store computer program code, wherein the computer program code comprises computer instructions; when the processor executes the computer instructions, the terminal device executes the method described in the first aspect and any possible design thereof.

[0030] In a third aspect, the present application provides a computer-readable storage medium comprising computer instructions, which, when executed on a terminal device, enables the terminal device to execute the method described in the first aspect and any possible design thereof.

[0031] In a fourth aspect, the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the method described in the first aspect and any possible design thereof.

[0032] In a fifth aspect, the present application provides a device, which is included in a terminal device, and the device has the function of implementing the behavior of the terminal device in any of the above aspects and possible implementation methods. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above function. For example, an allocation module or unit, a scanning module or unit, a recycling module or unit, a moving module or unit, and a storage module or unit, etc.

[0033] In a sixth aspect, an embodiment of the present application provides a chip system, which includes a processor and may also include a memory, for implementing any one of the methods provided in the first aspect. The chip system may be composed of a chip, or may include a chip and other discrete devices.

[0034] It can be understood that the terminal device described in the second aspect and any possible design method provided above, the computer-readable storage medium described in the third aspect, and the computer program product described in the fourth aspect are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A schematic diagram of a terminal device provided in an embodiment of the present application;

[0036] Figure 2 A schematic diagram of an air-interaction gesture control interface provided in an embodiment of the present application;

[0037] Figure 3 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application;

[0038] Figure 4 A schematic diagram of waking up a terminal device provided in an embodiment of the present application;

[0039] Figure 5 A flowchart of a method for identifying gestures in an air interaction manner provided in an embodiment of the present application;

[0040] Figure 6 A schematic diagram of a screen-off gesture provided in an embodiment of the present application;

[0041] Figure 7 A schematic diagram of waking up a terminal device provided in an embodiment of the present application;

[0042] Figure 8 A schematic diagram of a dynamic gesture change provided in an embodiment of the present application;

[0043] Fig. 9A schematic diagram of key points of a hand provided in an embodiment of the present application;

[0044] Fig.10 A schematic diagram of a flow chart of an air interaction gesture method provided in an embodiment of the present application;

[0045] Fig.11 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.

[0047] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0048] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0049] Before introducing the embodiments of the present application, the technologies involved in the embodiments of the present application are first introduced in detail.

[0050] 1. Motion detection (MD)

[0051] MD is also generally referred to as motion detection. Through MD technology, it is possible to detect whether the content of the image captured by the camera has changed. MD technology is often used for unattended surveillance recording and automatic alarm. For example, in a monitoring system, the surveillance camera can continuously collect images, and the processor of the monitoring system can use MD technology to detect whether the content of multiple frames of images has changed. Among them, someone walking in front of the surveillance camera or the lens of the surveillance camera being moved will cause the content of the image to change. When it is detected that the content of the multiple frames of images has changed, the processor of the monitoring system can make corresponding processing, such as ringing an alarm bell, issuing an alarm prompt message, etc.

[0052] In addition, mobile phones, tablets and other terminal devices can also use MD technology to assist shooting. For example, after detecting movement using MD technology, the mobile phone can adjust the shooting frame rate to take clear photos in scenes of shooting sports or fast-moving objects.

[0053] 2.mono format

[0054] The mono format is a grayscale image format. Mono format images (hereinafter referred to as mono images) can generally be output by devices such as single-channel cameras or black and white cameras. The mono format includes Mono8, Mono10, Mono10 Packed, Mono12, Mono12 Packed and other pixel formats.

[0055] In an embodiment of the present application, the terminal device can recognize air interaction gestures based on mono images.

[0056] In order to better understand the embodiments of the present application, the terminal device provided in the embodiments of the present application is first introduced.

[0057] like Figure 1 As shown, the terminal device 100 can be a mobile phone 11, a tablet computer 12, a smart screen 13, a laptop computer 14, a vehicle-mounted device, a wearable device (such as a smart watch), an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), an artificial intelligence device, and other terminal devices with air gesture interaction functions. The operating system installed in the terminal device 100 includes but is not limited to Or other operating systems. The embodiment of the present application does not limit the specific type of the terminal device 100 and the installed operating system.

[0058] In the following embodiments, the terminal device is described as a mobile phone as an example, but in actual applications, the terminal device can be any of the above devices.

[0059] Generally, different air interaction gestures correspond to different operations. For example, the air interaction gesture of sliding from bottom to top corresponds to the swipe up operation, the air interaction gesture of sliding from bottom to top corresponds to the swipe down operation, and the air interaction gesture of changing from five fingers open to five fingers clenched into a fist corresponds to the screenshot operation.

[0060] In some scenarios, the mobile phone can turn on or off the air gesture interaction function according to the user operation. In some examples, the setting application of the mobile phone includes an option for controlling the air gesture interaction function. In response to the user clicking the option, the interface for controlling the air gesture interaction function is displayed on the display screen of the mobile phone.

[0061] Among them, different air interaction gestures can correspond to different control interfaces. Figure 2 (a) shows a control interface for sliding the screen downward through the air, in which the operation guide information for sliding the screen downward through the air and the control buttons for the function of sliding the screen downward through the air are displayed. Figure 2 The air sliding screen function control button shown in (a) is in the turned-on state, indicating that the terminal device can recognize the user's downward sliding air interaction gesture in the gesture recognition area, and control the page displayed on the terminal device to slide downward after recognizing the downward sliding air interaction gesture. Figure 2 (b) shows a control page for remote screenshot, in which the operation guide information for remote screenshot and the control button for remote screenshot function are displayed. Figure 2 The air screenshot function control button shown in (b) is in the off state, indicating that the terminal device cannot recognize the air interaction gesture in the gesture recognition area.

[0062] When the air gesture interaction function of the above-mentioned mobile phone is turned on, the mobile phone in the bright screen state continuously collects images through the camera and identifies whether there is an air interaction gesture in each frame of the image to ensure that the mobile phone can timely recognize the air interaction gesture and perform the interactive action corresponding to the air interaction gesture. However, the mobile phone with the bright screen is not always in the state of air interaction with the user. For example, when the video playback application is playing a video, the screen of the mobile phone is always on. However, there is a scene where the user is not near the mobile phone at this time, that is, there is a situation where the mobile phone cannot recognize the air interaction gesture. For example, in the scene where the user is typing with the mobile phone, there is also a situation where the mobile phone will not recognize the air interaction gesture. In the above-mentioned scenarios, if the mobile phone is in the bright screen state and the air gesture interaction function is turned on, the mobile phone will continue to collect pictures and identify whether there is an air interaction gesture in each frame of the picture. In other words, the process used by the mobile phone to recognize the air interaction gesture is always in a running state. Since the recognition process of the air interaction gesture belongs to the recognition process of deep learning, the calculation process is relatively complex, which leads to high power consumption of the mobile phone and the recognition process will always occupy more computing resources.

[0063] In order to solve the above problems, an embodiment of the present application provides a method for identifying air-interaction gestures. When the terminal device is in a screen-on state, the terminal device first obtains a first picture at a lower frequency, and then identifies whether the first gesture exists in the first picture. When the terminal device identifies that there is a first gesture in the first picture, the terminal device then obtains and captures the second picture at a higher frequency, and identifies the air-interaction gesture in the second picture. Among them, the air-interaction gesture is a dynamic gesture. By identifying whether there is a first gesture in the first picture, the terminal device can determine whether the user has the intention to interact with the terminal device through the air. When it is determined that the user has the intention to interact with the terminal device through the air, the terminal device will obtain the second picture at a higher frequency and identify the air-interaction gesture included in the second picture. In this way, the recognition time of the air-interaction gesture can be effectively reduced, thereby reducing the power consumption of the terminal device and the occupied computing resources.

[0064] In some examples, the first gesture recognized by the above method is a static gesture. The static gesture recognition process has a lower requirement on the frame rate of image acquisition. Therefore, the terminal device obtains the first image at a lower frequency to ensure the accuracy of recognizing the first gesture based on the first image. In this way, the number of times the terminal device recognizes the first gesture can be reduced, thereby effectively reducing the power consumption of the terminal device in recognizing the starting gesture based on the first image.

[0065] In addition, when the terminal device recognizes the air interaction gesture, in order to ensure the accuracy of recognition, a higher frequency of image collection is required. Therefore, the second image collected at a higher frequency is obtained when the terminal device needs to recognize the air interaction gesture, which can ensure the recognition accuracy.

[0066] In some examples, the air interaction gesture is a dynamic gesture, and the first gesture can be the starting action of the air interaction gesture. Since the starting action takes a short time at the beginning of the dynamic gesture, the first gesture can be considered as the starting action of the air interaction gesture, which is a static action. In this way, the terminal device can recognize the first gesture and the air interaction gesture in the process of the user making a continuous dynamic gesture, which is convenient for the user to operate and brings a seamless user experience.

[0067] Among them, the first gesture is a static gesture, then the terminal device's recognition process of the first gesture is a binary classification judgment process. The air interaction gesture is a dynamic gesture, then the terminal device's recognition process of the air interaction gesture in the second picture is a process of deep learning recognition. The binary classification judgment process is relatively simple compared to the deep learning recognition process, so it can effectively reduce the amount of calculation of the terminal device, thereby reducing the computing resources occupied by the terminal device in the process of recognizing the air interaction gesture, and also reducing power consumption.

[0068] The following introduces the hardware structure of the terminal device provided in the embodiment of the present application.

[0069] Figure 3 The structure diagram of the terminal device 100 is shown. The terminal device 100 may include a processor 310, an external memory interface 320, an internal memory 321, a universal serial bus (USB) interface 330, a charging management module 340, a power management module 341, a battery 342, an antenna 1, an antenna 2, a mobile communication module 350, a wireless communication module 360, an audio module 370, a speaker 370A, a receiver 370B, a microphone 370C, an earphone interface 370D, a sensor module 380, a button 390, a motor 391, an indicator 392, a camera 393, a display screen 394, and a subscriber identification module (SIM) card interface 395.

[0070] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0071] The processor 310 may include one or more processing units, for example, the processor 310 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0072] In some embodiments, the processor 310 of the present application may include an AON (always on) ISP. The AON ISP can be understood as a low-power ISP for processing data fed back by the camera 393. In addition, the AON ISP can convert the low-resolution single-channel image data collected by the camera 393 into an image with normal exposure and quality.

[0073] NPU is a neural network computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input information and can also continuously self-learn. Through NPU, applications such as intelligent cognition of the terminal device 100 can be realized, such as: image recognition, face recognition, voice recognition, text understanding, etc.

[0074] In some embodiments, the processor 310 may include an eNPU, which is a low-power AI accelerator (eNPU) for supporting always-on audio, sensors, context data streams, and always-sensing cameras. Unlike ordinary NPUs, eNPUs are used to assist neural network models running on the processor. In some examples, the method provided in the embodiments of the present application runs in an eNPU, which can effectively reduce the power consumption of air-to-air interactive gesture recognition.

[0075] The controller may be the nerve center and command center of the terminal device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0076] The processor 310 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 310 is a cache memory. The memory may store instructions or data that the processor 310 has just used or cyclically used. If the processor 310 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 310, and thus improves the efficiency of the system.

[0077] In some embodiments, the processor 310 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0078] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 310 may include multiple groups of I2C buses. The processor 310 may be coupled to touch sensors, chargers, flashlights, cameras 393 and other devices through different I2C bus interfaces. For example: the processor 310 may be coupled to the camera 393 through the I2C interface, so that the processor 310 communicates with the camera 393 through the I2C bus interface, thereby realizing the air gesture interaction function of the terminal device 100.

[0079] The I2S interface can be used for audio communication. In some embodiments, the processor 310 can include multiple groups of I2S buses. The processor 310 can be coupled to the audio module 370 via the I2S bus to achieve communication between the processor 310 and the audio module 370. In some embodiments, the audio module 370 can transmit an audio signal to the wireless communication module 360 ​​via the I2S interface to achieve the function of answering a call through a Bluetooth headset.

[0080] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In some embodiments, the audio module 370 and the wireless communication module 360 ​​can be coupled via a PCM bus interface.

[0081] The UART interface is a universal serial data bus for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is generally used to connect the processor 310 and the wireless communication module 360.

[0082] The MIPI interface can be used to connect the processor 310 with peripheral devices such as the display screen 394 and the camera 393. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 310 and the camera 393 communicate via the CSI interface to implement the air gesture interaction function of the terminal device 100. The processor 310 and the display screen 394 communicate via the DSI interface to implement the display function of the terminal device 100.

[0083] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 310 with the camera 393, the display screen 394, the wireless communication module 360, the audio module 370, the sensor module 380, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0084] The USB interface 330 is an interface that complies with the USB standard specification, and specifically can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 330 can be used to connect a charger to charge the terminal device 100, and can also be used to transmit data between the terminal device 100 and peripheral devices. It can also be used to connect headphones to play audio through the headphones. The interface can also be used to connect other terminal devices, such as AR devices, etc.

[0085] It is understandable that the interface connection relationship between the modules illustrated in the embodiment of the present invention is only a schematic illustration and does not constitute a structural limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0086] The charging management module 340 is used to receive charging input from a charger. The charger may be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 340 may receive charging input from a wired charger through the USB interface 330. In some wireless charging embodiments, the charging management module 340 may receive wireless charging input through a wireless charging coil of the terminal device 100. While the charging management module 340 is charging the battery 342, it may also power the terminal device through the power management module 341.

[0087] The power management module 341 is used to connect the battery 342, the charging management module 340 and the processor 310. The power management module 341 receives input from the battery 342 and / or the charging management module 340, and supplies power to the processor 310, the internal memory 321, the external memory, the display screen 394, the camera 393, and the wireless communication module 360. The power management module 341 can also be used to monitor parameters such as battery capacity, battery cycle number, battery health status (leakage, impedance), etc. In some other embodiments, the power management module 341 can also be set in the processor 310. In other embodiments, the power management module 341 and the charging management module 340 can also be set in the same device.

[0088] The wireless communication function of the terminal device 100 can be implemented through antenna 1, antenna 2, mobile communication module 350, wireless communication module 360, modem processor and baseband processor.

[0089] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0090] The mobile communication module 350 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the terminal device 100. The mobile communication module 350 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 350 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 350 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1.

[0091] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be sent into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After the low-frequency baseband signal is processed by the baseband processor, it is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker 370A, a receiver 370B, etc.), or displays an image or video through a display screen 394.

[0092] The wireless communication module 360 ​​can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the terminal device 100. The wireless communication module 360 ​​can be one or more devices integrating at least one communication processing module. The wireless communication module 360 ​​receives electromagnetic waves via the antenna 2, modulates the frequency of the electromagnetic wave signal and performs filtering, and sends the processed signal to the processor 310. The wireless communication module 360 ​​can also receive the signal to be sent from the processor 310, modulate the frequency of it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0093] In some embodiments, the antenna 1 of the terminal device 100 is coupled to the mobile communication module 350, and the antenna 2 is coupled to the wireless communication module 360, so that the terminal device 100 can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).

[0094] The terminal device 100 implements the display function through a GPU, a display screen 394, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 394 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 310 may include one or more GPUs, which execute program instructions to generate or change display information.

[0095] The display screen 394 is used to display images, videos, etc. The display screen 394 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLED, Micro-OLED, quantum dot light-emitting diodes (QLED), etc. In some embodiments, the terminal device 100 may include 1 or N display screens 394, where N is a positive integer greater than 1.

[0096] The terminal device 100 can realize the shooting function through ISP, camera 393, video codec, GPU, display screen 394 and application processor.

[0097] The ISP is used to process the data fed back by the camera 393. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, the ISP can be set in the camera 393.

[0098] The camera 393 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the terminal device 100 may include 1 or N cameras 393, where N is a positive integer greater than 1.

[0099] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the terminal device 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0100] The video codec is used to compress or decompress digital video. The terminal device 100 may support one or more video codecs. In this way, the terminal device 100 can play or record videos in multiple coding formats.

[0101] The external memory interface 320 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external memory card communicates with the processor 310 through the external memory interface 320 to implement a data storage function.

[0102] The internal memory 321 may be used to store computer executable program codes, which include instructions. The processor 310 executes various functional applications and data processing of the terminal device 100 by running the instructions stored in the internal memory 321 .

[0103] The terminal device 100 can implement audio functions such as music playing and recording through the audio module 370, the speaker 370A, the receiver 370B, the microphone 370C, the headphone jack 370D, and the application processor.

[0104] The audio module 370 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signal. The audio module 370 can also be used to encode and decode audio signals. The speaker 370A is used to convert audio electrical signals into sound signals. The receiver 370B is used to convert audio electrical signals into sound signals. The microphone 370C is used to convert sound signals into electrical signals. The terminal device 100 can be provided with at least one microphone 370C. The headphone jack 370D is used to connect wired headphones.

[0105] Among them, the sensor module 380 may include a pressure sensor, a touch sensor, etc. The pressure sensor is used to sense the pressure signal and can convert the pressure signal into an electrical signal. The touch sensor is also called a "touch panel". The touch sensor can be set on the display screen 394, and the touch sensor and the display screen 394 form a touch screen, also called a "touch screen". The touch sensor is used to detect a touch operation acting on or near it. The touch sensor can pass the detected touch operation to the application processor to determine the type of touch event.

[0106] The key 390 includes a power key, a volume key, etc. The key 390 may be a mechanical key or a touch key. The terminal device 100 may receive key input and generate key signal input related to user settings and function control of the terminal device 100.

[0107] Motor 391 can generate vibration prompts. Motor 391 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback.

[0108] Indicator 392 may be an indicator light, which may be used to indicate charging status, power changes, messages, missed calls, notifications, etc.

[0109] The SIM card interface 395 is used to connect a SIM card. The SIM card can be connected to or disconnected from the terminal device 100 by inserting the SIM card interface 395 or removing the SIM card interface 395. The terminal device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1.

[0110] The technical solution in this application will be described below with reference to the accompanying drawings, taking a mobile phone as an example of a terminal device.

[0111] In the current air gesture interaction scheme, if the phone is in the off-screen state, it is generally assumed that the user has no need to interact with the phone through air gestures. Therefore, the phone in the off-screen state will not run the related processes of air gesture recognition. When the phone is in the off-screen state, if the user wants to control the phone through air gestures, he needs to wake up the phone first, that is, switch the phone from the off-screen state to the on-screen state. At present, the operation of waking up the phone is generally contact interaction, such as Figure 4 As shown, when the mobile phone in black screen state detects that the button 390 set on the side frame is pressed by the user, the mobile phone switches from black screen state to bright screen state and displays the unlocking interface. In this way, in the scenario where the user cannot touch the mobile phone, there is a situation where the mobile phone cannot be awakened, and then the mobile phone cannot recognize the air interaction gesture made by the user to achieve contactless interaction, which brings inconvenience to the user.

[0112] When the phone screen is off and the user cannot touch the phone, the following method can be used to wake up the phone through air gestures.

[0113] like Figure 5 As shown, when the air gesture interaction function of the mobile phone is turned on and the mobile phone is in the screen-off state, the mobile phone can recognize the screen-off gesture through the following steps (step S501-step S504).

[0114] Step S501: Detect whether there is any movement in the recognition area.

[0115] The recognition area is a detectable area. In some examples, if the terminal device detects through an image captured by a front camera located on the side of the mobile phone screen, the recognition area is the shooting area of ​​the front camera.

[0116] In some embodiments, the terminal device may use MD technology to detect mobile actions in the identification area.

[0117] In some examples, the mobile phone can use MD to detect whether the continuous frames captured by the camera are different. If the continuous frames are different, it indicates that there is movement in the recognition area. If the continuous frames are the same, it indicates that there is no movement in the recognition area. The terminal device can accurately detect the movement in the recognition area through MD.

[0118] The moving action may be caused by the movement of the camera or the appearance of a moving object in the recognition area. Therefore, if there is a moving action in the recognition area, it indicates that the user wants to wake up the phone by performing a screen-off gesture. Then, the phone can continue to determine whether there is a screen-off gesture in the recognition area (steps S502-S504).

[0119] Step S502 : in response to detecting that there is a moving action in the recognition area, obtaining a picture A of a size m1×n1 at a frame rate a.

[0120] Detecting movement in the identification area indicates that the user may need to interact with the terminal device. In this way, the terminal device can obtain picture A (i.e., the fourth picture mentioned above) to reduce the power consumption of the terminal device in obtaining picture A.

[0121] In some examples, the terminal device can obtain a picture A with a size of m1×n1 at a frame rate a (ie, the third frame rate mentioned above).

[0122] Frame rate refers to the number of frames transmitted, displayed or collected per second. The unit of frame rate is the number of frames transmitted per second (fps). When the terminal device captures dynamic actions and obtains multiple frames of pictures, the more frames captured per second, the smoother the dynamic actions displayed continuously in multiple frames. The size of a picture includes length and width. Generally, the length and width of a picture are in pixels. For example, a picture with a resolution of 640×480 is composed of 640 pixels horizontally and 480 pixels vertically. It can be seen that the larger the picture size, the more pixels are required to constitute it. Therefore, the larger the picture size, the higher its resolution.

[0123] In some embodiments, the specific values ​​of the frame rate a and size can be set according to actual applications. For example, the mobile phone can obtain a picture A of size 80×60 at a frame rate of 5fps (frames per second).

[0124] The mobile phone obtains pictures continuously collected by the camera, which can ensure that the mobile phone can recognize the gestures in the pictures in a timely manner and improve the sensitivity of gesture recognition.

[0125] In some embodiments, the image A can be a mono image, that is, a grayscale image. Compared with color images, grayscale images contain less information, so they occupy less memory and the mobile phone processes grayscale images faster. In addition, grayscale images have only one tone, so it is easier to extract texture features based on grayscale images, avoiding the interference of color on image recognition, thereby highlighting the target area. In this way, the terminal device can more accurately identify air interaction gestures based on grayscale images.

[0126] In some embodiments, the mobile phone can also pre-process the picture A, wherein the pre-processing can include denoising, enhancement, edge detection, background removal, etc. By performing the process of pre-processing the picture A, the mobile phone can eliminate irrelevant information in the picture A, restore useful real information, enhance the detectability of relevant information and simplify the data to the maximum extent, thereby improving the accuracy of the screen-off gesture recognition based on the picture A.

[0127] Step S503: Identify picture A and determine whether there is a screen-off gesture in picture A.

[0128] The screen-off gesture (i.e., the aforementioned third gesture) is used to instruct the mobile phone to switch from the screen-off state to the screen-on state. For example, when there is a screen-off gesture in picture A, it indicates that the user needs to wake up the mobile phone in the screen-off state, and the terminal device can switch from the screen-off state to the screen-on state. When there is no screen-off gesture in picture A, it indicates that the user does not need to wake up the mobile phone in the screen-off state at this time, and the mobile phone continues to remain in the screen-off state.

[0129] In some embodiments, the screen-off gesture is pre-set in the mobile phone. For example, the screen-off gesture can be pre-set before the mobile phone leaves the factory, or it can be a gesture input by the user when the air gesture interaction function of the mobile phone is turned on.

[0130] In some examples, the preset screen-off gesture is a palm hover gesture, such as Figure 6As shown in (a), the user stands in front of the screen of the phone that is off, that is, within the shooting range of the front camera, with the inside of the palm of the hand with five fingers extended facing the front camera. At this time, the mobile phone obtains the air interaction gesture based on image A as a palm hovering gesture, and can also determine that the gesture in image A is a screen off gesture. Figure 6 As shown in (b), the user points the hand making the OK gesture toward the front camera. At this time, the mobile phone obtains the air interaction gesture as the OK gesture based on picture A, and it can be determined that the gesture in picture A is not the screen-off gesture.

[0131] In some embodiments, the mobile phone may include a screen-off gesture recognition module, which is used to identify whether there is a screen-off gesture in picture A. Since the screen-off gesture is used to wake up the mobile phone, a static gesture is generally used as the screen-off gesture. Then, the screen-off gesture recognition module can use methods such as neural networks, convolutional neural networks, support vector machines (SVM), nearest neighbor algorithms, distributed local linear embedding, etc. to identify whether there is a screen-off gesture in picture A.

[0132] It should be noted that when the screen-off gesture is a dynamic gesture, the screen-off gesture recognition module can use a deep learning model to realize the recognition of the screen-off gesture.

[0133] In the above embodiment, the process (step S501) of the terminal device detecting whether there is a moving action in the identification area is simpler than the process (step S503) of the terminal device judging whether there is a screen-off gesture in the identification area. Therefore, the power consumption of the mobile phone when executing step S501 is low, and the computing resources used are relatively small. In this way, the mobile phone can determine whether the mobile phone needs to continue to judge whether there is a screen-off gesture in the identification area by identifying whether there is a moving action in the identification area, which can effectively reduce the power consumption of the mobile phone and reduce the computing resources occupied by the mobile phone, thereby improving the smoothness of the mobile phone operation.

[0134] In the above embodiment, when the terminal device is in the screen-off state, the terminal device can first detect whether there is a moving movement in the identification area through MD. If there is a moving movement in the identification area, the terminal device then obtains picture A and identifies whether there is a screen-off gesture in picture A.

[0135] In some other embodiments provided by the present application, a mobile phone in the screen-off state can also directly obtain picture A and identify whether there is a screen-off gesture in picture A. In some examples, the mobile phone obtains picture A and identifies whether there is a screen-off gesture in picture A. If there is no screen-off gesture in picture A, the mobile phone can directly obtain picture A again and identify whether there is a screen-off gesture in picture A. If there is a screen-off gesture in picture A, the mobile phone can execute subsequent steps.

[0136] Optionally, in order to ensure the accuracy of the screen-off gesture recognition result, the mobile phone can recognize multiple consecutive frames of pictures A, and when the screen-off gesture is recognized in the multiple consecutive frames of pictures A, the mobile phone switches from the screen-off state to the screen-on state. See the following step S504 and related embodiments for details.

[0137] Step S504: In response to the gesture in the first frame image A being a screen-off gesture, determine whether N1 consecutive frames after the first frame image A all have screen-off gestures.

[0138] Among them, the first frame of picture A is the picture A in which the screen-off gesture appears for the first time among the consecutive multiple frames of picture A. For example, the camera continuously acquires multiple frames of picture A. After the mobile phone acquires picture A, it sequentially identifies whether there is a screen-off gesture in picture A in the order in which picture A is acquired. If the mobile phone does not recognize the screen-off gesture in picture A, the mobile phone re-detects whether there is a moving action in the recognition area to re-acquire picture A, or the mobile phone can directly re-acquire picture A. If the mobile phone recognizes the screen-off gesture in picture A, the picture A is used as the first frame of picture A, and the mobile phone continues to determine whether there is a screen-off gesture in the consecutive N1 frames after the first frame of picture A.

[0139] When the mobile phone recognizes that there is a screen-off gesture in the first frame of picture A, it can continue to determine whether there are screen-off gestures in the consecutive N1 frames after the first frame of picture A to avoid misjudgment. In some examples, if there is at least one frame of picture A in the consecutive N1 frames after the first frame of picture A that does not include a screen-off gesture, the mobile phone can re-detect the recognition area to re-acquire picture A, or the mobile phone can directly re-acquire picture A. If the consecutive N1 frames of picture A after the first frame of picture A all include a screen-off gesture, it means that the mobile phone has detected a screen-off gesture made by the user in the recognition area.

[0140] In some examples, the user unconsciously makes a screen-off gesture within the recognition area of ​​the phone, but the user does not intend to wake up the phone for air interaction; or, the user originally wants to wake up the phone for air interaction through a screen-off gesture, but after making the screen-off gesture within the recognition area of ​​the phone, he suddenly changes his mind and changes the gesture. The picture A obtained by the mobile phone only includes part of the picture A where the screen-off gesture actually exists. Therefore, the mobile phone can detect picture A without the screen-off gesture in the consecutive N1 frames after the first frame of picture A.

[0141] For example, N1 takes a value of 5, and after the screen-off gesture appears in the first frame A, the screen-off gesture only appears in two consecutive frames A, that is, the screen-off gesture appears in the second and third frames A, but not in the fourth frame A. At this time, the mobile phone can determine that the screen-off gesture does not appear in all five consecutive frames after the first frame A.

[0142] It can be understood that in step S504, the value of N1 may affect the sensitivity and accuracy of the mobile phone in recognizing the screen-off gesture.

[0143] In some examples, the smaller the N1 value is, the shorter the time required for the mobile phone to finally recognize the screen-off gesture, and the higher the recognition sensitivity of the mobile phone. For example, when the acquisition frame rate of picture A is 5pfs and N1 is 5, the mobile phone can determine that the user has accurately and completely made the screen-off gesture and switch from the screen-off state to the screen-on state in at least 1 second when the user makes the screen-off gesture and keeps the screen-off gesture unchanged. For another example, when the acquisition frame rate of picture A is 5pfs and N1 is 10, the mobile phone can determine that the user has accurately and completely made the screen-off gesture in at least 2 seconds when the user makes the screen-off gesture and keeps the screen-off gesture unchanged. It can be seen that the smaller the value of N1, the higher the sensitivity of the mobile phone in recognizing the screen-off gesture.

[0144] In some examples, the larger the N1 value, the longer it takes for the mobile phone to finally recognize the screen-off gesture. The longer the user keeps the screen-off gesture after making it, the more the mobile phone can avoid being awakened by the screen-off gesture made unconsciously by the user, and the higher the accuracy of the mobile phone recognition. For example, when the acquisition frame rate of picture A is 5pfs and N1 is 5, the mobile phone can determine that the user has made a screen-off gesture and switch from the screen-off state to the screen-on state when the user makes a screen-off gesture and keeps the screen-off gesture unchanged within 1s. For another example, when the acquisition frame rate of picture A is 5pfs and N1 is 10, the mobile phone can keep the screen-off gesture unchanged within the first 1s after the user makes the screen-off gesture, and no longer makes the screen-off gesture within the next 1s. The mobile phone cannot determine that the user has accurately and completely made the screen-off gesture. In this way, the larger the value of N1, the higher the accuracy of the mobile phone in recognizing the screen-off gesture.

[0145] Step S505: In response to the presence of screen-off gestures in N1 consecutive frames after the first frame A, the mobile phone switches from the screen-off state to the screen-on state.

[0146] Based on the above embodiments, it can be known that in the scenario where the user clearly intends to make a screen-off gesture, the mobile phone can recognize the presence of a screen-off gesture in picture A, or the mobile phone can recognize the first frame of picture A in which the screen-off gesture exists, and recognize the screen-off gesture in the consecutive N1 frames after the first frame of picture A. Therefore, at this time, the mobile phone switches from the screen-off state to the screen-on state, which can ensure that the mobile phone in the screen-off state can avoid being awakened by the gesture made unintentionally by the user, thereby avoiding false touches, wasting power, etc.

[0147] It should be noted that when the user needs to perform contactless interaction with a mobile phone that is in the off state through a screen-off gesture, the mobile phone recognizes the presence of a screen-off gesture in the collected image A and can directly switch from the screen-off state to the screen-on state.

[0148] When the screen-off gesture is a palm hover gesture, and the user extends his fingers and the palm faces the camera in the recognition area of ​​the phone (such as Figure 7 As shown in (a) in the figure, the mobile phone can recognize the screen-off gesture and switch from the screen-off state to the screen-on state (as shown in Figure 7 As shown in (b) above, the mobile phone can recognize the air interaction gestures made by the user.

[0149] It is understandable that when the phone switches from the off screen state to the on screen state, the interface displayed by the phone may be the lock screen interface. At this time, the phone can be unlocked by non-contact unlocking methods such as facial recognition and voiceprint recognition, so that the user can perform subsequent air gesture operations. Optionally, the phone can also be unlocked by fingerprint unlocking, password input and contact unlocking methods. The specific unlocking method needs to be determined by the user's actual usage needs.

[0150] In some examples, after the phone switches from the off screen state to the on screen state, the phone displays an unlocking interface (such as Figure 7 In the example shown in (b), a prompt graphic 701 is displayed on the mobile phone display screen, and the prompt graphic 701 is used to prompt the user that the mobile phone is currently performing face recognition. At this time, the mobile phone can detect whether there is a user's facial image in the face recognition area, so that the mobile phone can realize face recognition. The face recognition area of ​​the mobile phone is the shooting area of ​​the camera. In this way, the process from waking up to unlocking the mobile phone is realized without any contact between the user and the mobile phone.

[0151] Step S506: Detect whether there is any movement in the recognition area.

[0152] The implementation method in step S506 is similar to the implementation method in step S501, which will not be described in detail here. For a specific description, please refer to the embodiment of the above step S501.

[0153] In some embodiments, when the mobile phone is in the screen-on state, the user does not necessarily have the intention to interact with the mobile phone through the air. Therefore, the mobile phone can use MD to detect whether there is a moving action in the identification area to preliminarily determine whether the user has the intention to interact through the air. Among them, the moving action can be caused by the movement of the camera and the appearance of a moving object in the identification area. Therefore, if the mobile phone uses MD to detect the presence of a moving action in the identification area, it means that the current user may have the intention to interact with the mobile phone through the air. However, if the mobile phone uses MD to detect the absence of a moving action in the identification area, it must indicate that the user has no intention to interact with the mobile phone through the air.

[0154] That is to say, in some cases, although the mobile phone in the screen-on state uses MD to detect the presence of moving movements in the identification area, for example, when the video playback software of the mobile phone is playing a touching video, the user is wiping away tears left by watching the video with his hands, the mobile phone can only use MD to detect the presence of moving movements in the identification area, and cannot determine whether the user is wiping tears or needs to operate the mobile phone remotely. Therefore, the mobile phone needs to further determine whether the gesture made by the user is an air interaction gesture.

[0155] In some embodiments, when the mobile phone detects movement in the identification area, step S507 may be executed.

[0156] Step S507 : In response to detecting that there is a moving action in the recognition area, a picture B of size m2×n2 is obtained at a frame rate b.

[0157] Detecting movement in the identification area indicates that when the terminal device is in the screen-on state, there may be a situation where the user needs to interact with the terminal device. In this way, the terminal device can obtain picture B (i.e., the first picture mentioned above) to reduce the power consumption of the terminal device in obtaining picture B.

[0158] In some embodiments, the value of frame rate b (i.e., the aforementioned first frame rate) is greater than the value of frame rate a, and the size of m2×n2 is greater than the size of m1×n1. The specific values ​​of frame rate b and size can be set according to the actual application, the value of frame rate a, and the size of picture A. For example, in step S502, if the mobile phone obtains picture A of size 80×60 captured by the camera at a frame rate of 5fps, then in step S507, the camera can obtain picture B of size 160×120 captured at a frame rate of 10fps.

[0159] In the above step S507, compared with step S502, the picture acquisition frequency is increased from 5fps to 10fps, which means that the mobile phone obtains more pictures B per unit time. Then, when there is action in picture B, the mobile phone can detect the action more smoothly, thus improving the accuracy of mobile phone gesture recognition.

[0160] In addition, the size of picture B is increased to 160×120 relative to the size of picture A, which means that the resolution of picture B is higher than that of picture A. Therefore, the gestures contained in picture B are clearer, and the mobile phone can achieve longer-distance and more accurate recognition based on picture B.

[0161] In some embodiments, the picture B may also be a mono picture, i.e., a grayscale image. In some embodiments, image preprocessing may also be performed on the picture B. Here, reference may be made to the specific embodiment described above for the picture A.

[0162] In some embodiments, when the mobile phone is in the screen-on state, the mobile phone can also directly obtain picture B.

[0163] Optionally, in some embodiments, considering the scenario where the mobile phone turns on the automatic screen-off function, if the mobile phone does not detect the user's operation within a certain period of time, it will switch from the screen-on state to the screen-off state. Therefore, before the mobile phone performs gesture recognition on the image, step S508 is performed to determine whether the mobile phone is currently in the screen-on state.

[0164] Step S508: determine whether the mobile phone is currently in a screen-on state.

[0165] If the mobile phone is currently in the screen-on state, the terminal device can continue to perform subsequent steps to identify image B. If the mobile phone is currently in the screen-off state, the mobile phone can perform the process in the above embodiment to realize the recognition of the screen-off gesture.

[0166] In some embodiments, before acquiring picture B, the mobile phone can determine whether the mobile phone is currently in a screen-on state. If the mobile phone is currently in a screen-on state, it indicates that the mobile phone can continue to acquire picture B and recognize gestures in picture B. If the mobile phone is currently in a screen-off state, it indicates that the mobile phone cannot interact with the user through the air, and there is no need to continue acquiring picture B.

[0167] Step S509: In response to the mobile phone being in the screen-on state, identifying picture B, and determining whether the gesture in picture B is an interaction start gesture.

[0168] The interaction start gesture (i.e. the aforementioned first gesture) is used to indicate that the user starts to make an air interaction gesture.

[0169] In some embodiments, the role of the air interaction gesture is to trigger the mobile phone to perform corresponding operations, and the user can use the air interaction gesture to interact with the mobile phone. The air interaction gesture is generally a dynamic gesture. For example, the air interaction gesture includes an upward swipe gesture, a downward swipe gesture, a left swipe gesture, a right swipe gesture, a press gesture, a palm flip gesture, a two-finger pinch gesture, etc. The interaction starting gesture in step S509 is the starting gesture when the user makes an air interaction gesture. For example, at the beginning of the upward swipe gesture, the user's fingers are facing downward, and the back of the hand is facing the camera of the mobile phone. Then, with the upward swipe action, the direction of the fingers slowly moves upward. After completing the upward swipe action, the fingers are facing upward, and the palm side is facing the camera of the mobile phone, such as Figure 8 As shown. Therefore, corresponding to the swipe up gesture, the interaction starting gesture may be the back of the hand facing the camera, and the fingers facing downward. For another example, corresponding to the swipe down gesture, the interaction starting gesture may be the palm of the hand, with the fingers facing upward; corresponding to the left swipe gesture, the interaction starting gesture may be the fingers facing right; corresponding to the right swipe gesture, the interaction starting gesture may be the fingers facing left; corresponding to the press gesture, the interaction starting gesture may be the palm of the hand; corresponding to the palm flip gesture, the interaction starting gesture may be the palm or the palm of the hand; corresponding to the two-finger pinch gesture, the interaction starting gesture may be the middle finger, ring finger, and little finger in a retracted state. It can be understood that the interaction starting gesture is set based on the air interaction gesture.

[0170] In some embodiments, the air interaction gesture and the interaction start gesture may be pre-set before the mobile phone leaves the factory. In other embodiments, the air interaction gesture may also be determined based on the user's input. For example, after the air interaction function of the mobile phone is turned on, the user is prompted to make a customized air interaction gesture in the recognition area. The mobile phone can obtain the user's customized air interaction gesture and determine the interaction start gesture.

[0171] The mobile phone recognizes the pictures B in the order in which the camera captures them. Since the camera captures pictures B continuously, the mobile phone also recognizes the interaction start gesture in multiple consecutive frames of pictures B. If the mobile phone does not recognize the interaction start gesture in one of the frames of picture B, it indicates that the air interaction gesture in the recognition area is not complete. At this time, the mobile phone can re-execute step S506. If the mobile phone recognizes the interaction start gesture in picture B, it indicates that the user intends to interact with the mobile phone through the air interaction gesture.

[0172] In some embodiments, since the mobile phone has acquired multiple consecutive frames of images B, the mobile phone can determine whether there is an interaction starting gesture in the multiple frames of images B (step S510) to ensure that the mobile phone determines whether to further recognize the air interaction gesture.

[0173] Step S510: determine whether the N2 consecutive frames after the first frame B all have an interaction start gesture.

[0174] The first frame of picture B is the picture B in which the interaction start gesture appears for the first time among the continuous multiple frames of pictures B. It can be understood that the value of N2 can be determined according to actual needs.

[0175] It can be understood that in this embodiment, the embodiment related to step S510 is similar to the embodiment related to the above step 504, and will not be repeated here. For details, please refer to the above embodiment.

[0176] When the mobile phone recognizes that there is an interaction start gesture in the first frame of picture B, it can continue to determine whether there are interaction start gestures in the consecutive N2 frames after the first frame of picture B to avoid misjudgment. For example, after the mobile phone determines that there is an interaction start gesture in the first frame of picture A, an air interaction prompt icon can be displayed in the display interface of the mobile phone to prompt the user that the air interaction gesture is currently being recognized. If the user does not actually want to interact with the mobile phone through the air, but makes an interaction start gesture in the recognition area of ​​the mobile phone for some reason, then after seeing the air interaction prompt icon displayed in the display interface, the user can stop the current hand movement in time to avoid misrecognition.

[0177] In some examples, if there is at least one frame of picture B in the consecutive N2 frames after the first frame of picture B that does not include the screen-off gesture, the mobile phone can re-use MD to detect and identify the area to re-acquire picture B, or the mobile phone can directly re-acquire picture B.

[0178] It is understandable that in step S510, the value of N2 may affect the sensitivity and accuracy of the mobile phone in recognizing the screen-off gesture. For details, please refer to the above embodiment, which will not be repeated here.

[0179] When there is an interaction starting gesture in N consecutive frames after the first frame B, it means that the user wants to continue to interact with the mobile phone through the air. At this time, the mobile phone can complete the recognition of the air interaction gesture through the following process (steps S511-S517) so as to perform corresponding actions based on the recognized air interaction gesture.

[0180] Step S511 : In response to the presence of an interaction start gesture in N2 consecutive frames after the first frame of the picture B, a picture C of a size of m3×n3 is obtained at a frame rate c.

[0181] Based on the above embodiments, it can be known that when the user clearly intends to make an air interaction gesture, the mobile phone can recognize the existence of an interaction start gesture in picture B, or the mobile phone can recognize the first frame of picture B with an interaction start gesture, and recognize the interaction start gesture in the consecutive N2 frames after the first frame of picture B. Therefore, at this time, the mobile phone can obtain picture C (that is, the aforementioned second picture) to recognize the air interaction gesture based on picture C. In this way, it can avoid the mobile phone from always executing the process of recognizing air interaction gestures, thereby reducing power consumption and reducing resource usage.

[0182] In some embodiments, the value of frame rate c (i.e., the aforementioned second frame rate) is greater than the value of frame rate b, and the size of m3×n3 is greater than the size of m2×n2. The specific values ​​of frame rate c and the size of picture C can be set according to the actual application, the value of frame rate b, and the size of picture B. For example, in step S502, when the mobile phone obtains picture A of size 80×60 captured by the camera at a frame rate of 5fps, in step S507, the mobile phone obtains picture B of size 160×120 captured by the camera at a frame rate of 10fps, then, in step S511, the mobile phone can obtain picture C of size 320×240 captured by the camera at a frame rate of 15fps.

[0183] In some embodiments, the image C may be a mono image, i.e., a grayscale image. In some embodiments, the mobile phone may also pre-process the image C. Here, reference may be made to the specific embodiment of the above description of the image A.

[0184] Step S512: Determine whether there is a hand recognition frame currently.

[0185] The hand recognition frame (i.e., the aforementioned palm frame) is used to indicate the position of the user's hand in the image C. In other words, the hand recognition frame is the position information of the hand. In some examples, the hand recognition frame is the minimum circumscribed rectangular frame determined based on the outer contour of the hand.

[0186] If there is no hand recognition frame currently, it means that the mobile phone has not recognized the hand in the previous frame of picture C, or the mobile phone has not processed picture C before. Then the mobile phone can recognize the hand information of the current frame of picture C (step S513). If there is a hand recognition frame currently, it means that the mobile phone recognizes that there is a hand in the previous frame of picture C. The mobile phone can process the current picture C based on the hand recognition frame obtained by processing the previous frame of picture C. In this way, the mobile phone does not need to recognize each frame of picture C, thereby achieving the effect of reducing power consumption and reducing computing resource usage.

[0187] Step S513: In response to the fact that there is no hand recognition frame currently, recognize the picture C to obtain a hand recognition frame.

[0188] In some embodiments, if the mobile phone does not obtain a hand recognition frame when executing step S513, it can be considered that there is no hand in picture C. At this time, the mobile phone can re-acquire picture B and determine whether there is an interaction start gesture in picture B, for example, re-execute step S506, or re-execute the process of acquiring picture B in step S507. If the hand recognition result obtained by the mobile phone when executing step S513 includes a hand recognition frame, it can be considered that there is a hand in picture C.

[0189] After the hand recognition frame is determined through the above step S512 or step S513, the mobile phone can process the image C based on the hand recognition frame.

[0190] Step S514: Process image C based on the hand recognition frame to obtain image C'.

[0191] In some embodiments, in the process of processing the image C based on the hand recognition frame, the image C can be cropped based on the hand recognition frame so that the cropped image C' (i.e., the aforementioned third image) includes as little background information as possible other than the hand. In this way, the mobile phone can more accurately recognize the user's gesture based on the image C'. In addition, since the size of the image C' becomes smaller after being cropped, the computing resources and power consumption occupied by the mobile phone in the subsequent step of recognizing the air interaction gesture in the image C' can also be effectively reduced.

[0192] In some embodiments, since the air interaction gesture is a dynamic gesture, after obtaining the image C', the mobile phone can use a dynamic gesture recognition method to recognize that the image C' is an air interaction gesture. In some examples, the following steps S515 to S517 are a dynamic gesture recognition method.

[0193] Step S515: Perform static gesture classification on the image C' to obtain the gesture classification result in each frame of the image C'.

[0194] The gesture classification result of static gesture classification of the image C' may include at least one of palm, back of hand, fist, pinch open, pinch closed, and other categories, wherein the identified categories all belong to other categories.

[0195] For example, in picture C' only the following Figure 6 In the case of the gesture shown in (a) of FIG. 1 , after the mobile phone has been subjected to static gesture classification, the gesture classification result obtained includes the palm. For another example, in the picture C ', only the palm Figure 6 In the case of the gesture shown in (b), the mobile phone undergoes static gesture classification and the gesture classification results obtained include pinch-to-close.

[0196] Step S516: perform hand key point detection on image C' to obtain hand key points.

[0197] The hand key points can also be understood as the joints of the hand skeleton, which are usually described by 21 3D key points. Each 3D key point has 3 degrees of freedom, so the dimension of the hand key points obtained in step S516 is 21*3. Therefore, we often use a 21*3 dimensional vector to describe it, such as Fig. 9 As shown. In this way, after the mobile phone detects the key points of the hand in the image C', it can judge the information such as the back of the hand and the direction of the fingers based on the obtained key points of the hand.

[0198] In some instances, after cropping the image C in step S514, the size of the cropped image may be further adjusted (resized) to obtain the image C'. In some examples, steps S515 and S516 perform static gesture classification and hand key point detection on the image C', which generally have requirements on the size of the image C', for example, requiring multiple images C' to have the same size. Therefore, the mobile phone may resize the image cropped based on the image C to obtain the image C' that meets the requirements, so as to ensure that the mobile phone can successfully complete steps S515 and S516 based on the image C'.

[0199] Step S517: determine the user's air interaction gesture based on the hand key points corresponding to each frame of image C and the gesture classification result.

[0200] As Figure 8 Take the air interaction gesture in the example as an example. When the detection in step S516 Figure 8 The interactive starting gesture on the left side of the middle, and after obtaining the key points, the mobile phone can judge based on the key points Figure 8 The back of the hand of the left side of the interaction start gesture faces upward and the fingers face downward. Figure 8 After static gesture classification of the interaction start gesture on the left side of the middle, the classification result can be obtained as the back of the hand. As the user's hand changes, the interaction gesture in picture C gradually changes to Figure 8 The air interaction gesture on the right side of the figure. In this process, the key points corresponding to each frame of the image and the hand classification results can be obtained. The mobile phone can then calculate the change trend of the gesture based on the key points corresponding to the two adjacent frames of the image. At the same time, it can make auxiliary judgments based on the hand classification results to determine the gesture. Figure 8 In the process shown, the finger gradually rises from a low position, the back of the hand gradually disappears from the recognition area, and the palm gradually appears completely in the recognition area. Then, based on the above-mentioned change process, it can be determined that the corresponding air interaction gesture is an upward swipe gesture.

[0201] In some embodiments, when the mobile phone is executing step S517, if the mobile phone cannot determine the air interaction gesture based on the hand key points, gesture classification results and hand position information corresponding to each frame of picture C, the mobile phone can re-acquire picture C.

[0202] In some examples, when the terminal device is in the screen-on state, such as Fig.10 As shown, the air interaction method provided in the embodiment of the present application may include the following steps:

[0203] Step S1001: Acquire a first picture at a first frame rate.

[0204] Step S1002: In response to the first picture including the first gesture, acquiring a second picture at a second frame rate, where the second frame rate is higher than the first frame rate.

[0205] Step S1003: Identify air interaction gestures based on the gesture classification and hand key points of each frame of the second image in multiple frames of the second image.

[0206] The specific implementation of the above steps can be found in the above embodiments, which will not be described again here.

[0207] In some embodiments, in the above process, if the mobile phone determines that the distance between the target gesture and the camera is far based on the key point information, the mobile phone can obtain a larger image C. For example, the user interacts with the mobile phone at a very long distance. For example, the detected key points can be used to determine that the user's hand is actually 1.2 meters away from the mobile phone camera. At this time, the effect of identifying the air interaction gesture through step S617 is not good, so the mobile phone can control the camera to obtain a picture C of size 480×360.

[0208] Understandably, the mobile phone can also determine whether it needs to control the camera to obtain a larger image C based on the relationship between the hand classification results of two adjacent frames. For example, in a scenario where the user is interacting with the mobile phone at a very long distance, the resolution of the image C obtained by the mobile phone is not high, resulting in the hand classification result of the first frame being the palm and the hand classification result of the second frame being the back of the hand. Since it is unlikely that a person's hand changes and can be flipped within 0.06 seconds, it can be determined that the classification result is unstable at this time, and the mobile phone can control the camera to obtain a larger image C.

[0209] In some embodiments, the air interaction gestures include an up swipe gesture, a down swipe gesture, a left swipe gesture, a right swipe gesture, a flip gesture, etc.

[0210] In some embodiments, after the mobile phone recognizes the air interaction gesture, it needs to make a corresponding response. In some examples, the response to the swipe up gesture is to swipe up the page or switch to the next video, the response to the swipe down gesture is to swipe down the page or switch to the previous video, the response to the left swipe gesture is to turn the page to the left or go back, the response to the right swipe gesture is to turn the page to the right or go forward, the response to the flip gesture is to return, etc.

[0211] It should be noted that the same air interaction gesture may produce different responses in different applications, and the specific response should be determined based on the operation of the application.

[0212] In some instances, the air interaction prompt icon can also be used to prompt the current recognition progress of the mobile phone.

[0213] In some embodiments, after the hand key points are obtained based on step S516, the hand recognition frame can be obtained through the following process (step S518-step S519):

[0214] Step S518: determine whether the hand in image C is blocked based on the hand key points.

[0215] If the identified hand key points are missing, the mobile phone will think that the hand in picture C is blocked. For example, 21 hand key points should be identified, but only 18 key points are actually identified, so the mobile phone thinks that the hand in picture C is blocked. At this time, the mobile phone can execute step S513.

[0216] Step S519: In response to the hand in picture C being not blocked, a hand recognition frame is obtained based on the hand key points.

[0217] Since the key points of the hand are 21 three-dimensional data, the mobile phone can calculate a frame of the outer contour of the palm and the position of the frame based on the key points, that is, obtain the hand recognition frame.

[0218] In some embodiments, through the above process, a hand recognition frame can be obtained based on the current frame. Then, after the mobile phone inputs the next frame of picture C through step S511, it can determine that a hand recognition frame currently exists through step S512. At this time, the mobile phone can skip step S513 and directly execute step S514, thereby reducing the computing resources and power consumption occupied by step S513. In the above embodiment, step S513 processes the larger picture C, while step S515 and step S516 both process the smaller picture C'. Therefore, through the above process, the number of executions of step S513 can be effectively reduced, thereby reducing power consumption.

[0219] In some examples, since the hand recognition frame is obtained by the terminal device based on the processing of the picture C, the hand recognition frame does not exist in the terminal device when the terminal device processes the picture C for the first time. At this time, the terminal device can crop the picture C' from each frame of the multiple frames of the picture C, and then obtain the gesture classification and hand key points based on the hand area in the picture C' (i.e., steps S513-S516).

[0220] In some embodiments, when the terminal device crops out the picture C' from each frame of the multiple frames of pictures C, it can be considered that the picture C' corresponding to the k-th frame of picture C is cropped based on the k-th frame of picture C.

[0221] After the terminal device starts processing the picture C, the terminal device can obtain the hand key points of the k-th frame picture C among the multiple frames of pictures C. Afterwards, the terminal device can detect whether the hand in the k-th frame picture C is complete based on the hand key points; k is a positive integer. In response to the hand in the k-th frame picture C being complete, the terminal device obtains a palm frame based on the hand key points in the k-th frame picture C, and crops the k+i-th frame picture C based on the palm frame to obtain picture C', where i is a positive integer.

[0222] Among them, the position of the hand in the k+i frame image C is the same as or similar to the position of the hand in the k frame image C, so the hand in the k frame image C is complete, indicating that the hand in the k+i frame image C is also complete, and the palm frame corresponding to the k frame image C can be directly used as the palm frame corresponding to the k+i frame image C. In this way, after the terminal device obtains the palm frame based on the hand key points in the k frame image C, the terminal device can obtain the k+i frame image C, and then when executing step S512, it is determined that the palm frame exists. Then, the terminal device can directly execute step S514.

[0223] In response to the incomplete hand in the k-th frame image C, it indicates that the hand in the k+i-th frame image C is not complete. Therefore,

[0224] The terminal device may directly crop the picture C' corresponding to the k+i-th frame picture C based on the k+i-th frame picture C (ie, execute step S513).

[0225] It should be noted that in the embodiment of the present application, the terminal device needs to obtain the user's consent when turning on the camera to obtain pictures, including but not limited to notifying and reminding the user to read the relevant user agreement and sign an agreement including authorization of relevant permissions before the user uses the air interaction function.

[0226] In the above embodiment, in the process of the mobile phone recognizing the air interaction gesture based on the picture C, the mobile phone needs to perform hand recognition, static gesture classification, hand key point detection and other processes on the picture C, and then determine the user's air interaction gesture based on the obtained hand key points, gesture classification results and hand position information. It can be seen that the process of the mobile phone recognizing based on the picture C is more complicated, and accordingly, the computing resources occupied and the power consumption consumed by the process of the mobile phone recognizing based on the picture C are both large. Therefore, before performing gesture recognition based on the picture C, the mobile phone can first recognize the interaction start gesture of the picture B, and when there is an interaction start gesture in the picture B, it further determines whether to further obtain the picture C and recognize the air interaction gesture based on the picture C. It is relatively simple to recognize the interaction start gesture of the picture B. Therefore, the power consumption and computing resource occupation of the entire recognition process can be effectively reduced through the above process.

[0227] In the above embodiment, the role of gesture recognition based on picture A by a mobile phone in the screen-off state is to trigger the mobile phone to switch from the screen-off state to the screen-on state, and the role of gesture recognition based on picture B by a mobile phone in the screen-on state is to determine whether it is necessary to further obtain picture C and perform a more complex gesture recognition process based on picture C. In other words, the mobile phone has the lowest requirements for the frequency and resolution of obtaining picture A, and the highest requirements for the frequency and resolution of obtaining picture C. In this way, as the above process proceeds, by successively increasing the frequency and size of picture acquisition, it can ensure accurate recognition of air interaction gestures based on picture C, while reducing the frequency and overhead of obtaining pictures A and B, or reducing the frequency of the mobile phone calling the camera to obtain images and the overhead of obtaining images, thereby achieving the effect of reducing power consumption.

[0228] Combination of the above Figure 3-Figure 9 The air interaction gesture recognition method provided by the embodiment of the present application is described in detail. Fig.11 The terminal device provided in the embodiment of the present application is described in detail.

[0229] In one possible design, Fig.11 This is a schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Fig.11 As shown, the terminal device 1100 may include: a display unit 1101, a processing unit 1102, and a transceiver unit 1103. The terminal device 1100 may be used to implement the functions of the terminal device involved in the above method embodiment.

[0230] Optionally, the display unit 1101 is used to support the terminal device 1100 to display interface content; and / or to support the terminal device 1700 to execute Figure 6 S609 in.

[0231] Optionally, the processing unit 1102 is used to support the terminal device 1100 to execute Figure 5 S501-S519 in.

[0232] Optionally, the transceiver unit 1103 is used to support the terminal device 1100 to execute Figure 5 S501-S519 in.

[0233] Among them, the transceiver unit may include a receiving unit and a sending unit, which may be implemented by a transceiver or a transceiver-related circuit component, and may be a transceiver or a transceiver module. The operations and / or functions of each unit in the terminal device 1100 are respectively to implement the corresponding process of the air-to-air interactive gesture recognition method described in the above method embodiment. All relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional unit. For the sake of brevity, they will not be repeated here.

[0234] Optionally, Fig.11 The terminal device 1100 shown may also include a storage unit ( Fig.11 (not shown), the storage unit stores a program or instruction. When the processing unit 1102 and the transceiver unit 1103 execute the program or instruction, Fig.11 The terminal device 1100 shown can execute the air-interaction gesture recognition method described in the above method embodiment.

[0235] Fig.11 The technical effects of the terminal device 1100 shown can refer to the technical effects of the air-interaction gesture recognition method described in the above method embodiment, and will not be repeated here.

[0236] In addition to being in the form of the terminal device 1100, the technical solution provided in the embodiment of the present application may also be a functional unit or chip in the terminal device, or a device used in conjunction with the terminal device.

[0237] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned terminal device, the terminal device executes each function or step in the above-mentioned method embodiment.

[0238] The embodiment of the present application also provides a computer program product, including a computer program. When the computer program is executed on a terminal device, the terminal device executes each function or step in the above method embodiment.

[0239] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0240] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0241] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0242] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0243] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0244] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for identifying gestures in air interaction, characterized in that: Applied to a terminal device, when the terminal device is in a screen-on state, the method includes: Acquire a first picture at a first frame rate; In response to the first picture including a first gesture, acquiring a second picture at a second frame rate, wherein the second frame rate is higher than the first frame rate; Based on the gesture classification and hand key points of each frame of the second image in multiple frames of the second image, the air interaction gesture is identified.

2. The method according to claim 1, characterized in that The size of the second picture is larger than the size of the first picture.

3. The method according to claim 1 or 2, characterized in that: The identifying the air interaction gesture based on the gesture classification and the hand key points of each frame of the second image in the plurality of frames of the second image includes: In response to a first frame of the second picture in the multiple frames of the second picture including a hand, an air gesture is identified based on gesture classification and hand key points of each frame of the second picture in the multiple frames of the second picture.

4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: In response to a first frame of the second picture in the plurality of frames of the second pictures not including a hand, returning to the step of acquiring the first picture at the first frame rate.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: A third picture is cropped from each of the second pictures in multiple frames of the second pictures, wherein the third picture includes a hand area, and gesture classification and hand key points are obtained based on the hand area in the third picture.

6. The method according to any one of claims 1 to 5, characterized in that After obtaining the hand key points of the kth second picture in the plurality of second pictures, the method further includes: Detecting whether the hand in the second picture of the kth frame is complete based on the hand key points; k is a positive integer; In response to the hand in the k-th frame of the second picture being complete, obtaining a palm frame based on the hand key points in the k-th frame of the second picture; Cropping a third picture from each of the second pictures in a plurality of frames of the second pictures comprises: The k+i-th frame of the second picture is cropped based on the palm frame to obtain a third picture, where i is a positive integer.

7. The method according to claim 6, characterized in that Cropping a third picture from each of the second pictures in a plurality of frames of the second pictures comprises: In response to the incomplete hand in the k-th frame second picture, a third picture corresponding to the k+i-th frame second picture is cropped based on the k+i-th frame second picture.

8. The method according to claim 6 or 7, characterized in that: Cropping a third picture from each of the second pictures in a plurality of frames of the second pictures comprises: A third picture corresponding to the k-th frame of the second picture is cropped based on the k-th frame of the second picture.

9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: In response to detecting a moving motion, a first picture is acquired at a first frame rate.

10. The method according to any one of claims 1 to 9, characterized in that The first picture has multiple frames, and determining that the first picture includes the first gesture includes: It is determined that N consecutive first pictures in the plurality of first pictures all include a first gesture, where N is a positive integer greater than 2.

11. The method according to any one of claims 1 to 10, characterized in that Before acquiring the first picture at the first frame rate, the method further includes: When the terminal device is in a screen-off state, a fourth image is acquired at a third frame rate, and in response to the fourth image including a second gesture, the terminal device switches to a screen-on state.

12. The method according to claim 11, characterized in that Determining that the fourth picture includes the second gesture includes: It is determined that M consecutive fourth pictures in the plurality of fourth pictures all include the second gesture, where M is a positive integer greater than 2.

13. A terminal device, characterized in that: The terminal device includes a display screen, a memory and one or more processors; the display screen, the memory and the processor are coupled; the display screen is used to display an image generated by the processor, and the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the terminal device executes the method described in any one of claims 1-12.

14. A computer-readable storage medium, characterized in that: The method comprises computer instructions, and when the computer instructions are executed on a terminal device, the terminal device executes the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Gesture recognition method and device and vehicle-mounted system

    CN106502570A

  • Shortcut function starting method and electronic equipment

    CN110058777A

  • Gesture recognition method and device, gesture control method and device, medium and terminal equipment

    CN111062312A

  • Gesture recognition method and electronic device and readable medium thereof

    CN114612928A

  • Gesture recognition object determination method and device

    CN115565241A