A method for recognizing a gesture for air interaction and a terminal device

By using a low frame rate to recognize the first gesture while the terminal device's screen is on and a high frame rate to recognize the air gesture after confirmation, the problem of high power consumption and high computational overhead under the air interaction function of the terminal device is solved, achieving more efficient recognition and more convenient user interaction.

CN119942628BActive Publication Date: 2026-03-31HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing terminal devices, after enabling the air interaction function, continuously identify each frame of image, resulting in significant power consumption and computational overhead.

Method used

When the terminal device is in the screen-on state, the first image is acquired at a lower frame rate to recognize the first gesture. If the first gesture is recognized, the second image is acquired at a higher frame rate to perform air gesture recognition. By combining gesture classification and hand key points from multiple frames, the image processing flow is optimized to reduce power consumption and computational overhead.

Benefits of technology

It effectively reduces the time and power consumption of air gesture recognition, improves the accuracy and robustness of recognition, and provides a more convenient user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942628B_ABST
    Figure CN119942628B_ABST
Patent Text Reader

Abstract

The application provides a method for recognizing a gesture for air interaction and a terminal device, and relates to the technical field of terminals. The method is applied to a terminal device. When the terminal device is in a bright screen state, the terminal device acquires a first picture at a first frame rate. In response to the first picture including a first gesture, the terminal device acquires a second picture at a second frame rate, the second frame rate is higher than the first frame rate, and the terminal device recognizes a gesture for air interaction based on gesture classification and hand key points of each second picture in multiple second pictures. The higher the frame rate, the more pictures acquired in a unit of time, and the higher the resource memory computing power required by the terminal device to calculate. Therefore, in the above process, whether the terminal device performs the process of acquiring a second picture and recognizing a gesture for air interaction based on the second picture is determined based on whether the first picture includes a first gesture. In this way, the process of the terminal device recognizing a gesture for air interaction can be reduced, thereby reducing power consumption and the cost of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to a method and terminal device for air gesture recognition. Background Technology

[0002] For scenarios where users find it inconvenient to touch or operate terminal devices, such as mobile phones and tablets, some devices offer interaction via air gestures. The terminal device can capture images using a camera and recognize air gestures within those images. When a specific air gesture is recognized, the terminal device can execute the corresponding action. In this way, users can achieve contactless interaction with the terminal device using air gestures.

[0003] In existing solutions utilizing air gesture interaction, the terminal device needs to perform air gesture recognition on each captured image frame after the air interaction function is enabled, which results in significant power consumption. Furthermore, the air gesture recognition process also incurs substantial computational overhead. Summary of the Invention

[0004] In view of this, this application provides a method and terminal device for recognizing air gestures, which can reduce the time for the terminal device to recognize air gestures, thereby reducing the power consumption and computational overhead of the terminal device in recognizing air gestures.

[0005] In a first aspect, this application provides a method for recognizing air gestures, applied to a terminal device. When the terminal device is in a screen-on state, the method includes: acquiring a first image at a first frame rate, wherein the first image can be captured by the terminal device's camera; in response to the first image including a first gesture, acquiring a second image at a second frame rate, where the second frame rate is higher than the first frame rate; and recognizing the air gesture based on gesture classification and hand key points in each frame of the multiple second images.

[0006] The initial action of the air gesture interaction is the first gesture. That is, when the terminal device recognizes the first gesture, it indicates that the user intends to interact with the terminal device air-to-air. Therefore, after recognizing the first gesture in the first image, the terminal device can acquire the second image and recognize the air gesture in the second image. In this way, the terminal device can reduce the time spent recognizing the air gesture in the second image, thereby reducing the power consumption and computational overhead of the terminal device in recognizing air gestures.

[0007] In some examples, the first gesture can be a gesture unrelated to the air interaction gesture. For example, if the air interaction gesture is an upward swipe, the first gesture is a clenched fist. In other examples, the air interaction gesture is a dynamic gesture, and the first gesture can be the initial action of making the air interaction gesture. For example, if the air interaction gesture is a downward swipe, the user's initial action is to point their palm upwards towards the camera, and the first gesture is correspondingly an upward-pointing palm. Thus, in the case of a user making a downward swipe gesture, the first image acquired by the terminal device includes the first gesture, and the subsequent second image includes the air interaction gesture, avoiding the need for the user to make two gestures and achieving a more convenient and faster air interaction.

[0008] Furthermore, the higher the frame rate of image acquisition, the faster the processing speed of the terminal device needs to ensure that the image can be processed. In the above implementation, the terminal device acquires the second image at a higher frame rate than the first image, ensuring that the terminal device processes the first image at a lower processing speed and the second image at a higher processing speed. By reducing the processing time of the second image, the power consumption of the terminal device can be further reduced.

[0009] After recognizing a gesture, the terminal device can execute the corresponding operation, enabling air interaction between the user and the terminal device.

[0010] In one possible implementation of the first aspect, the second image is larger than the first image. A larger image size indicates a higher resolution, meaning the second image is clearer than the first. The terminal device can identify the presence of a first gesture based on the lower-resolution first image, while accurately recognizing the air gesture based on the higher-resolution second image. This ensures the accuracy of the recognition results.

[0011] In another possible implementation of the first aspect, the air gesture is identified based on the gesture classification and hand key points of each frame of the multi-frame second image, including: in response to the first frame of the multi-frame second image including a hand, the air gesture is identified based on the gesture classification and hand key points of each frame of the multi-frame second image.

[0012] In the above implementation, gesture classification can represent the state of the gesture, and hand key points can represent the three-dimensional position of the hand in space. Therefore, the terminal device can determine the trend of gesture change based on each frame of the second image in multiple frames, and accurately determine the air interaction gesture.

[0013] In one possible implementation of the first aspect, in response to the fact that the first frame of the second image in a multi-frame set does not include a hand, the step of acquiring the first image at the first frame rate is returned. The fact that the first frame of the second image in a multi-frame set does not include a hand indicates a misidentification. Therefore, the terminal device can restart acquiring the first image, avoiding the increase in the time required for the terminal device to recognize the air gesture based on the second image due to misidentification, thereby reducing power consumption.

[0014] In one possible implementation of the first aspect, the method further includes: cropping a third image from each frame of the multi-frame second images, the third image including a hand region; subsequently, the terminal device can obtain gesture classification and hand key points based on the hand region in the third image. The smaller size of the third image cropped from the second images allows the terminal device to perform faster recognition based on the third size, thereby reducing recognition time and also achieving the effect of reducing power consumption.

[0015] In one possible implementation of the first aspect, after obtaining the hand keypoints of the k-th second image out of multiple frames of second images, the method further includes: detecting whether the hand in the k-th second image is complete based on the hand keypoints; where k is a positive integer. In response to the hand being complete in the k-th second image, a palm outline is obtained based on the hand keypoints in the k-th second image.

[0016] After obtaining the hand frame based on the hand key points in the second image of frame k, the terminal device can crop the second image of frame (k+i) based on the hand frame to obtain the third image, where i is a positive integer. The third image is obtained by cropping the second image of frame (k+i) based on the hand frame.

[0017] In the above implementation process, because the speed of hand movement is relatively slow compared to the speed of second image acquisition, the movement distance between hands in consecutive frames of the second image is small. In other words, the terminal device can consider the hand in the (k+i)th frame of the second image to be in the same or similar position as the hand in the kth frame. Therefore, if the hand in the kth frame is complete, the terminal device will also consider the hand in the (k+i)th frame to be complete, and the hand's position to be the same or similar to that in the (k+i)th frame. Thus, the terminal device can directly crop the (k+i)th frame of the second image based on the hand frame, thereby improving image processing speed.

[0018] In one possible implementation of the first aspect, in response to the incomplete hand in the second image of the k-th frame, the terminal device can crop out a third image corresponding to the second image of the (k+i)-th frame based on the second image of the k+i-th frame.

[0019] In the above implementation process, if the hand in the second image of the k-th frame is incomplete, the terminal device will assume that the hand in the second image of the (k+i)-th frame may also be incomplete. Since the position of the hand in the second image of the k-th frame is the same as or close to the position of the hand in the second image of the (k+i)-th frame, the terminal device can directly crop out the third image corresponding to the second image of the (k+i)-th frame, thereby improving the accuracy of image processing.

[0020] In another possible implementation of the first aspect, cropping the third image from each of the multiple frames of second images includes: the terminal device cropping the third image corresponding to the k-th frame of the second image based on the k-th frame of the second image. The terminal device can directly crop the third image corresponding to the k-th frame of the second image based on the k-th frame of the second image to improve the accuracy of the third image.

[0021] In one possible implementation of the first aspect, the method further includes: in response to detecting a movement, the terminal device acquires a first image at a first frame rate.

[0022] In one possible implementation of the first aspect, when the first image has multiple frames, determining that the first image includes the first gesture includes: determining that N consecutive frames of the first image all include the first gesture, where N is a positive integer greater than 2. Thus, by determining whether the first gesture is included in all N consecutive frames of the first image, the terminal device can identify errors and improve robustness in practical applications.

[0023] In one possible implementation of the first aspect, before the terminal device acquires the first image at a first frame rate, the method further includes: acquiring a fourth image at a third frame rate while the terminal device is in a screen-off state; and in response to the fourth image including a second gesture, the terminal device switches to a screen-on state, wherein the second gesture is a gesture used to wake up the terminal device. Thus, the terminal device can interact with the user to wake up the screen while in a screen-off state, allowing the user to interact with the terminal device using air gestures, improving the convenience of interaction.

[0024] In some designs, the third frame rate for acquiring the fourth image on the terminal device is lower than the first frame rate. When the terminal device is in a screen-off state, the user's need for air interaction with the terminal device is relatively small. Therefore, the terminal device can acquire the fourth image and recognize the second gesture within it at a slower speed, which can effectively reduce the terminal device's power consumption.

[0025] In one possible implementation of the first aspect, determining that the fourth image includes the second gesture includes: determining that M consecutive fourth images out of multiple frames all include the second gesture, where M is a positive integer greater than 2. Thus, by determining whether the second gesture is included in all M consecutive fourth images, the terminal device can identify errors and improve robustness in practical applications.

[0026] In some implementations, when the terminal device is in a screen-off state, in response to the detection of movement, the terminal device can acquire a fourth image at a third frequency. This method can reduce the power consumption of the terminal device.

[0027] In some implementations, the fourth image is smaller than the first image. A larger image size indicates higher resolution and higher memory usage. This allows the terminal device to process the smaller fourth image more quickly, and the first image is only processed if a second gesture is present in the fourth image. This improves the terminal device's recognition speed and reduces power consumption.

[0028] In some implementations, air gestures include swiping left, swiping right, flipping the palm, and pinching with two fingers.

[0029] In a second aspect, this application provides a terminal device, the terminal device including a display screen, a memory, and one or more processors; the display screen, the memory, and the processor are coupled; the display screen is used to display an image generated by the processor, the memory is used to store computer program code, the computer program code including computer instructions; when the processor executes the computer instructions, it causes the terminal device to perform the method as described in the first aspect and any possible design of the method.

[0030] Thirdly, this application provides a computer-readable storage medium including computer instructions that, when executed on a terminal device, cause the terminal device to perform the method described in the first aspect above and any possible design of the method thereof.

[0031] Fourthly, this application provides a computer program product that, when run on a terminal device, causes the terminal device to perform the method described in the first aspect above and any possible design of the method.

[0032] Fifthly, this application provides an apparatus included in a terminal device, which has the function of implementing the terminal device behavior in any of the above aspects and possible implementations. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above function. For example, an allocation module or unit, a scanning module or unit, a recycling module or unit, a moving module or unit, and a storage module or unit, etc.

[0033] Sixthly, embodiments of this application provide a chip system including a processor and potentially a memory, for implementing any of the methods provided in the first aspect above. The chip system may be composed of chips or may include chips and other discrete devices.

[0034] Understandably, the terminal device described in the second aspect and any possible design of the above-mentioned device, the computer-readable storage medium described in the third aspect, and the computer program product described in the fourth aspect are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here. Attached Figure Description

[0035] Figure 1 A schematic diagram of a terminal device provided in an embodiment of this application;

[0036] Figure 2 A schematic diagram of a gesture control interface for air interaction provided in an embodiment of this application;

[0037] Figure 3 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;

[0038] Figure 4 This application provides a schematic diagram of a terminal device wake-up method.

[0039] Figure 5 A flowchart illustrating a gesture recognition method for air interaction provided in an embodiment of this application;

[0040] Figure 6 A schematic diagram of a screen-off gesture provided in an embodiment of this application;

[0041] Figure 7 This application provides a schematic diagram of a terminal device wake-up method.

[0042] Figure 8 This is a schematic diagram of a dynamic gesture change provided in an embodiment of this application;

[0043] Figure 9A schematic diagram of key points of the hand provided in an embodiment of this application;

[0044] Figure 10 A flowchart illustrating a gesture-based air interaction method provided in an embodiment of this application;

[0045] Figure 11 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0046] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.

[0047] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0048] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0049] Before introducing the embodiments of this application, the technologies involved in the embodiments of this application will be described in detail.

[0050] 1. Motion detection (MD)

[0051] Motion detection (MD), also known as motion capture, is a technique used to detect changes in the content captured by a camera. MD technology is commonly used in unattended surveillance video recording and automatic alarm systems. For example, in a surveillance system, a camera continuously captures images, and the system's processor can use MD to detect changes in the content of multiple frames. For instance, someone walking past the camera or the camera lens being moved will cause changes in the image content. Once changes are detected in multiple frames, the system's processor can take appropriate action, such as sounding an alarm or issuing an alarm notification.

[0052] In addition, mobile phones, tablets, and other terminal devices can also use MD technology to assist in shooting. For example, after detecting movement using MD technology, the mobile phone can adjust the shooting frame rate to capture clear photos in scenes of moving or fast-moving objects.

[0053] 2. mono format

[0054] Mono format is a grayscale image format. Mono format images (hereinafter referred to as mono images) can generally be output by devices such as single-channel cameras or black and white cameras. Mono format includes various pixel formats such as Mono8, Mono10, Mono10 Packed, Mono12, and Mono12 Packed.

[0055] In this embodiment of the application, the terminal device can recognize air gestures based on mono images.

[0056] To better understand the embodiments of this application, the terminal device provided in the embodiments of this application will be introduced first.

[0057] like Figure 1 As shown, the terminal device 100 can specifically be a mobile phone 11, tablet computer 12, smart screen 13, laptop computer 14, in-vehicle equipment, wearable device (such as a smartwatch), ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), artificial intelligence device, or other terminal device with air gesture interaction function. The operating system installed on the terminal device 100 includes, but is not limited to, iOS®, Android®, Windows®, Linux®, or other operating systems. This application embodiment does not limit the specific type of terminal device 100 or the operating system installed on it.

[0058] In the following examples, the terminal device is a mobile phone, but in actual applications, the terminal device can be any of the above-mentioned devices.

[0059] Generally, different air gestures correspond to different operations. For example, swiping up from bottom to top corresponds to the swipe-up operation, swiping up from bottom to top corresponds to the swipe-down operation, and changing from spreading five fingers to making a fist corresponds to the screenshot operation.

[0060] In some scenarios, mobile phones can enable or disable air gesture interaction based on user input. In some examples, the phone's settings app includes options for controlling air gesture interaction; in response to the user tapping these options, the phone's screen displays an interface for controlling air gesture interaction.

[0061] Different air gestures can correspond to different control interfaces. Figure 2 Image (a) shows a control interface for a gesture-based downward swipe screen, displaying operation instructions and control buttons for the gesture-based downward swipe screen function. Figure 2 The control button for the air swipe screen function shown in (a) is in the on state, indicating that the terminal device can recognize the user's air swipe gesture in the gesture recognition area, and control the page displayed on the terminal device to swipe down after recognizing the air swipe gesture. Figure 2 Image (b) shows a control page for air screenshotting, which displays operation instructions for air screenshotting, as well as control buttons for the air screenshot function. Figure 2 The air screenshot function control button shown in (b) is in the off state, indicating that the terminal device cannot recognize the air interaction gesture in the gesture recognition area.

[0062] When the air gesture interaction function of the aforementioned mobile phone is enabled, the phone, with its screen on, continuously captures images through its camera and identifies the presence of air gestures in each frame to ensure timely recognition and execution of corresponding interactive actions. However, the phone is not always in a state of air interaction with the user while its screen is on. For example, when a video playback application is playing a video, the phone screen remains constantly lit, but there may be scenarios where the user is not near the phone, meaning the phone may not recognize the air gesture. Similarly, when the user is typing on the phone, there may be situations where the phone will not recognize the air gesture. In these scenarios, if the phone is on and the air gesture interaction function is enabled, the phone continuously captures images and identifies the presence of air gestures in each frame; that is, the process for recognizing air gestures is constantly running. Because the recognition process for air gestures involves deep learning, the computation is relatively complex, resulting in higher power consumption and the recognition process continuously consuming significant computing resources.

[0063] To address the aforementioned issues, this application provides a method for recognizing air gestures. When the terminal device is in a screen-on state, it first acquires a first image at a lower frequency and then identifies whether a first gesture exists in the first image. If the terminal device identifies the first gesture in the first image, it then acquires a second image at a higher frequency and recognizes the air gesture in the second image. The air gesture is a dynamic gesture; by identifying whether the first gesture exists in the first image, the terminal device can determine whether the user intends to interact with the terminal device air-to-ground. Only when the user's intention to interact air-to-ground is determined will the terminal device acquire the second image at a higher frequency and recognize the air gesture included in the second image. This effectively reduces the recognition time for air gestures, thereby reducing the power consumption and computing resources occupied by the terminal device.

[0064] In some examples, the first gesture identified by the above method is a static gesture. The static gesture recognition process has a lower frame rate requirement for image acquisition. Therefore, the accuracy of recognizing the first gesture based on the first image can be guaranteed by the terminal device acquiring the first image at a lower frequency. In this way, the number of times the terminal device recognizes the first gesture can be reduced, thereby effectively reducing the power consumption of the terminal device in recognizing the starting gesture based on the first image.

[0065] In addition, in order to ensure the accuracy of recognition, the terminal device requires a high image acquisition frequency when recognizing air gestures. Therefore, the second image, which is acquired at a high frequency, is only acquired when the terminal device needs to recognize air gestures, which can ensure the recognition accuracy.

[0066] In some examples, the air gestures are dynamic, and the first gesture can be the starting action of the air gesture. Since the starting action takes a short time at the beginning of the dynamic gesture, it can be considered the starting action of the air gesture, and is therefore a static action. In this way, the terminal device can recognize both the first gesture and the air gesture itself while the user is making a series of dynamic gestures, facilitating user operation and providing a seamless user experience.

[0067] In this case, the first gesture is a static gesture, so the terminal device's recognition process for the first gesture is a binary classification process. The air gesture is a dynamic gesture, so the terminal device's recognition process for the air gesture in the second image utilizes deep learning. The binary classification process is simpler than the deep learning recognition process, thus effectively reducing the computational load on the terminal device, thereby reducing the computational resources occupied by the terminal device in recognizing the air gesture, and also reducing power consumption.

[0068] The hardware structure of the terminal device provided in the embodiments of this application is described below.

[0069] Figure 3 A schematic diagram of the terminal device 100 is shown. The terminal device 100 may include a processor 310, an external memory interface 320, an internal memory 321, a universal serial bus (USB) interface 330, a charging management module 340, a power management module 341, a battery 342, antenna 1, antenna 2, a mobile communication module 350, a wireless communication module 360, an audio module 370, a speaker 370A, a receiver 370B, a microphone 370C, a headphone jack 370D, a sensor module 380, buttons 390, a motor 391, an indicator 392, a camera 393, a display screen 394, and a subscriber identification module (SIM) card interface 395, etc.

[0070] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0071] Processor 310 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. The different processing units may be independent devices or integrated into one or more processors.

[0072] In some embodiments, the processor 310 of this application may include an AON (always-on) ISP. An AON ISP can be understood as a low-power ISP used to process data fed back by the camera 393. Furthermore, the AON ISP can convert the low-resolution single-channel image data captured by the camera 393 into an image with proper exposure and quality.

[0073] NPU stands for Neural Network Processing Unit. By drawing inspiration from the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in terminal devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0074] In some embodiments, the processor 310 may include an eNPU, which is a low-power AI accelerator (eNPU) used to support always-on audio, sensors, contextual data streams, and always-sensing cameras. Unlike a regular NPU, the eNPU assists neural network models running on the processor. In some examples, the methods provided in this application embodiment run in an eNPU, which can effectively reduce the power consumption of air gesture recognition.

[0075] The controller can serve as the central nervous system and command center of the terminal device 100. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.

[0076] The processor 310 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 310 is a cache memory. This memory can store instructions or data that the processor 310 has just used or that are used repeatedly. If the processor 310 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 310, and thus improves the efficiency of the system.

[0077] In some embodiments, the processor 310 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0078] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 310 may include multiple I2C buses. The processor 310 can couple to devices such as touch sensors, chargers, flashlights, and cameras 393 through different I2C bus interfaces. For example, the processor 310 can couple to the camera 393 through the I2C interface, enabling the processor 310 and the camera 393 to communicate via the I2C bus interface, thereby realizing the air gesture interaction function of the terminal device 100.

[0079] The I2S interface can be used for audio communication. In some embodiments, the processor 310 may include multiple I2S buses. The processor 310 can be coupled to the audio module 370 via the I2S bus to enable communication between the processor 310 and the audio module 370. In some embodiments, the audio module 370 can transmit audio signals to the wireless communication module 360 ​​via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0080] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 370 and the wireless communication module 360 ​​can be coupled via the PCM bus interface.

[0081] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 310 and the wireless communication module 360.

[0082] The MIPI interface can be used to connect the processor 310 to peripheral devices such as the display screen 394 and the camera 393. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 310 and the camera 393 communicate via the CSI interface to enable air gesture interaction functionality of the terminal device 100. The processor 310 and the display screen 394 communicate via the DSI interface to enable the display function of the terminal device 100.

[0083] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 310 to a camera 393, a display screen 394, a wireless communication module 360, an audio module 370, a sensor module 380, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0084] USB port 330 is a USB standard compliant interface, which can be a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 330 can be used to connect a charger to charge terminal device 100, and can also be used for data transfer between terminal device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other terminal devices, such as AR devices.

[0085] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0086] The charging management module 340 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 340 receives charging input from the wired charger via a USB interface 330. In some wireless charging embodiments, the charging management module 340 receives wireless charging input via the wireless charging coil of the terminal device 100. While charging the battery 342, the charging management module 340 can also supply power to the terminal device via the power management module 341.

[0087] The power management module 341 connects the battery 342, the charging management module 340, and the processor 310. The power management module 341 receives input from the battery 342 and / or the charging management module 340, providing power to the processor 310, internal memory 321, external memory, display screen 394, camera 393, and wireless communication module 360. The power management module 341 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 341 may also be located within the processor 310. In other embodiments, the power management module 341 and the charging management module 340 may be housed in the same device.

[0088] The wireless communication function of the terminal device 100 can be implemented through antenna 1, antenna 2, mobile communication module 350, wireless communication module 360, modem processor and baseband processor, etc.

[0089] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0090] The mobile communication module 350 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the terminal device 100. The mobile communication module 350 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 350 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 350 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1.

[0091] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 370A, receiver 370B, etc.) or displays images or videos through the display screen 394.

[0092] The wireless communication module 360 ​​can provide solutions for wireless communication applications on the terminal device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 360 ​​can be one or more devices integrating at least one communication processing module. The wireless communication module 360 ​​receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 310. The wireless communication module 360 ​​can also receive signals to be transmitted from processor 310, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0093] In some embodiments, antenna 1 of terminal device 100 is coupled to mobile communication module 350, and antenna 2 is coupled to wireless communication module 360, enabling terminal device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0094] The terminal device 100 implements display functions through a GPU, a display screen 394, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 394 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. The processor 310 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0095] The display screen 394 is used to display images, videos, etc. The display screen 394 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a minimized display, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 100 may include one or N displays 394, where N is a positive integer greater than 1.

[0096] Terminal device 100 can perform shooting functions through ISP, camera 393, video codec, GPU, display 394 and application processor.

[0097] The ISP (Image Signal Processor) is used to process data fed back from the camera 393. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 393.

[0098] Camera 393 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the terminal device 100 may include one or N cameras 393, where N is a positive integer greater than 1.

[0099] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when terminal device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0100] Video codecs are used to compress or decompress digital video. Terminal device 100 may support one or more video codecs. Thus, terminal device 100 can play or record videos in various encoding formats.

[0101] The external memory interface 320 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external memory card communicates with the processor 310 through the external memory interface 320 to perform data storage functions.

[0102] Internal memory 321 can be used to store computer executable program code, which includes instructions. Processor 310 executes various functional applications and data processing of terminal device 100 by running the instructions stored in internal memory 321.

[0103] Terminal device 100 can implement audio functions, such as music playback and recording, through audio module 370, speaker 370A, receiver 370B, microphone 370C, headphone jack 370D, and application processor.

[0104] The audio module 370 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 370 can also be used for encoding and decoding audio signals. The speaker 370A is used to convert audio electrical signals into sound signals. The receiver 370B is used to convert audio electrical signals into sound signals. The microphone 370C is used to convert sound signals into electrical signals. The terminal device 100 may be equipped with at least one microphone 370C. The headphone jack 370D is used to connect wired headphones.

[0105] The sensor module 380 may include a pressure sensor, a touch sensor, etc. The pressure sensor senses pressure signals and converts them into electrical signals. The touch sensor, also known as a "touch panel," can be located on the display screen 394. The touch sensor and the display screen 394 together form a touchscreen, also known as a "touch screen." The touch sensor detects touch operations applied to or near it. The touch sensor can then transmit the detected touch operation to the application processor to determine the type of touch event.

[0106] Buttons 390 include a power button, volume buttons, etc. Buttons 390 can be mechanical buttons or touch-sensitive buttons. Terminal device 100 can receive button input and generate key signal inputs related to user settings and function control of terminal device 100.

[0107] Motor 391 can generate vibration alerts. Motor 391 can be used for incoming call vibration alerts or for touch vibration feedback.

[0108] Indicator 392 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0109] The SIM card interface 395 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 395 to achieve contact and separation with the terminal device 100. The terminal device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1.

[0110] The technical solution in this application will now be described with reference to the accompanying drawings, using a mobile phone as an example of the terminal device.

[0111] In current air gesture interaction solutions, if the phone is in a screen-off state, it's generally assumed that the user doesn't need to interact with the phone via air gestures. Therefore, the phone won't run the air gesture recognition process when the screen is off. If a user wants to control the phone via air gestures while the screen is off, they need to wake the phone first, that is, switch the phone from a screen-off state to a screen-on state. Currently, waking up the phone is generally done through contact interaction, such as... Figure 4 As shown, when a phone in a black screen state detects that the button 390 on the side bezel has been pressed by the user, the phone switches from a black screen state to a bright screen state and displays the unlock interface. Therefore, in scenarios where the user cannot physically touch the phone, it may be impossible to wake it up. Consequently, the phone cannot recognize the user's air gestures to achieve contactless interaction, thus causing inconvenience for the user.

[0112] When the phone screen is off and the user cannot touch the phone, the following methods can be used to wake the phone with air gestures.

[0113] like Figure 5 As shown, when the air gesture interaction function of the mobile phone is turned on and the mobile phone is in a screen-off state, the mobile phone can recognize the screen-off gesture through the following steps (steps S501-S504).

[0114] Step S501: Detect whether there is any movement in the recognition area.

[0115] The recognition area is the detectable area. In some examples, such as when the terminal device detects an image captured by a front-facing camera located on the side of the phone screen, the recognition area is the area captured by that front-facing camera.

[0116] In some embodiments, the terminal device may employ MD technology to detect movement within the recognition area.

[0117] In some examples, mobile phones use motion detection (MD) to detect whether multiple consecutive frames captured by the camera are different. If the frames are different, it indicates that there is movement in the detection area. If the frames are the same, it indicates that there is no movement in the detection area. The terminal device can accurately detect movement in the detection area using MD.

[0118] The movement could be caused by camera movement or a moving object appearing in the recognition area. Therefore, if movement is present in the recognition area, it indicates that the user intends to wake the phone using a screen-off gesture. The phone can then continue to determine whether a screen-off gesture exists in the recognition area (steps S502-S504).

[0119] Step S502: In response to detecting a movement in the recognition area, acquire an image A with size m1×n1 at frame rate a.

[0120] If movement is detected in the recognition area, it indicates that there may be a situation where the user needs to interact with the terminal device. In this case, the terminal device can acquire image A (i.e., the aforementioned fourth image) to reduce the power consumption of the terminal device in acquiring image A.

[0121] In some examples, the terminal device can acquire an image A of size m1×n1 at frame rate a (i.e., the aforementioned third frame rate).

[0122] Frame rate refers to the number of frames transmitted, displayed, or captured per second. The unit of frame rate is frames per second (fps). When a terminal device captures dynamic actions to obtain multiple frames, the more frames captured per second, the smoother the dynamic action will be when multiple frames are displayed continuously. Image size includes length and width, generally measured in pixels. For example, a 640×480 resolution image is composed of 640 pixels horizontally and 480 pixels vertically. It can be seen that the larger the image size, the more pixels are required to construct it; therefore, the larger the image size, the higher its resolution.

[0123] In some embodiments, the specific values ​​of frame rate 'a' and size can be set according to the actual application. For example, a mobile phone can obtain an image A with a size of 80×60 from the camera at a frame rate of 5fps (frames per second).

[0124] By capturing images continuously from the camera, the phone can ensure that it can recognize gestures in the images in a timely manner, thus improving the sensitivity of gesture recognition.

[0125] In some embodiments, image A can be a mono image, i.e., a grayscale image. Compared to color images, grayscale images contain less information, thus requiring less memory, and mobile phones can process grayscale images faster. Furthermore, grayscale images have only one tone, making it easier to extract texture features and avoid interference from color in image recognition, thereby highlighting the target area. Thus, terminal devices can more accurately recognize air gestures based on grayscale images.

[0126] In some embodiments, the mobile phone can also preprocess image A, wherein preprocessing may include noise reduction, enhancement, edge detection, background removal, etc. By performing the process of preprocessing image A, the mobile phone can eliminate irrelevant information in image A, restore useful real information, enhance the detectability of relevant information and simplify the data to the maximum extent, thereby improving the accuracy of screen-off gesture recognition based on image A.

[0127] Step S503: Identify image A and determine whether there is a screen-off gesture in image A.

[0128] The screen-off gesture (i.e., the aforementioned third gesture) is used to instruct the phone to switch from a screen-off state to a screen-on state. For example, when a screen-off gesture is present in image A, it indicates that the user needs to wake up the phone from its screen-off state, and the terminal device can switch from a screen-off state to a screen-on state. When a screen-off gesture is not present in image A, it indicates that the user does not need to wake up the phone from its screen-off state, and the phone remains in a screen-off state.

[0129] In some embodiments, the screen-off gesture is preset in the phone. For example, the screen-off gesture may be preset before the phone leaves the factory, or it may be a gesture entered by the user when the phone's air gesture interaction function is enabled.

[0130] In some examples, the preset screen-off gesture is a palm hover gesture, such as... Figure 6As shown in (a), when a user is in front of the screen of a phone that is in a screen-off state, i.e., within the field of view of the front-facing camera, they place the inside of their outstretched palm towards the front-facing camera. At this time, the phone, based on the air gesture obtained from image A, identifies the gesture as a palm-holding gesture, thus determining that the gesture in image A is a screen-off gesture. Figure 6 As shown in (b), the user points their hand, which makes an OK gesture, toward the front-facing camera. At this time, the air gesture obtained by the phone based on image A is the OK gesture, which means that the gesture in image A is not the screen-off gesture.

[0131] In some embodiments, the mobile phone may include a screen-off gesture recognition module, which is used to identify whether a screen-off gesture exists in image A. Since the screen-off gesture is used to wake up the phone, a static gesture is generally used as the screen-off gesture. Therefore, the screen-off gesture recognition module can use methods such as neural networks, convolutional neural networks, support vector machines (SVM), nearest neighbor algorithms, and distributed local linear embedding to identify whether a screen-off gesture exists in image A.

[0132] It should be noted that when the screen-off gesture is a dynamic gesture, the screen-off gesture recognition module can use a deep learning model to recognize the screen-off gesture.

[0133] In the above embodiments, the process of detecting whether there is a movement in the recognition area (step S501) is simpler than the process of determining whether there is a screen-off gesture in the recognition area (step S503). Therefore, the power consumption of the mobile phone is lower and the computing resources used are relatively less when performing step S501. Thus, by recognizing whether there is a movement in the recognition area, the mobile phone can determine whether it needs to continue to determine whether there is a screen-off gesture, which can effectively reduce the power consumption of the mobile phone and reduce the computing resources occupied by the mobile phone, thereby improving the smoothness of mobile phone operation.

[0134] In the above embodiments, when the terminal device is in a screen-off state, the terminal device can first detect whether there is a movement in the detection area through MD. If there is a movement in the detection area, the terminal device can then obtain image A and identify whether there is a screen-off gesture in image A.

[0135] In other embodiments provided in this application, a mobile phone in a screen-off state can also directly acquire image A and identify whether a screen-off gesture exists in image A. In some examples, the mobile phone acquires image A and identifies whether a screen-off gesture exists in image A. If no screen-off gesture exists in image A, the mobile phone can directly acquire image A again and identify whether a screen-off gesture exists in image A. If a screen-off gesture exists in image A, the mobile phone can perform subsequent steps.

[0136] Optionally, to ensure the accuracy of the screen-off gesture recognition result, the mobile phone can recognize multiple consecutive frames of images A. If the screen-off gesture is recognized in all of the consecutive frames of images A, the mobile phone switches from the screen-off state to the screen-on state. See step S504 below and related embodiments for details.

[0137] Step S504: In response to the gesture in the first frame image A being a screen-off gesture, determine whether the screen-off gesture exists in all N1 consecutive frames after the first frame image A.

[0138] In this system, the first frame, image A, is the image A in which the screen-off gesture first appears among a series of consecutive frames, image A. For example, the camera continuously acquires multiple frames, image A. After acquiring image A, the phone sequentially checks whether the screen-off gesture exists in each frame, following the order in which the images were acquired. If the phone does not recognize the screen-off gesture in image A, it re-detects whether there is movement in the recognition area to re-acquire image A; alternatively, the phone can directly re-acquire image A. If the phone recognizes the screen-off gesture in image A, that image A is designated as the first frame, and the phone continues to determine whether the screen-off gesture exists in all N1 consecutive frames following the first frame.

[0139] If the phone detects a screen-off gesture in the first frame (Image A), it can continue to determine if the screen-off gesture is present in all N1 consecutive frames following Image A to avoid false positives. In some examples, if at least one of the N1 consecutive frames following Image A does not contain a screen-off gesture, the phone can re-detect the recognition area to re-acquire Image A, or it can directly re-acquire Image A. If all N1 consecutive frames following Image A contain a screen-off gesture, it indicates that the phone has detected a user-made screen-off gesture in the recognition area.

[0140] In some examples, when a user unconsciously makes a screen-off gesture within the phone's recognition area, but the user's intention is not to wake the phone for air interaction, or when a user originally intended to wake the phone for air interaction with a screen-off gesture, but then changed their mind and gesture after making the gesture within the phone's recognition area, the image A obtained by the phone will only include a portion of image A in which the screen-off gesture actually exists. Therefore, the phone can detect images A without a screen-off gesture in the N1 consecutive frames following the first image A.

[0141] For example, if N1 is 5, and after the screen-off gesture appears in the first frame A, the gesture only appears in two consecutive frames A (the second and third frames A), but not in the fourth frame A. In this case, the phone can determine that the screen-off gesture does not appear in all five consecutive frames following the first frame A.

[0142] Understandably, in step S504, the value of N1 can affect the sensitivity and accuracy of the phone in recognizing screen-off gestures.

[0143] In some examples, the smaller the N1 value, the shorter the time required for the phone to recognize the screen-off gesture, and the higher the phone's recognition sensitivity. For instance, with image A's capture frame rate of 5 fps and N1 of 5, the phone can determine that the user has accurately and completely performed the screen-off gesture and switch from the screen-off state to the screen-on state within at least 1 second, provided the user maintains the gesture. Conversely, with image A's capture frame rate of 5 fps and N1 of 10, the phone can determine that the user has accurately and completely performed the screen-off gesture within at least 2 seconds. Therefore, it is evident that the smaller the N1 value, the higher the phone's sensitivity in recognizing screen-off gestures.

[0144] In some examples, a larger N1 value means a longer time is needed for the phone to recognize the screen-off gesture. The longer the user holds the gesture, the less likely the phone is to be woken up by an unconscious screen-off gesture, resulting in higher recognition accuracy. For instance, with image A captured at a frame rate of 5 fps and N1 of 5, the phone can determine that the user has performed a screen-off gesture and switch from screen-off to screen-on state if the user holds the gesture for 1 second. However, with image A captured at a frame rate of 5 fps and N1 of 10, the phone cannot determine that the user has accurately and completely performed the screen-off gesture if the user holds the gesture for the first second but doesn't perform it again in the second. Thus, a larger N1 value indicates higher accuracy in recognizing screen-off gestures.

[0145] Step S505: In response to the presence of a screen-off gesture in N1 consecutive frames following the first frame image A, the phone switches from the screen-off state to the screen-on state.

[0146] Based on the above embodiments, it is clear that in scenarios where the user intentionally makes a screen-off gesture, the phone can recognize the presence of a screen-off gesture in image A, or the phone can recognize the first frame image A containing the screen-off gesture, and recognize the screen-off gesture in all N1 consecutive frames following the first frame image A. Therefore, when the phone switches from a screen-off state to a screen-on state at this time, it can ensure that the phone in the screen-off state is not woken up by unintentional gestures made by the user, thereby avoiding accidental touches, wasted battery power, and other issues.

[0147] It should be noted that when a user needs to perform contactless interaction with a screen-off phone via a screen-off gesture, the phone can recognize the screen-off gesture in the captured image A and directly switch from the screen-off state to the screen-on state.

[0148] When the screen-off gesture is a palm-hold gesture, and the user makes a gesture with fingers outstretched and palm facing the camera within the phone's recognition area (e.g.) Figure 7 As shown in (a), the phone can recognize the screen-off gesture and switch from the screen-off state to the screen-on state (as shown in the image). Figure 7 After that (as shown in (b)), the mobile phone can recognize the user's air gestures.

[0149] Understandably, when a phone switches from screen-off to screen-on mode, the displayed interface can be the lock screen. At this time, the phone can be unlocked using contactless methods such as facial recognition or voiceprint recognition, allowing users to perform subsequent air gesture operations. Optionally, the phone can also be unlocked using fingerprint unlocking or password entry, depending on the user's specific needs.

[0150] In some examples, after the phone switches from a screen-off state to a screen-on state, the phone displays the unlock screen (e.g., Figure 7 As shown in (b), a prompt graphic 701 is displayed on the phone screen, indicating that the phone is currently performing facial recognition. At this time, the phone can detect the presence of the user's face image within the facial recognition area to facilitate facial recognition. The facial recognition area of ​​the phone is the camera's shooting area. Thus, the entire process from waking up to unlocking the phone is contactless with the user.

[0151] Step S506: Detect whether there is any movement in the recognition area.

[0152] The implementation method in step S506 is similar to that in step S501, and will not be repeated here. For a detailed description, please refer to the embodiment in step S501 above.

[0153] In some embodiments, when the phone screen is on, the user may not necessarily intend to interact with the phone via air gestures. Therefore, the phone can use magnetic motion detection (MD) to identify whether there is movement in the detection area to preliminarily determine whether the user intends to interact via air gestures. This movement could be caused by camera movement or the presence of moving objects in the detection area. Therefore, if the phone detects movement in the detection area using MD, it indicates that the user may intend to interact with the phone via air gestures. Conversely, if the phone detects no movement in the detection area using MD, it definitely indicates that the user does not intend to interact with the phone via air gestures.

[0154] In other words, in some situations, although a mobile phone with its screen on can detect and recognize movement in the area using MD detection, such as when a user is wiping away tears from watching a touching video, the phone can only detect and recognize movement in the area using MD detection. It cannot determine whether the user is wiping away tears or needs to operate the phone in the air. Therefore, the phone needs to further determine whether the gesture made by the user is an air interaction gesture.

[0155] In some embodiments, step S507 can be executed when the mobile phone detects movement in the recognition area.

[0156] Step S507: In response to detecting a movement in the recognition area, acquire an image B of size m2×n2 at frame rate b.

[0157] If movement is detected in the recognition area, it indicates that the user may need to interact with the terminal device while the screen is on. In this case, the terminal device can acquire image B (i.e., the aforementioned first image) to reduce the power consumption of acquiring image B.

[0158] In some embodiments, the frame rate b (i.e., the aforementioned first frame rate) is greater than the frame rate a, and the size of m2×n2 is greater than the size of m1×n1. The specific values ​​of frame rate b and size can be set according to the actual application, the value of frame rate a, and the size of image A. For example, in step S502, if the mobile phone acquires an image A of size 80×60 captured by the camera at a frame rate of 5fps, then in step S507, an image B of size 160×120 can be acquired by the camera at a frame rate of 10fps.

[0159] In step S507 above, the image acquisition frequency is increased from 5fps to 10fps compared to step S502, indicating that the mobile phone acquires more images B per unit time. Therefore, if there is action in image B, the mobile phone can detect the action more smoothly, thus improving the accuracy of the mobile phone in recognizing gestures.

[0160] Furthermore, the size of image B is increased to 160×120 compared to the size of image A, indicating that the resolution of image B is higher than that of image A. Therefore, the gestures contained in image B are clearer, and the mobile phone can achieve recognition from a greater distance and with greater accuracy based on image B.

[0161] In some embodiments, image B may also be a mono image, i.e., a grayscale image. In some embodiments, image preprocessing may also be performed on image B. See the specific embodiments described above for image A.

[0162] In some embodiments, the phone can directly acquire image B even when the screen is on.

[0163] Optionally, in some embodiments, considering scenarios where the phone's automatic screen-off function is enabled, the phone will switch from a screen-on state to a screen-off state if no user operation is detected within a certain period of time. Therefore, the phone executes step S508 before performing gesture recognition on the image to determine whether the phone is currently in a screen-on state.

[0164] Step S508: Determine whether the phone screen is currently on.

[0165] If the phone screen is currently on, the terminal device can continue to perform subsequent steps to recognize image B. If the phone screen is currently off, the phone can perform the process described in the above embodiment to recognize the off-screen gesture.

[0166] In some embodiments, before acquiring image B, the phone can determine whether the screen is currently on. If the screen is on, the phone can continue to acquire image B and recognize the gestures in image B. If the screen is off, the phone cannot interact with the user via air gestures, and therefore there is no need to acquire image B.

[0167] Step S509: In response to the phone being in a screen-on state, identify image B and determine whether the gesture in image B is an interaction start gesture.

[0168] The interaction initiation gesture (i.e. the aforementioned first gesture) is used to indicate that the user has begun to make a gesture for air interaction.

[0169] In some embodiments, the function of air gestures is to trigger the phone to perform corresponding operations, allowing users to interact with the phone using air gestures. Air gestures are generally dynamic gestures, such as swipe-up, swipe-down, swipe-left, swipe-right, press, palm flip, and pinch gestures. The interaction initiation gesture in step S509 is the initial gesture when the user performs an air gesture. For example, at the beginning of a swipe-up gesture, the user's fingers point downwards, and the back of their hand faces the phone's camera. Then, as the swipe continues upwards, the fingers gradually move upwards. After completing the swipe, the fingers point upwards, and the palm faces the phone's camera. Figure 8 As shown. Therefore, for an upward swipe gesture, the starting gesture can be the back of the hand facing the camera with fingers pointing downwards. Similarly, for a downward swipe gesture, the starting gesture can be the palm with fingers pointing upwards; for a left swipe gesture, the starting gesture can be fingers pointing to the right; for a right swipe gesture, the starting gesture can be fingers pointing to the left; for a press gesture, the starting gesture can be the palm; for a palm flip gesture, the starting gesture can be the palm or the center of the hand; and for a two-finger pinch gesture, the starting gesture can be the middle, ring, and little fingers in a contracted state. It can be understood that the starting gestures are based on the air gesture settings.

[0170] In some embodiments, the air gestures and the interaction initiation gesture can be pre-set before the phone leaves the factory. In other embodiments, the air gestures can also be determined based on user input. For example, after the air interaction function of the phone is enabled, the user is prompted to make a custom air gesture in the recognition area. The phone can obtain the user's custom air gesture and determine the interaction initiation gesture.

[0171] The phone sequentially identifies the interaction initiation gestures in the order the camera captures images B. Since the camera captures images B continuously, the phone also identifies the interaction initiation gestures in multiple consecutive frames of images B. If the phone does not recognize the interaction initiation gesture in one frame of image B, it indicates that the air gesture in the recognition area is incomplete. In this case, the phone can re-execute step S506. If the phone recognizes the interaction initiation gesture in image B, it indicates that the user intends to interact with the phone via air gestures.

[0172] In some embodiments, since the mobile phone has acquired a series of multiple frames of images B, the mobile phone can determine whether there is an interaction start gesture in the multiple frames of images B (step S510) to ensure that the mobile phone determines whether to further recognize the air interaction gesture.

[0173] Step S510: Determine whether the interaction start gesture exists in all N2 consecutive frames after the first frame image B.

[0174] Here, the first frame image B is the image B in which the interaction initiation gesture first appears among a series of consecutive frames of images B. Understandably, the value of N2 can be determined based on actual needs.

[0175] Understandably, in this embodiment, the embodiments related to step S510 are similar to those related to step 504 above, and will not be repeated here. For details, please refer to the above embodiments.

[0176] If the phone detects an interaction initiation gesture in the first frame (image B), it can continue to check if the interaction initiation gesture is present in all N2 consecutive frames following image B to avoid misjudgments. For example, after the phone detects an interaction initiation gesture in the first frame (image A), it can display a gesture recognition icon to indicate to the user that gesture recognition is in progress. If the user doesn't actually intend to interact with the phone but makes an interaction initiation gesture in the recognition area for some reason, seeing the gesture recognition icon will prompt the user to stop their hand movement, thus preventing misjudgment.

[0177] In some examples, if at least one of the N2 consecutive frames following the first image B does not include the screen-off gesture, the phone can re-use MD detection to identify the region to re-acquire image B, or the phone can directly re-acquire image B.

[0178] Understandably, in step S510, the value of N2 can affect the sensitivity and accuracy of the phone's recognition of screen-off gestures. For details, please refer to the above embodiments, which will not be repeated here.

[0179] If the interaction start gesture is present in all N consecutive frames after the first frame B, it indicates that the user wants to continue interacting with the phone via air gestures. At this time, the phone can complete the recognition of the air gestures through the following process (steps S511-S517) so as to execute the corresponding actions based on the recognized air gestures.

[0180] Step S511: In response to the presence of an interaction start gesture in all N2 consecutive frames following the first frame image B, obtain an image C with a size of m3×n3 at frame rate c.

[0181] As can be seen from the above embodiments, when a user makes a clear air gesture, the mobile phone can recognize the presence of an interaction initiation gesture in image B, or the mobile phone can recognize the first frame image B containing the interaction initiation gesture, and recognize the interaction initiation gesture in N2 consecutive frames after the first frame image B. Therefore, at this time, the mobile phone can acquire image C (i.e., the aforementioned second image) to recognize the air gesture based on image C. In this way, the mobile phone can avoid continuously performing the process of recognizing the air gesture, thereby reducing power consumption and resource consumption.

[0182] In some embodiments, the frame rate c (i.e., the aforementioned second frame rate) is greater than the frame rate b, and the size of m3×n3 is greater than the size of m2×n2. The specific values ​​of frame rate c and the size of image C can be set according to the actual application, the value of frame rate b, and the size of image B. For example, in step S502, if the mobile phone acquires an image A of size 80×60 captured by the camera at a frame rate of 5fps, and in step S507 the mobile phone acquires an image B of size 160×120 captured by the camera at a frame rate of 10fps, then in step S511, the mobile phone can acquire an image C of size 320×240 captured by the camera at a frame rate of 15fps.

[0183] In some embodiments, image C can be a mono image, i.e., a grayscale image. In some embodiments, the mobile phone can also preprocess image C. See the specific embodiments of image A described above.

[0184] Step S512: Determine whether a hand recognition frame exists.

[0185] The hand recognition bounding box (i.e., the aforementioned palm frame) is used to represent the position of the user's hand in image C; that is, the hand recognition bounding box is the position information of the hand. In some examples, the hand recognition bounding box is the smallest bounding rectangle determined based on the outer contour of the hand.

[0186] If no hand recognition box is currently present, it means the phone did not recognize a hand in the previous frame image C, or the phone did not process image C before. In this case, the phone can recognize the hand information in the current frame image C (step S513). If a hand recognition box is currently present, it means the phone recognized a hand in the previous frame image C, and the phone can process the current image C based on the hand recognition box obtained from processing the previous frame image C. In this way, the phone does not need to recognize every frame image C, thereby reducing power consumption and computational resource usage.

[0187] Step S513: In response to the absence of a hand recognition box, recognize image C to obtain the hand recognition box.

[0188] In some embodiments, if the phone does not obtain a hand recognition frame in step S513, it can be assumed that there is no hand in image C. In this case, the phone can re-acquire image B and determine whether there is an interaction initiation gesture in image B. For example, it can re-execute step S506 or re-execute the process of acquiring image B in step S507. If the hand recognition result obtained by the phone in step S513 includes a hand recognition frame, it can be assumed that there is a hand in image C.

[0189] After the hand recognition frame is determined through the above steps S512 or S513, the mobile phone can process image C based on the hand recognition frame.

[0190] Step S514: Process image C based on the hand recognition bounding box to obtain image C'.

[0191] In some embodiments, during the processing of image C based on the hand recognition bounding box, image C can be cropped based on the hand recognition bounding box to minimize the background information other than the hand in the cropped image C' (i.e., the aforementioned third image). This allows the mobile phone to more accurately recognize the user's gesture based on image C'. Furthermore, since the size of image C' is reduced after cropping, the computational resources and power consumption required for subsequent steps in recognizing the air gestures in image C' can be effectively reduced.

[0192] In some embodiments, since the air gesture is a dynamic gesture, after obtaining image C', the mobile phone can use a dynamic gesture recognition method to identify that the gesture in image C' is an air gesture. In some examples, steps S515-S517 below constitute a dynamic gesture recognition method.

[0193] Step S515: Perform static gesture classification on image C' to obtain the gesture classification results in each frame of image C'.

[0194] The gesture classification result of static gesture classification of image C' can include palm, back of hand, fist, open / closed, closed / closed, and at least one of other categories. Wherein, any category not identified belongs to the "other" category.

[0195] For example, image C' only includes... Figure 6 In the case of the gesture shown in (a), after static gesture classification, the mobile phone can obtain a gesture classification result that includes the palm. For example, in image C', only the following are included: Figure 6 In the case of the gesture shown in (b), the gesture classification results obtained by the mobile phone after static gesture classification include pinch, close, and shut.

[0196] Step S516: Perform hand key point detection on image C' to obtain hand key points.

[0197] Hand keypoints can also be understood as the joints of the hand skeleton, typically described using 21 3D keypoints. Each 3D keypoint has 3 degrees of freedom, so the hand keypoints obtained in step S516 have a dimension of 21*3. Therefore, we often use a 21*3 dimensional vector to describe them, such as... Figure 9 As shown. Thus, after the mobile phone performs hand key point detection on image C', it can determine information such as the orientation of the back of the hand and fingers based on the obtained hand key points.

[0198] In some instances, after cropping image C in step S514, the size of the cropped image can be further adjusted (resized) to obtain image C'. In some examples, steps S515 and S516 perform static gesture classification and hand keypoint detection on image C', which generally have requirements on the size of image C'. For example, multiple images C' are required to have the same size. Therefore, the mobile phone can resize the image cropped based on image C to obtain an image C' that meets the requirements, ensuring that the mobile phone can successfully complete steps S515 and S516 based on image C'.

[0199] Step S517: Determine the user's air interaction gestures based on the hand key points corresponding to each frame image C and the gesture classification results.

[0200] For example Figure 8 Take the air gesture interaction as an example. When the air gesture is detected through step S516... Figure 8 After obtaining the key points from the initial interaction gesture on the left side, the phone can make judgments based on these key points. Figure 8 The back of the hand in the interaction initiation gesture on the left side is facing up and the fingers are facing down. When step S516 is performed... Figure 8 After static gesture classification, the initial interactive gesture on the left side of image C is classified as the back of the hand. As the user's hand changes, the interactive gesture in image C gradually changes... Figure 8 The right side of the image shows the air gesture interaction. During this process, key points corresponding to each frame and hand classification results can be obtained. The phone can then calculate the trend of gesture changes based on the key points corresponding to adjacent frames. Simultaneously, based on the hand classification results, auxiliary judgment can be made to determine the gesture's direction. Figure 8 As shown in the process, the fingers gradually rise from a low position, the back of the hand gradually disappears from the recognition area, and the palm gradually appears completely in the recognition area. Then, based on the above changes, the corresponding air interaction gesture can be determined as an upward swipe gesture.

[0201] In some embodiments, when the mobile phone is performing step S517, if the mobile phone cannot determine the air interaction gesture based on the hand key points, gesture classification results and hand position information corresponding to each frame image C, the mobile phone can reacquire image C.

[0202] In some examples, when the terminal device is in a screen-on state, such as Figure 10 As shown, the air-to-air interaction method provided in this application embodiment may include the following steps:

[0203] Step S1001: Acquire the first image at the first frame rate.

[0204] Step S1002: In response to the first image including the first gesture, acquire the second image at a second frame rate, wherein the second frame rate is higher than the first frame rate.

[0205] Step S1003: Based on the gesture classification and hand key points of each frame of the second image in the multi-frame second image, identify air interaction gestures.

[0206] For details on the specific implementation of the above steps, please refer to the foregoing embodiments, which will not be repeated here.

[0207] In some embodiments, during the above process, if the mobile phone determines that the distance between the target gesture and the camera is far based on key point information, the mobile phone can acquire a larger image C. For example, if the user interacts with the mobile phone from a very far distance, such as when the detected key points indicate that the user's hand is actually 1.2 meters away from the mobile phone camera, the effect of recognizing the air gesture in step S517 is not good. Therefore, the mobile phone can control the camera to acquire an image C with a size of 480×360.

[0208] Understandably, the phone can also determine whether to control the camera to acquire a larger image C by analyzing the relationship between the hand classification results of two adjacent frames. For example, in a scenario where the user is interacting with the phone from a great distance, the resolution of the image C acquired by the phone is low, resulting in the hand being classified as the palm in the first frame and the back of the hand in the second frame. Since it is unlikely that a human hand can complete a flip within 0.06 seconds, it can be determined that the classification result is unstable, and the phone can control the camera to acquire a larger image C.

[0209] In some embodiments, air gestures include swipe up, swipe down, swipe left, swipe right, and flip gestures.

[0210] In some embodiments, after the mobile phone recognizes the air gesture, it needs to respond accordingly. In some examples, an upward swipe corresponds to swiping up the page or switching to the next video, a downward swipe corresponds to swiping down the page or switching to the previous video, a left swipe corresponds to turning the page left or going back, a right swipe corresponds to turning the page right or going forward, and a flip gesture corresponds to going back, etc.

[0211] It should be noted that the same air gesture may elicit different responses in different applications, and the specific response should be determined based on the operation of the application.

[0212] In some instances, the air interaction prompt icon can also be used to indicate the current recognition progress of the phone.

[0213] In some embodiments, after obtaining the hand key points based on step S516, the hand recognition box can be obtained through the following process (steps S518-S519):

[0214] Step S518: Determine whether the hand in image C is obscured based on the key points of the hand.

[0215] If any of the identified hand key points are missing, the phone will assume that the hand in image C is obscured. For example, if 21 hand key points should have been identified, but only 18 are actually identified, the phone will assume that the hand in image C is obscured. In this case, the phone can proceed to step S513.

[0216] Step S519: In response to the fact that the hand in image C is not occluded, obtain the hand recognition box based on the hand key points.

[0217] Since the key points of the hand are 21 3D data points, the mobile phone can calculate a bounding box of the outer contour of the palm and the position of the bounding box based on the key points, thus obtaining the hand recognition bounding box.

[0218] In some embodiments, the above process can obtain a hand recognition box based on the current frame. After the phone inputs the next frame image C in step S511, it can determine the existence of a hand recognition box in step S512. At this point, the phone can skip step S513 and directly execute step S514, thereby reducing the computing resources and power consumption of step S513. In the above embodiment, step S513 processes the larger image C, while steps S515 and S516 process the smaller image C'. Therefore, the above process effectively reduces the number of times step S513 is executed, thereby reducing power consumption.

[0219] In some examples, since the hand recognition box is obtained by the terminal device based on image C, the hand recognition box is not present in the terminal device when image C is processed for the first time. In this case, the terminal device can crop image C' from each frame of image C in the multi-frame image C. Then, based on the hand region in image C', gesture classification and hand key points can be obtained (i.e., steps S513-S516).

[0220] In some embodiments, when a terminal device crops an image C' from each of the multiple frames of images C, it can be considered that the image C' corresponding to the k-th frame image C is cropped based on the k-th frame image C.

[0221] After the terminal device begins processing image C, it can obtain the hand key points in the k-th frame of image C. Then, the terminal device can detect whether the hand in the k-th frame of image C is complete based on the hand key points; k is a positive integer. In response to the hand being complete in the k-th frame of image C, the terminal device obtains a palm frame based on the hand key points in the k-th frame of image C, and crops the (k+i)-th frame of image C based on the palm frame to obtain image C', where i is a positive integer.

[0222] In this case, the position of the hand in the (k+i)th frame image C is the same as or close to the position of the hand in the kth frame image C. Therefore, the hand in the kth frame image C is complete, indicating that the hand in the (k+i)th frame image C is also complete. The hand frame corresponding to the kth frame image C can be directly used as the hand frame corresponding to the (k+i)th frame image C. Thus, after the terminal device obtains the hand frame based on the key points of the hand in the kth frame image C, it can acquire the (k+i)th frame image C. Then, during step S512, it determines that a hand frame exists. Therefore, the terminal device can directly execute step S514.

[0223] The incomplete hand in the k-th frame image C indicates that the hand in the (k+i)-th frame image C is also incomplete. Therefore...

[0224] The terminal device can directly crop out the image C' corresponding to the (k+i)th frame image C based on the (k+i)th frame image C (i.e., execute step S513).

[0225] It should be noted that in this embodiment of the application, the terminal device needs to obtain the user's consent to turn on the camera to acquire images, including but not limited to notifying and reminding the user to read the relevant user agreement and sign the agreement that authorizes the relevant permissions before the user uses the air interaction function.

[0226] In the above embodiment, during the process of recognizing air gestures based on image C, the phone needs to perform hand recognition, static gesture classification, and hand key point detection on image C. Subsequently, it needs to determine the user's air gesture based on the obtained hand key points, gesture classification results, and hand position information. It can be seen that the process of recognizing gestures based on image C is relatively complex, and correspondingly, the computing resources and power consumption of this process are significant. Therefore, before recognizing gestures based on image C, the phone can first recognize the initial gesture of the interaction in image B. Only if the initial gesture exists in image B should it further determine whether to acquire image C and perform air gesture recognition based on image C. Recognizing the initial gesture in image B is relatively simple; therefore, the above process can effectively reduce the power consumption and computing resource consumption of the entire recognition process.

[0227] In the above embodiments, the purpose of gesture recognition based on image A when the phone is in a screen-off state is to trigger the phone to switch from a screen-off state to a screen-on state. Conversely, the purpose of gesture recognition based on image B when the phone is in a screen-on state is to determine whether further image C needs to be acquired and a more complex gesture recognition process performed based on image C. In other words, the phone has the lowest requirements for the frequency and resolution of acquiring image A, and the highest requirements for the frequency and resolution of acquiring image C. Thus, by sequentially increasing the acquisition frequency and size of the images, accurate recognition of air gestures based on image C can be ensured while reducing the acquisition frequency and overhead of images A and B, or reducing the frequency and overhead of the phone calling the camera to acquire images, thereby achieving the effect of reducing power consumption.

[0228] The above combination Figures 3-9 This application provides a detailed description of the air gesture recognition method provided in its embodiments. The following, in conjunction with... Figure 11 This application provides a detailed description of the terminal device provided in its embodiments.

[0229] In one possible design, Figure 11 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 11 As shown, the terminal device 1100 may include a display unit 1101, a processing unit 1102, and a transceiver unit 1103. The terminal device 1100 can be used to implement the functions of the terminal device involved in the above method embodiments.

[0230] Optionally, the display unit 1101 is used to support the display of interface content by the terminal device 1100, such as... Figure 2 The interface shown.

[0231] Optionally, the processing unit 1102 is configured to support the terminal device 1100 in performing operations. Figure 5 S501-S519 in the series.

[0232] Optionally, the transceiver unit 1103 is used to support the terminal device 1100 in performing... Figure 5 S501-S519 in the series.

[0233] The transceiver unit may include a receiving unit and a transmitting unit, and may be implemented by a transceiver or transceiver-related circuit components, and may be a transceiver or transceiver module. The operation and / or function of each unit in the terminal device 1100 are respectively to implement the corresponding process of the air-to-air gesture recognition method described in the above method embodiments. All relevant content of each step involved in the above method embodiments can be referred to the functional description of the corresponding functional unit, and for the sake of brevity, it will not be repeated here.

[0234] Optionally, Figure 11 The terminal device 1100 shown may also include a storage unit ( Figure 11 (Not shown in the image), this storage unit stores a program or instruction. When the processing unit 1102 and the transceiver unit 1103 execute the program or instruction, it causes... Figure 11 The terminal device 1100 shown can execute the air gesture recognition method described in the above method embodiments.

[0235] Figure 11 The technical effects of the terminal device 1100 shown can be referred to the technical effects of the air gesture recognition method described in the above method embodiments, and will not be repeated here.

[0236] In addition to being in the form of a terminal device 1100, the technical solutions provided in this application embodiment can also be functional units or chips in a terminal device, or devices used in conjunction with a terminal device.

[0237] This application also provides a computer-readable storage medium including computer instructions that, when executed on the terminal device, cause the terminal device to perform various functions or steps in the above method embodiments.

[0238] This application also provides a computer program product, including a computer program that, when run on a terminal device, causes the terminal device to perform various functions or steps in the above method embodiments.

[0239] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0240] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0241] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0242] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0243] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0244] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for recognizing gestures in air interaction, characterized in that, The method is applied to a terminal device, and when the terminal device is in a screen-on state, the method comprises: acquiring a first picture at a first frame rate; in response to the first picture including a first gesture, acquiring a second picture at a second frame rate, the second frame rate being higher than the first frame rate; when hand key points of the kth frame of the second picture are obtained, detecting whether a hand in the kth frame of the second picture is complete based on the hand key points; k is a positive integer; in response to the hand in the kth frame of the second picture being complete, obtaining a palm frame based on the hand key points in the kth frame of the second picture; cropping the k+i th frame of the second picture based on the palm frame to obtain a plurality of third pictures, the third pictures including a hand region, i is a positive integer; obtaining a gesture classification and hand key points based on the hand region in each of the third pictures; identifying a midair interaction gesture based on the gesture classification and the hand key points of each of the plurality of second pictures.

2. The method of claim 1, wherein, The size of the second picture is greater than the size of the first picture.

3. The method of claim 1, wherein, The identifying of the midair interaction gesture based on the gesture classification and the hand key points of each of the plurality of second pictures comprises: in response to the first frame of the plurality of second pictures including a hand, identifying a midair gesture based on the gesture classification and the hand key points of each of the plurality of second pictures.

4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: in response to the first frame of the plurality of second pictures not including a hand, returning to the step of acquiring the first picture at the first frame rate.

5. The method according to any one of claims 1-3, characterized in that, The method further comprises: in response to the hand in the kth frame of the second picture being incomplete, cropping the k+i th frame of the second picture to obtain a third picture corresponding to the k+i th frame of the second picture.

6. The method according to any one of claims 1-3, characterized in that, The method comprises: cropping the kth frame of the second picture to obtain a third picture corresponding to the kth frame of the second picture.

7. The method according to any one of claims 1-3, characterized in that, The method further comprises: in response to detecting a moving action, acquiring the first picture at the first frame rate.

8. The method of any one of claims 1-3, wherein, The first picture has a plurality of frames, and the determination that the first picture includes the first gesture comprises: determining that consecutive N frames of the first picture all include the first gesture, N being a positive integer greater than 2.

9. The method of any one of claims 1-3, wherein, Before acquiring the first picture at the first frame rate, the method further comprises: when the terminal device is in a screen-off state, acquiring a fourth picture at a third frame rate, and in response to the fourth picture including a second gesture, switching the terminal device to a screen-on state.

10. The method of claim 9, wherein, The determination that the fourth picture includes the second gesture comprises: determining that consecutive M frames of the fourth picture all include the second gesture, M being a positive integer greater than 2.

11. A terminal device, comprising: The terminal device comprises a display screen, a memory and one or more processors; the display screen, the memory and the processors are coupled; the display screen is configured to display images generated by the processors, the memory is configured to store computer program codes, the computer program codes comprise computer instructions; when the processors execute the computer instructions, the terminal device performs the method according to any one of claims 1-10.

12. A computer-readable storage medium, characterized in that, computer program comprising computer instructions which, when the computer instructions are executed on a terminal device, cause the terminal device to perform the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Air gesture recognition method, electronic equipment and storage medium

    CN116301363A

  • Gesture recognition method, electronic device, computer-readable storage medium, and chip

    US20220198836A1