Gesture recognition method, electronic device and storage medium

Through the Doppler frequency shift principle and noise filtering technology, multiple microphones are used to receive ultrasonic echoes and perform amplitude accumulation and normalization processing, the problems of short ultrasonic gesture recognition distance and poor anti-noise interference capabilities are solved, and gesture recognition with longer distances and higher accuracy is achieved.

CN115494935BActive Publication Date: 2025-08-29HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110679683.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-18
Publication Date
2025-08-29
Estimated Expiration
2041-06-18

AI Technical Summary

Technical Problem

The existing ultrasonic gesture recognition technology has the problems of short recognition distance and poor anti-noise interference capabilities, especially in long-distance and noise environments.

Method used

By using multiple microphones to receive ultrasonic echoes, and perform amplitude accumulation and normalization processing in the gesture recognition area, combining the Doppler shift principle and noise filtering technology, the distance and accuracy of gesture recognition are improved.

Benefits of technology

It effectively expands the distance of gesture recognition, improves the recognition efficiency and accuracy in noise environments, and enhances anti-interference ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115494935B_ABST
    Figure CN115494935B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a gesture recognition method, electronic device, and storage medium, relating to the field of communications technology. The method comprises: transmitting ultrasonic waves of a preset frequency; using multiple microphones to respectively receive ultrasonic echoes corresponding to the ultrasonic waves to obtain multiple first acoustic wave signals; accumulating the multiple first acoustic wave signals according to a gesture recognition area to obtain a second acoustic wave signal, wherein the gesture recognition area is a frequency area associated with the preset frequency, and the amplitude of the second acoustic wave signal in the gesture recognition area is greater than the amplitude of each first acoustic wave signal in the gesture recognition area; and performing gesture recognition based on the second acoustic wave signal. The method provided by embodiments of the present application can improve the recognition distance of gestures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of communication technologies, and in particular to a gesture recognition method, an electronic device, and a storage medium. Background Art

[0002] With the development of information technology, gestures are becoming an increasingly popular way to interact with devices such as computers, mobile phones, and speakers. Currently, two technologies are commonly used to achieve gesture recognition. The first is computer vision technology. While advances in computer vision enable efficient and accurate gesture recognition, it is sensitive to conditions such as lighting, is easily limited by the visual field, and is relatively complex to deploy. The second is ultrasonic gestures. This method uses a speaker in a device to emit ultrasonic waves. These waves, when reflected by moving objects (such as hand gestures) in the air within a certain distance, generate a Doppler shift. A microphone in the device receives the reflected signal and uses the Doppler shift to infer various gestures.

[0003] However, as ultrasonic waves propagate through a medium, their energy gradually weakens with increasing distance, a phenomenon known as ultrasonic attenuation. Long-distance recognition is poor primarily because the energy of the echo reflected by a moving object attenuates more significantly with increasing distance. Existing ultrasonic gesture recognition generally does not exceed a range of 2 meters, meaning it has a relatively short range. Furthermore, ultrasonic gesture recognition is susceptible to interference from full-band noise, single-frequency noise, and the device's own sound. Therefore, its noise immunity is relatively poor. Summary of the Invention

[0004] Embodiments of the present application provide a gesture recognition method, an electronic device, and a storage medium to provide a gesture recognition method, thereby improving the gesture recognition distance.

[0005] In a first aspect, an embodiment of the present application provides a gesture recognition method, which is applied to an electronic device, the electronic device including multiple microphones, including:

[0006] Send ultrasonic waves of a preset frequency; wherein the preset frequency can be the center frequency of the ultrasonic wave, for example, 24KHZ, 48KHz or 96KHz.

[0007] Multiple microphones are used to respectively receive ultrasonic echoes corresponding to the ultrasonic waves to obtain multiple first sound wave signals; wherein, if the ultrasonic wave is blocked by the palm of the user on its propagation path, the ultrasonic wave will generate an echo.

[0008] Accumulating and processing the plurality of first sound wave signals according to the gesture recognition area to obtain a second sound wave signal, wherein the gesture recognition area is a frequency area associated with a preset frequency, and an amplitude of the second sound wave signal in the gesture recognition area is greater than an amplitude of each first sound wave signal in the gesture recognition area.

[0009] Perform gesture recognition based on the second sound wave signal.

[0010] In the embodiment of the present application, by accumulating the amplitudes of multiple microphones in the gesture recognition area, the energy of the gesture recognition area can be increased, thereby effectively recognizing the user's gesture at a farther distance.

[0011] In one possible implementation, the gesture recognition area is a frequency area determined based on a Doppler shift principle and a preset frequency and gesture speed.

[0012] One possible implementation method also includes:

[0013] When the amplitude of the second sound wave signal in the gesture recognition area is greater than the preset amplitude, it is determined that the gesture is detected.

[0014] In the embodiment of the present application, gestures can be effectively recognized by determining the amplitude of the gesture recognition area, thereby improving the efficiency of gesture recognition.

[0015] In one possible implementation, the gesture recognition area includes a distant gesture recognition area located to the left of a preset frequency and / or a close gesture recognition area located to the right of the preset frequency; and further includes:

[0016] When the amplitude of the second sound wave signal in the away-gesture recognition area is greater than a preset amplitude, determining that a away gesture is detected;

[0017] When the amplitude of the second sound wave signal in the proximity gesture recognition area is greater than the preset amplitude, it is determined that the proximity gesture is detected.

[0018] In the embodiment of the present application, the flexibility of gesture recognition can be improved by distinguishing the types of gesture recognition areas and determining different types of gestures.

[0019] In one possible implementation, accumulating multiple first sound wave signals according to the gesture recognition area to obtain a second sound wave signal includes:

[0020] performing normalization processing on the plurality of first acoustic wave signals respectively to obtain a plurality of normalized first acoustic wave signals;

[0021] A second sound wave signal is obtained by accumulating multiple normalized first sound wave signals according to a gesture recognition area, wherein the gesture recognition area is a frequency area associated with a preset frequency, and an amplitude of the second sound wave signal in the gesture recognition area is greater than an amplitude of each normalized first sound wave signal in the gesture recognition area.

[0022] In the embodiment of the present application, by normalizing the amplitude of the gesture recognition area, the absolute value difference of the amplitude caused by the different positions of multiple microphones can be reduced, thereby improving the recognition accuracy.

[0023] In one possible implementation, performing gesture recognition based on the second sound wave signal specifically includes:

[0024] performing noise filtering on the second sound wave signal to obtain a noise-filtered second sound wave signal;

[0025] Gesture recognition is performed based on the second sound wave signal after noise filtering.

[0026] In the embodiment of the present application, by filtering the noise, the anti-interference capability can be improved, thereby improving the recognition accuracy.

[0027] In one possible implementation, the noise filtering process includes at least one of ambient noise filtering, single-frequency noise filtering, and electronic device's own sound filtering.

[0028] In one possible implementation, sending ultrasonic waves of a preset frequency specifically includes:

[0029] In response to a detected preset trigger event, ultrasonic waves of a preset frequency are transmitted, wherein the preset trigger event is any one of an alarm ringing, an incoming call, a timed reminder, and a light switch event.

[0030] One possible implementation method also includes:

[0031] If a gesture is detected, perform any of the following operations: silence the alarm, answer or reject the call, silence the timer reminder, or turn on or off the light.

[0032] In one possible implementation, the second sound wave signal further includes a non-gesture recognition area, and the amplitude of the non-gesture recognition area of ​​the second sound wave signal is determined by an average of the amplitudes of the non-gesture recognition areas of the plurality of first sound wave signals.

[0033] In one possible implementation, the electronic device is any one of a smart speaker, a mobile phone, a large screen, a tablet, and a smart switch.

[0034] In a second aspect, an embodiment of the present application provides a gesture recognition device, which is applied to an electronic device. The electronic device includes multiple microphones, including:

[0035] A sending module, used for sending ultrasonic waves of a preset frequency;

[0036] A receiving module, configured to use a plurality of microphones to respectively receive ultrasonic echoes corresponding to the ultrasonic waves to obtain a plurality of first sound wave signals;

[0037] an accumulation module, configured to accumulate the plurality of first sound wave signals according to the gesture recognition area to obtain a second sound wave signal, wherein the gesture recognition area is a frequency area associated with a preset frequency, and an amplitude of the second sound wave signal in the gesture recognition area is greater than an amplitude of each first sound wave signal in the gesture recognition area;

[0038] The recognition module is used to perform gesture recognition based on the second sound wave signal.

[0039] In one possible implementation, the gesture recognition area is a frequency area determined based on a Doppler shift principle and a preset frequency and gesture speed.

[0040] In one possible implementation, the apparatus further includes:

[0041] The first determining module is configured to determine that a gesture is detected when the amplitude of the second sound wave signal in the gesture recognition area is greater than a preset amplitude.

[0042] In one possible implementation, the gesture recognition area includes a distant gesture recognition area located to the left of a preset frequency and / or a close gesture recognition area located to the right of the preset frequency, and the apparatus further includes:

[0043] a second determining module, configured to determine that a distance gesture is detected when the amplitude of the second sound wave signal in the distance gesture recognition area is greater than a preset amplitude;

[0044] When the amplitude of the second sound wave signal in the proximity gesture recognition area is greater than the preset amplitude, it is determined that the proximity gesture is detected.

[0045] In one possible implementation, the accumulation module is further configured to perform normalization processing on the plurality of first sound wave signals respectively to obtain a plurality of normalized first sound wave signals;

[0046] A second sound wave signal is obtained by accumulating multiple normalized first sound wave signals according to a gesture recognition area, wherein the gesture recognition area is a frequency area associated with a preset frequency, and an amplitude of the second sound wave signal in the gesture recognition area is greater than an amplitude of each normalized first sound wave signal in the gesture recognition area.

[0047] In one possible implementation, the recognition module is further configured to perform noise filtering on the second sound wave signal to obtain a noise-filtered second sound wave signal;

[0048] Gesture recognition is performed based on the second sound wave signal after noise filtering.

[0049] In one possible implementation, the noise filtering process includes at least one of ambient noise filtering, single-frequency noise filtering, and electronic device's own sound filtering.

[0050] In one possible implementation, the sending module is further used to send ultrasonic waves of a preset frequency in response to a detected preset trigger event, wherein the preset trigger event is any one of an alarm ringing, an incoming call, a timed reminder, and a light switch event.

[0051] In one possible implementation, the apparatus further includes:

[0052] The execution module is used to execute any one of the operations of turning off the alarm, answering or rejecting a call, turning off a timer reminder, or turning on or off a light when it is determined that a gesture is detected.

[0053] In one possible implementation, the second sound wave signal further includes a non-gesture recognition area, and the amplitude of the non-gesture recognition area of ​​the second sound wave signal is determined by an average of the amplitudes of the non-gesture recognition areas of the plurality of first sound wave signals.

[0054] In one possible implementation, the electronic device is any one of a smart speaker, a mobile phone, a large screen, a tablet, and a smart switch.

[0055] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0056] A memory, wherein the memory is used to store computer program code, the computer program code includes instructions, and the electronic device includes multiple microphones. When the electronic device reads the instructions from the memory, the electronic device performs the following steps:

[0057] Sending ultrasonic waves of preset frequency;

[0058] Using multiple microphones to receive ultrasonic echoes corresponding to the ultrasonic waves respectively to obtain multiple first sound wave signals;

[0059] Accumulating the plurality of first sound wave signals according to the gesture recognition area to obtain a second sound wave signal, wherein the gesture recognition area is a frequency area associated with a preset frequency, and an amplitude of the second sound wave signal in the gesture recognition area is greater than an amplitude of each first sound wave signal in the gesture recognition area;

[0060] Perform gesture recognition based on the second sound wave signal.

[0061] In one possible implementation, the gesture recognition area is a frequency area determined based on a Doppler shift principle and a preset frequency and gesture speed.

[0062] In one possible implementation, when the instruction is executed by the electronic device, the electronic device further performs the following steps:

[0063] When the amplitude of the second sound wave signal in the gesture recognition area is greater than the preset amplitude, it is determined that the gesture is detected.

[0064] In one possible implementation, the gesture recognition area includes a distant gesture recognition area located to the left of a preset frequency and / or a close gesture recognition area located to the right of the preset frequency. When the above instruction is executed by the above electronic device, the above electronic device further performs the following steps:

[0065] When the amplitude of the second sound wave signal in the away-gesture recognition area is greater than a preset amplitude, determining that a away gesture is detected;

[0066] When the amplitude of the second sound wave signal in the proximity gesture recognition area is greater than the preset amplitude, it is determined that the proximity gesture is detected.

[0067] In one possible implementation, when the instruction is executed by the electronic device, the electronic device performs accumulation processing on the plurality of first sound wave input signals according to the gesture recognition area to obtain the second sound wave signal, including:

[0068] performing normalization processing on the plurality of first acoustic wave signals respectively to obtain a plurality of normalized first acoustic wave signals;

[0069] A second sound wave signal is obtained by accumulating multiple normalized first sound wave signals according to a gesture recognition area, wherein the gesture recognition area is a frequency area associated with a preset frequency, and an amplitude of the second sound wave signal in the gesture recognition area is greater than an amplitude of each normalized first sound wave signal in the gesture recognition area.

[0070] In one possible implementation, when the instruction is executed by the electronic device, the electronic device performs the step of performing gesture recognition based on the second sound wave signal, including:

[0071] performing noise filtering on the second sound wave signal to obtain a noise-filtered second sound wave signal;

[0072] Gesture recognition is performed based on the second sound wave signal after noise filtering.

[0073] In one possible implementation, the noise filtering process includes at least one of ambient noise filtering, single-frequency noise filtering, and electronic device's own sound filtering.

[0074] In one possible implementation, when the instruction is executed by the electronic device, causing the electronic device to execute the step of transmitting ultrasonic waves of a preset frequency includes:

[0075] In response to a detected preset trigger event, ultrasonic waves of a preset frequency are transmitted, wherein the preset trigger event is any one of an alarm ringing, an incoming call, a timed reminder, and a light switch event.

[0076] In one possible implementation, when the instruction is executed by the electronic device, the electronic device further performs the following steps:

[0077] If a gesture is detected, perform any of the following operations: silence the alarm, answer or reject the call, silence the timer reminder, or turn on or off the light.

[0078] In one possible implementation, the second sound wave signal further includes a non-gesture recognition area, and the amplitude of the non-gesture recognition area of ​​the second sound wave signal is determined by an average of the amplitudes of the non-gesture recognition areas of the plurality of first sound wave signals.

[0079] In one possible implementation, the electronic device is any one of a smart speaker, a mobile phone, a large screen, a tablet, and a smart switch.

[0080] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer-readable storage medium is run on a computer, the computer executes the method described in the first aspect.

[0081] In a fifth aspect, an embodiment of the present application provides a computer program, which, when executed by a computer, is used to execute the method described in the first aspect.

[0082] In one possible design, the program in the fifth aspect may be stored in whole or in part on a storage medium packaged with the processor, or may be stored in whole or in part on a memory not packaged with the processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 A schematic diagram of an application scenario is provided for the embodiment of the present application;

[0084] Figure 2 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0085] Figure 3A flowchart of a gesture recognition method provided in an embodiment of the present application;

[0086] Figure 4 A schematic diagram of the gesture recognition area division provided in an embodiment of the present application;

[0087] Figure 5 Provides a schematic diagram of gesture recognition area accumulation for the embodiment of the present application;

[0088] Figure 6a and Figure 6b A schematic diagram of environmental noise filtering provided in an embodiment of the present application;

[0089] Figure 7 A schematic diagram of single-frequency noise filtering provided in an embodiment of the present application;

[0090] Figure 8 A schematic diagram of the gesture recognition effect provided in an embodiment of the present application;

[0091] Figure 9 This is a schematic diagram of the structure of a gesture recognition device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0092] The following describes the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " represents "or." For example, A / B can represent A or B. "And / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, or B exists alone.

[0093] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0094] Computer vision recognition requires a camera installed on the device and a certain level of data processing capabilities, which is costly. Ultrasonic gesture recognition, on the other hand, is universal and can be implemented on most devices with limited computing power. For example, it can utilize the device's built-in speaker and microphone for gesture recognition without requiring any modifications to the device's existing hardware.

[0095] However, as ultrasonic waves propagate through a medium, their energy gradually weakens with increasing distance, a phenomenon known as ultrasonic attenuation. Long-distance recognition is poor primarily because the energy of the echo reflected by a moving object attenuates more significantly with increasing distance. Existing ultrasonic gesture recognition generally does not exceed a range of 2 meters, meaning it has a relatively short range. Furthermore, ultrasonic gesture recognition is susceptible to interference from full-band noise, single-frequency noise, and the device's own sound. Therefore, its noise immunity is relatively poor.

[0096] Based on the above problems, an embodiment of the present application proposes a gesture recognition method that can improve the distance of ultrasonic gesture recognition.

[0097] Now combined Figures 1-8 The gesture recognition method provided in the embodiment of the present application is described as follows. Figure 1 The following is an application scenario provided by the embodiment of the present application, refer to Figure 1 The above application scenario includes a device 10 and a user 20. The device 10 may include one or more speakers and multiple microphones. The speakers may transmit ultrasonic waves, and the microphones may receive echoes of the ultrasonic waves after they are reflected by the user 20. For example, the device 10 may include four or six speakers and four or six microphones.

[0098] The device 10 may be a smart device with a speaker and a microphone, such as a smart speaker. It should be understood that the smart speaker is merely an example. In some embodiments, the device 10 may be any other smart device with a speaker and a microphone, including but not limited to a mobile phone, tablet, large screen, smart robot, smart switch, etc. The embodiments of this application do not specifically limit the specific form of the smart device 10.

[0099] The following combination Figure 2 First, an exemplary electronic device provided in the following embodiments of the present application is introduced. Figure 2 A structural schematic diagram of an electronic device 100 is shown, and the electronic device 100 may be the above-mentioned device 10.

[0100] The electronic device 100 may include a processor 110 , an external memory interface 120 , an internal memory 121 , a universal serial bus (USB) interface 130 , a charging management module 140 , a power management module 141 , a battery 142 , an audio module 150 , a speaker 150A, a microphone 150B, and a wireless communication module 160 .

[0101] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0102] The processor 110 may include one or more processing units, for example: the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors. Among them, the controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals based on instruction opcodes and timing signals to complete the control of instruction fetching and execution.

[0103] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0104] The execution of the gesture recognition method provided in the embodiment of the present application can be controlled by the processor 110 or completed by calling other components, such as calling the processing program of the embodiment of the present application stored in the internal memory 121 to realize gesture recognition of the user.

[0105] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0106] The USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as augmented reality devices.

[0107] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0108] The wireless communication module 160 can provide wireless communication solutions for the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module.

[0109] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0110] The internal memory 121 can be used to store computer executable program codes, which include instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor.

[0111] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 150, speaker 150A, microphone 150B, and application processor. In some embodiments, the electronic device 100 may include one or S speakers 150A; in addition, the electronic device 100 may also include S microphones 150B, where S is a positive integer greater than 1.

[0112] like Figure 3 The figure shows a flow chart of an embodiment of the gesture recognition method provided by the embodiment of the present application, including:

[0113] Step 301: The device 10 plays ultrasonic waves.

[0114] Specifically, the ultrasonic wave can be played through the speaker in device 10. The ultrasonic wave played by device 10 can be a preset ultrasonic audio file stored in the memory of device 10. The frequency (or center frequency) of the ultrasonic wave can be pre-set, for example, 24 kHz, 48 kHz, or 96 kHz. It will be understood that the embodiments of the present application do not specifically limit the preset frequency of the ultrasonic wave. In a specific implementation, the trigger for the device 10 to play the ultrasonic wave can be an event. For example, taking an alarm clock as an example, if a user sets an alarm on device 10, when the alarm of device 10 goes off, device 10 can simultaneously play the ultrasonic wave. That is, the playback of the ultrasonic wave by device 10 can be triggered by the alarm ringing event. Therefore, if a user makes a gesture near device 10, device 10 can recognize the user's gesture by receiving the ultrasonic wave echo, determine whether the gesture corresponds to a gesture to turn off the alarm, and determine whether to turn off the alarm. It will be understood that the above example only illustrates the alarm clock scenario and does not constitute a limitation of the embodiments of the present application. In some embodiments, the device 10 can also be triggered to play the ultrasonic wave by other scenarios. For example, in an incoming call scenario, when device 10 receives an incoming call, the device 10 can trigger the playback of the ultrasonic wave based on the incoming call event. Another example is a timed reminder scenario, when a timed reminder is issued on device 10, the device 10 can trigger the playback of the ultrasonic wave based on the timed reminder event. Another example is a light on / off scenario, when device 10 turns a light on or off, the device 10 can trigger the playback of the ultrasonic wave based on the light on / off event.

[0115] In step 302 , the device 10 receives ultrasonic echoes, normalizes the amplitudes of the ultrasonic echoes, and accumulates the normalized amplitudes of the gesture recognition area to obtain normalized ultrasonic echoes.

[0116] Specifically, when the speaker of device 10 plays ultrasonic waves, if the ultrasonic waves are blocked by the palm of user 20 along their propagation path, the ultrasonic waves will generate echoes. At this point, the multiple microphones in device 10 can collect these ultrasonic echoes, thereby obtaining echo data from each microphone. This echo data can be used to recognize the user's gestures.

[0117] It is understood that the above echo data is raw audio data, that is, time domain data. Therefore, the above time domain data can be subjected to a Fast Fourier Transform (FFT) to obtain frequency domain data. The above example only illustrates how to obtain frequency domain data through FFT and does not constitute a limitation of the embodiments of the present application. In some embodiments, frequency domain data can also be obtained through other conversion methods.

[0118] After acquiring the frequency domain data corresponding to the time domain microphone 12 data collected by each microphone, the amplitude of the frequency domain data corresponding to the time domain data collected by each microphone in the gesture recognition area can be accumulated, thereby obtaining an accumulated ultrasonic echo. The specific meaning of the gesture recognition area will be explained in detail in the subsequent sections. By accumulating the amplitudes of the gesture recognition area corresponding to the frequency domain data collected by multiple microphones, the amplitudes of the gesture recognition area can be superimposed. In other words, the frequency domain energy of the gesture recognition area can be expanded, making it easier for device 10 to recognize gestures based on the frequency domain data. This can improve the device 10's recognition distance for gestures, that is, the device 10 can recognize gestures at a greater distance.

[0119] Preferably, after obtaining the frequency domain data corresponding to the time domain data collected by each microphone, the amplitude of the frequency domain data corresponding to the time domain data collected by each microphone can be normalized to obtain the normalized amplitude of the frequency domain data corresponding to the time domain data collected by each microphone. Subsequently, the normalized amplitudes of the gesture recognition areas of all microphones can be accumulated to obtain a normalized accumulated ultrasonic echo.

[0120] It should be noted that gestures can only be recognized by recognizing the frequency domain data in the gesture recognition area, and the frequency domain data in the gesture recognition areas of multiple microphones can be accumulated to expand the frequency domain energy of the gesture recognition area. Therefore, the frequency domain data outside the gesture recognition area is not accumulated. This can prevent the frequency domain data in other areas from interfering with the frequency domain data in the gesture recognition area after the energy is enhanced.

[0121] For example, the normalized amplitude accumulation method can be implemented by the following formula:

[0122] Among them, E sum is the normalized amplitude accumulated in the gesture recognition area, E iis the normalized amplitude in the gesture recognition area of ​​the i-th microphone, Ei,max is the maximum amplitude in the frequency domain data corresponding to the time domain data collected by the i-th microphone, and n is the total number of microphones in device 10. In a specific implementation, the maximum amplitude Ei,max can be the amplitude corresponding to the center frequency of each microphone. In other words, the maximum amplitude Ei,max is the maximum amplitude in the entire frequency domain of the frequency domain data corresponding to the time domain data collected by each microphone. After summing the normalized amplitudes of the gesture recognition areas of all the microphones, the normalized amplitudes of the gesture recognition areas of the multiple microphones can be combined into a normalized amplitude accumulation value Esum. In other words, the normalized amplitude of the gesture recognition area is amplified by summing the normalized amplitudes of the multiple microphones.

[0123] It should be noted that the frequency domain data outside the gesture recognition area can be obtained by averaging the frequency domain data outside the gesture recognition area of ​​multiple microphones. For example, the amplitude of the area outside the gesture recognition area of ​​multiple microphones can be calculated to obtain the average value. Taking two microphones as an example, E 11 is the amplitude of the area outside the gesture recognition area of ​​the first microphone, E 12 is the amplitude of the area outside the gesture recognition area of ​​the second microphone, then the average value of the amplitudes outside the gesture recognition areas of the multiple microphones can be (E 11 +E 12 ) / 2.

[0124] It is understandable that each microphone is located at a different position in the device 10 (for example, arranged in a ring), so that each microphone is at a different distance from the user. Ultrasonic waves attenuate quickly in the air, resulting in a large difference in the absolute value of the amplitude of the ultrasonic echo received by each microphone. This is the value obtained after amplitude normalization. Since the value obtained after normalization represents the percentage of the amplitude collected by each microphone relative to the amplitude of the ultrasonic center frequency, it can avoid the direct accumulation of the absolute values ​​of the amplitudes, and thus avoid the differences in the absolute values ​​of the amplitudes caused by differences between microphones. This can make the statistics more accurate, more in line with the principles of the gesture recognition algorithm, and better extend the gesture recognition distance.

[0125] Now combined Figure 4 The gesture recognition area in ultrasonic echo is described. Figure 4As shown, the frequency domain waveform 400 is the frequency domain data obtained after normalization of the ultrasonic echo collected by any microphone. Among them, the horizontal axis of the frequency domain waveform 400 is frequency, and the vertical axis is amplitude. The frequency domain waveform 400 includes a slow gesture area 410, a gesture recognition area 420 and a center frequency point 430. It can be understood that the center frequency point 430 is the center frequency of the ultrasonic echo. Usually, the energy near the center frequency is the highest, so the amplitude near the center frequency point 430 is the highest (such as Figure 4 The peak shown in FIG. 4 is a graph showing the frequency of the slow gesture region 410. The frequency points of the slow gesture region 410 correspond to abnormal gestures, such as those in slow-moving scenes such as walking. In actual applications, the device 10 may mistakenly identify walking as a gesture when detecting a user walking. Walking is a scene that does not require recognition for gesture recognition. Therefore, it is necessary to distinguish between normal user gestures and walking. Therefore, embodiments of the present application can filter the slow gesture region 410 and only perform gesture recognition in the gesture recognition region 420.

[0126] The above-mentioned method of determining the gesture recognition area 420 can be based on the Doppler formula. For example, when the user approaches the device 10, the Doppler formula is as follows:

[0127]

[0128] When the user is far away from the device 10, the Doppler equation is as follows:

[0129]

[0130] Among them, f r is the receiving frequency of the microphone, f t is the original frequency of the speaker, that is, the center frequency of the played ultrasound, C is the speed of sound, and v is the user's hand speed.

[0131] In the specific implementation, Figure 4 For example, when the user approaches the device 10, assuming that the user moves slowly at a speed v1=1m / s, then substituting v1 into the above Doppler formula (1) can obtain the frequency point fr1; then, assuming that the user's maximum hand speed is v2=6m / s, then substituting v2 into the above Doppler formula (2) can obtain the frequency point fr2; after obtaining the above frequency points fr1 and fr2, it can be determined that the frequency domain interval between the frequency points fr1 and fr2 is the gesture recognition area 420, and the gesture recognition area 420 includes the frequency domain corresponding to the user's normal hand speed (for example, the normal hand speed is between v1 and v2). Within the gesture recognition area 420, the user's gesture can be recognized.

[0132] Next, when the user moves away from the device 10, v1 and v2 are substituted into the Doppler formula (2) to obtain frequency points fr3 and fr4. After obtaining the above frequency points fr3 and fr4, it can be determined that the frequency domain interval between frequency points fr3 and fr1 is a slow gesture area 410. The slow gesture area 410 includes the frequency domain corresponding to a hand speed less than the moving speed v1. Within the slow gesture area 410, there is no need to recognize the user's gesture; the frequency domain interval between frequency points fr3 and fr4 is a gesture recognition area 420.

[0133] Next, combine Figure 5 , take two microphones as an example (for example, microphone 1 and microphone 2) for explanation, Figure 5 Figure 5 is a schematic diagram of normalized amplitude accumulation. Frequency domain waveform 510 is the frequency domain data obtained by normalizing the ultrasonic echo collected by microphone 1. Frequency domain waveform 510 includes a gesture recognition region 511. Frequency domain waveform 520 is the frequency domain data obtained by normalizing the ultrasonic echo collected by microphone 2. Frequency domain waveform 520 includes a gesture recognition region 521. By summing the amplitudes of the gesture recognition regions of microphone 1 and microphone 2, frequency domain waveform 500 is obtained. It can be seen that the normalized amplitude of gesture recognition region 501 in frequency domain waveform 500 is higher than that of both the gesture recognition regions of microphone 1 and microphone 2. In other words, by accumulating the normalized amplitudes of the ultrasonic echoes from multiple microphones in the gesture recognition regions, the ultrasonic echoes can be enhanced, thereby facilitating gesture recognition.

[0134] It should be noted that the following description uses normalized amplitude as an example, but the present application does not limit the normalization of ultrasonic echoes. In other words, normalization may not be performed, but the amplitude of the ultrasonic echo may be directly accumulated and then the noise may be filtered out in the following steps.

[0135] In step 303 , the device 10 filters the ambient noise.

[0136] Specifically, in order to accurately recognize the user's gesture, the device 10 can filter the ambient noise in the gesture recognition area, thereby improving the anti-interference ability and making the gesture recognition more accurate. The filtering of the ambient noise can be performed using the following formula:

[0137] E i '=E sum -E noise Among them, E i ' is the amplitude of the gesture recognition area after filtering the ambient noise; E noiseThe average amplitude of the ambient noise floor is E. It is understood that the average amplitude of the ambient noise floor can be the average value of the normalized amplitude of the ambient noise floor. noise It can be a gesture recognition area (such as Figure 4 The average amplitude of M frequency points on both sides outside the area 420).

[0138] It should be noted that the value of M can be pre-set. In specific implementation, if the value of M is too small, it may lead to insufficient sample size. If the value of M is too large, it may lead to excessive calculation. Therefore, the value of M can be an empirical value.

[0139] Next, combine Figure 6a and Figure 6b For example, Figure 6a As shown, the frequency domain waveform 600 is a frequency domain waveform obtained by normalizing and accumulating the frequency domain data collected by multiple microphones in the gesture recognition area. The frequency domain waveform 600 includes a slow gesture area 610, a gesture recognition area 620, an ambient noise area 630, and a center frequency point 640. Next, the gesture recognition area 620 can be filtered for ambient noise, wherein the filtering of the ambient noise can be performed by subtracting the average amplitude of the ambient noise from the normalized amplitude accumulation value in the gesture recognition area 620. Figure 6a After filtering the ambient noise in the gesture recognition area 620 shown in FIG. Figure 6b The frequency domain waveform 601 after filtering the ambient noise is shown. By filtering the ambient noise, the anti-interference capability can be improved.

[0140] In step 304 , the device 10 filters single-frequency noise.

[0141] Optionally, during the gesture recognition process, noise that continuously affects one or more frequencies may occur, for example, the noise caused by the fan in a device near the device 10 (such as a computer). Since the fan usually operates at a fixed frequency, it will cause noise at a fixed frequency. Therefore, the device 10 can also filter these single-frequency noises, thereby improving the accuracy of gesture recognition. The above-mentioned single-frequency noise filtering can be performed by an Alpha filtering method. It will be understood that the above example only illustrates the filtering of single-frequency noise by Alpha filtering, and does not constitute a limitation of the embodiments of the present application. In some embodiments, single-frequency noise can also be filtered by other methods.

[0142] In specific implementation, since the characteristics of the above-mentioned single-frequency noise are that it only affects specific frequencies, and the amplitude at certain frequencies is large and persists, a single-frequency reference noise can be preset first. The initial amplitude of the single-frequency reference noise can be the normalized amplitude accumulation value E after the earliest m accumulated noises are collected. sum The earliest acquisition time may be after the device 10 is started (e.g., powered on), at which time there is usually no gesture. It is understandable that when the microphone of the device 10 acquires the ultrasonic echo, it may sequentially acquire the audio frames of the ultrasonic echo. Therefore, the microphone of the device 10 may sequentially acquire the earliest m audio frames acquired, and may calculate the normalized amplitude accumulation value E corresponding to each audio frame. sum Then, the normalized amplitude accumulation value E corresponding to the above m audio frames is further calculated. sum The average value is calculated to obtain the single-frequency reference noise amplitude. Since the first three audio frames generally do not contain any gesture information, m can preferably be 3. It should be noted that m being 3 is only a preferred embodiment of the present application and does not constitute a limitation of the present application. In some embodiments, m can also be other values.

[0143] Then, the single frequency noise can be filtered using the following formula:

[0144] E′=E sum -E alpha ; Where, E' is the amplitude after filtering the single frequency noise, E alpha is the single frequency reference noise amplitude.

[0145] Optionally, the above single frequency reference noise amplitude E can be further alpha The above single frequency reference noise amplitude E alpha The update method can be: if no gesture is currently recognized, it can be considered that the current frequency domain data contains the latest single frequency point noise, and E alpha The above update method can be performed using the following formula:

[0146] E′ alpha =E alpha *α+E sum *(1-α); where E alpha ' is the updated single-frequency reference noise amplitude, and α is a preset coefficient. By updating the single-frequency reference noise amplitude, the first device 10 can filter single-frequency noise based on the latest reference noise amplitude, thereby improving the accuracy of gesture recognition.

[0147] It should be noted that the above judgment result of identifying whether a gesture exists can be determined based on E'. For example, if E' is less than or equal to a preset threshold, it can be determined that no gesture exists in the currently collected frequency domain data.

[0148] Figure 7 This is the filtering effect diagram of single frequency noise. Figure 7 As shown, waveform 710 is the frequency domain waveform before filtering the single frequency point noise. Waveform 710 contains noise near the specific frequency point X. Waveform 720 is the frequency domain waveform after filtering the single frequency point noise. Figure 7 After the single-frequency noise is filtered, the noise near the specific frequency X is successfully filtered out. Therefore, by filtering the single-frequency noise, the anti-interference capability can be improved.

[0149] It is understandable that the execution order of step 204 and step 203 can be in any order. In other words, step 204 can be executed before step 203, after step 203, or simultaneously with step 203.

[0150] In step 305 , the device 10 filters the sound emitted by itself.

[0151] Specifically, when the speaker of device 10 emits sound, it affects the frequency domain data collected by the microphone, generating a number of irregular and erratic peaks. For example, when device 10 plays ultrasound while ringing, the ringing sound will cause a number of irregular and erratic peaks to be generated in the ultrasonic echo data collected by the microphone. It is understandable that the above-mentioned ringing sound is a low-frequency pulse wave, and these pulse waves are irregularly and erratically distributed on the frequency domain data of the entire frequency band. These irregular and erratic peaks will affect the frequency domain data of the gesture recognition area, thereby affecting the recognition of gestures. Therefore, device 10 can also filter the sound emitted by device 10 itself through echo cancellation or high-pass filtering, thereby improving the anti-interference ability. It should be noted that since the ringing sound is low-frequency data, the noise generated by the above-mentioned ringing can be filtered out through high-pass filtering.

[0152] It is understandable that the execution order of step 205, step 203, and step 204 may not be specific. In other words, step 205 may be executed before step 203 or step 204, after step 203 or step 204, or simultaneously with step 203 and / or step 204.

[0153] In step 306 , the device 10 performs energy amplification.

[0154] Specifically, after the device 10 filters the above-mentioned ultrasonic echo for ambient noise and / or single-frequency noise and / or the own sound, it can also perform energy amplification on the frequency domain data after filtering the above-mentioned ambient noise and / or single-frequency noise and / or the own sound. In specific implementation, the above-mentioned energy amplification method can be achieved by the following formula:

[0155] E″=E1*factor; wherein, E″ is the amplitude after amplification, E1 is the amplitude of the frequency domain data after filtering the ambient background noise and / or filtering the single-frequency noise and / or filtering the own sound, and factor is the preset amplification factor.

[0156] In step 307 , the device 10 performs gesture recognition.

[0157] Specifically, in order to avoid misjudgment, the first device 10 can obtain a sufficient number of audio frames for identification, and the number of the above-mentioned audio frames can be pre-set. In specific implementation, the device 10 can collect audio data within a preset time length, and the device 10 can also collect a preset number of audio frames to ensure that the device 10 can collect a certain amount of audio data. The above-mentioned audio data can be the echo data of the above-mentioned ultrasound, and multiple audio frames can be obtained by sampling the above-mentioned audio data. It can be understood that the embodiment of the present application does not specifically limit the method of collecting a preset amount of audio data.

[0158] After the device 10 acquires a certain number of audio frames, it can process the above audio frames through steps 202 to 206, thereby obtaining a frequency domain waveform after normalization, accumulation and noise filtering. Then, gesture recognition can be performed based on the above waveform to determine whether the user's gesture is present.

[0159] In a specific implementation, the gesture recognition area in the audio frame can be detected to determine whether there are frequency points exceeding a preset amplitude threshold. If a frequency point exceeding the preset amplitude threshold is detected in the gesture recognition area in the audio frame, it can be determined that a user gesture is present. If a frequency point exceeding the preset amplitude threshold is detected in the gesture recognition area in the audio frame, it can be determined that no user gesture is present. The user gesture can be a gesture to turn off the ringer, a gesture to answer or turn off an incoming call, a gesture to turn off a timed reminder, or a gesture to turn a light on or off.

[0160] If the device 10 confirms the presence of a gesture (e.g., a gesture to turn off the ringer), it can perform a corresponding operation based on the currently recognized gesture. For example, if the device 10 detects a gesture from the user to turn off the ringer while the device is ringing, the device 10 can turn off the ringer. If the device 10 does not recognize a gesture, the device 10 can continue to collect ultrasonic echo data and further recognize the user's gesture based on the ultrasonic echo data.

[0161] Figure 8 This is a schematic diagram of the gesture recognition effect. Figure 8 As shown, frequency domain waveform 800 is the frequency domain waveform of the ultrasonic echo collected by a single microphone. When the user gesture is far away from the device 10, the amplitude of the gesture recognition area 810 of the frequency domain waveform 800 does not exceed the preset amplitude threshold. Therefore, the single microphone cannot recognize the user gesture at a distance. Frequency domain waveform 801 is the frequency domain waveform after accumulation, noise filtering and energy amplification of multiple microphones. Since the amplitude of the gesture recognition area 811 of the frequency domain waveform 801 has exceeded the preset amplitude threshold, it can be confirmed that the user gesture exists. In other words, through the above-mentioned multiple microphone accumulation, noise filtering and energy amplification, the user gesture can be recognized at a distance, thereby improving the distance of gesture recognition through the technical solution of this application.

[0162] It can be understood that in the above embodiment, steps 201 to 207 are all optional steps. This application only provides a feasible embodiment, which may also include more or fewer steps than steps 201 to 207. This application does not limit this.

[0163] Figure 9 This is a structural diagram of an embodiment of the gesture recognition device of the present application. Figure 9 As shown, the gesture recognition device 90 is applied to an electronic device, which includes multiple microphones and may include: a sending module 91, a receiving module 92, an accumulation module 93 and a recognition module 94; wherein,

[0164] A sending module 91 is used to send ultrasonic waves of a preset frequency;

[0165] The receiving module 92 is configured to use a plurality of microphones to respectively receive ultrasonic echoes corresponding to the ultrasonic waves to obtain a plurality of first sound wave signals;

[0166] an accumulation module 93 for accumulating the plurality of first sound wave signals according to the gesture recognition area to obtain a second sound wave signal, wherein the gesture recognition area is a frequency area associated with a preset frequency, and the amplitude of the second sound wave signal in the gesture recognition area is greater than the amplitude of each first sound wave signal in the gesture recognition area;

[0167] The recognition module 94 is configured to perform gesture recognition based on the second sound wave signal.

[0168] In one possible implementation, the gesture recognition area is a frequency area determined based on a Doppler shift principle and a preset frequency and gesture speed.

[0169] In one possible implementation, the device 90 further includes:

[0170] The first determining module 95 is configured to determine that a gesture is detected when the amplitude of the second sound wave signal in the gesture recognition area is greater than a preset amplitude.

[0171] In one possible implementation, the gesture recognition area includes a distant gesture recognition area located to the left of a preset frequency and / or a close gesture recognition area located to the right of the preset frequency. The apparatus 90 further includes:

[0172] A second determining module 96 is configured to determine that a distance gesture is detected when the amplitude of the second sound wave signal in the distance gesture recognition area is greater than a preset amplitude;

[0173] When the amplitude of the second sound wave signal in the proximity gesture recognition area is greater than the preset amplitude, it is determined that the proximity gesture is detected.

[0174] In one possible implementation, the accumulation module 93 is further configured to perform normalization processing on the plurality of first sound wave signals respectively to obtain a plurality of normalized first sound wave signals;

[0175] A second sound wave signal is obtained by accumulating multiple normalized first sound wave signals according to a gesture recognition area, wherein the gesture recognition area is a frequency area associated with a preset frequency, and an amplitude of the second sound wave signal in the gesture recognition area is greater than an amplitude of each normalized first sound wave signal in the gesture recognition area.

[0176] In one possible implementation, the identification module 94 is further configured to perform noise filtering on the second sound wave signal to obtain a noise-filtered second sound wave signal.

[0177] Gesture recognition is performed based on the second sound wave signal after noise filtering.

[0178] In one possible implementation, the noise filtering process includes at least one of ambient noise filtering, single-frequency noise filtering, and electronic device's own sound filtering.

[0179] In one possible implementation, the sending module 91 is also used to send ultrasonic waves of a preset frequency in response to a detected preset trigger event, wherein the preset trigger event is any one of an alarm ringing, an incoming call, a timed reminder, and a light switch event.

[0180] In one possible implementation, the device 90 further includes:

[0181] The execution module 97 is used to execute any one of the operations of turning off the alarm, answering or rejecting a call, turning off a timer reminder, turning on or off a light when it is determined that a gesture is detected.

[0182] In one possible implementation, the second sound wave signal further includes a non-gesture recognition area, and the amplitude of the non-gesture recognition area of ​​the second sound wave signal is determined by an average of the amplitudes of the non-gesture recognition areas of the plurality of first sound wave signals.

[0183] In one possible implementation, the electronic device is any one of a smart speaker, a mobile phone, a large screen, a tablet, and a smart switch.

[0184] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0185] It is understandable that, in order to realize the above functions, the above-mentioned electronic device 100 and the like include hardware structures and / or software modules corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0186] In the embodiment of the present application, the functional modules of the electronic device 100 and the like can be divided according to the above-mentioned method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.

[0187] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0188] The functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0189] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.

[0190] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A gesture recognition method, applied to an electronic device, characterized in that: The electronic device includes a plurality of microphones, and the method includes: Sending ultrasonic waves of preset frequency; Using the plurality of microphones to respectively receive ultrasonic echoes corresponding to the ultrasonic waves to obtain a plurality of first sound wave signals; Accumulating the plurality of first sound wave signals according to a gesture recognition area to obtain a second sound wave signal, wherein the gesture recognition area is a frequency area associated with the preset frequency, and an amplitude of the second sound wave signal in the gesture recognition area is greater than an amplitude of each of the first sound wave signals in the gesture recognition area; Perform gesture recognition according to the second sound wave signal.

2. The method according to claim 1, characterized in that The gesture recognition area is a frequency area determined based on the Doppler frequency shift principle and according to the preset frequency and gesture speed.

3. The method according to claim 2, characterized in that The method further comprises: If the amplitude of the second sound wave signal in the gesture recognition area is greater than a preset amplitude, it is determined that a gesture is detected.

4. The method according to claim 2, characterized in that The gesture recognition area includes a distant gesture recognition area located to the left of the preset frequency, and / or a close gesture recognition area located to the right of the preset frequency; the method further includes: When the amplitude of the second sound wave signal in the moving-away gesture recognition area is greater than a preset amplitude, determining that a moving-away gesture is detected; If the amplitude of the second sound wave signal in the approach gesture recognition area is greater than the preset amplitude, it is determined that the approach gesture is detected.

5. The method according to any one of claims 1 to 4, characterized in that The step of accumulating the plurality of first sound wave signals according to the gesture recognition area to obtain a second sound wave signal includes: performing normalization processing on the plurality of first acoustic wave signals respectively to obtain a plurality of normalized first acoustic wave signals; Accumulating the plurality of normalized first sound wave signals according to the gesture recognition area to obtain the second sound wave signal, wherein the gesture recognition area is a frequency area associated with the preset frequency, and an amplitude of the second sound wave signal in the gesture recognition area is greater than an amplitude of each of the normalized first sound wave signals in the gesture recognition area.

6. The method according to any one of claims 1 to 5, characterized in that The performing gesture recognition according to the second sound wave signal specifically includes: performing noise filtering on the second sound wave signal to obtain a noise-filtered second sound wave signal; Perform gesture recognition based on the second sound wave signal after noise filtering.

7. The method according to claim 6, characterized in that The noise filtering process includes at least one of ambient noise filtering, single frequency noise filtering, and electronic device's own sound filtering.

8. The method according to any one of claims 1 to 7, characterized in that The sending of ultrasonic waves of a preset frequency specifically includes: In response to a detected preset trigger event, ultrasonic waves of a preset frequency are transmitted, wherein the preset trigger event is any one of an alarm ringing, an incoming call, a timed reminder, and a light on / off event.

9. The method according to claim 8, characterized in that The method further comprises: If a gesture is detected, perform any of the following operations: silence the alarm, answer or reject the call, silence the timer reminder, or turn on or off the light.

10. The method according to any one of claims 1 to 9, characterized in that The second sound wave signal further includes a non-gesture recognition area, and the amplitude of the non-gesture recognition area of ​​the second sound wave signal is determined by an average of the amplitudes of the non-gesture recognition areas of the plurality of first sound wave signals.

11. The method according to any one of claims 1 to 10, characterized in that The electronic device is any one of a smart speaker, a mobile phone, a large screen, a tablet, and a smart switch.

12. An electronic device, characterized in that: The electronic device comprises a memory for storing computer program codes, wherein the computer program codes include instructions. When the electronic device reads the instructions from the memory, the electronic device executes the method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that The method comprises computer instructions, which, when executed on the electronic device, cause the electronic device to execute the method according to any one of claims 1 to 11.

14. A computer program product, characterized in that When the computer program product is run on a computer, the computer is caused to perform the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Ultrasonic-based gesture recognition method and system

    CN107943300A

  • Gesture recognition method and terminal

    CN109857245A