Gesture determination method and device, storage medium and electronic equipment
By transmitting and receiving sound wave signals through the speaker and microphone of an electronic device, extracting speed and distance feature information, and combining it with a lightweight neural network model to recognize gestures, the high hardware cost and privacy exposure problems of traditional air gesture control are solved, achieving low-cost and accurate air gesture recognition.
Patent Information
- Application Number
- CN202410635312.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional air gesture control solutions rely on sensors and cameras that are expensive, inconvenient to wear, and expose user privacy.
By using speakers and microphones in electronic devices to transmit and receive sound wave signals, and extracting motion feature information from the sound wave signals, including speed and distance features, a lightweight neural network model is used to recognize gestures.
It achieves low-cost, no-additional-hardware-required air gesture recognition, reduces privacy exposure, and improves the accuracy and adaptability of gesture recognition.
Smart Images

Figure CN120994047A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a gesture recognition method and apparatus, storage medium and electronic device. Background Technology
[0002] With the rapid development of computer vision and machine learning technologies, air gesture control has been widely used, including but not limited to virtual reality, augmented reality, smart home control, and game interaction.
[0003] Traditional gesture control solutions mostly rely on devices such as sensor gloves and cameras to capture user gestures. These methods are often limited by high hardware costs and inconvenience of wearing them. Furthermore, using cameras to capture user gestures can expose user privacy to some extent. Summary of the Invention
[0004] In view of the above, embodiments of this disclosure provide a gesture determination method and apparatus, a storage medium and an electronic device.
[0005] According to a first aspect of this disclosure, a gesture determination method is proposed, the method comprising:
[0006] After receiving a sound wave signal through the microphone, motion feature information is extracted from the received sound wave signal; the sound wave signal is a signal that is transmitted to the microphone after being propagated by a specific sound wave signal emitted by the speaker to determine the gesture; the motion feature information includes at least speed feature information, which is used to characterize the speed of the gesture execution.
[0007] Determine the gesture information corresponding to the motion feature information.
[0008] In conjunction with any of the embodiments provided in this disclosure,
[0009] The motion feature information also includes distance feature information; the distance feature information is used to characterize the distance between the gesture and the target device.
[0010] In any embodiment provided by this disclosure, the specific acoustic signal consists of a single-frequency signal and a multi-frequency signal;
[0011] The single-frequency signal is used to extract the velocity feature information, and the multi-frequency signal is used to extract the distance feature information. The target frequency corresponding to the single-frequency signal is different from the frequency in the frequency range corresponding to the multi-frequency signal.
[0012] In any embodiment provided by this disclosure, the received acoustic signal is a time-domain signal sequence;
[0013] The extraction of motion feature information from the received acoustic signal includes:
[0014] After performing a Fourier transform operation on the time-domain signal sequence to obtain a frequency-domain signal sequence, the target frequency located to the left of the target frequency and the target frequency located to the right of the target frequency in the frequency-domain signal sequence are determined.
[0015] The difference between the amplitudes of the target frequency located to the right of the target frequency and the target frequency located to the left of the target frequency is determined as the velocity characteristic information.
[0016] In conjunction with any embodiment provided in this disclosure, the step of extracting motion feature information from the received acoustic signal includes:
[0017] A filtering operation is performed on the received acoustic signal to obtain a filtered signal within the frequency range corresponding to the multi-frequency signal;
[0018] The reflection duration is determined based on the filtered signal; the reflection duration is used to characterize the time required for the specific sound wave signal to be emitted from the speaker and reflected to the microphone via a gesture.
[0019] The distance feature information is determined by half the product of the reflection time and the propagation speed of the specific sound wave signal in the air.
[0020] In conjunction with any embodiment provided in this disclosure, determining the reflection duration based on the filtered signal includes:
[0021] The first sampling point in the filtered signal is determined based on the first cross-correlation result of the multi-frequency signal and the filtered signal; the first sampling point is the sampling point corresponding to when the specific sound wave signal is directly transmitted from the speaker to the microphone;
[0022] Using the first sampling point as the starting sampling point, a signal of a preset length is extracted from the filtered signal to obtain a sub-filtered signal;
[0023] The second sampling point in the sub-filtered signal is determined based on the second cross-correlation result of the multi-frequency signal and the sub-filtered signal; the second sampling point is the sampling point corresponding to when the specific sound wave signal is emitted by the speaker, reflected by a gesture, and then transmitted to the microphone;
[0024] The reflection duration is determined based on the second sampling point and the preset sampling frequency.
[0025] In conjunction with any embodiment provided in this disclosure, determining the second sampling point in the sub-filtered signal based on the second cross-correlation result of the multi-frequency signal and the sub-filtered signal includes:
[0026] Perform a Hilbert transform operation on the second cross-correlation result of the multi-frequency signal and the sub-filtered signal to obtain the first echo characteristic corresponding to the second cross-correlation result;
[0027] Perform a difference operation on the first echo feature based on the target echo feature to obtain a first-order difference value;
[0028] Calculate the modulus of the first-order difference value, and determine the sampling point corresponding to the maximum modulus of the first-order difference value as the second sampling point.
[0029] In conjunction with any embodiment provided in this disclosure, determining the gesture information corresponding to the motion feature information includes:
[0030] The motion feature information is input into the gesture recognition model to obtain the gesture information output by the gesture recognition model that corresponds to the motion feature information.
[0031] In conjunction with any embodiment provided in this disclosure, the gesture recognition model is trained in the following manner:
[0032] Acquire training data; the training data includes at least one set of motion feature information labeled with gesture information, the gesture information being used to characterize the gesture corresponding to the motion feature information;
[0033] The first motion feature information in the training data is input into the gesture recognition model to be trained, and the predicted gesture information output by the gesture recognition model corresponding to the first motion feature information is obtained.
[0034] The network parameters of the gesture recognition model are adjusted based on the difference between the predicted gesture information and the gesture information in the training data corresponding to the first motion feature information.
[0035] In conjunction with any embodiment provided in this disclosure, before extracting motion feature information from the received sound wave signal after receiving it through the microphone, the method further includes:
[0036] In response to receiving a gesture recognition activation command, the specific sound wave signal is emitted through the speaker.
[0037] According to a second aspect of this disclosure, a gesture recognition device is provided, the device comprising:
[0038] The motion feature information extraction module is used to extract motion feature information from the received sound wave signal after receiving the sound wave signal through the microphone; the sound wave signal is a signal transmitted to the microphone after being propagated by a specific sound wave signal emitted by the speaker to determine the gesture; the motion feature information includes at least speed feature information, which is used to characterize the speed of the gesture execution.
[0039] The gesture information determination module is used to determine the gesture information corresponding to the motion feature information.
[0040] According to a third aspect of this disclosure, a computer-readable storage medium is provided, the machine-readable storage medium storing machine-readable instructions, which, when invoked and executed by a processor, cause the processor to implement a gesture determination method according to any embodiment of this disclosure.
[0041] According to a fourth aspect of this disclosure, an electronic device is provided, comprising:
[0042] processor;
[0043] Memory used to store processor-executable instructions;
[0044] The processor is configured to perform a gesture determination method according to any embodiment of the present disclosure.
[0045] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0046] The gesture determination method, apparatus, storage medium, and electronic device provided in this disclosure consider the differences in speed between different gesture operations. This method extracts speed feature information characterizing the speed of gesture execution based on received sound wave signals, and determines the corresponding gesture information based on this speed feature information. In this method, the device can utilize existing speakers and microphones to emit and receive sound wave signals, and determine the user's gesture based on these signals, eliminating the need for additional hardware deployment and resulting in low hardware costs. Furthermore, compared to methods that capture user gestures using a camera, this method does not expose user privacy.
[0047] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0048] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0049] Figure 1 This is a flowchart illustrating a gesture determination method according to an exemplary embodiment of the present disclosure;
[0050] Figure 2 This is a flowchart illustrating another gesture determination method according to an exemplary embodiment of the present disclosure;
[0051] Figure 3 This is a flowchart illustrating another gesture determination method according to an exemplary embodiment of the present disclosure;
[0052] Figure 4 This is a flowchart illustrating another gesture determination method according to an exemplary embodiment of the present disclosure;
[0053] Figure 5 This is a flowchart illustrating another gesture determination method according to an exemplary embodiment of the present disclosure;
[0054] Figure 6 This is a schematic diagram of the structure of a gesture determination device according to an exemplary embodiment of the present disclosure;
[0055] Figure 7 This is a schematic diagram of the structure of a gesture determination device according to an exemplary embodiment of the present disclosure;
[0056] Figure 8 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0057] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. In the following description relating to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements.
[0058] With the rapid development of computer vision and machine learning technologies, air gesture control has been widely used, including but not limited to virtual reality, augmented reality, smart home control, and game interaction.
[0059] The advantage of air gesture control is that it allows people to completely free their hands when using electronic devices, eliminating the need to touch the screen, mouse, keyboard, or other operating media. This control method is particularly convenient in certain scenarios, such as when the user is wearing gloves or has wet hands.
[0060] Traditional gesture control solutions mostly rely on devices such as sensor gloves and cameras to capture user gestures. These methods are often limited by high hardware costs and inconvenience of wearing them. Furthermore, using cameras to capture user gestures can expose user privacy to some extent.
[0061] To address the aforementioned issues, this disclosure provides a gesture determination method. In this method, the device can utilize its existing speaker and microphone to emit and receive sound wave signals, and determine the user's gesture based on these signals. No additional hardware deployment is required, resulting in low hardware costs. Furthermore, compared to methods that capture user gestures using a camera, this method has relatively fewer privacy concerns.
[0062] The gesture determination method of this disclosure will now be described in detail with reference to the accompanying drawings.
[0063] Figure 1 This is a flowchart illustrating a gesture determination method according to an exemplary embodiment of the present disclosure. This method can be applied to electronic devices such as smartphones, tablets, and smart TVs. Figure 1 As shown, the exemplary embodiment method may include the following steps:
[0064] In step 101, after receiving the sound wave signal through the microphone, motion feature information is extracted from the received sound wave signal.
[0065] The aforementioned acoustic signal is a specific acoustic signal emitted by the speaker to determine the gesture, which is then transmitted to the microphone after propagation. In practical applications, the speaker of an electronic device can continuously emit the specific acoustic signal to promptly capture the user's gesture. Alternatively, the electronic device can control the speaker to emit the specific acoustic signal only after receiving a gesture recognition activation command from the user, thereby reducing power consumption.
[0066] After a specific sound wave signal is emitted through the speaker of an electronic device and received by the microphone of the electronic device after the specific sound wave signal has been propagated, motion feature information can be further extracted from the received sound wave signal. The motion feature information includes at least speed feature information, which is used to characterize the speed of the gesture execution.
[0067] It should be noted that this example is based on the premise that the electronic device has one speaker and one microphone. In practical applications, the electronic device may have two or more speakers and microphones, and this disclosure does not limit this.
[0068] In step 102, gesture information corresponding to the motion feature information is determined.
[0069] In this example, the gesture information corresponding to the aforementioned speed feature information is determined.
[0070] Considering the speed differences of different gesture operations, the gesture determination method provided in this embodiment can extract speed feature information that characterizes the speed of user gesture execution from the received sound wave signal, and determine the corresponding gesture information based on the speed feature information, thereby accurately identifying the user's gesture operation in various scenarios.
[0071] In one embodiment, the aforementioned motion feature information may further include distance feature information, which is used to characterize the distance between the gesture and the target device.
[0072] The target device is the electronic device corresponding to the loudspeaker that emits a specific sound wave signal. Similarly, it can be understood as the electronic device corresponding to the microphone that receives the sound wave signal.
[0073] In this example, the subsequent step of determining the gesture information corresponding to the motion feature information can be performed only when the distance value indicated by the aforementioned distance feature information is less than or equal to a preset distance value. Conversely, when the distance value indicated by the aforementioned distance feature information is greater than the preset distance value, it is considered that the user is passing through a scene, and gesture recognition is not required, that is, the subsequent step of determining the gesture information corresponding to the motion feature information is not required.
[0074] In the gesture determination method provided in this embodiment, distance feature information can also be extracted from the received sound wave signal, and the position of the gesture operation can be constrained based on the distance feature information, thereby effectively reducing the problem of misrecognition of gestures by users in scenarios such as passing by.
[0075] In one embodiment, the specific sound wave signal emitted by the aforementioned loudspeaker can be composed of a single-frequency signal and a multi-frequency signal. The single-frequency signal is used to extract velocity feature information, and the multi-frequency signal is used to extract distance feature information. Furthermore, the target frequency corresponding to the single-frequency signal is different from the frequency within the frequency range corresponding to the multi-frequency signal.
[0076] Taking the emission of specific sound wave signals through two loudspeakers as an example, where the specific sound wave signal emitted by the first loudspeaker is denoted as x1(t) and the specific sound wave signal emitted by the second loudspeaker is denoted as x2(t), as follows:
[0077] x1(t)=cos(2πf M t)+x0(t)
[0078] x2(t)=cos(2πf N t)+x0(t)
[0079] in,
[0080]
[0081]
[0082] Where t is the time when the signal is generated, f is the frequency, and A i It is a real number weight, F s Let L be the sampling frequency of the acoustic signal, L be the number of sampling points in one period of the signal x0(t), and M, N, K0, K1 be four selected integer constants that satisfy:
[0083] M≠N
[0084]
[0085] In other words, the specific sound wave signal emitted by the first loudspeaker is composed of a single-frequency signal cos(2πf) M The second loudspeaker emits a specific sound wave signal consisting of a single-frequency signal cos(2πf) and a multi-frequency signal x0(t). N It consists of a single-frequency signal cos(2πf) and a multi-frequency signal x0(t). M The target frequency f corresponding to t) M Frequency range corresponding to multi-frequency signals The frequencies in the signal are different; a single-frequency signal cos(2πf) N The target frequency f corresponding to t) N Frequency range corresponding to multi-frequency signals The frequencies also differ.
[0086] In subsequent feature extraction, velocity feature information can be extracted based on the single-frequency signal in the specific sound wave signal emitted by each speaker, and distance feature information can be extracted based on the multi-frequency signal in the specific sound wave signal emitted by each speaker. For details, please refer to the description of the following embodiments, which will not be repeated here.
[0087] In the gesture determination method provided in this embodiment, the specific sound wave signal emitted by the loudspeaker is designed to consist of a single-frequency signal and a multi-frequency signal. In this way, speed feature information can be extracted based on the single-frequency signal in the specific sound wave signal, and distance feature information can be extracted based on the multi-frequency signal in the specific sound wave signal, which is highly practical.
[0088] The following is a detailed description of the process of extracting velocity feature information from single-frequency signals in the aforementioned specific acoustic signals, and extracting distance feature information from multi-frequency signals in the aforementioned specific acoustic signals.
[0089] In one embodiment, for step 101 described above, the received acoustic signal is a time-domain signal sequence. At this time, as... Figure 2 As shown, velocity feature information can be extracted from the single-frequency signal in the received acoustic signal through the following steps:
[0090] In step 201, after performing a Fourier transform operation on the time-domain signal sequence to obtain a frequency-domain signal sequence, the target frequency located to the left of the target frequency and the target frequency located to the right of the target frequency in the frequency-domain signal sequence are determined.
[0091] Taking the reception of sound wave signals through two microphones as an example, the time-domain signal sequence received by the first microphone is y1(t), and the time-domain signal sequence received by the second microphone is y2(t).
[0092] For the time-domain signal sequence y1(t) received by the first microphone, a Fourier transform operation is first performed on it to obtain the frequency-domain signal sequence y1(j), where j = 1, 2, ..., L, which refers to the j-th frequency.
[0093] For example, it can be shown in Table 1 below:
[0094] Table 1
[0095]
[0096] Before explaining how to extract velocity feature information based on this frequency domain signal sequence, it should be noted that in the specific sound wave signal emitted by the first loudspeaker, besides the frequency f... M , Except for the frequency f, the amplitude corresponding to all other frequencies is 0. In the specific sound wave signal emitted by the second speaker, except for the frequency f... N , Apart from that, the amplitudes corresponding to other frequencies are all 0. However, after being reflected by the gesture, the frequency f... M f N , A Doppler frequency shift will occur, causing changes in the amplitude corresponding to other frequencies. This example method determines velocity characteristic information based on this shift. Specifically, in this example method, a single-frequency signal cos(2πf) can be used. M t), cos(2πf) N The frequency f corresponding to t) M f N The Doppler frequency shift is used to determine the velocity characteristic information.
[0097] Specifically, after obtaining the aforementioned frequency domain signal sequence y1(j), 2C+1 items can be selected from this sequence. For example, these can be items with frequencies ranging from MC to M+C, where C is a positive integer. That is, items located at the target frequency f in the frequency domain signal sequence are selected. M The frequency of item C on the left and the target frequency f M The frequency of item C on the right is taken as the frequency of the target item.
[0098] Taking C=1 as an example, the frequencies of the target terms are X2 and X3 respectively. Taking C=2 as an example, the frequencies of the target terms are X1, X2, X3 and X4 respectively.
[0099] As can be seen from the previous examples, the method in this example is based on the single-frequency signal cos(2πf) M The frequency f corresponding to t) M The Doppler frequency shift determines the velocity characteristic information; therefore, the frequency f corresponding to the single-frequency signal should be removed. N The frequencies corresponding to multi-frequency signals The impact of this. Specifically, when determining the aforementioned C, it should be ensured that the selected target frequency corresponds to the frequency f of the single-frequency signal. N and the frequencies within the frequency range corresponding to multi-frequency signals. They do not overlap.
[0100] Similarly, for the time-domain signal sequence y2(t) received by the second microphone, a Fourier transform operation is first performed on it to obtain the frequency-domain signal sequence y2(j), where j = 1, 2, ..., L, which refers to the j-th frequency.
[0101] For example, it can be shown in Table 2 below:
[0102] Table 2
[0103]
[0104] Then, 2C+1 items can be selected from this frequency domain signal sequence. For example, these can be items with frequencies ranging from NC to N+C, where C is a positive integer. That is, items located at the target frequency f in the frequency domain signal sequence... N The frequency of item C on the left and the target frequency f N The frequency of item C on the right is taken as the frequency of the target item.
[0105] Taking C=1 as an example, the frequencies of the target terms are X3 and X4 respectively. Taking C=2 as an example, the frequencies of the target terms are X2, X3, X4 and X5 respectively.
[0106] In step 202, the difference between the amplitudes corresponding to the target frequency located to the right of the target frequency and the target frequency located to the left of the target frequency is determined as the velocity characteristic information.
[0107] The velocity feature information extracted from the sound wave signal received by the first microphone, denoted as v1(t), is obtained through the following formula 1:
[0108]
[0109] Where |Y1(M+j)| is the frequency f (M+j) The corresponding amplitude, |Y1(Mj)| is the frequency f(M-j) The corresponding amplitude.
[0110] In the example where C is 1, the difference between the amplitude corresponding to frequency X3 and the amplitude corresponding to frequency X2 can be determined as the speed characteristic information of the first microphone.
[0111] Similarly, the velocity feature information extracted from the sound wave signal received by the second microphone, denoted as v2(t), is obtained through the following formula 2:
[0112]
[0113] Where |Y2(N+j)| is the frequency f (N+j) The corresponding amplitude, |Y2(Nj)| is the frequency f (N-j) The corresponding amplitude.
[0114] In the example where C is 2, the difference between the amplitudes corresponding to frequencies X4 and X5 and the amplitudes corresponding to frequencies X2 and X3 can be determined as the speed characteristic information of the second microphone.
[0115] Thus, in this example, two velocity characteristic information can be obtained based on the sound wave signals received by the two microphones.
[0116] It should be noted that in the aforementioned example, frequency f M and between, with f N Taking a phase difference of two frequencies as an example, in practical applications, f M and between, with f N The frequency difference can be more than two frequencies, but the specific standard is to ensure that the selected target frequency does not overlap with the frequency range corresponding to the multi-frequency signal.
[0117] The gesture determination method provided in this embodiment can extract speed feature information based on a single-frequency signal in the received sound wave signal, which is simple to implement and highly usable.
[0118] In one embodiment, such as Figure 3 As shown, distance feature information can be extracted from the multi-frequency signals in the received acoustic signal through the following steps:
[0119] In step 301, a filtering operation is performed on the received acoustic signal to obtain a filtered signal within the frequency range corresponding to the multi-frequency signal.
[0120] For any time-domain signal sequence y received by a microphone i(t), firstly, a bandpass filter is applied to filter out the frequency range corresponding to the multi-frequency signal. External noise is used to obtain the frequency range corresponding to the multi-frequency signal. The filtered signal within.
[0121] In step 302, the reflection duration is determined based on the filtered signal.
[0122] The reflection duration is used to characterize the time required for a specific sound wave signal to be emitted from the speaker and reflected to the microphone via a gesture.
[0123] In practical applications, electronic devices can extract a segment of length L from the aforementioned filtered signal, denoted as z. i (t j (j = 1, 2, ..., L). Then, the aforementioned multi-frequency signal x0(t) is calculated based on the following formula 3. j (j = 1, 2, ..., L) and z i (t j The first cross-correlation result of (j = 1, 2, ..., L) is denoted as r. i (k), as follows:
[0124]
[0125] in, It is x0(t) j The average value of (j = 1, 2, ..., L). It is z i (t j The average value of (j = 1, 2, ..., L), where k is the time delay, specifically representing a delay of k sampling points.
[0126] Furthermore, it can make r i The largest k among (k) (k = 0, 1, ..., L-1) is denoted as k. i,max In other words, at a delay of k i,max After sampling points, the filtered signal z i (t j (j = 1, 2, ..., L) and multi-frequency signal x0(t) j The correlation is highest for the sequence (j = 1, 2, ..., L). Based on practical considerations, we can determine the k-th... i,max The k-th sampling point is the sampling point corresponding to a specific sound wave signal being directly transmitted from the speaker to the microphone. In this example, the k-th sampling point mentioned above can be... i,max One sampling point is used as the first sampling point.
[0127] Then, using the aforementioned first sampling point as the starting sampling point, a segment of length L can be extracted from the filtered signal, denoted as u. i (t j (j = 1, 2, ..., L), where u i (t j )=z i (t j+ki,max And calculate the multi-frequency signal x0(t). j (j = 1, 2, ..., L) and u i (t j The second cross-correlation result of (j = 1, 2, ..., L) is denoted as s. i (k)(k=0,1,...,L-1) will make the second cross-correlation result s i The largest k is denoted as k. i,max That is to say, at a delay of k i,max After 'sampling points, the filtered signal u i (t j (j = 1, 2, ..., L) and multi-frequency signal x0(t) j The correlation is highest for the sequence (j = 1, 2, ..., L). Based on practical considerations, we can determine the k-th... i,max The k-th sampling point is the sampling point corresponding to when a specific sound wave signal is emitted by the speaker, reflected by a gesture, and then transmitted to the microphone. In this example, the k-th sampling point can be... i,max One sampling point is used as the second sampling point.
[0128] As can be seen from the previous information, the second sampling point is the k-th point in the filtered signal. i,max With 'sampling points,' the electronic device can calculate k i,max The quotient of the frequency and the preset sampling frequency yields the aforementioned reflection duration.
[0129] The aforementioned preset sampling frequency can be, for example, 48kHz or 96kHz, and can be set by relevant personnel based on actual conditions. This disclosure does not limit this setting.
[0130] In an optional example, to improve the accuracy of the determined reflection duration, after calculating the aforementioned second cross-correlation result s... i After (k) (k = 0, 1, ..., L-1), a Hilbert transform operation can be performed on the second cross-correlation result to obtain the first echo characteristic f corresponding to the second cross-correlation result. i (k)(k=0,1,...,L-1).
[0131] Then, the electronic device can use the echo feature obtained in the previous moment as the target echo feature, and perform a difference between the first echo feature at the current moment and the target echo feature at the previous moment to obtain the first difference value Δf. i (k)(k=0,1,...,L-1).
[0132] Furthermore, the first-order difference value Δf is calculated. i The modulus of (k) (k = 0, 1, ..., L-1) is determined, and the sampling point corresponding to the k that maximizes the modulus of the first difference is determined as the aforementioned second sampling point.
[0133] In the gesture determination method provided in this example, a Hilbert transform operation is performed on the second cross-correlation result to convert the real signal into a complex signal, thereby providing more dimensional information features and better extracting the amplitude and phase information of the sound wave signal. Furthermore, in this scheme, the difference between the first echo feature at the current moment and the target echo feature at the previous moment can cancel out the signal reflected by the static object, leaving only the signal obtained from the gesture reflection, thus improving the accuracy of the determined reflection duration.
[0134] In step 303, half of the product of the reflection time and the propagation speed of the specific sound wave signal in the air is determined as the distance feature information.
[0135] In this example, the distance feature information d can be determined based on the following formula 4. i (t):
[0136]
[0137] Where v is the speed of propagation of a specific sound wave signal in the air, and t is the reflection time.
[0138] Based on the aforementioned method, two distance feature information can be extracted from the sound wave signals received by the two microphones respectively.
[0139] Considering that the cross-correlation value of multi-frequency signals at a certain sampling point will be very high and the sidelobes will be low when performing cross-correlation calculations, the gesture determination method provided in this embodiment can be designed to consist of a single-frequency signal and a multi-frequency signal emitted by the speaker, and distance feature information can be extracted based on the cross-correlation calculation results of the multi-frequency signals to improve the accuracy of the determined distance feature information.
[0140] In one embodiment, such as Figure 4 As shown, a gesture recognition model can be pre-trained based on the following steps:
[0141] In step 401, training data is acquired.
[0142] The training data includes at least one set of motion feature information labeled with gesture information, the gesture information being used to characterize the gesture corresponding to the motion feature information. The motion feature information includes velocity feature information and distance feature information. Furthermore, for each set of motion feature information, the number of velocity feature information and the number of distance feature information can be one or more, and this disclosure does not limit this.
[0143] It should be noted that the aforementioned training data includes not only positive examples but also negative examples, namely examples where the labeled gesture information is "no gesture" (corresponding to the aforementioned scenarios such as users passing by).
[0144] In step 402, the first motion feature information in the training data is input into the gesture recognition model to be trained, and the predicted gesture information output by the gesture recognition model corresponding to the first motion feature information is obtained.
[0145] The gesture recognition model to be trained can be a lightweight neural network model.
[0146] In step 403, the network parameters of the gesture recognition model are adjusted based on the difference between the predicted gesture information and the gesture information in the training data corresponding to the first motion feature information.
[0147] The gesture determination method provided in this disclosure can use a lightweight neural network model to train the constructed training data to obtain a neural network model for gesture recognition, thereby improving the accuracy of gesture recognition. The use of a lightweight neural network model also facilitates deployment on devices such as smartphones and tablets, simplifying implementation.
[0148] Figure 5 This is a flowchart illustrating another gesture determination method according to an exemplary embodiment of this disclosure. In this embodiment, the same steps as in the foregoing embodiments will be briefly described and will not be detailed further. For details, please refer to any of the foregoing embodiments. Figure 5 As shown, the exemplary embodiment method may include the following steps:
[0149] In step 501, in response to receiving a gesture recognition activation command, a specific sound wave signal is emitted through the speaker of the electronic device.
[0150] The specific sound wave signal emitted by the loudspeaker consists of a single-frequency signal and a multi-frequency signal, and the target frequency corresponding to the single-frequency signal is different from the frequency in the frequency range corresponding to the multi-frequency signal.
[0151] In step 502, after receiving the sound wave signal through the microphone of the electronic device, motion feature information is extracted from the received sound wave signal.
[0152] The sound wave signal is a specific sound wave signal emitted by the speaker to determine the gesture, which is then transmitted to the microphone after propagation. The aforementioned motion feature information includes speed feature information and distance feature information. The speed feature information is used to characterize the speed of the gesture execution, and the distance feature information is used to characterize the distance between the gesture and the electronic device.
[0153] Specifically, velocity feature information can be extracted from single-frequency signals in a specific acoustic signal, and distance feature information can be extracted from multi-frequency signals in a specific acoustic signal.
[0154] The processes for extracting velocity feature information based on single-frequency signals and for extracting distance feature information based on multi-frequency signals can be found in the description of the foregoing embodiments, and will not be repeated here.
[0155] In step 502, the motion feature information is input into the gesture recognition model to obtain the gesture information output by the gesture recognition model corresponding to the motion feature information.
[0156] In the gesture determination method provided in this disclosure, an electronic device can extract distance and velocity feature information from a received acoustic signal, and input this information into a pre-trained gesture recognition model to obtain gesture information output by the model corresponding to the aforementioned distance and velocity feature information. This method can constrain the position of gesture operations based on distance feature information, thereby effectively reducing misrecognition of gestures by users in scenarios such as passing by. Furthermore, considering the speed differences of different gesture operations, this method can accurately identify the user's actual gestures in various scenarios.
[0157] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should know that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps may be performed in other orders or simultaneously.
[0158] Corresponding to the aforementioned application function implementation method embodiments, this disclosure also provides embodiments of application function implementation apparatus and corresponding terminals.
[0159] Figure 6 This is a schematic diagram illustrating the structure of a gesture recognition device according to an exemplary embodiment of the present disclosure, as shown below. Figure 6 As shown, the gesture recognition device may include:
[0160] The motion feature information extraction module 61 is used to extract motion feature information from the received sound wave signal after receiving the sound wave signal through the microphone; the sound wave signal is a signal transmitted to the microphone after being propagated by a specific sound wave signal emitted by the speaker to determine the gesture; the motion feature information includes at least speed feature information, which is used to characterize the speed of the gesture execution.
[0161] The gesture information determination module 62 is used to determine the gesture information corresponding to the motion feature information.
[0162] Optionally, the motion feature information may also include distance feature information; the distance feature information is used to characterize the distance between the gesture and the target device.
[0163] Optionally, the specific acoustic signal consists of a single-frequency signal and a multi-frequency signal.
[0164] The single-frequency signal is used to extract the velocity feature information, and the multi-frequency signal is used to extract the distance feature information. The target frequency corresponding to the single-frequency signal is different from the frequency in the frequency range corresponding to the multi-frequency signal.
[0165] Optionally, the received acoustic signal is a time-domain signal sequence.
[0166] The motion feature information extraction module 61, when used to extract motion feature information from the received acoustic signal, includes:
[0167] After performing a Fourier transform operation on the time-domain signal sequence to obtain a frequency-domain signal sequence, the target frequency located to the left of the target frequency and the target frequency located to the right of the target frequency in the frequency-domain signal sequence are determined.
[0168] The difference between the amplitudes of the target frequency located to the right of the target frequency and the target frequency located to the left of the target frequency is determined as the velocity characteristic information.
[0169] Optionally, the motion feature information extraction module 61, when used to extract motion feature information from the received acoustic signal, includes:
[0170] A filtering operation is performed on the received acoustic signal to obtain a filtered signal within the frequency range corresponding to the multi-frequency signal.
[0171] The reflection duration is determined based on the filtered signal; the reflection duration is used to characterize the time required for the specific sound wave signal to be emitted from the speaker and reflected to the microphone via a gesture.
[0172] The distance feature information is determined by half the product of the reflection time and the propagation speed of the specific sound wave signal in the air.
[0173] Optionally, the motion feature information extraction module 61, when used to determine the reflection duration based on the filtered signal, includes:
[0174] The first sampling point in the filtered signal is determined based on the first cross-correlation result of the multi-frequency signal and the filtered signal; the first sampling point is the sampling point corresponding to when the specific sound wave signal is directly transmitted to the microphone by the speaker.
[0175] Using the first sampling point as the starting sampling point, a signal of a preset length is extracted from the filtered signal to obtain a sub-filtered signal.
[0176] The second sampling point in the sub-filtered signal is determined based on the second cross-correlation result of the multi-frequency signal and the sub-filtered signal; the second sampling point is the sampling point corresponding to when the specific sound wave signal is emitted by the speaker and transmitted to the microphone after being reflected by a gesture.
[0177] The reflection duration is determined based on the second sampling point and the preset sampling frequency.
[0178] Optionally, the motion feature information extraction module 61, when determining the second sampling point in the sub-filtered signal based on the second cross-correlation result of the multi-frequency signal and the sub-filtered signal, includes:
[0179] Perform a Hilbert transform operation on the second cross-correlation result of the multi-frequency signal and the sub-filtered signal to obtain the first echo feature corresponding to the second cross-correlation result.
[0180] Perform a difference operation on the first echo feature based on the target echo feature to obtain a first-order difference value.
[0181] Calculate the modulus of the first-order difference value, and determine the sampling point corresponding to the maximum modulus of the first-order difference value as the second sampling point.
[0182] Optionally, the gesture information determination module 62, when determining gesture information corresponding to the motion feature information, includes:
[0183] The motion feature information is input into the gesture recognition model to obtain the gesture information output by the gesture recognition model that corresponds to the motion feature information.
[0184] Optionally, the gesture recognition model is trained in the following manner:
[0185] Acquire training data; the training data includes at least one set of motion feature information labeled with gesture information, the gesture information being used to characterize the gesture corresponding to the motion feature information.
[0186] The first motion feature information in the training data is input into the gesture recognition model to be trained, and the predicted gesture information output by the gesture recognition model corresponding to the first motion feature information is obtained.
[0187] The network parameters of the gesture recognition model are adjusted based on the difference between the predicted gesture information and the gesture information in the training data corresponding to the first motion feature information.
[0188] Optional, such as Figure 7 As shown, in Figure 6 Based on the module shown, the gesture determination device may further include:
[0189] The sound wave signal transmitting module 71 is used to transmit the specific sound wave signal through the speaker in response to receiving a gesture recognition activation command.
[0190] For the apparatus embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to in the description of the method embodiment.
[0191] Figure 8 This is a schematic diagram illustrating the structure of an electronic device 800 according to an exemplary embodiment. For example, the electronic device 800 can be an electronic device such as a smartphone, tablet computer, smart TV, or computer.
[0192] Reference Figure 8 The electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0193] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0194] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of this data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0195] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.
[0196] Multimedia component 808 includes a screen that provides an output interface between the aforementioned electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0197] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0198] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0199] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0200] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 4G or 5G, 4G LTE, 5G NR, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the aforementioned communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0201] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0202] In an exemplary embodiment, a non-transitory computer-readable storage medium is also provided, such as a memory 804 including instructions, which, when executed by a processor 820 of an electronic device 800, enables the electronic device 800 to perform the gesture determination method of any embodiment of the present disclosure.
[0203] The non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0204] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0205] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A gesture recognition method, characterized in that, The method includes: After receiving a sound wave signal through the microphone, motion feature information is extracted from the received sound wave signal; the sound wave signal is a signal that is transmitted to the microphone after being propagated by a specific sound wave signal emitted by the speaker to determine the gesture; the motion feature information includes at least speed feature information, which is used to characterize the speed of the gesture execution. Determine the gesture information corresponding to the motion feature information.
2. The method according to claim 1, characterized in that, The motion feature information also includes distance feature information; the distance feature information is used to characterize the distance between the gesture and the target device.
3. The method according to claim 2, characterized in that, The specific acoustic signal consists of single-frequency signals and multi-frequency signals; The single-frequency signal is used to extract the velocity feature information, and the multi-frequency signal is used to extract the distance feature information. The target frequency corresponding to the single-frequency signal is different from the frequency in the frequency range corresponding to the multi-frequency signal.
4. The method according to claim 3, characterized in that, The received acoustic signal is a time-domain signal sequence; The extraction of motion feature information from the received acoustic signal includes: After performing a Fourier transform operation on the time-domain signal sequence to obtain a frequency-domain signal sequence, the target frequency located to the left of the target frequency and the target frequency located to the right of the target frequency in the frequency-domain signal sequence are determined. The difference between the amplitudes of the target frequency located to the right of the target frequency and the target frequency located to the left of the target frequency is determined as the velocity characteristic information.
5. The method according to claim 3, characterized in that, The extraction of motion feature information from the received acoustic signal includes: A filtering operation is performed on the received acoustic signal to obtain a filtered signal within the frequency range corresponding to the multi-frequency signal; The reflection duration is determined based on the filtered signal; the reflection duration is used to characterize the time required for the specific sound wave signal to be emitted from the speaker and reflected to the microphone via a gesture. The distance feature information is determined by half the product of the reflection time and the propagation speed of the specific sound wave signal in the air.
6. The method according to claim 5, characterized in that, Determining the reflection duration based on the filtered signal includes: The first sampling point in the filtered signal is determined based on the first cross-correlation result of the multi-frequency signal and the filtered signal; the first sampling point is the sampling point corresponding to when the specific sound wave signal is directly transmitted from the speaker to the microphone; Using the first sampling point as the starting sampling point, a signal of a preset length is extracted from the filtered signal to obtain a sub-filtered signal; The second sampling point in the sub-filtered signal is determined based on the second cross-correlation result of the multi-frequency signal and the sub-filtered signal; the second sampling point is the sampling point corresponding to when the specific sound wave signal is emitted by the speaker, reflected by a gesture, and then transmitted to the microphone; The reflection duration is determined based on the second sampling point and the preset sampling frequency.
7. The method according to claim 6, characterized in that, Determining the second sampling point in the sub-filtered signal based on the second cross-correlation result of the multi-frequency signal and the sub-filtered signal includes: Perform a Hilbert transform operation on the second cross-correlation result of the multi-frequency signal and the sub-filtered signal to obtain the first echo characteristic corresponding to the second cross-correlation result; Perform a difference operation on the first echo feature based on the target echo feature to obtain a first-order difference value; Calculate the modulus of the first-order difference value, and determine the sampling point corresponding to the maximum modulus of the first-order difference value as the second sampling point.
8. The method according to claim 2, characterized in that, The determination of gesture information corresponding to the motion feature information includes: The motion feature information is input into the gesture recognition model to obtain the gesture information output by the gesture recognition model that corresponds to the motion feature information.
9. The method according to claim 8, characterized in that, The gesture recognition model is trained in the following manner: Acquire training data; the training data includes at least one set of motion feature information labeled with gesture information, the gesture information being used to characterize the gesture corresponding to the motion feature information; The first motion feature information in the training data is input into the gesture recognition model to be trained, and the predicted gesture information output by the gesture recognition model corresponding to the first motion feature information is obtained. The network parameters of the gesture recognition model are adjusted based on the difference between the predicted gesture information and the gesture information in the training data corresponding to the first motion feature information.
10. The method according to claim 1, characterized in that, Before extracting motion feature information from the received sound wave signal after receiving it through the microphone, the method further includes: In response to receiving a gesture recognition activation command, the specific sound wave signal is emitted through the speaker.
11. A gesture recognition device, characterized in that, The device includes: The motion feature information extraction module is used to extract motion feature information from the received sound wave signal after receiving the sound wave signal through the microphone; the sound wave signal is a signal transmitted to the microphone after being propagated by a specific sound wave signal emitted by the speaker to determine the gesture; the motion feature information includes at least speed feature information, which is used to characterize the speed of the gesture execution. The gesture information determination module is used to determine the gesture information corresponding to the motion feature information.
12. A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method of any one of claims 1-10.
13. An electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to perform the steps of the method according to any one of claims 1-10.
Citation Information
Patent Citations
Head-mounted display equipment, gesture recognition method and device thereof and storage medium
CN114118152A