A gesture recognition system and method based on sound waves and sensors

Through a gesture recognition system based on sound waves and sensors, combined with channel estimation and gyroscope data, gesture recognition is achieved using neural network models, which solves the problems of the prior art being susceptible to environmental interference and high equipment costs, and improves recognition accuracy and user experience.

CN114692671BActive Publication Date: 2025-06-03NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210078481.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-06-03
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

Existing gesture recognition technology is susceptible to interference from lighting environments and external noise, and IMU-based methods require expensive smart devices, which are difficult to promote and popularize.

Method used

Using a gesture recognition system based on sound waves and sensors, channel information changed due to gesture movement is extracted through channel estimation technology, and the motion information recorded by the gyroscope is fused to realize gesture recognition using a neural network model.

Benefits of technology

Improves the accuracy of gesture recognition, provides a new input method that is not affected by lighting conditions, reduces dependence on device performance, and reduces deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114692671B_ABST
    Figure CN114692671B_ABST
Patent Text Reader

Abstract

The present invention discloses a gesture recognition system and method based on sound waves and sensors. The system includes a microphone and a smart watch. The smart watch is worn on the wrist and is used to transmit modulated ultrasonic signals and collect gesture motion information in the sensor. The microphone is responsible for receiving the ultrasonic signals emitted by the watch. When the microphone receives the signals, the channel changes caused by the gesture motion are obtained through channel estimation technology. At the same time, the sensor information recorded in the smart watch is combined. The two kinds of information are fused and input into the designed neural network model to output the probability values of each gesture, and the gesture with the maximum probability is selected as the recognized gesture. The present invention can enrich the input mode of users, improve the user experience, effectively improve the gesture recognition accuracy, and provide a new input mode for intelligent devices under limited devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless sensing, and particularly relates to a gesture recognition system and method based on sound waves and sensors. Background Art

[0002] The technology of using CIR (Channel Impulse Response) to estimate the channel state mainly involves transmitting a pre-generated synchronization code, which is then reflected and received by the receiver. By correlating the transmitted signal and the received signal, the corresponding channel information can be estimated.

[0003] Through this technology, the channel information changed due to gesture movement can be obtained, so this feature can be used for gesture recognition.

[0004] At present, gesture recognition can be mainly divided into three types:

[0005] (1) Using a vision-based method to obtain the corresponding gesture category by analyzing the video. However, this method is easily affected by the lighting environment and has high requirements for equipment.

[0006] (2) Using a sound-wave-based method to modulate the transmitted sound wave, calculate the distance feature, and extract the gesture movement information. However, these methods are easily interfered by external noise.

[0007] (3) Using an IMU (Inertial Measurement Unit)-based method to record the movement characteristics of the gesture through the sensors built into the smart device. However, these methods often require wearing special smart devices, which are often expensive and difficult to popularize. Summary of the Invention

[0008] Object of the Invention: The object of the present invention is to provide a gesture recognition system and method based on sound waves and sensors, which utilize channel estimation technology to extract the channel information changed due to gesture movement, fuse the values recorded by the gyroscope, and use a neural network model to achieve gesture recognition, so as to provide a new input method for smart devices.

[0009] Technical Solution: The gesture recognition system based on sound waves and sensors of the present invention includes a signal transmitter and a signal receiver; the signal transmitter is a smart watch worn on the wrist, which is worn on the user's wrist, transmits modulated ultrasonic waves, and records gyroscope data; the signal receiver is a microphone fixed in position; when the microphone receives the corresponding sound wave signal, channel estimation is performed by the cross-correlation method, that is, the dynamic information of the gesture; the gesture movement information recorded by the gyroscope is fused; a neural network model is used to extract and fuse the two gesture features and complete the final gesture classification.

[0010] Furthermore, the formula for transmitting modulated ultrasonic waves is:

[0011] s(t) = a(t)cos(2πf c t)

[0012] where a(t) is the Barker code sequence and f c is the modulation carrier;

[0013] The corresponding acoustic wave signal received by the microphone is expressed as:

[0014]

[0015] τ i represents the time delay corresponding to the i-th path, and n represents the number of all paths;

[0016] Demodulate the received acoustic wave signal, and the formula is:

[0017]

[0018] where R(t) is the received signal. After passing through a low-pass filter, the real part and the imaginary part of the signal are obtained. SingalI represents the real signal after demodulation, and SingalQ represents the imaginary signal after demodulation;

[0019] Finally, using the cross-correlation calculation formula, the channel information is obtained:

[0020]

[0021] c(m) = R xy (m - N) m ∈ [1, 2N - 1]

[0022] where y and x respectively represent the received signal after demodulation and the transmitted signal, c(m) is the output, and N is the signal length;

[0023] In order to obtain the dynamic characteristics of the gesture, subtract the previous frame from the current frame:

[0024] h(t) = h(t) - h(t - 1)

[0025] where h(t) is the channel information at the current moment, and h(t - 1) is the channel information at the previous moment;

[0026] By calculating the channel information within a fixed time period and combining the motion information recorded in the smart watch as gesture features; design a neural network model based on CNN. After offline training, it can recognize the current gesture.

[0027] The present invention also discloses an identification method for a gesture recognition system based on acoustic waves and sensors, including the following steps:

[0028] (1) Gesture feature extraction based on channel estimation: Extract the channel features changed due to gesture movement through channel estimation technology;

[0029] (2) Feature extraction based on gyroscope: By recording the three-axis features of the gyroscope built in the watch, the motion features of the gesture can be obtained after Kalman filtering;

[0030] (3) Gesture classification based on neural network model: After collecting data, train the network model to obtain the corresponding offline model, and then classify the extracted gesture features;

[0031] (4) Gesture interaction: Enrich the system input by recognizing gestures to improve the user experience.

[0032] Furthermore, in step (1), the transmitter emits modulated Barker code. After the receiver receives the signal, filter the signal and then perform IQ demodulation. Use the transmitted signal and the received signal to perform cross-correlation operation to obtain the channel information at each moment, and then use the method of subtracting the previous and the next frames to obtain the dynamic features changed due to gesture movement. The specific steps of step (1) include:

[0033] First, the formula for the watch to emit the modulated acoustic wave signal:

[0034] s(t) = a(t)cos(2πf c t)

[0035] where a(t) is the Barker code sequence and f c is the modulated carrier;

[0036] The acoustic wave signal received by the microphone is expressed as:

[0037]

[0038] τ i represents the time delay corresponding to the i-th path, and n represents the number of all paths;

[0039] Demodulate the received acoustic wave signal, and the formula is:

[0040]

[0041] where R(t) is the received signal. After passing through the low-pass filter, the real part and the imaginary part of the signal are obtained. SingalI represents the demodulated real signal, and SingalQ represents the demodulated imaginary signal;

[0042] Finally, use the cross-correlation calculation formula to obtain the channel information:

[0043]

[0044] c(m) = R xy (m - N) where m ∈ [1, 2N - 1]

[0045] where y and x represent the demodulated received signal and the transmitted signal respectively, c(m) is the output, and N is the signal length;

[0046] To obtain the dynamic features of the gesture, subtract the previous frame from the current frame:

[0047] h(t) = h(t) - h(t - 1)

[0048] where h(t) is the channel information at the current moment, and h(t - 1) is the channel information at the previous moment.

[0049] Furthermore, in step (2), locate the position of the gesture movement through an energy detection method, as shown in the formula:

[0050] E(t) = x(t) 2 + y(t) 2 + z(t) 2

[0051] where x(t), y(t), and z(t) are the values of the three axes of the gyroscope respectively; find the position exceeding the threshold as the starting point of the gesture movement by setting a threshold, and then use Kalman filtering to smooth the data and filter out the noise interference caused by the jitter of the hand itself.

[0052] Furthermore, in step (3), classify the extracted gesture features by using a convolutional neural network model. The method is as follows: After extracting the channel features, regard them as an image. Then further extract features by setting multiple convolutional kernels, as shown in the formula:

[0053]

[0054] where X, H, and Y represent the input, convolutional kernel, and output respectively. And k 0 , k 1 , k 2 represent the width, length, and number of channels of the input features respectively, and * represents the convolutional operation; after convolution, select the activation function:

[0055] ReLU(x) = max(x, 0)

[0056] where x is the input to enhance the fitting ability of the network; regard the extracted gyroscope features as a three-channel picture and extract features using a similar method; fuse the two features, then use a single fully connected layer as the output, and finally through the formula:

[0057]

[0058] where z i the output of the i-th neuron, S i is the final probability of the i-th class, and k is the number of classes; the final gesture class probability output is obtained, and the class with the maximum probability is selected as the recognized gesture.

[0059] Classification is performed using a convolutional neural network model. The two features are regarded as one-channel and three-channel images respectively, and then a convolutional layer is designed to automatically extract features. The two features are fused by a merging method, and finally a fully connected layer is used to achieve probability output.

[0060] In step (4), by recognizing the user's gesture, the device is controlled, realizing a more flexible input method and greatly improving the user experience.

[0061] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: By using channel estimation technology, the present invention extracts the channel information changed due to gesture movement, and at the same time, by fusing the motion information recorded in the gyroscope, multi-dimensional features of gesture movement can be extracted. The two features are complementary, which can effectively improve the gesture recognition accuracy and provide a new input method for intelligent devices. Different from the currently widely used vision-based solutions, the present invention is not limited by the influence of lighting conditions and reduces the dependence on device performance at the same time. The present invention only needs to utilize the microphones, speakers and IMU sensors commonly equipped in commercial devices, which can effectively reduce the deployment cost and improve the practicability of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is the system flowchart of the present invention;

[0063] Figure 2 is the baseband signal diagram of the present invention;

[0064] Figure 3 is the system schematic diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0065] The technical solution of the present invention will be further described below with reference to the drawings.

[0066] The present invention is a gesture recognition system based on sound waves and sensors, including a smart watch and a microphone, wherein:

[0067] The smart watch is worn on the user's wrist, emits modulated ultrasonic waves, and records gyroscope data.

[0068] The microphone at a fixed position is responsible for receiving the sound wave signal.

[0069] The present invention uses a 13-bit Barker code as the baseband signal. In fact, to increase the signal energy, the present invention performs an upsampling operation on it. After replicating each point multiple times, it is filtered using a low-pass filter. At the same time, to avoid inter-frame interference, zeros are appended at the end. To avoid interfering with the human ear, the signal is up-converted using a high-frequency carrier. Finally, a band-pass filter is used to smooth the data and avoid spectral leakage.

[0070] When the microphone receives the corresponding acoustic wave signal, channel estimation is performed through the method of cross-correlation, that is, the dynamic information of the gesture; the gesture motion information recorded by the gyroscope is fused; two gesture features are extracted and fused using a neural network model, and the final gesture classification is completed.

[0071] The gesture recognition method based on the above system includes the following steps:

[0072] Gesture feature extraction based on channel estimation: Through channel estimation technology, the channel features changed due to gesture motion are extracted; a modulated Barker code is transmitted by the transmitter. After being received by the receiver, the signal is filtered and then IQ demodulated. The cross-correlation operation is performed on the transmitted signal and the received signal to obtain the channel information at each moment, and then the dynamic features changed due to gesture motion are obtained by subtracting two consecutive frames.

[0073] (1) The formula for the acoustic wave signal modulated and transmitted by the watch:

[0074] s(t) = a(t)cos(2πf c t)

[0075] where a(t) is the Barker code sequence and f c is the modulated carrier

[0076] The received signal is expressed as

[0077]

[0078] τ i represents the time delay corresponding to the i-th path, and n represents all paths.

[0079] Demodulate the received signal, and the formula is:

[0080]

[0081] Then, after low-pass filtering, the real part and the imaginary part of the signal are obtained.

[0082] Finally, using the channel estimation calculation formula, the channel information is obtained:

[0083]

[0084] c(m) = Rxy (m - N) where m ∈ [1, 2N - 1]

[0085] Among them, y and x respectively represent the received signal after demodulation and the transmitted signal, c(m) is the output, and N is the signal length.

[0086] In order to obtain the dynamic features of the gesture, subtract the previous and next frames:

[0087] h(t) = h(t) - h(t - 1)

[0088] Among them, h(t) is the channel information at the current moment, and h(t - 1) is the channel information at the previous moment.

[0089] By calculating the channel information within a fixed time, obtain the dynamic information of a gesture once

[0090] (2) Feature extraction based on gyroscope: Locate the position of the gesture movement through an energy detection method, such as the formula:

[0091] E(t) = x(t) 2 + y(t) 2 + z(t) 2

[0092] Among them, x(t), y(t), and z(t) are the values of the three axes of the gyroscope respectively. By setting a sliding window of a fixed size, detect the standard deviation of the energy in each window, such as the formula:[[]]

[0093]

[0094] Among them, window represents the window size, and represents the average value of the current window. Specifically, find the position exceeding the threshold as the starting point of the gesture movement by setting a threshold. Then use Kalman filtering to smooth the data and filter out the noise interference caused by the jitter of the hand itself.

[0095] After extracting the channel information and gyroscope information, transmit them to the server through the local area network for gesture position correction. The gesture position extracted by the gyroscope can obtain the corresponding time point, and find the same position in the channel information, so as to effectively locate the gesture features.

[0096] (3) Gesture recognition based on neural network model: Use a convolutional neural network model for classification, and the method is as follows:[[]]

[0097] After extracting the channel features, regard them as an image that arrives together. Then further extract features by setting multiple convolutional kernels, such as the formula:[[]]

[0098]

[0099] Among them, X, H, and Y respectively represent the input, convolution kernel, and output. And k 0 , k 1 , k 2 respectively represent the width, length, and number of channels of the input feature map, and * represents the convolution operation. After convolution, an activation function is selected:

[0100] ReLU(x) = max(x, 0)

[0101] where x is the input, enhancing the fitting ability of the network.

[0102] Similarly, the extracted gyroscope features are regarded as three-channel pictures, and similar methods are used to extract features.

[0103] The two features are fused, and then a fully connected layer is used as the output. Finally, through the formula:

[0104]

[0105] where z i is the output of the i-th neuron, S i is the final probability of the i-th class, and k is the number of classes. The final gesture class probability output is obtained, and the class with the maximum probability is selected as the recognized gesture.

[0106] (4) Gesture interaction: By recognizing the user's gestures, the device is controlled, realizing a more flexible input method and greatly improving the user experience.

[0107] Embodiment

[0108] The present invention uses an existing smart watch and a commercial microphone as a signal reflector and a signal receiver respectively to preliminarily implement and verify the system. The speaker and microphone of the smart watch both support a sampling rate of 48 kHz. After the microphone receives the signal, the effective frequency band is extracted through a band-pass filter, and then IQ demodulation is used to obtain the corresponding real and imaginary part information. The channel estimation is obtained through cross-correlation operation with the transmitted signal. By fusing the motion information recorded by the gyroscope, it is input into the neural network model for gesture classification.

Claims

1. A gesture recognition system based on sound waves and sensors, characterized in that, it includes a signal transmitter and a signal receiver; the signal transmitter is a smart watch worn on the wrist, the smart watch is worn on the user's wrist, emits modulated ultrasonic waves, and records gyroscope data; the signal receiver is a microphone at a fixed position; when the microphone receives the corresponding sound wave signal, channel estimation is performed by the cross-correlation method, that is, the dynamic information of the gesture; the gesture motion information recorded by the gyroscope is fused; a neural network model is used to extract and fuse the two gesture features and complete the final gesture classification; The recognition method of the gesture recognition system based on sound waves and sensors includes the following steps: (1) Gesture feature extraction based on channel estimation: Through channel estimation technology, extract the channel features changed due to gesture motion; (2) Feature extraction based on gyroscope: By recording the three-axis features of the gyroscope built into the watch, the motion features of the gesture can be obtained after Kalman filtering; (3) Gesture classification based on neural network model: After collecting data, perform network model training to obtain the corresponding offline model, and then classify the extracted gesture features; (4) Gesture interaction: Enrich the system input by recognizing gestures and improve the user experience.

2. The gesture recognition system based on sound waves and sensors according to claim 1, characterized in that, the formula for emitting modulated ultrasonic waves is: s(t) = a(t)cos(2πf c t) where a(t) is the Barker code sequence, f c is the modulation carrier frequency; the corresponding sound wave signal received by the microphone is expressed as: τ i represents the time delay corresponding to the i-th path, and n represents the total number of paths; demodulate the received sound wave signal, and the formula is: where R(t) is the received signal, and after low-pass filtering, the real part and imaginary part of the demodulated signal are obtained. SingalI represents the real signal of demodulation, and SingalQ represents the imaginary signal of demodulation; Finally, use the cross-correlation calculation formula to obtain the channel information: c(m) = R xy (m - N) where m ∈ [1, 2N - 1] where y and x respectively represent the demodulated received signal and transmitted signal, c(m) is the final output, and N is the signal length; In order to obtain the dynamic characteristics of the gesture, subtract the previous and next frames: h(t) = h(t) - h(t - 1) where h(t) is the channel information at the current moment, and h(t - 1) is the channel information at the previous moment; By calculating the channel information within a fixed time and combining the motion information recorded in the smart watch as gesture features; Design a neural network model based on CNN, and after offline training, recognize the current gesture.

3. The gesture recognition system based on sound waves and sensors according to claim 1, characterized in that, in the recognition method step (2), a method of energy detection is used to locate the position of the gesture motion, and the formula is as follows: E(t) = x(t) 2 + y(t) 2 + z(t) 2 where x(t), y(t), and z(t) are the values of the three axes of the gyroscope respectively; by setting a threshold, find the position exceeding the threshold as the starting point of the gesture motion, and then use Kalman filtering to smooth the data and filter out the noise interference caused by the shaking of the hand itself.

4. The gesture recognition system based on sound waves and sensors according to claim 1, characterized in that, In step (3) of the recognition method, the classification of the extracted gesture features is performed using a convolutional neural network model, and the method is as follows: After extracting the channel features, they are regarded as an image with one channel, and then features are further extracted by setting multiple convolutional kernels. The formula is as follows: where X, H, and Y represent the input, convolutional kernel, and output respectively, and k 0 , k 1 , k 2 represent the width, length, and number of channels of the input features respectively, * represents the convolution operation; after convolution, an activation function is selected: ReLU(x) = max(x, 0) where x is the input to enhance the fitting ability of the network; the extracted gyroscope features are regarded as a three-channel picture, and features are extracted using a similar method; the two types of features are fused, and then a single fully connected layer is used as the output. Finally, through the formula: where z i is the output of the i-th neuron, S i is the final probability of the i-th class, and k is the number of classes; obtain the final gesture class probability output and select the class with the maximum probability as the recognized gesture.