Wearable Audio Gesture Detection via Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Wearable devices, such as frame-based augmented reality (AR) devices, face challenges in detecting user inputs without adding bulky and unattractive physical sensors, which can compromise the device's form factor and user experience.
Innovation Solution
The use of two or more internal microphones in the wearable device to generate audio signals based on user interactions, which are then processed to predict or generate user inputs using a machine-learned model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If physical sensors (e.g., touch pads) are added to detect user inputs, then user input detection capability is improved, but device form factor and aesthetics are compromised
Solution Approach 1:
The patent replaces mechanical touch sensors with an acoustic field-based detection system. Microphones capture sound waves generated by user interactions (taps, swipes, pinches) and machine learning models interpret these acoustic signatures to identify gestures, eliminating the need for physical touch pads while maintaining gesture detection functionality
Solution Approach 2:
The patent introduces sound waves as an intermediary medium between user input and device detection. Instead of direct mechanical contact with sensors, user gestures generate acoustic signals that propagate through air to microphones, enabling indirect but accurate detection of interaction patterns
2Ease of operation
If physical sensors are added to detect user inputs, then user input detection capability is improved, but device weight and bulk are increased
Solution Approach 1:
The patent substitutes heavy mechanical touch sensor components with lightweight acoustic sensing elements (microphones). The detection mechanism transitions from mechanical contact sensing to acoustic wave detection, significantly reducing the weight and bulk of input detection hardware while maintaining functionality
3Measurement precision
If more microphones are added to improve gesture detection accuracy, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent makes the microphone array multi-functional by using the same acoustic sensors for both voice command recognition and gesture detection. The machine learning system distinguishes between different input types (speech vs. gestures) based on acoustic characteristics, allowing existing microphones to serve dual purposes and reducing the need for additional specialized sensors
Solution Approach 2:
The patent changes the parameters used for analyzing microphone signals to enable gesture detection. By extracting features such as sound intensity, frequency spectrum, temporal patterns, and inter-microphone time delays from acoustic signals, the system transforms voice-oriented microphone data into gesture-detectable parameters without adding hardware complexity
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution allows for intuitive user input detection without the need for physical sensors, preserving the device's aesthetics and functionality while enabling accurate gesture recognition and coordinate prediction.
Implementation Method 1
generating two or more audio signals (e.g., by two or more microphones) based on a user interaction with the wearable device
Implementation Method 2
The device includes a phase offset layer and the microphones are embedded internally in the frame, wherein the coordinate is based on phase delays caused by the phase offset layer
Data Source
AI summary
A method including generating two or more audio signals (e.g., by two or more microphones) based on a user interaction with a wearable device, generating an audio signature based on the two or more audio signals, and identifying at least one of a coordinate or a gesture based on an output of a machine-learned model given the audio signature.


