Machine-Learning Gesture Recognition for Wearable Audio Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Wearable devices like earbuds without touch sensors face challenges in detecting touch inputs and gestures due to size, power, and manufacturing constraints, limiting their ability to recognize user interactions.
Innovation Solution
Implementing a machine-learning based gesture recognition system using non-touch sensors such as accelerometers, optical sensors, and microphones, which feed data into a convolutional neural network model to predict gestures without the need for a touch sensor, enabling detection of taps, swipes, and other gestures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If touch sensors are used for gesture detection, then gesture recognition accuracy is improved, but device size and manufacturing cost increase
Solution Approach 1:
The patent replaces the mechanical/electrical touch sensor system with a machine learning-based system that processes data from existing non-touch sensors (accelerometers, optical sensors, microphones). The convolutional neural network model analyzes patterns from these sensors to detect gestures, eliminating the need for dedicated touch sensors and reducing device complexity while maintaining gesture recognition capability
Solution Approach 2:
The patent makes existing multi-functional sensors serve dual purposes: their primary functions (motion detection, light detection, audio capture) plus gesture detection. The machine learning model integrates data from these sensors to perform gesture recognition, allowing the same hardware to serve multiple functions without adding dedicated touch sensor components
2Measurement precision
If touch sensors are used for gesture detection, then gesture recognition accuracy is improved, but power consumption increases
Solution Approach 1:
The patent replaces the power-intensive touch sensor system with a machine learning-based processing approach that utilizes data from existing sensors. The convolutional neural network model processes sensor data to detect gestures, eliminating the continuous power consumption associated with dedicated touch sensors while maintaining accurate gesture recognition
Solution Approach 2:
The patent enables existing sensors to serve themselves by having the machine learning model directly process their output data for gesture detection. The system uses the data these sensors already collect for their primary functions, extracting additional gesture information without requiring separate power-intensive sensing hardware
3Measurement precision
If touch sensors are used for gesture detection, then gesture recognition accuracy is improved, but manufacturing cost increases
Solution Approach 1:
The patent replaces expensive touch sensor hardware with a software-based machine learning solution that processes data from existing sensors. The convolutional neural network model provides accurate gesture recognition without requiring costly touch sensor components, simplifying manufacturing and reducing bills of materials
Solution Approach 2:
The patent creates a virtual model of touch interaction through machine learning by training the convolutional neural network on sensor data patterns that correspond to various gestures. This software-based copy of touch sensor functionality achieves similar recognition accuracy without the physical hardware cost
Data Source
AI summary
The subject technology receives, from a first sensor of a device, first sensor output of a first type. The subject technology receives, from a second sensor of the device, second sensor output of a second type, the first and second sensors being non-touch sensors. The subject technology provides the first sensor output and the second sensor output as inputs to a machine learning model, the machine learning model having been trained to output a predicted touch-based gesture based on sensor output of the first type and sensor output of the second type. The subject technology provides a predicted touch-based gesture based on output from the machine learning model. Further, the subject technology adjusts an audio output level of the device based on the predicted gesture, and where the device is an audio output device.


