In-Ear Audio Event Classification for Real-Time Silent Commands
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems struggle to accurately and efficiently classify a wide variety of nonverbal audio events produced by humans, particularly those captured from inside the ear canal, due to processing power limitations and real-time communication constraints, which are essential for applications like health monitoring, artifact removal, and silent speech interfaces.
Innovation Solution
A system and method for training a classification module using an in-ear microphone to capture audio signals, extract features, and associate them with specific nonverbal events, employing machine learning algorithms like SVM, GMM, and MLP to accurately identify events such as teeth clicking, blinking, and throat clearing without requiring extensive processing power.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If filtering methods are used to detect and identify nonverbal audio events, then measurement precision is improved, but productivity deteriorates due to time-consuming processing and extensive processing power requirements
Solution Approach 1:
The patent replaces traditional mechanical filtering methods with a machine learning-based classification system. The system uses trained classifiers (such as support vector machines, neural networks, or decision trees) that have been pre-trained on labeled audio data to directly classify nonverbal audio events, eliminating the need for real-time filtering operations and significantly improving processing efficiency while maintaining detection accuracy
Solution Approach 2:
The patent applies preliminary action by pre-training classification models offline using extensive labeled audio data before deployment. The training process extracts features and learns patterns in advance, so that during real-time operation, the system can quickly classify events without performing computationally intensive processing, thus resolving the contradiction between accuracy and efficiency
2Adaptability or versatility
If multiple nonverbal audio event types are detected in real time, then adaptability is improved, but device complexity increases due to processing power requirements
Solution Approach 1:
The patent implements a universal classification system that can detect and classify multiple types of nonverbal audio events (such as coughing, sneezing, throat clearing, laughter, etc.) using a single machine learning model. The system extracts general audio features and uses a trained classifier to identify various event types, eliminating the need for separate detection mechanisms for each event type and reducing overall device complexity
Solution Approach 2:
The patent changes the approach from using multiple specialized filters for different event types to using a single classification model that processes audio features and outputs multiple event type predictions. This parameter change from filter-based to model-based classification allows the system to handle diverse event types with unified processing logic, reducing computational complexity
Data Source
AI summary
A system and method for training a classification module of nonverbal audio events and a classification module for use in a variety of nonverbal audio event monitoring, detection and command systems. The method comprises capturing an in-ear audio signal from an occluded ear and defining at least one nonverbal audio event associated to the captured in-ear audio signal. Then sampling and extracting features from the in-ear audio signal. Once the extracted features are validated, associating the extracted features to the at least one nonverbal audio event and updating the classification module with the association. The nonverbal audio event comprises one or a combination of user-induced or externally-induced nonverbal audio events such as teeth clicking, tongue clicking, blinking, eye closing, teeth grinding, throat clearing, saliva noise, swallowing, coughing, talking, yawning with inspiration, yawning with expiration, respiration, heartbeat and head or body movement, wind, earpiece insertion or removal, degrading parts, etc.


