ASR Triggering via Acoustic and Bone Conduction Signal Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition systems based on verbal commands or physical taps often fail to function seamlessly in noisy environments and are prone to false triggers, leading to device power drainage and user frustration.
Innovation Solution
An automatic speech recognition (ASR) triggering system that combines acoustic and non-acoustic signals from a microphone and accelerometer, respectively, to generate a trigger signal, using a processor to integrate these signals and prevent false activations by ensuring simultaneous detection of both acoustic and non-acoustic voice activity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech recognition is triggered by verbal commands or physical taps, then the system can be activated easily, but false triggers occur in noisy environments leading to device power drainage
Solution Approach 1:
The patent combines acoustic signal detection (microphone) with non-acoustic vibration signal detection (accelerometer) to create a dual-channel triggering system. The processor analyzes both acoustic and vibration signals simultaneously, requiring coincidence between the two channels to generate an ASR trigger. This merging of detection modalities resolves the contradiction by maintaining ease of activation through voice commands while eliminating false triggers from ambient noise or unintended physical contacts.
2Speed
If the system uses acoustic signals alone for triggering, then the response is fast, but bystander speech can cause false triggers
Solution Approach 1:
The vibration signal from the accelerometer acts as an intermediary verification mechanism. When the acoustic detector identifies a potential key phrase, the system checks for corresponding vibration signals that would indicate the user is physically interacting with the device (e.g., tapping or holding it). This intermediary vibration detection layer prevents false triggers from bystander speech while maintaining fast response times, as the vibration channel provides rapid confirmation of legitimate user intent.
3Productivity
If the system continuously monitors for triggers, then activation is immediate, but device power is drained
Solution Approach 1:
The triggering system is segmented into two independent detection channels (acoustic and vibration) that operate in parallel but are processed cooperatively. Each channel can independently detect its respective signal type, and the processor combines their outputs to make the final trigger decision. This segmentation allows the system to maintain continuous monitoring capability for immediate activation while optimizing power management by requiring both channels to confirm a trigger event, thereby reducing unnecessary full-system activations and conserving battery power.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system effectively reduces false triggers by confirming that the user is the source of the command, enhancing the reliability and efficiency of speech recognition in noisy conditions while conserving device power.
Implementation Method 1
an accelerometer to generate a non-acoustic signal representing a bone conduction vibration
Data Source
AI summary
An automatic speech recognition (ASR) triggering system, and a method of providing an ASR trigger signal, is described. The ASR triggering system can include a microphone to generate an acoustic signal representing an acoustic vibration and an accelerometer worn in an ear canal of a user to generate a non-acoustic signal representing a bone conduction vibration. A processor of the ASR triggering system can receive an acoustic trigger signal based on the acoustic signal and a non-acoustic trigger signal based on the non-acoustic signal, and combine the trigger signals to gate an ASR trigger signal. For example, the ASR trigger signal may be provided to an ASR server only when the trigger signals are simultaneously asserted. Other embodiments are also described and claimed.


