Voice Event Detection Using Vibration Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice activity detection (VAD) systems cannot distinguish human speech from other sounds, leading to unnecessary power consumption in computing systems as they trigger downstream processes based on amplitude alone, regardless of the audio signal type.
Innovation Solution
A voice event detection apparatus and method that converts input audio signals into vibration data, using a vibration to digital converter and a computing unit to sum vibration counts over a number of frames, allowing for the correct triggering of downstream modules and distinguishing speech from non-speech by analyzing vibration patterns and rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If VAD classifies audio signal by comparing amplitude with threshold, then speech detection can be performed, but human speech cannot be distinguished from other sounds leading to unnecessary power consumption
Solution Approach 1:
The patent changes the detection parameter from simple amplitude comparison to analyzing vibration rate characteristics. By examining the rate of vibration in the audio signal, the system can distinguish human speech from other sounds with similar amplitudes, thereby improving detection precision while avoiding unnecessary downstream processing and reducing power consumption.
2Productivity
If downstream module is triggered by large amplitude audio signal, then speech processing can be initiated, but unnecessary processes are activated for non-speech sounds
Solution Approach 1:
The patent applies preliminary action by performing vibration rate analysis before triggering the downstream speech processing module. This preliminary check examines whether the audio signal exhibits characteristics of human speech (specific vibration rate patterns) before committing resources to full speech processing, thus avoiding wasted energy on non-speech sounds while maintaining efficient speech processing when needed.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Effectively reduces power consumption by accurately identifying voice events and wake phonemes, preventing unnecessary activation of downstream processes and optimizing power usage in computing systems.
Implementation Method 1
The vibration to digital converter is configured to convert an input audio signal into vibration data
Data Source
AI summary
A voice event detection apparatus is disclosed. The apparatus comprises a vibration to digital converter and a computing unit. The vibration to digital converter is configured to convert an input audio signal into vibration data. The computing unit is configured to trigger a downstream module according to a sum of vibration counts of the vibration data for a number X of frames. In an embodiment, the voice event detection apparatus is capable of correctly distinguishing a wake phoneme from the input vibration data so as to trigger a downstream module of a computing system. Thus, the power consumption of the computing system is saved.


