Rotating Capacitive Sampler for Low Power Voice Activity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice activity detection (VAD) methods are slow and consume high power due to the need for multiple overlapping fast Fourier transforms (FFTs), which increases complexity and power consumption.
Innovation Solution
A multiphase differential output rotating capacitive sampler is used to achieve frequency down conversion across multiple frequency bands, sampling signals synchronously with a chirp multiplied by a window function, and detecting energy changes to indicate speech presence with low power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple overlapping FFTs are used to perform voice activity detection, then the detection accuracy is improved, but the power consumption and device complexity increase
Solution Approach 1:
The patent segments the frequency spectrum into multiple bands and processes each band independently using separate FFTs. This segmentation allows the system to achieve accurate voice activity detection across the full spectrum while managing computational complexity by dividing the problem into smaller, parallelizable units that can be efficiently implemented in hardware.
Solution Approach 2:
The patent transitions from time-domain analysis to frequency-domain analysis by applying FFT transforms. This dimensional change from time to frequency domain enables the system to identify voice activity patterns that are not apparent in the time domain, improving detection accuracy while the hardware implementation optimizes power consumption.
2Measurement precision
If multiple overlapping FFTs are used to perform voice activity detection, then the detection accuracy is improved, but the device complexity increases
Solution Approach 1:
The patent segments the frequency spectrum into multiple bands and processes each band independently using separate FFTs. This segmentation allows the system to achieve accurate voice activity detection across the full spectrum while managing computational complexity by dividing the problem into smaller, parallelizable units that can be efficiently implemented in hardware.
Solution Approach 2:
The patent combines multiple FFT outputs from different frequency bands to make a unified voice activity detection decision. By merging the results from parallel FFT processors in a coordinated manner, the system achieves high detection accuracy without proportionally increasing overall system complexity, as the combination logic is simpler than the individual FFT processing units.
3Productivity
If the FFT processing time is reduced to increase detection speed, then the productivity is improved, but the measurement precision deteriorates
Solution Approach 1:
The patent segments the frequency spectrum into multiple bands and processes each band independently using separate FFTs. This segmentation allows the system to achieve accurate voice activity detection across the full spectrum while managing computational complexity by dividing the problem into smaller, parallelizable units that can be efficiently implemented in hardware.
Solution Approach 2:
The patent implements continuous overlapping FFT processing where each subsequent FFT window overlaps with the previous one. This continuous processing ensures that voice activity detection occurs without gaps, maintaining high detection speed while the overlapping windows ensure that no transient voice signals are missed, preserving measurement precision.
Data Source
AI summary
An apparatus and method for voice activity detection. A multiphase differential output rotating capacitive sampler achieves a frequency down conversion over as many specific frequency bands as are required for analysis. A chirp is created in the rotating sampler as the sum of arbitrary frequencies across the desired analysis band multiplied by a window function. The chirp is sampled at a rate of rotation synchronous with the last state of burst of the chirp, allowing a non-phase synchronous pattern in the coefficient values and allowing a high-Q and arbitrary decomposition of the signal. After the sample is taken, the next clock signal to the sampler is used to define the output voltage of the sampler by shorting the output, which is entirely capacitive, to ground. Processing occurs in the analog domain rather than digitally, avoiding the need for FFTs and allowing for greater speed and lower power consumption.


