Low Bit Rate Voice Encoding Using Zero Crossing Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional low bit rate vocoders face challenges in reproducing authentic speech quality, particularly for non-English speakers, due to limitations in capturing resonant frequencies and accurate voice/unvoiced decision and pitch estimation.
Innovation Solution
The system employs zero crossings of the first formant, dividing by two and sampling at specific frequencies to use half the bit rate for excitation and the remainder for short-term spectrum analysis, eliminating voice/unvoiced pitch tracking, and utilizing Hanning modified sawtooth and spectral flattening to produce natural-sounding speech for all languages and speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional low bit rate vocoders use band-pass filters and multiple coders tuned to different frequencies, then speech can be reproduced at reduced bandwidth, but the speech quality becomes mechanical and lacks authentic resonant frequencies
Solution Approach 1:
The patent extracts only the essential excitation information (zero crossings of first formant) rather than transmitting complete spectral data through multiple band-pass filters. This extraction approach reduces bit rate while preserving the fundamental characteristics needed for natural speech reproduction.
Solution Approach 2:
The patent performs preliminary analysis to identify zero crossings of the first formant before encoding. This preliminary extraction of critical excitation features allows the system to operate at lower bit rates while maintaining speech quality, as the essential resonant frequency information is captured in advance.
2Ease of operation
If conventional vocoders use voice/unvoiced decision and pitch tracking algorithms, then excitation can be determined, but accuracy deteriorates for non-English speakers and various languages
Solution Approach 1:
The patent uses the speech signal itself to determine excitation characteristics through zero crossing detection of the first formant, rather than relying on external algorithms trained for specific languages. This self-service approach makes the system universally applicable across different languages and speakers without requiring language-specific tuning.
Solution Approach 2:
The patent replaces complex voice/unvoiced decision algorithms and pitch tracking mechanisms with a simpler zero crossing detection method. This substitution eliminates the need for language-specific training data and improves accuracy across diverse speakers and languages.
3Manufacturing precision
If standard systems record speech from 300 Hz to 4 kHz at 64 kbit/s, then full speech spectrum is captured, but bandwidth consumption is excessive for reduced bandwidth applications
Solution Approach 1:
The patent extracts only the critical first formant zero crossings rather than transmitting the complete speech spectrum. This selective extraction reduces bandwidth requirements from 64 kbit/s to significantly lower rates while preserving the essential resonant frequency information needed for natural speech reproduction.
Solution Approach 2:
The patent focuses computational and transmission resources on the specific frequency region of the first formant (300-1100 Hz) where zero crossings contain the most critical excitation information. This localized approach allows reduced bandwidth consumption while maintaining speech quality through selective spectral analysis.
Data Source
AI summary
A voice encoder/decoder (vocoder) may provide receiving a voice sample and generating zero crossings of the voice sample in response to voice excitation in a first formant and creating a corresponding output signal. Additional operations may include dividing the output signal by two, and sampling the output signal at a predefined frequency such that a resulting combination uses half of a bit rate for an excitation and a remainder for short term spectrum analysis.


