Speech Coding Selection via Pitch Lag Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech coding technologies face challenges in effectively classifying between time-domain coding and frequency-domain coding, particularly at high bit rates, leading to suboptimal performance for signals with short pitch or varying voicing periodicity, resulting in perceptual distortion and quality degradation.
Innovation Solution
A method and apparatus that select between time-domain coding and frequency-domain coding based on bit rate and pitch lag detection, where frequency-domain coding is chosen for high bit rates with short pitch signals and time-domain coding is preferred for low bit rates or strong voicing periodicity, ensuring optimal coding for specific speech characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If frequency domain coding is used for high bit rates, then coding efficiency is improved, but perceptual distortion increases for short pitch signals
Solution Approach 1:
The patent implements dynamic selection between time-domain and frequency-domain coding methods based on real-time detection of pitch lag and voicing periodicity characteristics. The coding method is not fixed but adapts dynamically to the specific signal characteristics, allowing the system to switch between domains to optimize both coding efficiency and perceptual quality for different speech segments.
Solution Approach 2:
The system changes the fundamental parameter of coding domain (time vs frequency) based on detected signal characteristics such as pitch lag and voicing periodicity. By monitoring these parameters and selecting the appropriate coding domain, the system resolves the contradiction between coding efficiency and perceptual quality for short pitch signals at high bit rates.
2Manufacturing precision
If time domain coding is used for short pitch signals, then perceptual quality is maintained, but coding efficiency decreases at high bit rates
Solution Approach 1:
The system dynamically selects the coding domain based on pitch lag detection and voicing periodicity analysis. For short pitch signals, the system can dynamically switch to time-domain coding to maintain perceptual quality, while for other signals it uses frequency-domain coding for better efficiency.
Solution Approach 2:
The coding domain parameter is changed based on detected signal characteristics. When short pitch is detected, the system changes from frequency-domain to time-domain coding approach, optimizing the balance between quality and efficiency for specific signal conditions.
3Device complexity
If a single coding method is used for all bit rates, then device complexity is reduced, but performance deteriorates for varying speech characteristics
Solution Approach 1:
The patent implements a dynamic classification system that detects pitch lag and voicing periodicity to determine the appropriate coding method. This dynamic approach maintains relatively simple individual coding methods while using intelligent classification to achieve high performance across varying speech characteristics and bit rates.
Solution Approach 2:
The system changes the coding method parameter based on detected speech characteristics such as pitch lag and voicing periodicity. This allows the system to maintain low complexity for each individual coding method while achieving high overall performance through parameter-based selection.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
A method for processing speech signals prior to encoding a digital signal comprising audio data includes selecting frequency domain coding or time domain coding based on a coding bit rate to be used for coding the digital signal and a short pitch lag detection of the digital signal.