CELP Speech Energy Estimation via Partial Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for determining speech energy in CELP-encoded bit streams require full decoding, which is computationally expensive and inefficient, especially in scenarios like voice conferencing where only active speakers need to be identified.
Innovation Solution
Estimating speech energy based on a set of four or fewer CELP parameters extracted from a partially decoded CELP-encoded bit stream, without calculating the linear prediction coding (LPC) filter response energy, allowing for quick and resource-efficient estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If full decoding of CELP-encoded bit stream is performed to determine speech energy, then measurement precision is improved, but device complexity and computational load increase
Solution Approach 1:
The patent extracts only the necessary CELP parameters (excitation energy, LPC synthesis filter energy, and their product) from the encoded bit stream without performing full decoding. This extraction approach obtains sufficient statistics for speech energy estimation while avoiding the computational complexity of complete decoding operations.
Solution Approach 2:
The patent creates a simplified model of the CELP decoding process that copies only the essential parameter relationships needed for energy estimation. By working with extracted parameters rather than reconstructing the full audio signal, the system achieves measurement precision equivalent to full decoding with significantly reduced complexity.
2Measurement precision
If full decoding is performed to measure speech energy, then measurement precision is improved, but processing time increases
Solution Approach 1:
The patent extracts only the necessary CELP parameters (excitation energy, LPC synthesis filter energy, and their product) from the encoded bit stream without performing full decoding. This extraction approach obtains sufficient statistics for speech energy estimation while avoiding the computational complexity of complete decoding operations.
Solution Approach 2:
The patent applies partial action by performing only the minimal decoding operations required to extract energy-related parameters, rather than completing the full decoding process. This partial decoding approach provides sufficient information for speech energy measurement with substantially reduced processing time.
3Measurement precision
If LPC filter response energy is calculated for each frame, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent pre-calculates and stores the LPC synthesis filter energy values during the encoding process, making them available for direct extraction during decoding. This preliminary computation eliminates the need to recalculate filter response energy for each frame, reducing real-time computational complexity while maintaining measurement precision.
Solution Approach 2:
The patent extracts pre-computed energy values (excitation energy and LPC synthesis filter energy) directly from the encoded bit stream parameters without performing real-time filter response calculations. This extraction of pre-prepared data eliminates complex real-time computations while preserving frame energy estimation accuracy.
Data Source
AI summary
Methods, systems, and non-transitory computer readable media for estimating speech energy of an encoded bit stream based on coding parameters extracted from the partially-decoded bit stream are disclosed. In an embodiment, a disclosed method includes receiving a CELP-encoded bit stream, partially decoding the bit stream, and estimating the speech energy of the bit stream based a set of four or fewer CELP parameters extracted from the partially decoded bit stream. In another embodiment, a disclosed method includes receiving a CELP-encoded bit stream, partially decoding the bit stream, extracting at least one CELP parameter from the partially decoded bit stream, and estimating the speech energy of the bit stream based on the extracted at least one CELP parameter without calculating a linear prediction coding (LPC) filter response energy.


