Audio Encoding Method Using Spectral Sparseness for Complexity Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio encoding methods in voice communications systems are complex and lack accuracy due to the high operation complexity of hybrid encoders, which compare and select between speech and non-speech signal encoders.
Innovation Solution
An audio encoding method that determines the sparseness of energy distribution on the spectrum of audio frames to choose between a time-frequency transform-based encoding method and a linear prediction-based method, reducing encoding complexity while maintaining high accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a hybrid encoder with two sub encoders is used to encode speech and non-speech signals, then the encoding accuracy is improved, but the operation complexity increases significantly
Solution Approach 1:
The patent applies preliminary action by calculating sparseness parameters (spectral flatness, spectral entropy, zero-crossing rate) before the encoding decision is made. These parameters are computed in advance to predict whether the signal is speech or non-speech, allowing the system to select the appropriate sub-encoder without requiring complex real-time comparison and iteration, thus reducing operation complexity while maintaining encoding accuracy
Solution Approach 2:
The patent introduces sparseness parameters as intermediary variables that mediate between the raw audio signal and the encoder selection decision. These parameters (spectral flatness, spectral entropy, zero-crossing rate) serve as intermediate measurements that simplify the complex task of signal classification, enabling accurate speech/non-speech distinction without direct complex signal comparison
2Measurement precision
If a closed-loop encoding method with comparison and selection is used, then the optimum sub encoder is selected, but the encoding complexity increases
Solution Approach 1:
The patent extracts the essential characteristics of speech and non-speech signals by computing specific sparseness parameters (spectral flatness, spectral entropy, zero-crossing rate) separately. Instead of comparing entire complex signals in a closed loop, the system extracts these key features and makes encoding decisions based on the extracted parameter values, significantly simplifying the encoding process while maintaining selection accuracy
Solution Approach 2:
The patent transforms the complex signal comparison problem into a simpler parameter-based decision problem by changing from direct signal comparison to comparison of derived parameters (spectral flatness threshold, spectral entropy threshold, zero-crossing rate threshold). This parameter transformation reduces encoding complexity while preserving the ability to accurately select the optimum sub-encoder
Data Source
AI summary
An audio encoding method includes dividing an energy spectrum of a current audio frame into P FFT energy spectrum coefficients; determining a minimum bandwidth of distribution, on spectrum, of first-preset-proportion energy of the current audio frame according to the energy of the P FFT energy spectrum coefficients of the current audio frame, wherein the minimum bandwidth of distribution, on spectrum, of first preset proportion energy of the current audio frame indicates sparseness of distribution, on the spectrum, of energy of the current audio frame; and determining to use a linear-prediction-based encoding method to encode the current audio frame in response to the minimum bandwidth of distribution is greater than a first preset value.
