Audio Encoding Method Using Spectral Sparseness for Complexity Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio encoding methods in voice communications systems are complex and lack accuracy due to the high operation complexity of hybrid encoders, which compare and select between speech and non-speech signal encoders.

Innovation Solution

An audio encoding method that determines the sparseness of energy distribution on the spectrum of audio frames to choose between a time-frequency transform-based encoding method and a linear prediction-based method, reducing encoding complexity while maintaining high accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a hybrid encoder with two sub encoders is used to encode speech and non-speech signals, then the encoding accuracy is improved, but the operation complexity increases significantly

Engineering Contradiction:
Improveencoding accuracyVSAvoidoperation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by calculating sparseness parameters (spectral flatness, spectral entropy, zero-crossing rate) before the encoding decision is made. These parameters are computed in advance to predict whether the signal is speech or non-speech, allowing the system to select the appropriate sub-encoder without requiring complex real-time comparison and iteration, thus reducing operation complexity while maintaining encoding accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces sparseness parameters as intermediary variables that mediate between the raw audio signal and the encoder selection decision. These parameters (spectral flatness, spectral entropy, zero-crossing rate) serve as intermediate measurements that simplify the complex task of signal classification, enabling accurate speech/non-speech distinction without direct complex signal comparison

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a closed-loop encoding method with comparison and selection is used, then the optimum sub encoder is selected, but the encoding complexity increases

Engineering Contradiction:
Improveencoder selection accuracyVSAvoidencoding complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the essential characteristics of speech and non-speech signals by computing specific sparseness parameters (spectral flatness, spectral entropy, zero-crossing rate) separately. Instead of comparing entire complex signals in a closed loop, the system extracts these key features and makes encoding decisions based on the extracted parameter values, significantly simplifying the encoding process while maintaining selection accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the complex signal comparison problem into a simpler parameter-based decision problem by changing from direct signal comparison to comparison of derived parameters (spectral flatness threshold, spectral entropy threshold, zero-crossing rate threshold). This parameter transformation reduces encoding complexity while preserving the ability to accurately select the optimum sub-encoder

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11074922B2Hybrid encoding method and apparatus for encoding speech or non-speech frames using different coding algorithms
Publication Date: 2021.07.27 TOP QUALITY TELEPHONY LLC
  • US11074922B2 patent drawing

AI summary

An audio encoding method includes dividing an energy spectrum of a current audio frame into P FFT energy spectrum coefficients; determining a minimum bandwidth of distribution, on spectrum, of first-preset-proportion energy of the current audio frame according to the energy of the P FFT energy spectrum coefficients of the current audio frame, wherein the minimum bandwidth of distribution, on spectrum, of first preset proportion energy of the current audio frame indicates sparseness of distribution, on the spectrum, of energy of the current audio frame; and determining to use a linear-prediction-based encoding method to encode the current audio frame in response to the minimum bandwidth of distribution is greater than a first preset value.