Wideband Speech Encoder Inactive Frame Bit Rate Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech coding technologies face challenges in efficiently encoding wideband speech signals while maintaining perceptual quality, particularly in reducing bit rates for inactive frames without degrading sound quality and supporting extended frequency ranges beyond traditional PSTN limits.

Innovation Solution

A method and apparatus for encoding speech signals that involve producing encoded frames with varying bit lengths based on frame activity, using a speech activity detector to differentiate active and inactive frames, and employing a coding scheme selector to adjust bit rates and coding modes, including descriptions of spectral envelopes over different frequency bands to optimize bit allocation and quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a lower bit rate is used to encode inactive frames, then the average bit rate is reduced, but the perceptual quality may be degraded

Engineering Contradiction:
Improvebit rateVSAvoidperceptual quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent applies parameter changes by dynamically adjusting the bit rate allocation based on frame activity status. Inactive frames are encoded at a lower bit rate (e.g., 16 bits) while active frames use a higher bit rate (e.g., 171 bits). The coding scheme selector changes coding parameters (wideband vs. narrowband mode) based on the activity detection, optimizing the balance between bit rate reduction and perceptual quality maintenance.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system dynamically adapts the encoding parameters based on the detected speech activity. The speech activity detector continuously monitors incoming frames and dynamically switches between different coding modes (wideband for active frames, narrowband for inactive frames). This dynamic adaptation allows the system to optimize bit rate usage while maintaining perceptual quality by applying the appropriate encoding strategy in real-time based on the current frame's characteristics.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If wideband frequency range is supported, then intelligibility and presence are improved, but the bit rate requirement increases

Engineering Contradiction:
ImproveintelligibilityVSAvoidbit rate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent applies local quality by providing different frequency band representations for different frame types. Wideband spectral envelope descriptions (covering extended frequency ranges up to 7-8 kHz) are applied locally to active frames where speech content requires high intelligibility. For inactive frames, only narrowband spectral descriptions are used, reducing bit rate while maintaining sufficient quality for background noise representation. This localized application of wideband encoding optimizes the trade-off between intelligibility and bit rate.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The speech signal is segmented into active and inactive frames based on speech activity detection. This segmentation allows the system to apply different encoding strategies: wideband encoding for active frames containing speech content and narrowband encoding for inactive frames containing only background noise. By segmenting the signal and applying appropriate encoding to each segment, the system achieves high intelligibility when needed while minimizing bit rate during silence periods.

Inventive Principle:
Principle #1Segmentation

3Quantity of substance

If different coding modes are used for active and inactive frames, then the average bit rate is reduced, but the device complexity increases

Engineering Contradiction:
Improveaverage bit rateVSAvoidcoding scheme complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The speech activity detector performs preliminary analysis of each incoming frame to determine whether it contains speech or only background noise before encoding. This preliminary classification allows the coding scheme selector to pre-select the appropriate encoding mode (wideband or narrowband) based on the frame type. By performing this classification in advance, the system avoids complex real-time decision-making during encoding and simplifies the overall processing architecture while achieving significant bit rate reduction.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The coding scheme selector automatically selects the appropriate coding mode based on the speech activity detection results without requiring external control or complex optimization algorithms. The system serves itself by using the speech activity information to directly determine the encoding parameters, eliminating the need for additional complexity in mode selection logic. This self-service approach simplifies the device architecture while achieving efficient bit rate optimization.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9324333B2Systems, methods, and apparatus for wideband encoding and decoding of inactive frames
Publication Date: 2016.04.26 QUALCOMM INC
  • US9324333B2 patent drawing
  • US9324333B2 patent drawing
  • US9324333B2 patent drawing

AI summary

Speech encoders and methods of speech encoding are disclosed that encode inactive frames at different rates. Apparatus and methods for processing an encoded speech signal are disclosed that calculate a decoded frame based on a description of a spectral envelope over a first frequency band and the description of a spectral envelope over a second frequency band, in which the description for the first frequency band is based on information from a corresponding encoded frame and the description for the second frequency band is based on information from at least one preceding encoded frame. Calculation of the decoded frame may also be based on a description of temporal information for the second frequency band that is based on information from at least one preceding encoded frame.