Speech Encoder Embedded Silence Noise Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current digital speech communication systems face challenges in efficiently managing bandwidth, particularly during silence and noise periods, leading to network congestion and quality degradation, as they lack effective methods for embedded silence and noise compression.

Innovation Solution

A speech encoder that generates embedded active and inactive speech bitstreams using a voice activity detector (VAD) to differentiate between active and inactive speech, employing discontinuous transmission (DTX) for intermittent updates of silence/noise information, ensuring smooth bandwidth continuity and efficient resource allocation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If continuous transmission of silence and noise information is used, then the decoded speech quality is maintained, but the network bandwidth consumption increases

Engineering Contradiction:
Improvedecoded speech qualityVSAvoidnetwork bandwidth consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent implements discontinuous transmission (DTX) where silence and noise information is transmitted periodically rather than continuously. During inactive speech periods, only intermittent updates of background noise parameters are sent, significantly reducing bandwidth consumption while maintaining adequate speech quality through periodic refreshes of the noise model at the decoder.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The patent extracts and separates the transmission of silence and noise information from active speech data. By using a voice activity detector to identify inactive speech periods, the system transmits only essential noise parameters during silence periods, removing unnecessary data transmission while preserving the ability to reconstruct background noise accurately at the decoder.

Inventive Principle:
Principle #2Taking out (Extraction)

2Productivity

If bandwidth reduction is implemented for inactive speech, then network capacity increases, but the ability to maintain bandwidth continuity decreases

Engineering Contradiction:
Improvenetwork capacityVSAvoidbandwidth continuity
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

The patent employs dynamic bandwidth allocation where the transmission rate adapts based on speech activity. During active speech, full bandwidth is used; during inactive speech, bandwidth is reduced to only essential noise parameters. The system dynamically switches between these modes using voice activity detection, allowing network capacity to handle more channels while maintaining perceived bandwidth continuity through smooth transitions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes transmission parameters based on speech activity state. During inactive speech, the system transmits reduced-parameter representations of background noise (such as spectral envelope parameters) rather than full-bandwidth audio data. This parameter reduction maintains network capacity while the decoder reconstructs the full bandwidth experience, preserving continuity of the audio output.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If embedded structure is used in bitstream, then network congestion is reduced, but the decoder complexity increases

Engineering Contradiction:
Improvenetwork congestion reductionVSAvoiddecoder complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the bitstream into embedded layers with a core layer containing essential speech and noise parameters, and enhancement layers containing additional quality information. During inactive speech, only the core layer with compressed noise parameters is transmitted, reducing network congestion. The decoder is designed to process this segmented structure efficiently, using the core layer to reconstruct acceptable quality without requiring complex processing of enhancement layers.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8195450B2Decoder with embedded silence and background noise compression
Publication Date: 2012.06.05 NYTELL SOFTWARE LLC
  • US8195450B2 patent drawing
  • US8195450B2 patent drawing
  • US8195450B2 patent drawing

AI summary

There is provided a method for use by a speech encoder to encode an input speech signal. The method comprises receiving the input speech signal; determining whether the input speech signal includes an active speech signal or an inactive speech signal; low-pass filtering the inactive speech signal to generate a narrowband inactive speech signal; high-pass filtering the inactive speech signal to generate a high-band inactive speech signal; encoding the narrowband inactive speech signal using a narrowband inactive speech encoder to generate an encoded narrowband inactive speech; generating a low-to-high auxiliary signal by the narrowband inactive speech encoder based on the narrowband inactive speech signal; encoding the high-band inactive speech signal using a wideband inactive speech encoder to generate an encoded wideband inactive speech based on the low-to-high auxiliary signal from the narrowband inactive speech encoder; and transmitting the encoded narrowband inactive speech and the encoded wideband inactive speech.