Speech Encoding Bit Allocation via Audible Significance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech/audio encoding methods fail to accurately encode audibly significant frequency domain regions independently of subbands, leading to suboptimal sound quality due to insufficient bit allocation in these regions.

Innovation Solution

A speech/audio encoding apparatus that identifies audibly significant frequency domain regions using linear prediction coefficients, repositions them, and determines bit allocation accordingly, allowing for high-accuracy encoding of these regions without influencing non-significant frequency domains.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If bits are allocated based on subband energy distribution, then encoding efficiency is improved, but encoding accuracy of audibly significant frequency regions is not necessarily improved

Engineering Contradiction:
Improveencoding efficiencyVSAvoidencoding accuracy of audibly significant frequency regions
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent applies local quality by differentiating bit allocation based on audible significance rather than uniform subband treatment. Specifically, the system identifies frequency regions where human ears are sensitive (audibly significant regions) and allocates more bits to these regions, while allocating fewer bits to less sensitive regions. This creates non-uniform bit allocation across different frequency regions within subbands, improving encoding accuracy where it matters most to human perception.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the parameter basis for bit allocation from subband energy distribution to audible significance assessment. By introducing audible significance as a new parameter and using it to determine bit allocation strategy, the system optimizes encoding accuracy for perceptually important frequency regions while maintaining overall encoding efficiency.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If additional information indicating bit allocation to frequency domain regions is required, then encoding precision is improved, but device complexity increases

Engineering Contradiction:
Improveencoding precisionVSAvoiddevice complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies self-service by enabling the decoding device to independently determine bit allocation information without requiring explicit transmission from the encoding device. The decoding device uses its own audible significance assessment (based on received linear prediction coefficients) to reconstruct the same bit allocation decisions made during encoding. This eliminates the need for additional bit allocation information in the bitstream while maintaining encoding precision.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The linear prediction coefficients serve multiple functions: they are used for speech parameter encoding and simultaneously serve as the basis for both encoding and decoding side audible significance assessments. This multi-functionality allows the system to derive bit allocation information implicitly without requiring separate information channels.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10446159B2Speech/audio encoding apparatus and method thereof
Publication Date: 2019.10.15 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US10446159B2 patent drawing
  • US10446159B2 patent drawing
  • US10446159B2 patent drawing

AI summary

A speech/audio encoding device for selectively allocating bits for higher precision encoding. The speech/audio encoding device receives a time-domain speech/audio input signal, transforms the speech/audio input signal into a frequency domain, and quantizes an energy envelope corresponding to an energy level for a frequency spectrum of the speech/audio input signal. The speech/audio encoding device further groups quantized energy envelopes into a plurality of groups, determines a perceptual significant group including one or more significant bands and a local-peak frequency, and allocates bits to a plurality of subbands corresponding to the grouped quantized energy envelopes, in which each of the subbands is obtained by splitting the frequency spectrum of the speech/audio input signal. The speech/audio encoding device encodes the frequency spectrum using the bits allocated to the subbands.