Speech Coding System Reducing Metallic Artifacts via Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Low-bit rate speech coders, such as sinusoidal coders, produce metallic-sounding artifacts due to an inadequate sparse signal representation, which are not mitigated by higher bit rates or compensation for transmission losses, delays, and jitter.

Innovation Solution

A system that generates an artificial mixed signal by extracting features from the decoded or encoded audio signal and mixing it with the decoded signal to enhance the frequency band, effectively reducing metallic artifacts without requiring additional bit transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a low-bit rate speech coder is used, then the bit rate is reduced, but metallic artifacts are introduced due to inadequate sparse signal representation

Engineering Contradiction:
Improvebit rateVSAvoidmetallic artifacts
Core Design Contradiction:
Quantity of substanceVSObject-generated harmful factors

Solution Approach 1:

The patent introduces an artificial mixed signal as an intermediary component that bridges the gap between the sparse encoded signal and the desired full speech signal. This artificial signal, generated by mapping extracted features, acts as a mediator that fills in the missing speech structures without requiring additional transmission bits, thereby reducing metallic artifacts while maintaining low bit rate operation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a copy of the speech signal in the form of an artificial mixed signal that is synthesized from extracted features. This copied representation captures the missing speech structures and when mixed with the decoded signal, reconstructs the absent speech components, effectively eliminating metallic artifacts without increasing the transmitted bit rate

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If more information is added to the transmitted signal to capture speech structure, then the quality improves, but the bit rate increases

Engineering Contradiction:
Improvespeech structure captureVSAvoidbit rate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent extracts essential speech features from the encoded or decoded signal and uses only these extracted features to generate the artificial mixed signal. By taking out only the necessary feature information rather than transmitting the complete speech signal, the system captures speech structures effectively while maintaining low bit rate operation

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of transmitting additional speech information, the patent creates a synthetic copy of the missing speech structures by mapping extracted features to an artificial mixed signal. This copied representation provides the necessary speech structure information without increasing the transmitted bit rate

Inventive Principle:
Principle #26Copying

3Reliability

If speech frames are stretched or concealment frames are inserted to compensate for jitter, then transmission issues are handled, but metallic artifacts are introduced

Engineering Contradiction:
Improvejitter compensationVSAvoidmetallic artifacts
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The artificial mixed signal serves as an intermediary that masks the artifacts introduced by jitter compensation operations. By mixing this synthesized signal with the decoded signal, the system covers up the metallic artifacts that arise from frame stretching and concealment operations, maintaining reliability while reducing harmful artifacts

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8069049B2Speech coding system and method
Publication Date: 2011.11.29 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8069049B2 patent drawing
  • US8069049B2 patent drawing
  • US8069049B2 patent drawing

AI summary

A system for enhancing a signal regenerated from an encoded audio signal. The system comprises a decoder arranged to receive the encoded audio signal and produce a decoded audio signal, a feature extraction means arranged to receive at least one of the decoded and encoded audio signal and extract at least one feature from at least one of the decoded and encoded audio signal, a mapping means arranged to map the at least one feature to an enhancement signal and operable to generate and output the enhancement signal, whereby the enhancement signal has a frequency band that is within the decoded audio signal frequency band, and a mixing means arranged to receive the decoded audio signal and the enhancement signal and mix the enhancement signal with the decoded audio signal.