Speech Coding System Reducing Metallic Artifacts via Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Low-bit rate speech coders, such as sinusoidal coders, produce metallic-sounding artifacts due to an inadequate sparse signal representation, which are not mitigated by higher bit rates or compensation for transmission losses, delays, and jitter.
Innovation Solution
A system that generates an artificial mixed signal by extracting features from the decoded or encoded audio signal and mixing it with the decoded signal to enhance the frequency band, effectively reducing metallic artifacts without requiring additional bit transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a low-bit rate speech coder is used, then the bit rate is reduced, but metallic artifacts are introduced due to inadequate sparse signal representation
Solution Approach 1:
The patent introduces an artificial mixed signal as an intermediary component that bridges the gap between the sparse encoded signal and the desired full speech signal. This artificial signal, generated by mapping extracted features, acts as a mediator that fills in the missing speech structures without requiring additional transmission bits, thereby reducing metallic artifacts while maintaining low bit rate operation
Solution Approach 2:
The patent creates a copy of the speech signal in the form of an artificial mixed signal that is synthesized from extracted features. This copied representation captures the missing speech structures and when mixed with the decoded signal, reconstructs the absent speech components, effectively eliminating metallic artifacts without increasing the transmitted bit rate
2Manufacturing precision
If more information is added to the transmitted signal to capture speech structure, then the quality improves, but the bit rate increases
Solution Approach 1:
The patent extracts essential speech features from the encoded or decoded signal and uses only these extracted features to generate the artificial mixed signal. By taking out only the necessary feature information rather than transmitting the complete speech signal, the system captures speech structures effectively while maintaining low bit rate operation
Solution Approach 2:
Instead of transmitting additional speech information, the patent creates a synthetic copy of the missing speech structures by mapping extracted features to an artificial mixed signal. This copied representation provides the necessary speech structure information without increasing the transmitted bit rate
3Reliability
If speech frames are stretched or concealment frames are inserted to compensate for jitter, then transmission issues are handled, but metallic artifacts are introduced
Solution Approach 1:
The artificial mixed signal serves as an intermediary that masks the artifacts introduced by jitter compensation operations. By mixing this synthesized signal with the decoded signal, the system covers up the metallic artifacts that arise from frame stretching and concealment operations, maintaining reliability while reducing harmful artifacts
Data Source
AI summary
A system for enhancing a signal regenerated from an encoded audio signal. The system comprises a decoder arranged to receive the encoded audio signal and produce a decoded audio signal, a feature extraction means arranged to receive at least one of the decoded and encoded audio signal and extract at least one feature from at least one of the decoded and encoded audio signal, a mapping means arranged to map the at least one feature to an enhancement signal and operable to generate and output the enhancement signal, whereby the enhancement signal has a frequency band that is within the decoded audio signal frequency band, and a mixing means arranged to receive the decoded audio signal and the enhancement signal and mix the enhancement signal with the decoded audio signal.


