Audio Signal Processing Apparatus Dynamic Coding Mode Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies fail to provide consistent performance when encoding and decoding signals that combine speech and music, as they are optimized for individual types of signals rather than mixed audio signals.
Innovation Solution
An audio signal processing apparatus that classifies input signals into speech, music, or mixed signals and selects an appropriate coding scheme using a signal classifying unit, employing linear prediction modeling for speech signals, psychoacoustic modeling for music signals, and a combined approach for mixed signals, including linear prediction and frequency transformation to optimize coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If perceptual audio coding is used for music signals, then coding efficiency for music is improved, but performance on speech and mixed signals deteriorates
Solution Approach 1:
The patent implements dynamic switching between different coding modes (speech coding, music coding, and mixed coding) based on the detected characteristics of the input audio signal. The coding apparatus adapts its processing method in real-time according to the signal type, thereby achieving high coding efficiency for each specific signal type while maintaining versatility across different audio contents.
Solution Approach 2:
The patent changes coding parameters such as the use of linear prediction coefficients, quantization methods, and synthesis filter settings according to the detected audio signal type. By adjusting these parameters dynamically, the system optimizes coding performance for speech, music, and mixed signals separately, resolving the contradiction between specialized efficiency and general adaptability.
2Productivity
If linear prediction based coding is used for speech signals, then coding efficiency for speech is improved, but performance on music and mixed signals deteriorates
Solution Approach 1:
The system dynamically selects between linear prediction based coding and perceptual coding based on the input signal characteristics. For speech signals, linear prediction coding is activated to achieve high efficiency, while for music and mixed signals, the system switches to appropriate alternative coding methods, thus maintaining both specialized performance and general adaptability.
Solution Approach 2:
The patent modifies coding parameters including the prediction order, analysis window size, and synthesis filter characteristics according to the detected signal type. These parameter adjustments enable the system to optimize speech coding performance when speech is detected while maintaining acceptable performance on other signal types through appropriate mode switching.
3Device complexity
If a single coding scheme is used for all audio signals, then device complexity is reduced, but coding efficiency for specific signal types deteriorates
Solution Approach 1:
The patent segments the audio signal processing into distinct coding paths for speech, music, and mixed signals. Each segment uses optimized coding parameters and methods tailored to its specific signal type. This segmentation allows the system to maintain relatively simple individual coding schemes while achieving high overall coding efficiency through appropriate selection and combination of different coding approaches.
Data Source
Figure 1
Figure 2
Figure 3(a)~4
AI summary
An apparatus for processing an encoded signal and method thereof are disclosed, by which an audio signal can be compressed and reconstructed in higher efficiency. An audio signal processing method includes the steps of identifying whether a type of an audio signal is a music using first type information, if the type of the audio signal is not the music signal, identifying whether the type of the audio signal is a speech signal or a mixed signal using second type information, and if the type of the audio signal is determined as either the speech signal or the mixed signal, reconstructing the audio signal according to a coding scheme applied per frame using coding identification information. If the type of the audio signal is the music signal, the first type information is received only. If the type of the audio signal is the speech signal or the mixed signal, both of the first type information and the second type information are received. Accordingly, various kinds of audio signals can be encoded/decoded in higher efficiency.