Unified Speech Audio Codec Module Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech and audio codecs fail to provide optimal performance when integrated, as they are optimized for specific signal characteristics, leading to suboptimal processing of unified speech/audio signals in communication and broadcasting services.
Innovation Solution
A unified codec apparatus and method that selects between speech and audio encoding/decoding modules based on input signal characteristics, using a combination of Code Excitation Linear Prediction (CELP) and Modified Discrete Cosine Transform (MDCT) operations, with initialization and multiplexing to generate and decode bitstreams without distortion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a unified codec uses a single encoding/decoding module optimized for speech signals, then speech processing performance is improved, but audio signal processing performance deteriorates
Solution Approach 1:
The codec dynamically switches between speech-optimized module and audio-optimized module based on the characteristics of the input signal. The system determines whether each input signal is a speech signal or an audio signal and selects the appropriate module accordingly, enabling adaptive optimization for different signal types rather than using a static single-module design
Solution Approach 2:
The unified codec achieves multi-functionality by integrating both speech-optimized and audio-optimized encoding/decoding modules within a single system. This allows the codec to handle both speech signals and audio signals effectively, providing universal processing capability across different signal types while maintaining optimal performance for each
2Manufacturing precision
If a unified codec uses a single encoding/decoding module optimized for audio signals, then audio processing performance is improved, but speech signal processing performance deteriorates
Solution Approach 1:
The codec dynamically switches between audio-optimized module and speech-optimized module based on the characteristics of the input signal. The system determines whether each input signal is an audio signal or a speech signal and selects the appropriate module accordingly, enabling adaptive optimization for different signal types rather than using a static single-module design
Solution Approach 2:
The unified codec achieves multi-functionality by integrating both audio-optimized and speech-optimized encoding/decoding modules within a single system. This allows the codec to handle both audio signals and speech signals effectively, providing universal processing capability across different signal types while maintaining optimal performance for each
3Adaptability or versatility
If the codec switches between different encoding/decoding modules for each frame, then signal processing adaptability is improved, but signal distortion increases due to module structure changes
Solution Approach 1:
The codec performs preliminary determination of whether the current frame is a speech frame or an audio frame before selecting the encoding/decoding module. This advance identification allows the system to pre-select the appropriate module structure, avoiding mid-processing switches that would cause distortion and ensuring continuous processing with a single consistent module structure for each frame
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided is an apparatus for integrally encoding and decoding a speech signal and an audio signal. An encoding apparatus for integrally encoding a speech signal and an audio signal, may include: a module selection unit to analyze a characteristic of an input signal and to select a first encoding module for encoding a first frame of the input signal; a speech encoding unit to encode the input signal according to a selection of the module selection unit and to generate a speech bitstream; an audio encoding unit to encode the input signal according to the selection of the module selection unit and to generate an audio bitstream; and a bitstream generation unit to generate an output bitstream from the speech encoding unit or the audio encoding unit according to the selection of the module selection unit.