Audio Simplification Unit for IVAS Codec Format Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice and video codecs face complexity and cost issues in supporting a wide range of audio capture and rendering formats, making it impractical for Immersive Voice and Audio Services (IVAS) codecs to address all formats directly, leading to increased complexity and expense.
Innovation Solution
A simplification unit within audio devices converts captured audio signals into a limited number of formats, such as mono, stereo, and spatial, that can be processed by an IVAS codec, reducing complexity and enabling deployment on various devices regardless of their capture capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If an IVAS codec supports a wide range of audio capture and rendering formats directly, then adaptability and versatility are improved, but device complexity increases
Solution Approach 1:
The patent introduces a simplification unit as an intermediary component that sits between the audio capture device and the IVAS codec. This simplification unit converts various audio formats (mono, stereo, spatial, immersive) into a standardized internal representation that the codec can process uniformly. By using this intermediary, the codec doesn't need to directly support all audio formats, thereby reducing codec complexity while maintaining adaptability across different capture devices
Solution Approach 2:
The patent segments the audio processing system into distinct functional components: capture devices that generate audio in various formats, a simplification unit that standardizes these formats, and the IVAS codec that processes the standardized format. This segmentation allows each component to specialize in specific tasks, with the simplification unit handling format conversion and the codec handling encoding/decoding, thereby reducing the complexity burden on the codec itself
2Ease of manufacture
If an IVAS codec is designed to be simple and cost-effective, then ease of manufacture is improved, but adaptability to different audio capture formats deteriorates
Solution Approach 1:
The simplification unit serves as a format conversion intermediary that enables simple IVAS codecs to work with complex audio capture formats. By converting all input formats into a unified internal representation before the codec processes them, the system allows the codec to remain simple and cost-effective while still being adaptable to various audio capture devices through the format conversion layer
3Device complexity
If all audio capture formats are converted to a limited number of formats, then device complexity is reduced, but processing time for format conversion increases
Solution Approach 1:
The simplification unit performs format conversion as a preliminary action before the audio data reaches the IVAS codec. By pre-converting all audio formats into the standardized internal representation at the point of capture, the system avoids the need for time-consuming format conversion during the encoding/decoding process itself, thereby reducing overall processing time while maintaining reduced codec complexity
Data Source
AI summary
The disclosed embodiments enable converting audio signals captured in various formats by various capture devices into a limited number of formats that can be processed by an audio codec (e.g., an Immersive Voice and Audio Services (IVAS) codec). In an embodiment, a simplification unit of the audio device receives an audio signal captured by one or more audio capture devices coupled to the audio device. The simplification unit determines whether the audio signal is in a format that is supported/not supported by an encoding unit of the audio device. Based on the determining, the simplification unit, converts the audio signal into a format that is supported by the encoding unit. In an embodiment, if the simplification unit determines that the audio signal is in a spatial format, the simplification unit can convert the audio signal into a spatial “mezzanine” format supported by the encoding.


