Audio Stream Mixing via Spectral Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conferencing systems face challenges in achieving a balance between audio quality, bandwidth, and computational complexity, particularly when processing multiple audio signals, leading to increased delay and hardware requirements, and existing solutions do not effectively utilize advanced coding techniques like AAC-ELD to minimize quantization noise and reduce re-quantization steps.
Innovation Solution
The method involves determining an input data stream by comparing spectral information from multiple input streams and copying relevant information to the output stream, omitting re-quantization and using psycho-acoustic models to reduce computational complexity and noise, allowing direct processing in the frequency domain without transforming back to the time domain.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If the number of quantization levels is increased to reduce quantization noise and improve audio quality, then the quality of the reconstructed audio signal is improved, but the amount of data to be transmitted increases, violating bandwidth restrictions
Solution Approach 1:
The patent segments the audio signal processing into frequency domain operations, where spectral information is processed independently across different frequency bins. This allows selective quantization and mixing operations that reduce the overall data requirement while maintaining audio quality through frequency-based optimization rather than uniform time-domain processing.
Solution Approach 2:
The patent transforms the audio signal from the time domain to the frequency domain using Fast Fourier Transform (FFT). This dimensional transformation enables mixing operations to be performed on spectral components rather than time samples, reducing the number of quantization levels needed and consequently the data transmission requirements while maintaining or improving audio quality.
2Manufacturing precision
If modern coding techniques like AAC-ELD are implemented to improve audio quality, then the quality of reconstructed audio signal is improved, but the system complexity and processing overhead increase
Solution Approach 1:
The patent performs preliminary frequency domain transformation and spectral analysis before the actual mixing operation. By pre-processing the audio signals into the frequency domain and identifying dominant spectral components in advance, the system can simplify subsequent mixing operations and reduce the computational complexity required for real-time processing.
Solution Approach 2:
The patent extracts and identifies the dominant input data stream based on spectral comparison, then copies only the relevant spectral information from that stream to the output. This extraction approach eliminates the need to process all input streams equally, reducing computational complexity while maintaining audio quality through selective copying of dominant spectral components.
3Adaptability or versatility
If multiple input audio signals are processed to create a mixed output signal, then the functionality and versatility of the system is improved, but the computational complexity and hardware requirements increase
Solution Approach 1:
The patent merges multiple input audio signals into a single output signal by operating on their frequency domain representations. Spectral information from multiple inputs is combined through comparison and selection operations in the frequency domain, then transformed back to the time domain. This merging approach maintains versatility for processing multiple signals while reducing hardware requirements through efficient spectral processing.
Solution Approach 2:
The patent copies spectral information from the dominant input data stream to the output stream rather than performing complex mixing operations on all inputs. By identifying which input stream has the dominant spectral content and copying its relevant frequency components, the system achieves versatile mixing functionality with reduced computational complexity and lower hardware requirements.
4Ease of manufacture
If audio signals are mixed in the time domain by superimposing signals, then the mixing process is simple, but the delay introduced by processing outside the time-domain increases
Solution Approach 1:
The patent moves the mixing operation from the time domain to the frequency domain by applying Fast Fourier Transform. This allows mixing to be performed on spectral components simultaneously across all frequency bins, dramatically reducing processing delay while maintaining simplicity. The frequency domain operations can be performed in parallel and require minimal computational overhead compared to time-domain signal processing.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An apparatus (500) for mixing a plurality of input data streams (510) is described, wherein the input data streams (510) each comprise a frame (540) of audio data in the spectral domain, a frame (540) of an input data stream (510) comprising spectral information for a plurality of spectral components. The apparatus comprises a processing unit (520) adapted to compare the frames (540) of the plurality of input data streams (510). The processing unit (520) is further adapted to determine, based on the comparison, for a spectral component of an output frame (550) of an output data stream (530), exactly one input data stream (510) of the plurality of input data streams (510). The processing unit (520) is further adapted to generate the output data stream (530) by copying at least a part of an information of a corresponding spectral component of the frame of the determined data stream (510) to describe the spectral component of the output frame (550) of the output data stream (530). Further or alternatively, the control value of the frames (540) of the first input data stream (510-1) and the second input data stream (510-2) may be compared to yield a comparison result and, if the comparison result is positive, the output data stream (530) comprising an output frame(550) may be generated such that the output frame (550) comprises a control value equal to that of the first and second input data streams (510) and payload data derived from the payload data of the frames of the first and second input data streams by processing the audio data in the spectral domain.