Audio Object Extraction via Cross-Correlation Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting audio objects from multiple microphones positioned at different distances suffer from reduced reliability and increased latency due to the need for complex training of neural networks and independent correlation calculations, which often amplify noise and disrupt the extraction process, especially in dynamic environments.
Innovation Solution
A method utilizing a trained neural network with two operators for transforming and synchronizing audio input signals, followed by iterative optimization and compensation for acoustic effects, enhances signal separation quality and reduces latency by up to 40 ms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If neural networks are trained for all possible microphone distances to synchronize audio signals, then synchronization accuracy is improved, but training complexity and time consumption increase significantly
Solution Approach 1:
The patent changes the approach from training neural networks for all possible distances to using a single trained operator that works with analytical cross-correlation. The parameter being changed is the training scope - instead of training for multiple distance scenarios, the system trains once and then uses analytical calculation adapted to any distance configuration, thereby reducing training complexity while maintaining synchronization accuracy
Solution Approach 2:
The patent replaces the neural network-based synchronization mechanism with an analytical cross-correlation calculation. This substitution eliminates the need for extensive neural network training while achieving the same synchronization function, thereby reducing training complexity and time consumption while maintaining or improving synchronization accuracy
2Productivity
If analytical cross-correlation is used to synchronize audio input signals, then processing speed is improved, but extraction reliability deteriorates due to noise amplification
Solution Approach 1:
The patent introduces a trained operator as an intermediary between the analytical cross-correlation calculation and the audio extraction process. This trained operator processes the cross-correlation results to suppress noise amplification while preserving the speed advantage of analytical calculation, thereby maintaining processing speed while improving extraction reliability
Solution Approach 2:
The system uses a trained operator that incorporates feedback mechanisms to adjust the processing of cross-correlation results. By analyzing the correlation output and applying learned transformations, the system can suppress noise artifacts while maintaining accurate synchronization, thus improving extraction reliability without sacrificing processing speed
3Area of stationary object
If microphones are positioned at different distances from the audio object, then spatial coverage is improved, but temporal alignment becomes more complex and extraction reliability decreases
Solution Approach 1:
The patent applies preliminary synchronization using a trained operator that calculates cross-correlation and applies appropriate time shifts before the extraction process. By pre-aligning the audio signals from microphones at different distances, the system maintains spatial coverage while eliminating temporal misalignment issues that would otherwise degrade extraction reliability
Data Source
Figure 1~2
Figure 3~5
Figure 6
AI summary
The invention relates to a method for extracting at least one audio object from at least two audio input signals each containing the audio object. According to the invention, the following steps are provided: synchronising the second audio input signal with the first audio input signal obtaining a synchronised second audio input signal; extracting the audio object by applying at least one trained model to the first audio signal and to the synchronised second audio input signal and outputting the audio object. The invention further provides that the method step of synchronising the second audio input signal with the first audio input signal comprises the following method steps: generating audio signals; analytically calculating a correlation between the audio signals; optimising the correlation vector; and determining the synchronised second audio input signal with the aid of the optimised correlation vector. The invention also provides a system having a control unit which is designed to perform the method according to the invention. A computer program containing program code means is also provided, the program being designed to perform the steps of the method according to the invention.