Audio Signal Directional Processing via Time-Frequency Factorization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies face challenges in accurately separating a signal of interest from multiple sources using closely spaced microphones, particularly in noisy environments, due to limitations in computation capacity at user devices and inefficiencies in data transmission for further processing.
Innovation Solution
A method involving time-frequency directional processing using non-negative matrix or tensor factorization, where user devices compute direction estimates and spectral characteristics, forming a data structure for server-based processing, reducing data transmission and leveraging greater computational resources for enhanced signal separation and speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If beamforming is used to separate signals from multiple sources, then directional sensitivity is improved, but microphone separation distance must be large (worsening compactness)
Solution Approach 1:
The patent transitions from spatial domain beamforming (requiring physical microphone separation) to time-frequency domain processing. By applying short-time Fourier transform and analyzing signals in the time-frequency plane, the system achieves source separation without relying on large physical distances between microphones, effectively moving the separation mechanism to a different dimensional space.
Solution Approach 2:
The patent replaces the mechanical beamforming approach (which requires specific physical arrangements of microphones) with a computational signal processing approach. Instead of using physical microphone geometry to achieve directional sensitivity, the system uses algorithms operating on time-frequency representations of the signals, substituting mechanical spatial filtering with computational processing.
2Speed
If all signal processing is performed at the user device, then processing speed is improved, but computation capacity is exceeded
Solution Approach 1:
The patent divides the signal processing task into two segments: initial processing at the user device and further processing at the server. The user device performs time-frequency analysis and extracts relevant features, then transmits these processed features to the server for additional analysis and source separation. This segmentation allows computationally intensive operations to be distributed, reducing the burden on the user device while maintaining processing efficiency.
Solution Approach 2:
The user device performs preliminary signal processing (time-frequency transformation and feature extraction) before transmitting data to the server. This preliminary action reduces the amount of raw data that needs to be processed centrally, allowing the server to focus on more complex analysis tasks with already-preprocessed input, thereby improving overall processing efficiency and reducing communication bandwidth requirements.
3Measurement precision
If all acquired data is transmitted to the server, then processing accuracy is improved, but data transmission volume increases
Solution Approach 1:
The patent extracts only the essential time-frequency features and directional information from the raw acoustic signals at the user device, rather than transmitting all raw data. By identifying and extracting only the relevant characteristics needed for source separation, the system reduces transmission volume while preserving the information necessary for accurate server-side processing.
Solution Approach 2:
The patent transforms the raw signal data into a different parameter space (time-frequency domain with directional characteristics). This parameter transformation compresses the information representation, converting large volumes of raw time-domain data into a more compact time-frequency representation that contains the essential information for source separation but occupies less transmission bandwidth.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
An approach to processing of acoustic signals acquired at a user's device include one or both of acquisition of parallel signals from a set of closely spaced microphones, and use of a multi-tier computing approach in which some processing is performed at the user's device and further processing is performed at one or more server computers in communication with the user's device. The acquired signals are processed using time versus frequency estimates of both energy content as well as direction of arrival. In some examples, a non-negative matrix or tensor factorization approach is used to identify multiple sources each associated with a corresponding direction of arrival of a signal from that source. In some examples, data characterizing direction of arrival information is passed from the user's device to a server computer where direction-based processing is performed.