Time-Frequency Audio Processing for Closely Spaced Microphones
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies face challenges in accurately separating signals from multiple sources using closely spaced microphones, particularly in noisy environments, due to limitations in beamforming and computation capacity of user devices.
Innovation Solution
A multi-tier computing approach is employed, where user devices compute time-dependent spectral characteristics and direction estimates, and a non-negative matrix or tensor factorization is used to identify sources, with further processing performed on a server to selectively process signals from specific sources, reducing data transmission and leveraging greater computational resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If beamforming is used to separate signals from multiple sources, then directional sensitivity is improved, but microphone separation distance must be large (worsening compactness)
Solution Approach 1:
The patent transitions from spatial domain beamforming (requiring physical microphone separation) to time-frequency domain processing. By transforming the signal separation problem into the time-frequency domain using STFT and spectral factorization, the system achieves directional sensitivity without relying on large physical microphone spacing, thus resolving the contradiction between directional sensitivity and compact microphone array.
Solution Approach 2:
The patent replaces the mechanical beamforming approach (which requires physically spaced microphones) with a computational signal processing approach. Instead of using physical microphone separation to achieve directional sensitivity, the system uses time-frequency spectral analysis and matrix factorization to separate sources computationally, eliminating the need for large microphone separation distances.
2Speed
If complete signal processing is performed at the user device, then processing speed is improved, but device complexity and power consumption increase
Solution Approach 1:
The patent divides the signal processing task into two segments: time-frequency transformation and spectral characteristic computation are performed at the user device for speed, while the computationally intensive non-negative matrix factorization and source separation are performed at the server. This segmentation allows the user device to maintain low complexity while achieving fast initial processing.
Solution Approach 2:
The patent performs preliminary time-frequency transformation and spectral analysis at the user device before transmitting data to the server. This preliminary processing reduces the amount of raw data that needs to be transmitted and prepared the signal in advance for the more complex server-side factorization operations, optimizing the overall processing workflow.
3Measurement precision
If all processed data is transmitted to the server, then processing accuracy is improved, but data transmission requirements increase
Solution Approach 1:
The patent extracts only the essential time-frequency spectral characteristics and direction estimates from the raw audio signal at the user device, rather than transmitting the complete processed dataset. By extracting only the necessary features needed for server-side factorization, the system maintains source separation accuracy while significantly reducing data transmission volume.
4Volume of moving object
If closely spaced microphones are used, then device compactness is improved, but signal separation performance deteriorates
Solution Approach 1:
The patent replaces the mechanical solution of widely spaced microphones with a computational signal processing approach. By using time-frequency spectral analysis and non-negative matrix factorization, the system achieves effective signal separation using closely spaced microphones, thus maintaining device compactness while restoring signal separation performance through computational methods rather than physical spacing.
Data Source
AI summary
An approach to processing of acoustic signals acquired at a user's device include one or both of acquisition of parallel signals from a set of closely spaced microphones, and use of a multi-tier computing approach in which some processing is performed at the user's device and further processing is performed at one or more server computers in communication with the user's device. The acquired signals are processed using time versus frequency estimates of both energy content as well as direction of arrival. In some examples, a non-negative matrix or tensor factorization approach is used to identify multiple sources each associated with a corresponding direction of arrival of a signal from that source. In some examples, data characterizing direction of arrival information is passed from the user's device to a server computer where direction-based processing is performed.


