Sound Separation on Personal Devices Using AR Glasses
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In environments with multiple sound sources, existing technologies face challenges in identifying and distinguishing individual sound sources from mixed sound streams, particularly in determining the direction and proximity of sounds like bird songs, vehicle sounds, and natural sounds such as wind and rain.
Innovation Solution
A method and system using augmented reality glasses with multiple microphones and machine learning algorithms for sound separation, employing time-frequency decomposition and temporal regularity techniques to classify and localize sound sources, allowing users to select and record specific sounds by displaying categorized icons on a user interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sound separation techniques are applied to identify individual sound sources from mixed sound streams, then sound localization accuracy and sound source identification are improved, but device complexity and processing requirements increase
Solution Approach 1:
The patent segments the mixed sound stream into individual sound sources by applying sound separation techniques that decompose the aggregate sound into distinct components. Each sound source is then independently analyzed and localized using binaural cues and spectral information, enabling precise identification of individual sounds like bird songs, vehicle sounds, and natural sounds from complex mixtures.
Solution Approach 2:
The patent introduces an intermediary processing layer that applies sound separation algorithms between the microphone input and final sound localization. This intermediary step extracts and separates individual sound sources from the mixed stream, making subsequent localization and identification tasks more accurate and manageable despite increased processing requirements.
2Reliability
If multiple microphones are used to capture sound streams from multiple sources, then sound source separation capability is improved, but device complexity increases
Solution Approach 1:
The patent combines multiple microphone inputs to capture sound streams from multiple sources simultaneously. By merging the signals from multiple microphones, the system creates a comprehensive audio capture that enables subsequent sound separation and identification of individual sources within the mixed sound stream.
Solution Approach 2:
The patent implements a multi-functional audio processing system where the same microphone array serves multiple purposes: capturing aggregate sound streams, separating individual sound sources, localizing sound directions, and identifying sound types. This universal approach maximizes the utility of the hardware while managing complexity through integrated processing.
3Measurement precision
If sound separation and classification algorithms are implemented, then sound identification accuracy is improved, but processing time and energy consumption increase
Solution Approach 1:
The patent applies preliminary sound separation and classification algorithms to the captured sound stream before final localization and identification. By pre-processing the audio data to separate and categorize sound sources in advance, the system reduces the computational burden during real-time processing and enables faster response times for sound identification and user interaction.
Data Source
AI summary
The method provides for one or more processor receiving on a personal device, a mixture of sounds within a sound stream from multiple sources. The one or more processors identifying one or more sounds of the mixture of sounds from the multiple sources, based on a sound separation technique. The one or more processors displaying on a user interface of the personal device an icon corresponding respectively to a classification of the one or more sounds identified from the multiple sources. The one or more processors receiving a selection of a sound from the mixture of the multiple sounds, based on an action by a user of the personal device selecting the icon displayed on the user interface of the personal device, and the one or more processors recording the sound from the mixture of the multiple sounds selected by the user.


