Sound Source Mapping for Selective Voice Separation in Hearables
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hearing aids and hearables have limited control over which voices are amplified, making it difficult for users to selectively hear desired sounds and reduce unwanted noise, leading to increased cognitive load and reduced user experience in noisy environments.
Innovation Solution
A method and system that utilizes acoustic fingerprints and speech separation techniques to identify and separate voices in a noisy environment, allowing users to control which sounds are amplified or muted through a user interface, and provides a map view to manage sound sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If hearing aids and hearables amplify all sounds in a noisy environment, then the user can hear all voices, but the cognitive load increases and unwanted noise is not reduced
Solution Approach 1:
The patent segments the audio environment by identifying and separating individual sound sources (voices) from the noisy background. The system divides the continuous audio stream into discrete voice segments that can be independently controlled, allowing users to select which voices to hear clearly while filtering out unwanted noise and conversations.
Solution Approach 2:
The patent applies different processing qualities to different sound sources. Desired voices receive enhanced processing (speech separation, noise reduction, amplification) while unwanted sounds receive different treatment. This local quality enhancement allows clear hearing of selected voices without unnecessarily processing all sounds, reducing cognitive load.
2Adaptability or versatility
If hearing aids and hearables provide no control over voice selection, then the device complexity is low, but the user cannot selectively hear desired sounds
Solution Approach 1:
The patent introduces a map view interface as an intermediary between the user and the audio processing system. This visual map displays detected sound sources and their locations, allowing users to easily select which voices to hear by interacting with the map. The intermediary interface simplifies the complexity of voice selection by providing an intuitive visual representation and control mechanism.
Solution Approach 2:
The patent replaces traditional mechanical or simple electronic volume controls with an intelligent software-based system that uses acoustic fingerprinting, speech separation algorithms, and directional audio processing. This substitution enables sophisticated voice selection and noise reduction capabilities while maintaining user-friendly control through the map interface.
3Measurement precision
If the system processes all audio signals equally, then the processing is simple, but the audibility of desired sounds is reduced by unwanted noise
Solution Approach 1:
The patent performs preliminary actions by first identifying and fingerprinting different voice sources before the user needs to hear them. The system pre-processes the audio environment to detect, classify, and tag different sound sources, creating a structured representation of the acoustic environment. This preliminary identification enables rapid and accurate voice selection when the user interacts with the map view, without requiring complex real-time processing during the selection moment.
Data Source
AI summary
A method, system and product includes displaying to a user, via a mobile device, a map view depicting locations of at least a portion of a plurality of people relative to a location of the user, wherein the user and the plurality of people are located in an environment, the user having the mobile device used for obtaining user input, the user having a hearable device used for providing audio output to the user; receiving, via the mobile device, an activation selection of a target person from the map view; capturing a noisy audio signal from the environment; processing the noisy audio signal by applying speech separation on the target person, whereby generating an enhanced audio signal; and outputting to the user via the hearable device the enhanced audio signal.


