Wearable Audio Voice Control System for Natural Self-Voice Perception

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Legacy wearable audio devices cause users to perceive their voice differently than expected, leading to an unnatural experience due to the mixture of sound emissions and vibrations, resulting in undesirable sound perception.

Innovation Solution

The implementation of a User Voice Control (UVC) system using sensors and algorithms to detect the user's voice, allowing for real-time adjustments to the audio stream, such as removing the user's voice or amplifying it, to mitigate this difference, and incorporating a Self-Voice Activity Detect (SVAD) function to enhance voice clarity in noisy environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the wearable audio device plays back the user's voice through the speaker, then the user can hear their own voice, but the user perceives the voice as unnatural and different from expected

Engineering Contradiction:
Improvevoice playback functionalityVSAvoidvoice perception accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent extracts the user's voice from the audio stream using voice activity detection and separation algorithms. By identifying and isolating the user's voice components from the mixed audio signal, the system can selectively remove or modify only those portions while preserving other sounds, thereby eliminating the unnatural self-voice perception without affecting overall audio playback functionality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary processing stage between the audio capture and playback components. This intermediary layer includes voice detection modules, voice separation algorithms, and mixing controls that mediate between the raw audio signal and the final output, enabling natural voice perception by filtering out the problematic self-voice components before playback.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system removes the user's voice from the audio stream, then the user perceives their voice more naturally, but the user may lose the ability to hear their own voice in certain contexts

Engineering Contradiction:
Improvevoice perception accuracyVSAvoidvoice hearing flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic voice removal control that adapts to different usage scenarios. The system can adjust the degree of voice removal in real-time based on detected conditions such as phone call mode, audio recording mode, or general audio playback mode. This dynamic adjustment allows the system to maintain versatility by enabling voice removal only when necessary while preserving voice hearing capabilities in other contexts.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent utilizes parameter changes to control the voice removal effect. By adjusting parameters such as voice removal intensity, frequency bands to be modified, and temporal characteristics, the system can create different playback experiences. This allows the same hardware to adapt to multiple use cases - from complete voice removal for natural perception to partial modification for hearing assistance, thereby maintaining versatility.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If the system processes the audio stream in real-time to remove or amplify voice, then the user experiences improved voice clarity, but the device complexity increases

Engineering Contradiction:
Improvevoice clarityVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the audio processing task into distinct functional modules: voice activity detection module, voice separation module, voice removal module, and audio mixing module. Each module handles a specific aspect of the processing independently, which simplifies the overall system architecture compared to a monolithic approach. This segmentation allows for more manageable complexity while achieving real-time voice clarity improvement.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent designs the audio processing system with multi-functionality, where the same processing pipeline can perform multiple operations including voice removal, voice amplification, and noise reduction depending on the active mode. This universal approach reduces the need for separate dedicated hardware for each function, thereby controlling device complexity while maintaining versatile voice clarity enhancement capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If the system uses sensors and algorithms to detect user voice, then the user benefits from enhanced voice clarity in noisy environments, but the device complexity and energy consumption increase

Engineering Contradiction:
Improvevoice detection accuracyVSAvoidenergy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent implements periodic voice activity detection rather than continuous full-processing mode. The system periodically samples the audio stream to detect voice presence, and only activates the full voice separation and removal algorithms when voice activity is detected. This periodic action significantly reduces energy consumption during idle periods while maintaining high voice detection accuracy when needed, effectively resolving the contradiction between detection reliability and energy usage.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS11557307B2User voice control system
Publication Date: 2023.01.17 PERSON-AIZ AS
  • US11557307B2 patent drawing
  • US11557307B2 patent drawing
  • US11557307B2 patent drawing

AI summary

Embodiments include techniques and objects related to a wearable audio device that includes a microphone to detect a plurality of sounds in an environment in which the wearable audio device is located. The wearable audio device further includes a non-acoustic sensor to detect that a user of the wearable audio device is speaking. The wearable audio device further includes one or more processors communicatively to alter, based on an identification by the non-acoustic sensor that the user of the wearable audio device is speaking, one or more of the plurality of sounds to generate a sound output. Other embodiments may be described or claimed.