Wearable Audio Source Identification via Phase Shift Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional wearable electronic devices require speech training and the use of specific trigger phrases for voice recognition, limiting their ability to respond to multiple users and making them impractical for devices like smart glasses used by multiple wearers without prior training.

Innovation Solution

The implementation of a speech analysis manager that uses multiple audio sensors to determine phase shifts in audio signals, allowing the device to identify speech generated by the wearer without the need for training or trigger phrases, by comparing measured phase shifts to expected values and eliminating non-wearer generated audio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech training techniques are used to recognize user voice, then voice recognition accuracy is improved, but device complexity and ease of operation deteriorate due to requiring training procedures and trigger phrases

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidoperational simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs self-identification of the wearer by automatically analyzing audio phase shifts without requiring external training data or user-provided voice samples. The device serves itself by using its own audio sensors to detect and analyze the acoustic characteristics of speech originating from the wearer's position, eliminating the need for users to complete training procedures.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the conventional speech recognition system (which relies on trained voice models and trigger phrases) with a physics-based acoustic localization system. Instead of using machine learning models trained on user voice, the system uses phase shift analysis of audio waves captured by multiple sensors to geometrically determine the origin of speech, substituting a physical measurement approach for a data-driven recognition approach.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If speech training is required for voice recognition, then recognition reliability is improved, but adaptability deteriorates for multi-user devices like smart glasses

Engineering Contradiction:
Improvevoice recognition reliabilityVSAvoidmulti-user compatibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The phase shift-based speech origin detection system is universal and works for any wearer without requiring device-specific customization. The same hardware (multiple audio sensors) and algorithm (phase shift analysis) serve all users equally, making the device inherently adaptable to multiple users while maintaining consistent reliability across different wearers.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

Instead of having the system adapt to each user through training (user-specific approach), the patent inverts the approach by having the system identify which user is speaking based on physical acoustic properties (phase shift) that are independent of user identity. This reversal allows the device to work with any user immediately without requiring adaptation or training procedures.

Inventive Principle:
Principle #13The other way round (Inversion)

3Adaptability or versatility

If multiple audio sensors are used to determine phase shifts, then speech source identification capability is improved, but device complexity increases

Engineering Contradiction:
Improvespeech source identification capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent divides the audio detection function into multiple independent audio sensors positioned at different locations on the device. Each sensor independently captures audio signals, and the system processes each sensor's output separately through phase shift analysis. This segmentation enables spatial differentiation of speech sources while keeping each sensor's function simple and well-defined.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces phase shift analysis as an intermediary processing step between audio capture and speech recognition. Instead of directly comparing raw audio signals from multiple sensors (which would be complex), the system uses phase shift as an intermediate physical quantity that simplifies the determination of speech origin. This intermediary approach transforms a complex multi-sensor problem into a manageable calculation based on wave propagation physics.

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables effortless voice detection for any user without the need for training or trigger phrases, making the technology suitable for multi-user applications and enhancing the natural interaction with wearable electronic devices.

Implementation Method 1

determine a phase shift between a first sample of first audio captured at a first audio sensor and a second sample of the first audio captured at the second audio sensor

Methodology Applied
Scientific EffectPhase shift:

Data Source

PatentUS10522160B2Methods and apparatus to identify a source of speech captured at a wearable electronic device
Publication Date: 2019.12.31 INTEL CORP
  • US10522160B2 patent drawing
  • US10522160B2 patent drawing
  • US10522160B2 patent drawing

AI summary

Methods, systems and articles of manufacture for a wearable electronic device having an audio source identifier are disclosed. Example audio source identifiers disclosed herein include first and second audio sensors disposed at first and second locations, respectively, on a wearable electronic device. Such audio source identifiers also include a phase shift determiner to determine a phase shift between a first sample of first audio captured at the first audio sensor and a second sample of the first audio captured at the second audio sensor. The first audio includes first speech generated by a first speaker wearing the wearable electronic device. Example audio source identifiers further include a speaker identifier to determine, based on the phase shift determined by the phase shift determiner, whether second audio includes speech generated by a second speaker wearing the wearable electronic device.