Keyword-Based Audio Localization with Beamforming for Noisy Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Devices struggle to distinguish between a person speaking and other ambient sounds, leading to difficulty in understanding voice commands due to background noise or multiple speakers.

Innovation Solution

Implementing multiple listening zones using beamformers to detect a keyword, determine the direction and distance of the speaker, and form active acoustic beams to enhance speech recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the device listens for keywords from multiple directions using multiple listening zones, then the device can detect keywords spoken by anyone in the environment, but the device cannot distinguish between the speaker and background noise

Engineering Contradiction:
Improvekeyword detection coverageVSAvoidspeaker identification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent divides the listening environment into multiple acoustic zones (first acoustic zone, second acoustic zone, etc.) with each zone having associated beamformers that point in specific directions. This segmentation allows the system to simultaneously monitor multiple directions for keywords while also being able to identify which zone the keyword was spoken in, enabling both broad detection coverage and directional speaker identification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a spatial dimension by using beamformers that create directional acoustic zones in three-dimensional space. Each beamformer points in a specific direction and creates an acoustic zone that can be distinguished from other zones. This adds the dimension of spatial location to the keyword detection process, allowing the system to not only detect keywords from multiple directions but also to identify the specific direction from which the keyword was spoken.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If the device uses multiple beamformers pointing in various directions to detect keywords, then the device can identify the direction of the speaker, but the device complexity increases

Engineering Contradiction:
Improvespeaker location accuracyVSAvoidbeamformer configuration
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the beamformers multi-functional by using them for both keyword detection and speaker location identification. The same beamformers that detect keywords in specific acoustic zones also provide the spatial information needed to determine speaker direction. This eliminates the need for separate systems for detection and localization, reducing overall device complexity while maintaining precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent performs preliminary action by pre-configuring multiple beamformers to point in various directions before keyword detection begins. These beamformers are set up in advance to create overlapping acoustic zones that cover the entire listening environment. This preliminary configuration allows the system to immediately begin both keyword detection and directional identification without additional processing complexity during operation.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the device forms active acoustic beams directed toward the speaker after keyword detection, then the device can enhance speech recognition, but the device must switch between different listening modes

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidlistening mode switching
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements dynamic switching between different listening modes: a first listening mode for keyword detection using multiple acoustic zones, and a second listening mode for enhanced speech recognition using active acoustic beams directed at the identified speaker location. The system dynamically transitions between these modes based on whether keyword detection is needed or enhanced speech recognition is required, allowing the device to adapt its behavior to different operational contexts.

Inventive Principle:
Principle #15Dynamics

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances speech recognition by focusing on the speaker's voice, reducing background noise interference and improving command understanding.

Implementation Method 1

the device may implement multiple listening zones, such as using one or more beamformers pointing in various directions around a horizontal plane and/or a vertical plane

Methodology Applied
Scientific EffectAcoustic beamforming:

Implementation Method 2

form one or more active acoustic beams directed toward the person speaking

Methodology Applied
Scientific EffectAcoustic beam focusing: Focusing

Data Source

PatentEP3816993B1Keyword-based audio source localization
Publication Date: 2025.07.16 COMCAST CABLE COMM LLC
  • EP3816993B1 patent drawingFigure 1
  • EP3816993B1 patent drawingFigure 2
  • EP3816993B1 patent drawingFigure 3

AI summary

Systems, apparatuses, and methods are described for determining a direction associated with a detected spoken keyword, forming an acoustic beam in the determined direction, and listening for subsequent speech using the acoustic beam in the determined direction.