Adaptive Microphone Beamforming for Speaker Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing microphone beamforming technologies face challenges in accurately distinguishing a speaker's speech from ambient noise due to the speaker's continuous movement, making it difficult to adaptively change the beamforming direction effectively.

Innovation Solution

A method and apparatus that recognize a speaker's speech, search for a corresponding image using a camera, and adjust microphone beamforming based on the speaker's position, amplifying speech from the speaker's location while reducing noise from other areas.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional microphone beamforming is used with a fixed direction, then the device complexity is low, but the adaptability to speaker movement is poor

Engineering Contradiction:
Improveadaptability to speaker positionVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines multiple functions into an integrated system: speech recognition unit identifies speaker identity, image searching unit retrieves speaker images, camera searches for speaker position, and beamforming unit adjusts microphone directions. This merging of units enables adaptive beamforming while managing complexity through functional integration.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The beamforming directions are made dynamic by continuously tracking speaker position through camera-based image processing. The system adaptively changes beamforming directions based on real-time speaker location, transforming a static system into a dynamic one that responds to speaker movement.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If the beamforming direction is continuously adjusted to track speaker movement, then the speech recognition accuracy is improved, but the loss of time for position tracking and beamforming adjustment increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtime for position tracking
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-recognizing speaker identity through speech analysis and pre-searching for speaker images before initiating the tracking process. This preparation reduces the time required for position tracking, as the system already has speaker identification and reference images ready when the speaker moves into view.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback loops where the camera continuously monitors speaker position, the position recognizing unit processes this information, and the beamforming performing unit adjusts directions based on the feedback. This closed-loop feedback enables efficient tracking by only making adjustments when position changes are detected, reducing unnecessary processing time.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If the system uses multiple units for speech recognition, image searching, and position tracking, then the adaptability to speaker position is improved, but the device complexity increases

Engineering Contradiction:
Improveadaptive beamforming capabilityVSAvoidnumber of processing units
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system achieves multi-functionality through integrated units that perform multiple tasks: the speech recognizing unit identifies both speaker identity and speech content, the image searching unit retrieves and processes speaker images, the camera serves both as a position detector and a verification tool, and the beamforming unit adjusts directions while maintaining speech quality. This universality reduces the need for separate dedicated components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system segments functionality into distinct but coordinated units: speech recognition unit, image searching unit, speaker searching unit, position recognizing unit, and beamforming performing unit. This segmentation allows each unit to specialize in specific tasks while working together as an integrated system, managing complexity through modular functional division.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9330673B2Method and apparatus for performing microphone beamforming
Publication Date: 2016.05.03 SAMSUNG ELECTRONICS CO LTD
  • US9330673B2 patent drawing
  • US9330673B2 patent drawing
  • US9330673B2 patent drawing

AI summary

A method and apparatus for performing microphone beamforming. The method includes recognizing a speech of a speaker, searching for a previously stored image associated with the speaker, searching for the speaker through a camera based on the image, recognizing a position of the speaker, and performing microphone beamforming according to the position of the speaker.