Multi-Microphone Beamforming for Moving-User Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition performance deteriorates when a user is moving due to fixed beamforming directions, necessitating a method to adapt the electronic device's location and direction based on the user's location.

Innovation Solution

An electronic device equipped with multiple microphones that identify a designated call word, perform speech recognition, and adjust its location using phase differences and beamforming to enhance target voice signals while reducing noise and echo.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If beamforming is used to track user location, then speech recognition accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the speech recognition process into two stages: first identifying call words using a single microphone to reduce complexity, then performing full speech recognition using multiple microphones with beamforming only when necessary. This segmentation allows the system to achieve high accuracy when needed while maintaining low complexity during normal operation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts the number of microphones active in beamforming based on whether call words have been identified. When call words are detected, the system transitions from single-microphone mode to multi-microphone beamforming mode, adapting the device complexity to the current operational requirements and improving accuracy at the appropriate moment.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If multiple microphones are used for speech recognition, then speech recognition accuracy is improved, but computational load increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcomputational load
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent segments the speech recognition task into call word identification (using one microphone, low computational load) and continuous word recognition (using multiple microphones, high computational load). By separating these functions, the system performs heavy computations only when necessary, reducing overall power consumption while maintaining high accuracy when needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary call word identification using a single microphone before engaging the full multi-microphone beamforming system. This preliminary action filters out unnecessary computations by identifying trigger words first, then activating the computationally intensive speech recognition only for relevant segments, thereby reducing total computational load.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If beamforming direction is fixed, then device complexity is reduced, but speech recognition performance deteriorates when user moves

Engineering Contradiction:
Improvedevice complexityVSAvoidspeech recognition performance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements dynamic beamforming that adjusts its direction and orientation based on detected user location and call word identification. The system transitions from a fixed beamforming configuration to an adaptive one that tracks user movement, improving reliability while managing complexity through conditional activation only when call words are detected.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from call word identification and user location detection to dynamically adjust beamforming parameters. By continuously monitoring user position and call word detection results, the system modifies its beamforming direction and orientation in real-time, ensuring optimal performance while maintaining manageable complexity through feedback-driven adaptation.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Improves speech recognition accuracy by selectively processing voice signals from a single microphone, reducing computational load, and accurately tracking the user's location for enhanced performance.

Implementation Method 1

track a location of a user associated with the plurality of voice signals by using a phase difference between the plurality of voice signals

Methodology Applied
Scientific EffectPhase difference:

Implementation Method 2

identify a second voice signal, which is obtained by enhancing a target voice corresponding to an utterance of the user, from the plurality of voice signals obtained by performing beamforming toward the location of the user

Methodology Applied
Scientific EffectBeamforming:

Data Source

PatentUS20250273229A1Electronic device and method thereof
Publication Date: 2025.08.28 HYUNDAI MOTOR CO LTD
  • US20250273229A1 patent drawing
  • US20250273229A1 patent drawing
  • US20250273229A1 patent drawing

AI summary

An electronic device may include a plurality of microphones, a speaker, a processor, and a memory. The processor may be configured to obtain a plurality of voice signals by using the plurality of microphones, to identify a designated call word, by performing speech recognition on a first voice signal obtained from a first microphone, from among the plurality of voice signals according to a passage of time in obtaining the plurality of voice signals, to obtain a second time point, which precedes an utterance time corresponding to the designated call word from a first time point at which the designated call word is identified completely, and to perform speech recognition on one portion, which is obtained from the second time point, from among the plurality of voice signals.