Microphone Array Wave Beam Selection via Voice Energy Metrics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for selecting an output wave beam of a microphone array rely on pre-stored speaker information, require wake word recognition, and are susceptible to high volume noise interference and low volume unstable signal interference, making them inefficient for resource-constrained devices.

Innovation Solution

A method that calculates the overall voice signal energy of each wave beam by converting the wave beam output signal from time domain to frequency domain, obtaining frequency and power spectrum vectors, and then selecting the wave beam with the maximal overall voice signal energy as the output wave beam, without relying on pre-stored information or wake word recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If pre-stored speaker information or wake word recognition is used for wave beam selection, then the direction of arrival can be recognized, but the system cannot handle unstored speakers and increases computational complexity

Engineering Contradiction:
Improvedirection of arrival recognition accuracyVSAvoidcomputational complexity and resource requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the essential characteristics needed for wave beam selection (voice existence probability and energy metrics) from the complex wake word recognition system, eliminating the need for pre-stored speaker information while maintaining directional accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces expensive, complex wake word recognition models with simpler, computationally efficient voice detection algorithms that consume fewer resources and can be executed on resource-constrained devices

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Device complexity

If only energy information of each beam is used for wave beam selection, then the system is simple to implement, but it cannot distinguish between human voice signals and non-human voice signals, making it susceptible to high volume noise interference

Engineering Contradiction:
Improvesystem implementation simplicityVSAvoidnoise interference susceptibility
Core Design Contradiction:
Device complexityVSObject-affected harmful factors

Solution Approach 1:

The patent changes the selection criterion from simple energy magnitude to a composite metric incorporating voice existence probability and energy, allowing discrimination between voice and noise while maintaining implementation feasibility

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces voice existence probability as an intermediary metric that bridges simple energy detection and complex voice recognition, enabling noise rejection without requiring full wake word recognition

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If wave beam selection is based only on voice existence probability, then the system can identify potential voice sources, but it is susceptible to interference from low volume unstable signals

Engineering Contradiction:
Improvevoice signal detection capabilityVSAvoidunstable signal interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent merges voice existence probability with energy information into a unified selection metric, combining the strengths of both approaches: probability-based detection and energy-based filtering, to overcome各自的 weaknesses

Inventive Principle:
Principle #5Merging (Combining)

4Reliability

If wake word engine is used to calculate existence probability before direction of arrival judgment, then voice signals can be identified, but the computational complexity increases making it impractical for resource-constrained devices

Engineering Contradiction:
Improvevoice signal identification accuracyVSAvoidprocessing efficiency on resource-constrained devices
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces computationally expensive wake word engines with simpler voice detection algorithms that provide sufficient accuracy for wave beam selection while being feasible for deployment on resource-constrained IoT devices

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentUS12223976B2Method for selecting output wave beam of microphone array
Publication Date: 2025.02.11 ESPRESSIF SYST SHANGHAI
  • US12223976B2 patent drawing
  • US12223976B2 patent drawing
  • US12223976B2 patent drawing

AI summary

A method for estimating a direction of arrival of sound signals from a microphone array, comprising: receiving sound signals from the microphone array, and performing beamforming on the sound signals to obtain wave beams and corresponding wave beam output signals; performing the following operation on each wave beam: converting the wave beam output signal of a current wave beam to frequency domain from time domain to obtain a frequency spectrum vector and a power spectrum vector; calculating comprehensive voice signal energy of the current wave beam, wherein the comprehensive voice signal energy is the product of comprehensive energy indicating the energy level of the wave beam output signal and a comprehensive voice existence probability indicating an existence probability of voice in the wave beam output signal; and selecting the wave beam with a maximal comprehensive voice signal energy value as the output wave beam.