Microphone Array Voice Pickup with Source Localization and Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional microphone array systems face challenges in far-field voice pickup, particularly in high noise environments, where it's difficult to determine the active speaker and initiate recording without user intervention, especially when multiple speakers are present.

Innovation Solution

A microphone array based pickup method that performs voice activation detection, locates the voice source, enhances the signal, conducts voice wakeup detection, and uses a pickup indicator lamp to direct the user on the active speaker, allowing for efficient voice signal capture and processing in far-field environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If microphone array recording is used to control noise and enhance target voice signals in far-field pickup, then the voice signal quality is improved, but the device complexity increases

Engineering Contradiction:
Improvevoice signal qualityVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The voice pickup process is segmented into distinct stages: voice activation detection stage, voice source locating stage, and voice enhancement stage. Each stage processes only when necessary, dividing the complex far-field pickup task into manageable segments that reduce overall system complexity while maintaining signal quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Voice activation detection is performed preliminarily before full voice enhancement processing. The system detects voice activation signals in advance, determines voice source directions, and only then initiates the computationally intensive beam-forming enhancement, avoiding unnecessary processing when no voice is present.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If continuous recording is performed in far-field environment to capture all possible voices, then the voice capture capability is improved, but the loss of time and energy increases

Engineering Contradiction:
Improvevoice capture capabilityVSAvoidrecording time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs periodic voice activation detection on incoming audio signals to identify when voice signals are present. Recording and enhancement operations are activated only during periods when voice activation is detected, rather than running continuously, thus reducing time loss while maintaining capture capability.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The microphone array system automatically detects voice activation signals and initiates voice source locating and enhancement processes without external intervention. This self-service mechanism ensures voice capture capability is maintained while avoiding unnecessary continuous recording operations.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If voice source locating is performed using time difference of signals from multiple microphones, then the voice source direction accuracy is improved, but the computational workload increases

Engineering Contradiction:
Improvevoice source direction accuracyVSAvoidcomputational workload
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

Voice source locating based on time difference of arrival is performed preliminarily to determine the direction of the voice source. This preliminary localization enables subsequent beam-forming enhancement to focus computational resources only on the identified voice direction, reducing overall computational workload while maintaining direction accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts only the necessary computational operations for voice source locating (time difference calculation) from the full signal processing chain. By separating the localization step from the enhancement step, the system performs minimal necessary computation for direction finding while deferring more intensive processing to when needed.

Inventive Principle:
Principle #2Taking out (Extraction)

4Object-affected harmful factors

If beam-forming technology is used to enhance voice signals from specific directions, then the noise control capability is improved, but the device complexity increases

Engineering Contradiction:
Improvenoise control capabilityVSAvoidsignal processing complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

Beam-forming enhancement is applied locally to specific voice source directions identified by the voice source locating unit, rather than processing all directions uniformly. This localized approach enhances noise control capability for target voices while reducing signal processing complexity by excluding other directions from intensive enhancement.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system applies beam-forming enhancement partially, focusing computational resources only on directions where voice activation is detected. Rather than enhancing all incoming signals, the system performs partial enhancement on selected voice sources, improving noise control while managing processing complexity.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11302341B2Microphone array based pickup method and system
Publication Date: 2022.04.12 YUTOU TECH HANGZHOU
  • US11302341B2 patent drawing
  • US11302341B2 patent drawing
  • US11302341B2 patent drawing

AI summary

The present invention relates to a microphone array based pickup method, comprising: performing voice activation detection using one channel voice signal among multichannel voice signals picked up and output by a microphone array, and determining if a voice activation signal occurs; locating the voice source by using the multichannel voice signals output by the microphone array to obtain the voice source locating direction; enhancing a voice signal in the voice source locating direction to obtain an enhanced voice signal; conducting voice wakeup detection on the enhanced voice signal and determining if a voice wakeup is detected; picking up and outputting the multichannel voice signals by the microphone array; Step 6: processing the multichannel voice signals picked up by the microphone array into one channel enhanced voice, and outputting the one channel enhanced voice as a finally picked up voice.