Voice Control With Microphone Array Echo Cancellation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice interaction technologies face low recognition accuracy and poor user experience due to echo interference and environmental noise in remote usage scenarios, particularly in audio playing states, leading to ineffective voice command recognition and delayed responses.

Innovation Solution

A method and device with a microphone array that analyze interference sounds before and after detecting a wake-up word to adjust voice enhancement modes, employing beamforming and adaptive cancellation to eliminate echo and noise, and processing voice commands in two stages to enhance recognition accuracy and user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice recognition is performed during audio playing state, then user can issue commands while audio plays, but echo interference from audio playback reduces voice command recognition accuracy

Engineering Contradiction:
Improvevoice command recognition availabilityVSAvoidvoice command recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the audio playing process into distinct states: before wake-up word detection and after wake-up word detection. In the first state, voice enhancement focuses on echo cancellation to enable command recognition. In the second state, voice enhancement focuses on noise and reverberation cancellation. This segmentation allows different voice enhancement strategies to be applied at different times, resolving the contradiction between maintaining audio playback and achieving accurate voice recognition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts the voice enhancement mode based on the current state. The system transitions from an audio-playing-state voice enhancement mode (optimized for echo cancellation) to a non-audio-playing-state voice enhancement mode (optimized for noise and reverberation cancellation) when a wake-up word is detected. This dynamic adaptation allows the system to optimize voice recognition accuracy for each specific operating condition.

Inventive Principle:
Principle #15Dynamics

2Duration of action of stationary object

If audio playing continues during voice command detection, then uninterrupted audio experience is maintained, but environmental noise and reverberation interfere with command recognition

Engineering Contradiction:
Improveaudio playing continuityVSAvoidcommand word recognition accuracy
Core Design Contradiction:
Duration of action of stationary objectVSMeasurement precision

Solution Approach 1:

The patent performs preliminary voice enhancement processing on the user's voice signal before it reaches the voice recognition system. By pre-processing the voice signal to remove echo, noise, and reverberation interference, the system ensures that the voice command can be accurately recognized even while audio continues to play. This preliminary action resolves the contradiction by preparing the voice signal in advance to withstand the interfering environment.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If single voice enhancement mode is used, then system complexity is reduced, but voice enhancement effect varies poorly across different sound environments

Engineering Contradiction:
Improvevoice enhancement system simplicityVSAvoidvoice enhancement effectiveness across environments
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent changes the parameters of the voice enhancement system by switching between different enhancement modes based on the audio playing state. The system uses two distinct voice enhancement modes: one optimized for audio-playing state (echo cancellation focused) and another for non-audio-playing state (noise and reverberation cancellation focused). This parameter change approach allows the system to adapt to different sound environments without requiring a completely complex reconfiguration, resolving the contradiction between simplicity and adaptability.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Improves voice command recognition accuracy and user experience by adapting voice enhancement modes to specific sound environments, ensuring timely and accurate detection of user commands, and aligning with user habits by stopping audio playback upon wake-up word detection.

Implementation Method 1

echo interference and environmental noise and reverberant interference

Methodology Applied
Scientific EffectEcho: Echo

Implementation Method 2

echo interference and environmental noise and reverberant interference

Methodology Applied
Scientific EffectReverberation: Reverberation

Implementation Method 3

employing beamforming and adaptive cancellation to eliminate echo and noise

Methodology Applied
Scientific EffectBeamforming:

Implementation Method 4

employing beamforming and adaptive cancellation to eliminate echo and noise

Methodology Applied
Scientific EffectAdaptive cancellation:

Data Source

PatentUS10453457B2Method for performing voice control on device with microphone array, and device thereof
Publication Date: 2019.10.22 LITTLE BIRD CO LTD
  • US10453457B2 patent drawing
  • US10453457B2 patent drawing
  • US10453457B2 patent drawing

AI summary

A method and device for performing voice control on a device with a microphone array are disclosed. The method includes the following steps. It is confirmed that the device is in an audio playing state. An interference sound interfering the device in the audio playing state is analyzed. A voice enhancement mode adopted by the device is selected according to a feature of the interference sound. A user's voice is detected in real time for a wake-up word, and when the wake-up word is detected, the device is controlled to stop audio playing. An interference sound interfering the device after playing audios is stopped is analyzed, and the voice enhancement mode adopted by the device is adjusted according to a feature of the interference sound. A command word from a user is acquired to control the device to execute a corresponding function, to respond to the user.