Multi-zone Voice Control with Echo Cancellation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems in vehicles face interference from multiple talkers and multimedia sources, leading to degraded performance in identifying and processing voice commands accurately in a multi-talker and multimedia environment.

Innovation Solution

A multi-zone speech recognition system with interference and echo cancellation methods that determine the location of each talker in a vehicle cabin, process each speech command to isolate the desired voice, and remove echo and interference, allowing for accurate voice control through adaptive filtering and signal processing techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional speech recognition systems are used in vehicles, then the system can process voice commands, but the performance is degraded by interference from multiple talkers and multimedia sources

Engineering Contradiction:
Improvevoice command identification accuracyVSAvoidinterference from multiple talkers and multimedia sources
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The vehicle cabin is divided into multiple acoustic zones, each with dedicated microphones and processing. The speech recognition system segments the acoustic environment into distinct zones, allowing independent processing of speech signals from each zone. This segmentation enables the system to isolate and identify speech commands from specific zones while filtering out interference from other zones and multimedia sources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each acoustic zone is equipped with specialized processing tailored to its specific characteristics and interference profile. The system applies zone-specific echo cancellation, interference suppression, and speech enhancement techniques. This local quality approach allows optimal processing for each zone's unique acoustic environment and interference conditions, improving overall voice command identification accuracy.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If multiple microphones are used to capture speech from different zones, then the system can identify speech sources, but the complexity of processing multiple signals increases

Engineering Contradiction:
Improvespeech source location detectionVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The microphone array and processing system are segmented into zone-specific groups. Each zone has dedicated microphones and processing channels, which simplifies the overall processing architecture by dividing the complex multi-microphone problem into manageable zone-specific sub-problems. This segmentation reduces cross-zone interference in processing and makes the system more tractable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediate processing stages including zone-specific echo cancellers, interference suppressors, and speech enhancers that act as mediators between the raw microphone signals and the final speech recognition engine. These intermediaries simplify the processing by progressively cleaning and organizing signals before they reach the recognition system, reducing the overall processing complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If echo cancellation and interference removal are applied to each zone, then the voice control reliability improves, but the processing time and computational load increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary processing of speech signals in each zone before speech recognition, including echo cancellation, interference suppression, and speech enhancement. By preparing and cleaning the signals in advance, the recognition engine receives pre-processed, high-quality input that requires less computational effort and time to accurately recognize. This preliminary action reduces the overall processing time despite the additional preprocessing steps.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Processing is segmented into parallel zone-specific channels, allowing simultaneous echo cancellation and interference removal for multiple zones. This parallel processing approach reduces total processing time compared to sequential processing, as multiple zones are handled concurrently rather than one after another.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3678135B1Voice control in a multi-talker and multimedia environment
Publication Date: 2023.10.25 BLACKBERRY LTD
  • EP3678135B1 patent drawingFigure 1
  • EP3678135B1 patent drawingFigure 2
  • EP3678135B1 patent drawingFigure 3~4

AI summary

Voice control in a multi-talker and multimedia environment is disclosed. In one aspect, there is provided a method comprising: receiving a microphone signal for each zone in a plurality of zones of an acoustic environment; generating a processed microphone signal for each zone in the plurality of zones of the acoustic environment, the generating including removing echo caused by audio transducers in the acoustic environment from each of the microphone signals, and removing interference from each of the microphone signals; and performing speech recognition on the processed microphone signals.