Multi-Microphone Wake-Word Detection for In-Car Voice Isolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice recognition systems in vehicles fail to accurately detect the direction of a wake-up word in noisy environments, leading to delayed activation of vehicular devices due to the inability to differentiate between voices from different directions.

Innovation Solution

A voice processing device with multiple microphones arranged across seats, a storing unit, a word detection unit, a microphone determining unit, and a voice processing unit that stores voice signals, detects the presence of a wake-up word, determines the speaker's position, and suppresses other voices to isolate the speaker's voice.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple microphones are used to detect voice direction in noisy vehicle environments, then voice recognition accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidmicrophone arrangement complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The vehicle interior is divided into multiple detection zones, each assigned to a specific microphone. The microphones are strategically positioned at different locations (front, rear, left, right) to segment the monitoring space, allowing independent analysis of voice signals from each direction while maintaining overall system accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Multiple microphone signals are merged and processed together through signal processing algorithms. The system combines the output from all microphones to create a comprehensive voice detection result, leveraging the collective information from all sensors to improve recognition accuracy while managing complexity through integrated processing.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If voice signals from all directions are processed equally, then no voice is lost, but activation speed decreases due to inability to prioritize relevant voices

Engineering Contradiction:
Improvevoice detection completenessVSAvoidactivation speed
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

Different processing priorities are assigned to different microphone signals based on their spatial location and relevance. The system applies local quality control by giving higher weight to microphones detecting voices from the driver's direction or front seats, while still monitoring all other directions to ensure no important voice is missed.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system continuously monitors voice signals from all microphones and provides feedback to adjust processing priorities in real-time. When a wake-up word is detected from a specific direction, the system feedback-adjusts the processing focus to that direction while maintaining awareness of other directions, enabling fast activation without losing comprehensive monitoring.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables quick detection and extraction of the wake-up word's direction, enhancing the accuracy and speed of voice recognition, and allowing for timely activation of vehicular devices.

Implementation Method 1

a plurality of different microphones MC1-MC6 corresponding to a plurality of respective seats

Methodology Applied
Scientific EffectSound wave detection: Sound

Implementation Method 2

a voice processing unit which outputs a voice uttered by the speaker while suppressing a voice uttered by a passenger other than the speaker

Methodology Applied
Scientific EffectAcoustic signal processing: Sound

Data Source

PatentUS12118990B2Voice processing device, voice processing method and voice processing system
Publication Date: 2024.10.15 PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
  • US12118990B2 patent drawing
  • US12118990B2 patent drawing
  • US12118990B2 patent drawing

AI summary

A voice processing device includes plural microphones arranged so as to correspond to a plurality of positions. The voice processing device includes at least one memory that stores instructions and voice signals from the plural microphones, and a processor. The voice signals collected by the plural microphones, respectively, during a prescribed period before a present time, are repeatedly stored in the at least one memory as buffered voice signals. The processor detects whether a prescribed word is uttered by a speaker based on the voice signals collected by the plural microphones, determines a microphone corresponding to the speaker by referring to the buffered voice signals, and suppresses the voice signals collected by the plural microphones other than the microphone corresponding to the speaker.