Dual-Microphone Audio Buffer Catch-Up for Low-Latency Voice Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing microphone systems experience processing delays due to buffering, which can lead to user dissatisfaction in voice-activated applications where quick signal processing is necessary.

Innovation Solution

The use of two microphones, where one operates in a low power mode with buffering and the other provides real-time data once a trigger phrase is detected, allowing for immediate activation and stitching together of complete phrases by a processor to minimize delays.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If microphones use buffers to aid in voice activity detection, then voice detection capability is improved, but processing delays occur

Engineering Contradiction:
Improvevoice detection capabilityVSAvoidprocessing delays
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The audio signal processing is segmented into two parallel paths: a first microphone processes audio through a buffer for reliable voice activity detection, while a second microphone processes audio in real-time without buffering. This segmentation allows the system to benefit from both buffered detection reliability and unbuffered processing speed simultaneously.

Inventive Principle:
Principle #1Segmentation

2Use of energy by moving object

If microphones operate in low power mode with buffering, then power consumption is reduced, but signal processing speed deteriorates

Engineering Contradiction:
Improvepower consumptionVSAvoidsignal processing speed
Core Design Contradiction:
Use of energy by moving objectVSSpeed

Solution Approach 1:

The system segments the microphone functions: the first microphone operates in low-power buffered mode for detection, while the second microphone operates in high-speed unbuffered mode for real-time processing. This allows the system to achieve both low power consumption and high processing speed through functional segmentation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses a second microphone as a copy of the first microphone's function but operates it without buffering to provide real-time audio data. This copy provides the fast processing path while the original buffered microphone handles power-efficient detection.

Inventive Principle:
Principle #26Copying

3Ease of operation

If voice commands are processed quickly, then user satisfaction is improved, but buffering requirements are compromised

Engineering Contradiction:
Improveuser satisfactionVSAvoidbuffering delays
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The dual-microphone architecture segments the processing into detection (buffered) and execution (unbuffered) paths, enabling the system to provide both reliable voice detection and fast command processing, thereby improving user satisfaction without sacrificing buffering benefits.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10121472B2Audio buffer catch-up apparatus and method with two microphones
Publication Date: 2018.11.06 KNOWLES ELECTRONICS LLC
  • US10121472B2 patent drawing
  • US10121472B2 patent drawing
  • US10121472B2 patent drawing

AI summary

A first microphone is operated in a low power sensing mode, and a buffer at the first microphone is used to temporarily store at least some of the phrase. Subsequently the first microphone is deactivated, then the first microphone is re-activated to operate in normal operating mode where the buffer is no longer used to store the phrase. The first microphone forms first data that does not include the entire phrase. A second microphone is maintained in a deactivated mode until the trigger portion is detected in the first data, and when the trigger portion is detected, the second microphone is caused to operate in normal operating mode where no buffer is used. The second microphone forms second data that does not include the entire phrase. A first electronic representation of the phrase as received at the first microphone and a second electronic representation of the phrase as received at the second microphone are formed from selected portions of the first data and the second data.