Voice Interface Audio Buffering for Seamless Command Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech processing systems require users to speak a keyword to initiate commands, leading to friction and a less seamless user experience, especially when additional information is needed or when transitioning between commands.

Innovation Solution

The system allows devices to continuously send audio data to servers without detecting a spoken keyword, even after a command has been processed, to capture potential subsequent commands or ongoing conversations, reducing the need for explicit keyword detection in certain contexts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the system requires users to speak a keyword to initiate commands, then system control and accuracy are improved, but user experience and interaction seamlessness deteriorate

Engineering Contradiction:
Improvesystem controlVSAvoiduser experience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system performs preliminary actions by continuously capturing and pre-processing audio data in the background before a keyword is spoken. Audio buffering and preliminary analysis occur proactively, so when a keyword is detected, the system can immediately process the command without interruption, eliminating the need to pause for keyword detection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system maintains continuous audio capture and processing operations without interruption. Audio data is continuously buffered and pre-analyzed in the background, ensuring that the system is always ready to process commands seamlessly. This continuous operation eliminates gaps between keyword detection and command processing.

Inventive Principle:
Principle #20Continuity of useful action

2Productivity

If the system continuously sends audio data to servers without keyword detection, then responsiveness and efficiency are improved, but energy consumption and processing overhead increase

Engineering Contradiction:
ImproveresponsivenessVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by selectively transmitting only relevant audio segments to the server. Instead of continuously sending all audio data, the system buffers audio locally and only transmits segments containing potential commands or relevant information, reducing network traffic and energy consumption while maintaining responsiveness.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary local processing and filtering of audio data before transmission. Audio segments are pre-analyzed on-device to identify relevant content, and only these filtered segments are sent to the server. This preliminary action reduces the volume of data transmitted and the computational load on the server.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the system requires explicit keyword prompts for each command, then command accuracy is improved, but interaction friction and time delay increase

Engineering Contradiction:
Improvecommand accuracyVSAvoidinteraction delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary audio buffering and continuous speech recognition in the background before a command is fully articulated. By preparing the audio processing pipeline in advance and continuously analyzing speech patterns, the system can accurately identify commands without requiring explicit keyword prompts, reducing interaction delay while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from continuous speech recognition to dynamically adjust its command processing. By continuously monitoring speech patterns and providing real-time feedback on recognized words and intent, the system can confirm command accuracy without requiring additional keyword prompts, thereby reducing interaction time while maintaining precision.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10847149B1Speech-based attention span for voice user interface
Publication Date: 2020.11.24 AMAZON TECH INC
  • US10847149B1 patent drawing
  • US10847149B1 patent drawing
  • US10847149B1 patent drawing

AI summary

Techniques for enabling a device to send to a speech processing server further input audio data following a completed utterance dialog to prevent the need for subsequent keywords to be spoken to invoke subsequent commands are described. A system receives input audio data corresponding to an utterance from a device upon the device detecting speech corresponding to a keyword. The system performs speech processing on the input audio data to determine a command. The system determines output data responsive to the command and sends same to the device, thus completing operations regarding the utterance. The system may also send an instruction to the device to: send to the system further input audio data corresponding to further input audio without the device first detecting a wake command.