Voice-Controlled Device Speech Pausing via Audio and Vision Sensors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice-controlled devices face challenges in determining when to initiate or stop conversations, as they often lack the ability to accurately assess whether speech input is directed at them or if it is appropriate to speak in the presence of humans or other devices, leading to inefficient and unnatural interactions.

Innovation Solution

The system employs a combination of audio and vision sensors, along with natural language classification, to determine the number of people present and their attention, allowing the device to decide whether to speak without an attention word, pause speech during ambient noise, and resume when conditions are suitable, thereby enhancing human-device conversation naturalness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the voice-controlled device continuously monitors audio input to detect speech and pause appropriately, then the naturalness of conversation improves, but the device complexity increases due to multiple sensors and processing requirements

Engineering Contradiction:
Improveconversation naturalnessVSAvoidsensor and processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The voice-controlled device integrates multiple sensors (audio sensors, vision sensors) and combines them with existing speech recognition capabilities to create a multi-functional system that can detect both speech content and speaker presence, enabling natural conversation pausing without requiring entirely separate detection systems

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses natural language classification as an intermediary process that analyzes audio input to determine whether the detected speech is directed at the device or is ambient conversation, enabling intelligent decision-making about when to pause without complex direct control logic

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the device pauses synthesized speech upon detecting ambient noise above a threshold, then conversation naturalness improves, but the loss of information increases due to potential interruption of important output

Engineering Contradiction:
Improveconversation naturalnessVSAvoidspeech output interruption
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system continuously monitors audio input during synthesized speech output and uses this feedback to dynamically adjust speech delivery, pausing only when ambient noise exceeds thresholds that would interfere with communication effectiveness, thereby maintaining information delivery while adapting to environmental conditions

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The pause decision mechanism dynamically adjusts based on real-time audio analysis, comparing detected noise levels against configurable thresholds and considering the context of the conversation, allowing the system to maintain speech output when appropriate and pause only when environmental conditions warrant interruption

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If the device uses vision sensors to detect person presence and attention, then the accuracy of speech direction detection improves, but the use of energy increases due to additional sensor operation

Engineering Contradiction:
Improvespeech direction detection accuracyVSAvoidsensor energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The vision sensor operates periodically rather than continuously, activating specifically when audio speech detection triggers a need to verify speaker identity or attention direction, thereby reducing overall energy consumption while maintaining detection accuracy when needed

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system performs preliminary audio analysis to identify potential speech directed at the device before activating vision sensors for confirmation, using the lower-energy audio detection as a filter to reduce unnecessary vision sensor activation and associated energy consumption

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10923101B2Pausing synthesized speech output from a voice-controlled device
Publication Date: 2021.02.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10923101B2 patent drawing
  • US10923101B2 patent drawing
  • US10923101B2 patent drawing

AI summary

A system, a computer program product, and method for controlling synthesized speech output on a voice-controlled device. The voice-controlled device recognized that speech input is being received. The voice-controlled device outputs synthesized speech based on the speech input. While outputting synthesized speech based on the audio is captured. The voice-controlled device recognized the audio input as speech and pausing the outputting of synthesized speech. Otherwise, in response to the captured audio not being recognized as speech and above a settable background noise threshold, pausing the outputting of synthesized speech. The paused output of speech based on the synthesized speech input is resumed after the pausing of the output of synthesized speech being within a settable pause timeframe.