Voice Input Handling for Speech Mode Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current electronic computing devices lack the ability to detect whether users are providing speech input using 'press and release' (P & R) or 'push to talk' (PTT) modes, leading to ambiguity in determining when a speech recognition session begins and ends.

Innovation Solution

A speech input device with a mechanism to detect the state of an input indicator, such as a button or switch, and a microphone to differentiate between P & R and PTT modes by analyzing the duration of input mechanism activation and sound levels, determining the speech input mode through a process that includes recording sounds upon activation, filtering for human voice frequencies, and determining the mode based on whether sounds are detected and the duration of input mechanism selection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a device uses a physical or virtual button to indicate speech recognition session initiation, then the device can accept voice input, but the device cannot determine whether the user intends P & R or PTT mode

Engineering Contradiction:
Improvespeech input mode detectionVSAvoidinput mechanism complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system monitors the duration of button press and sound level feedback to automatically determine whether the user intends P & R or PTT mode, eliminating the need for explicit user selection while improving adaptability

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The device automatically detects and adapts to the user's intended speech input mode through analysis of input duration and sound characteristics, allowing the system to self-determine the appropriate mode without additional user input

Inventive Principle:
Principle #25Self-service

2Ease of operation

If the device uses PTT mode requiring button depression for the duration of speech input, then continuous speech input is enabled, but users who expect P & R mode experience confusion

Engineering Contradiction:
Improvespeech input operationVSAvoiduser intent ambiguity
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system provides feedback by monitoring input duration and sound levels to infer user intent, allowing the device to adapt to either P & R or PTT expectations based on actual usage patterns rather than requiring explicit mode selection

Inventive Principle:
Principle #23Feedback

3Speed

If the device uses P & R mode with brief button press, then quick session initiation is achieved, but continuous speech input sessions cannot be maintained

Engineering Contradiction:
Improvesession initiation speedVSAvoidspeech recognition session duration
Core Design Contradiction:
SpeedVSDuration of action of moving object

Solution Approach 1:

The system dynamically adjusts the interpretation of button press duration based on sound level detection, allowing the same physical action to initiate either brief or continuous speech sessions depending on user behavior patterns

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9767797B2Voice input handling
Publication Date: 2017.09.19 DISH TECHNOLOGIES LLC
  • US9767797B2 patent drawing
  • US9767797B2 patent drawing
  • US9767797B2 patent drawing

AI summary

In an input device it is determined that an input indicator mechanism is selected for a predetermined period of time. Speech is recorded based on the input indicator mechanism being selected.