Speech Dialog Input Far-Field Near-Field Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech-based systems face challenges in accurately determining user intent through two-way speech dialogs, particularly in noisy environments, due to differences in microphone coverage and interference handling between stationary base devices and portable handheld devices.

Innovation Solution

A speech-based system comprising a base device with omnidirectional or directional microphones and a handheld device with a push-to-talk button, both communicating via a network interface, engages in dialog turns using automatic speech recognition and natural language understanding, with the system determining user intent and responding through speech, leveraging cloud-based services for processing and noise filtering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If the system uses stationary base devices with omnidirectional microphones for speech input, then the coverage area is improved, but the accuracy of intent determination deteriorates in noisy environments

Engineering Contradiction:
Improvemicrophone coverage areaVSAvoidintent determination accuracy
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The system segments the speech input source into two distinct categories: far-field sources (stationary base device) and near-field sources (handheld device). This segmentation allows the system to apply different processing strategies and noise filtering techniques appropriate to each source type, thereby maintaining intent determination accuracy while preserving the broad coverage capability of omnidirectional microphones.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary classification mechanism that identifies whether speech originates from a far-field or near-field source. This intermediary step enables the system to selectively apply noise filtering and processing parameters based on the identified source type, resolving the contradiction between broad coverage and accurate intent determination in noisy environments.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If the system uses portable handheld devices with push-to-talk buttons for speech input, then the noise filtering capability is improved, but the ease of operation deteriorates due to additional user actions required

Engineering Contradiction:
Improvenoise interferenceVSAvoidspeech input convenience
Core Design Contradiction:
Object-affected harmful factorsVSEase of operation

Solution Approach 1:

The system dynamically adapts its noise filtering and processing behavior based on the identified speech source type. When near-field sources (handheld devices) are detected, enhanced noise filtering is applied. When far-field sources (stationary devices) are detected, the system uses alternative noise reduction strategies. This dynamic adaptation allows the system to maintain ease of operation with stationary devices while achieving effective noise filtering when handheld devices are used.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters based on the speech source identification. Different noise filtering thresholds, gain settings, and signal processing parameters are applied depending on whether the speech originates from a far-field or near-field source. This parameter adaptation enables the system to optimize noise filtering performance without permanently compromising the ease of operation.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If the system integrates both stationary base devices and portable handheld devices for speech dialogs, then the adaptability to different environments is improved, but the device complexity increases

Engineering Contradiction:
Improveenvironmental adaptabilityVSAvoidsystem configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements a universal speech processing framework that can handle both far-field (stationary device) and near-field (handheld device) speech inputs through a single integrated architecture. The core speech recognition and natural language understanding components remain unchanged, while only the noise filtering and source identification modules adapt to different device types. This universal approach enables environmental adaptability without proportionally increasing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs self-service mechanisms through automatic source identification and adaptive parameter selection. The system automatically detects whether speech originates from a stationary or handheld device and configures appropriate processing parameters without requiring manual user configuration. This self-service capability enables the system to adapt to different environments while keeping the user interface simple and the operational complexity low.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9792901B1Multiple-source speech dialog input
Publication Date: 2017.10.17 AMAZON TECH INC
  • US9792901B1 patent drawing
  • US9792901B1 patent drawing
  • US9792901B1 patent drawing

AI summary

A speech system may be configured to operate in conjunction with a stationary base device and a handheld remote device to receive voice commands from a user. A user may direct speech either to the base device or to the handheld device. In order to direct speech to the base device, the user first speaks a keyword. In order to direct speech to the handheld device, the user presses a talk control on the handheld device. A dialog may be conducted with the user in multiple turns, where each turn comprises user speech and a speech response by the speech system. The user speech in any given dialog turn may be provided from the base device and/or the handheld device.