Wearable Motion Sensor Wake Command for Speech Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech recognition systems face challenges in noisy environments and privacy concerns due to the need for continuous audio transmission and high computational resources, especially when dealing with low signal-to-noise ratios and the desire to process all incoming audio, which can be inefficient and invasive.

Innovation Solution

A local wearable device equipped with motion sensors detects user movements to serve as a wake command, allowing for the selective capture and processing of audio data only when intended, using a combination of speech recognition and natural language understanding to interpret both spoken and gestural inputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If continuous audio transmission is used for speech recognition, then speech processing capability is improved, but energy consumption and computational resources increase

Engineering Contradiction:
Improvespeech processing capabilityVSAvoidenergy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary wake word detection on audio data before initiating full speech processing. The processor detects whether the audio data includes a wake word, and only proceeds with comprehensive speech recognition and natural language understanding when the wake word is detected, thereby avoiding continuous high-energy processing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of processing all audio data continuously, the system applies partial processing by first analyzing audio for wake words only. Full speech processing is applied selectively only when necessary (when wake word is detected), reducing overall computational load and energy consumption while maintaining speech processing capability

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If all incoming audio is processed, then speech recognition accuracy is improved, but processing efficiency decreases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary filtering by detecting wake words in audio data before initiating full speech recognition processing. This preliminary action identifies which audio segments warrant detailed processing, improving overall processing efficiency while maintaining accuracy for relevant inputs

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts and processes only the relevant portions of audio data - specifically segments containing wake words - rather than processing all incoming audio continuously. This extraction approach maintains speech recognition accuracy for meaningful inputs while significantly improving processing efficiency by eliminating unnecessary processing of irrelevant audio segments

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If audio data is transmitted for processing, then speech recognition is enabled, but privacy concerns increase

Engineering Contradiction:
Improvespeech recognitionVSAvoidprivacy concerns
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary wake word detection locally on the wearable device before transmitting audio data. This preliminary action serves as a privacy gate, ensuring that only audio segments following wake word detection (indicating user intent to speak) are transmitted for processing, thereby reducing unnecessary privacy intrusion while maintaining speech recognition functionality

Inventive Principle:
Principle #10Preliminary action

4Productivity

If motion sensors are added to detect wake commands, then system efficiency is improved, but device complexity increases

Engineering Contradiction:
Improvesystem efficiencyVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The motion sensors serve multiple functions: detecting wake commands, detecting user gestures for input, and potentially tracking device orientation. This multi-functionality justifies the added complexity by providing several capabilities from a single sensor addition, improving system efficiency through gesture-based wake and input methods

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10692489B1Non-speech input to speech processing system
Publication Date: 2020.06.23 AMAZON TECH INC
  • US10692489B1 patent drawing
  • US10692489B1 patent drawing
  • US10692489B1 patent drawing

AI summary

A system and method for incorporating motion into a speech processing system. A wearable device that is capable of both capturing spoken utterances and capturing motion data may be used to interact with a speech processing system. In certain circumstances, such as when voice communication are unreliable (due to noise) or when controlling the system by motion is desired, motion of a device may be used to provide input to a speech processing system. For example, sensor data or gesture data resulting from movement of a device may be processed and input into a natural language system as representative of a spoken command portion or other input. The motion information may be interpreted to provide prompts to the system (e.g., “yes,”“no,” etc.), to perform certain commands (skip, forward, back, cancel) or to otherwise control the system.