Distributed Keyphrase Recognition for Noisy Voice Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing keyphrase recognition systems are resource-intensive and inefficient, particularly in noisy environments, making it difficult to reliably control apparatuses using spoken commands, especially when noise levels are high or when dealing with accented speech.
Innovation Solution
A method and system that utilize a keyphrase recognition model based on multiple keyphrases, integrated into each apparatus, which detects keyphrases in speech and communicates them wirelessly, combining a domain-specific language model with a wide-vocabulary model to enhance recognition accuracy and reduce computational resources, allowing for reliable voice control in noisy conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine-learning techniques with substantial data collection are used to train neural networks for keyphrase recognition, then recognition capability is improved, but computing resources and human resources become excessively intensive
Solution Approach 1:
The patent segments the keyphrase recognition task into two distinct components: a noise-robust speech-to-text conversion stage and a separate keyphrase detection stage. This segmentation allows each component to be optimized independently, reducing the overall computational burden while maintaining recognition accuracy. The speech-to-text module handles noisy input efficiently, and the keyphrase detector operates on the converted text with lower resource requirements.
Solution Approach 2:
The patent introduces an intermediary speech-to-text conversion process that transforms noisy speech input into text before keyphrase detection. This intermediary step acts as a mediator that separates the noise handling function from the keyphrase recognition function, allowing the system to achieve accurate keyphrase detection without directly processing noisy audio through complex neural networks.
2Ease of operation
If traditional keyphrase recognition systems are used in noisy environments, then speech input can be processed, but noise from apparatuses interferes with accurate recognition
Solution Approach 1:
The system segments the processing into noise-robust speech-to-text conversion and separate keyphrase detection. The speech-to-text module is specifically designed to handle noisy audio environments, extracting meaningful text content while filtering out apparatus noise. This segmentation enables reliable operation in noisy environments by dedicating specific processing stages to handle different aspects of the input signal.
Solution Approach 2:
The speech-to-text conversion process serves as an intermediary that mediates between noisy speech input and keyphrase detection. This intermediary layer converts audio signals to text in a noise-robust manner, protecting the subsequent keyphrase detection stage from noise interference and enabling reliable voice control in challenging acoustic environments.
3Adaptability or versatility
If existing keyphrase recognition systems are deployed for controlling multiple apparatuses, then speech control is enabled, but the systems are burdensome to generate and modify
Solution Approach 1:
The patent segments the system into modular components: a speech-to-text engine and a configurable keyphrase detection module. This modular architecture allows the keyphrase detection component to be independently configured and modified for different apparatuses without regenerating the entire system. The segmentation enables easy adaptation to multi-apparatus control by simply adding or modifying keyphrase patterns in the detection module.
Solution Approach 2:
The speech-to-text conversion engine serves as a universal component that handles speech input for multiple different apparatuses. This universal module can be paired with different keyphrase detection configurations to control various devices, enabling the system to adapt to multi-apparatus control scenarios without requiring separate speech processing systems for each device.
Data Source
AI summary
Technologies are provided for a distributed spoken language interface for speech control of multiple apparatuses. In some aspects, a first apparatus can receive an audio signal representative of speech. The first apparatus can detect, based on applying a keyphrase recognition model to the speech, a keyphrase. The keyphrase can include a first string of characters defining an identifier corresponding to at least one second apparatus and also includes a second string of characters defining a command. The first apparatus can cause, based on the identifier, a communication unit integrated in the first apparatus to send the keyphrase to the at least one second apparatus. The at least one second apparatus can receive the keyphrase, and can cause one or more components to execute the command.


