Voice Command Processing via Local-Remote Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Advanced premises security systems require more sophisticated control and automation capabilities, including voice interaction, to manage various devices and services efficiently, while existing systems often lack seamless integration of local and remote command processing and user tracking.
Innovation Solution
A voice interactive system that processes control commands by receiving speech phrases on a local device, comparing them to local commands, and sending unmatched phrases to a remote server for further processing, while using audio beam forming and steering to determine user location and emotions, enabling interactive applications and device control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech phrases are sent to a remote server for processing, then the system can handle complex commands and interactives applications, but the response time increases and system complexity increases
Solution Approach 1:
The patent segments command processing into two categories: local control commands processed immediately by the local device, and interactive application commands sent to the remote server. This segmentation allows common commands to be executed locally without network latency, while complex interactions still benefit from remote processing capabilities.
Solution Approach 2:
The local device acts as an intermediary between the user and the remote server, filtering and pre-processing speech phrases before forwarding them. This intermediary role reduces the burden on the remote server and enables faster local response for straightforward commands.
2Measurement precision
If audio beam forming and steering systems are implemented, then user location tracking is improved, but device complexity increases
Solution Approach 1:
The audio beam forming and steering system serves multiple functions: it determines user spatial location for avatar control, identifies which user is speaking for context tracking, and enhances speech recognition accuracy. This multi-functionality justifies the added complexity by delivering multiple benefits from a single system.
3Reliability
If multiple local devices are used for speech detection, then coverage and reliability are improved, but system complexity and coordination difficulty increase
Solution Approach 1:
The system uses audio beam forming feedback to determine which user is speaking and their spatial location. This feedback mechanism allows the system to dynamically route speech phrases to the appropriate processing path and track user context across multiple devices without complex manual coordination.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables efficient control of premises devices, provides enhanced user interaction through voice commands, and improves speech recognition by analyzing context and emotions, offering a scalable and flexible solution for advanced security and automation needs.
Implementation Method 1
detecting the speech phrases via a beam forming and steering system of the local device that determines a spatial location of the user from the speech phrases
Data Source
AI summary
A system and method for processing user speech commands in a voice interactive system is disclosed. Users issue speech phrases on a local device in a premises network, and the local devices first determine if the speech phrases match any commands in a set of local control commands. The control commands, in examples, can activate and deactivate premises devices such as “smart” televisions and simpler lighting devices connected to home automation hubs. In the event of a command match, local actions associated with the commands are executed directly on the premises devices in response. When no match is found on the local device, the speech phrases are sent in messages to a remote server over a network cloud such as the Internet for further processing. This can save on bandwidth and cost as compared to current voice recognition systems.


