IoT Multi-VA Voice Command Capture for Moving Users
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice assistant systems struggle to capture and process complete voice commands when a user is moving within an environment with multiple devices, leading to incomplete command execution, privacy and power consumption issues, and limitations in handling long voice commands.
Innovation Solution
A system and method that utilizes a multi-VA environment to receive voice commands by determining a user's breathing pattern and movement, selecting appropriate VAs based on signal strength and voiceprint, merging audio from multiple devices to form a complete command, and dynamically adjusting listening times to ensure seamless execution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a single voice assistant device is used to capture voice commands, then the device can process commands within its listening range, but the user cannot move freely and parts of the utterance are missed when moving
Solution Approach 1:
The system segments the voice command capture task across multiple voice assistant devices distributed in the environment. Each device captures the portion of the utterance it receives, and these segments are then combined to form the complete command, allowing the user to move freely while maintaining command completeness.
Solution Approach 2:
The system merges audio inputs from multiple voice assistant devices that have captured different portions of the user's utterance. By combining these audio segments and removing overlaps, the system reconstructs the complete voice command, enabling both user mobility and full command capture.
2Ease of operation
If multiple voice assistants capture overlapping portions of the utterance, then the user can move freely, but the command cannot be fully understood due to overlapping captures
Solution Approach 1:
The system extracts and removes the overlapping portions from the audio captures of multiple voice assistants. By identifying and eliminating duplicate segments, the system preserves only the unique portions of each capture, ensuring accurate command recognition while maintaining user mobility benefits.
Solution Approach 2:
The system introduces an intermediary processing layer that receives audio from multiple VAs, identifies overlaps, and synthesizes a non-redundant complete utterance. This intermediary process resolves the conflict between multiple captures and accurate recognition by mediating the integration of audio segments.
3Ease of operation
If all voice assistants remain in always listening mode to capture moving users, then command capture is seamless, but privacy and power consumption issues increase
Solution Approach 1:
Instead of continuous listening mode, the system uses periodic wake-up triggers where voice assistants activate only when a wake word or trigger is detected. This periodic activation allows the system to maintain command capture capability while significantly reducing power consumption and addressing privacy concerns by remaining dormant otherwise.
Solution Approach 2:
The system performs preliminary detection of wake words or trigger signals before activating full listening mode. This preliminary action allows voice assistants to remain in a low-power state and only activate when needed, balancing seamless command capture with reduced energy consumption and improved privacy.
Data Source
AI summary
A system and method of receiving a voice command while a user is moving in an Internet of Things (IoT) environment comprising a plurality of virtual assistants (VAs) is provided. The method includes receiving a first audio uttered by a user at a first VA. The user's intention to speak more (ISM) is determined based on the user's breathing pattern. Furthermore, the method includes selecting a second VA to initiate listening to a second audio uttered by the user, based on detecting a movement of the user with respect to the plurality of VAs in the IoT environment. The method includes merging the first audio and the second audio to determine a voice command corresponding to the user's complete utterance, for execution within the IoT environment.


