Management Layer for Multi-Service Voice Assistant Interference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face inconvenience when interacting with multiple intelligent personal assistant (IPA) services, as they typically require separate devices and are not able to issue voice commands to multiple IPAs simultaneously without interference, limiting the natural and conversational experience.
Innovation Solution
A management layer for smart devices that detects activation phrases, extracts query content, and transmits voice commands to multiple IPA services, allowing for sequential responses without interference, enabling users to issue voice commands to multiple IPAs via a single device with more conversational syntax.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a user subscribes to multiple IPA services, then the user can access diverse assistant services, but the user requires multiple separate devices to access each IPA service
Solution Approach 1:
The patent combines multiple IPA service clients into a single smart device, creating a unified interface that manages multiple IPA services simultaneously. This merging eliminates the need for separate devices for each IPA service while maintaining full functionality of each service.
Solution Approach 2:
The smart device is configured to perform multiple functions by supporting various IPA services through a universal client application. This single device can interact with multiple different IPA services, making it a multi-functional platform that replaces multiple specialized devices.
2Adaptability or versatility
If a single device supports multiple IPA services, then device versatility improves, but response interference occurs when multiple IPAs receive commands simultaneously
Solution Approach 1:
The system performs preliminary actions by detecting activation phrases and identifying the intended target IPA service before processing the voice command. This preliminary identification and sequential management of commands prevents response interference by ensuring that only one IPA service processes a command at a time, maintaining reliable operation.
3Ease of operation
If a default IPA service is configured, then system operation is simplified, but switching to another IPA service is cumbersome and time-consuming
Solution Approach 1:
The system implements dynamic service switching by detecting activation phrases that indicate the user's intent to change the target IPA service. Rather than requiring manual reconfiguration, the system dynamically adapts to user preferences in real-time, making service switching as simple as speaking a different activation phrase.
4Ease of operation
If voice commands are transmitted to multiple IPAs simultaneously, then user convenience improves, but response interference and system reliability deteriorate
Solution Approach 1:
The system performs preliminary identification of the target IPA service through activation phrase detection before transmitting voice commands. This preliminary action enables the system to route commands sequentially to the appropriate IPA service, maintaining user convenience while preventing response interference through proper command management.
Data Source
AI summary
Performing speech recognition in a multi-device system includes receiving a first audio signal that is generated by a first microphone in response to a verbal utterance, and a second audio signal that is generated by a second microphone in response to the verbal utterance; dividing the first audio signal into a first sequence of temporal segments; dividing the second audio signal into a second sequence of temporal segments; comparing a sound energy level associated with a first temporal segment of the first sequence to a sound energy level associated with a first temporal segment of the second sequence; based on the comparing, selecting, as a first temporal segment of a speech recognition audio signal, one of the first temporal segment of the first sequence and the first temporal segment of the second sequence; and performing speech recognition on the speech recognition audio signal.


