AI Agent Servers for Cross-Device Voice Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems struggle to organically manage interactive utterances across multiple devices with different voice recognition service providers and to provide seamless voice recognition services between devices with and without user interfaces.
Innovation Solution
The implementation of two AI agent servers, one for devices capable of providing a user interface and another for those incapable, which communicate through a natural language processing server to analyze and process voice commands, ensuring consistent and interactive voice recognition services across diverse devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple devices with different voice recognition service providers are used, then device diversity and functionality are improved, but organic management of utterances and service consistency deteriorate
Solution Approach 1:
The patent introduces a TV device as an intermediary hub that receives utterances from multiple peripheral devices (remote controller, IoT speaker, smartphone) and coordinates with multiple AI agent servers. The TV device's controller manages the complex interactions by routing utterances to appropriate AI agents and integrating responses, thereby reducing the management burden on individual peripheral devices and enabling seamless cross-device voice recognition services.
2Adaptability or versatility
If devices with and without user interfaces coexist, then system versatility is improved, but interactive service provision deteriorates
Solution Approach 1:
The patent creates a universal interaction framework where the TV device serves multiple functions: it acts as a display interface for devices without screens (like IoT speakers), as a coordination hub for multi-device interactions, and as a fallback interface when peripheral devices lack display capabilities. This multi-functional approach enables seamless interactive services across diverse device types with different interface capabilities.
Solution Approach 2:
The TV device functions as an intermediary that bridges devices with and without user interfaces. When a peripheral device without a display (e.g., IoT speaker) receives a voice command, the TV device displays the relevant information and coordinates the response, thereby providing a unified interactive experience across devices with varying interface capabilities.
3Reliability
If independent utterance management is implemented per device, then device autonomy is improved, but service integration and interaction deteriorate
Solution Approach 1:
The patent merges the autonomous capabilities of multiple devices with centralized coordination through the TV device. Each peripheral device maintains its independent voice recognition and initial processing capabilities, but the TV device consolidates the management of utterances across devices and integrates responses from multiple AI agent servers, creating a unified service experience that leverages both autonomy and integration.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An artificial intelligence device may receive first voice data corresponding to first voice uttered by a user from a first peripheral device, acquire a first intention corresponding to the first voice data, transmit a first search result corresponding to the first intention to the first peripheral device, receive second voice data corresponding to second voice uttered by the user from a second peripheral device, acquire a second intention corresponding to the received second voice data, and transmit a search result corresponding to the second intention to the second peripheral device depending on whether the second intention is an interactive intention associated with the first intention.