Interactive Agent Control via Gaze and Communication State Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current interactive agent systems face challenges in efficiently processing user requests through a combination of voice and gesture recognition, particularly in varying communication states and environments, leading to suboptimal user interaction and information transmission.
Innovation Solution
The interactive agent system employs a dual interaction engine approach, utilizing both local and cloud interaction engines based on communication state, allowing seamless switching between them to ensure effective request processing and information transmission, and dynamically adjusts the agent's mode and interaction style based on user and vehicle states.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single interaction engine is used for processing user requests, then the device complexity is low, but the adaptability to different communication states and environments deteriorates
Solution Approach 1:
The system dynamically switches between local and cloud interaction engines based on communication state. When communication is available, the cloud engine handles complex tasks; when communication is unavailable, the local engine maintains basic functionality. This dynamic adaptation resolves the contradiction by adjusting system configuration according to environmental conditions.
Solution Approach 2:
The interaction system is designed with multi-functionality by incorporating both local and cloud interaction engines that can handle different types of requests. The local engine processes simple tasks independently, while the cloud engine handles complex tasks requiring more resources, providing universal adaptability across various communication states.
2Productivity
If cloud interaction engine is used for all processing, then the productivity is high, but the reliability in varying communication states deteriorates
Solution Approach 1:
Different parts of the system have different functional qualities tailored to their capabilities. The local interaction engine is optimized for basic, low-resource tasks that can be executed reliably without communication, while the cloud interaction engine is optimized for high-productivity complex tasks when communication is available. This local quality differentiation resolves the contradiction between productivity and reliability.
Solution Approach 2:
The system introduces a communication state detection mechanism that acts as an intermediary between the local and cloud interaction engines. This mediator determines which engine should handle each request based on current communication conditions, ensuring both high productivity when possible and reliable operation when communication is unavailable.
3Adaptability or versatility
If voice and gesture recognition are used simultaneously, then the adaptability to user interactions is improved, but the difficulty of detecting and measuring increases
Solution Approach 1:
The system implements partial action by activating only the necessary recognition modes based on current interaction needs. Instead of continuously running both voice and gesture recognition, the system selectively engages appropriate modes, reducing detection complexity while maintaining adaptability when needed.
Data Source
AI summary
A control apparatus controls an agent apparatus functioning as a user interface of a request processing apparatus that acquires a request indicated by at least one of a voice and a gesture of a user and performs a process corresponding to the request. The control apparatus includes a gaze point specifying section specifying a gaze point of the user, and a face control section controlling an orientation of a face or line of sight of an agent used to transmit information to the user. The face control section controls the orientation of the face or line of sight of the agent such that the face or line of sight of the agent becomes oriented toward the user, if the gaze point is positioned at (i) a portion of the agent or (ii) a portion of an image output section that displays or projects an image of the agent.


