Vehicle Agent Image Orientation Control for Natural Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional display systems in vehicles with multiple occupants often fail to display results in a position where the operator can easily recognize them, leading to unnatural agent behavior.
Innovation Solution
An agent device with a microphone, speaker, interpreter, and display system that adjusts the agent image's face direction based on audio input interpretation, prioritizing non-driver occupants and changing direction when necessary to ensure the agent can perform natural interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If the agent image maintains a fixed face direction, then the display is simple, but the agent behavior becomes unnatural when multiple occupants interact
Solution Approach 1:
The system uses audio input from microphones to detect which occupant is speaking, and this feedback is used to dynamically adjust the agent image's face direction. This closed-loop feedback mechanism enables natural behavior adaptation while maintaining relatively simple system architecture.
2Device complexity
If the agent image is displayed in a fixed position, then the display system is simple, but the display result may not be visible to the occupant who performed the operation
Solution Approach 1:
The agent image's display position and orientation are made dynamic based on the detected conversation target. The system adjusts the agent image's position and facing direction to ensure visibility and appropriate interaction with the active speaker, while maintaining a relatively simple display system architecture.
Data Source
AI summary
An agent device includes a microphone which collects audio in a vehicle cabin, a speaker which outputs audio to the vehicle cabin, an interpreter which interprets the meaning of the audio collected by the microphone, a display provided in the vehicle cabin, and an agent controller which displays an agent image in a form of speaking to an occupant in a region of the display and causes the speaker to output audio by which the agent image speaks to at least one occupant, and the agent controller changes the face direction of the agent image to an direction different from an direction of the occupant who is a conversation target in a case that an utterance with respect to the face direction is interpreted by the interpreter after the agent image is displayed on the display.


