Embodied Negotiation Agent Multimodal Interaction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current human-agent negotiation systems are limited to text-based interactions and do not support multi-lateral negotiations between humans and software agents, lacking the naturalness of human-human interactions, which restricts the feasibility of practical human-machine negotiation.
Innovation Solution
The development of an embodied negotiation agent and platform that transcribes human speech and non-speech behavioral traces, allowing humans to interact with multiple software agents using speech and non-verbal cues, such as gestures, to facilitate more natural and multi-modal negotiations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If text-based communication interface is used for human-agent negotiation, then system complexity is reduced, but interaction naturalness deteriorates
Solution Approach 1:
The patent replaces text-based mechanical input with voice-based acoustic input. The system captures human speech signals and converts them to text through speech-to-text conversion, allowing users to communicate naturally through voice instead of typing. This substitution maintains system simplicity while dramatically improving interaction naturalness.
Solution Approach 2:
The system integrates multiple communication modalities (voice, text, gestures) into a single negotiation platform. By supporting multiple input methods simultaneously, the system achieves multi-functionality that enhances naturalness without requiring separate specialized systems for each modality.
2Ease of operation
If speech-to-text conversion is implemented, then interaction naturalness is improved, but processing time increases
Solution Approach 1:
The system performs speech-to-text conversion in real-time as speech is captured, rather than processing after the entire conversation is complete. This preliminary action allows the text representation to be available immediately for subsequent negotiation processing, minimizing delays in the interaction flow.
3Ease of operation
If multi-modal input (speech and behavioral traces) is accepted, then interaction naturalness is improved, but system complexity increases
Solution Approach 1:
The patent merges speech signals and non-speech behavioral traces into a unified input representation. By combining these different modalities into a single integrated data structure that includes both audio and behavioral information, the system achieves multi-modal input without requiring separate complex processing pipelines for each modality.
4Ease of operation
If avatars with visual representations are introduced, then interaction naturalness is improved, but computational requirements increase
Solution Approach 1:
The system uses visual avatars as simplified copies or representations of the software agents rather than implementing full graphical interfaces. These avatar copies provide visual feedback and representation with minimal computational overhead, allowing natural visual interaction without the heavy processing requirements of complex graphical systems.
Data Source
AI summary
Human speech signals that are uttered within an environment are transcribed; the environment includes one or more avatars representing one or more software agents; the human speech signals are directed to at least one of the avatars. At least one non-speech behavioral trace is obtained within the environment; the trace is representative of non-speech behavior directed to the at least one of the avatars. The transcribed human speech signals and the at least one non-speech behavioral trace are forwarded to the one or more software agents. A proposed act is obtained from at least one of the agents; responsive thereto, a command is issued to cause the avatar corresponding to the software agent from which the proposed act is obtained to emit synthesized speech and to act visually in accordance with the proposed act.


