Voice Presentation System Using Dual-Agent Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice output technologies, such as those for reading newspaper articles, fail to explicitly convey the essential points, making it difficult for users to understand the main information, especially when emotions are involved.
Innovation Solution
An information presenting apparatus that includes an acquirer, a generator, and a presenter, which acquire and analyze presentation information to generate support information, synthesizing voices for both a main and sub-agent to clearly convey essential points and emotions, with the sub-agent providing additional context or questions to deepen user understanding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a single agent outputs voice information, then the device complexity is low, but the user cannot easily understand the essential points and emotions
Solution Approach 1:
The information presentation system is segmented into multiple agents: a main agent that presents primary information and a sub-agent that presents support information including essential points and emotions. This segmentation allows the system to provide comprehensive information presentation while maintaining manageable complexity through clear division of roles between agents.
2Loss of information
If multiple agents are used to present information, then the user understanding of essential points improves, but the device complexity increases
Solution Approach 1:
The system segments information presentation into two distinct agents: a main agent for primary content and a sub-agent for support information. This segmentation ensures that essential points are not lost while keeping the number of agents minimal (only two), thus balancing information clarity with system simplicity.
Solution Approach 2:
The sub-agent serves multiple functions simultaneously: it presents support information, emphasizes essential points, and conveys emotional context. This multi-functionality allows the system to achieve comprehensive information presentation with minimal additional agents, reducing the complexity increase that would otherwise result from having separate specialized agents for each function.
3Ease of operation
If emotional expressions are added to voice output, then user engagement improves, but the monotony of standard reading is reduced
Solution Approach 1:
Emotional expression is segmented and assigned specifically to the sub-agent rather than being implemented across the entire system. The sub-agent uses emotional words and expressions to present support information, while the main agent maintains standard reading tone. This segmentation increases user engagement through emotional content while limiting emotion synthesis complexity to only the sub-agent.
Data Source
AI summary
There is provided an information presenting apparatus to make it easy for the user to understand presentation information that is output as a voice. The information presenting apparatus includes an acquirer that acquires presentation information to be presented to a user, a generator that generates support information to be presented, together with the presentation information, to the user, based on the acquired presentation information, and synthesizes a voice that corresponds to the generated support information, and a presenter that presents the synthesized voice that corresponds to the support information, as an utterance of a first agent. The present technology is applicable to a robot, a signage apparatus, a car navigation apparatus, a watch-over system, a moving-image reproduction apparatus, or the like, for example.


