Voice Presentation System Using Dual-Agent Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice output technologies, such as those for reading newspaper articles, fail to explicitly convey the essential points, making it difficult for users to understand the main information, especially when emotions are involved.

Innovation Solution

An information presenting apparatus that includes an acquirer, a generator, and a presenter, which acquire and analyze presentation information to generate support information, synthesizing voices for both a main and sub-agent to clearly convey essential points and emotions, with the sub-agent providing additional context or questions to deepen user understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a single agent outputs voice information, then the device complexity is low, but the user cannot easily understand the essential points and emotions

Engineering Contradiction:
Improveuser understanding of essential pointsVSAvoidagent structure complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The information presentation system is segmented into multiple agents: a main agent that presents primary information and a sub-agent that presents support information including essential points and emotions. This segmentation allows the system to provide comprehensive information presentation while maintaining manageable complexity through clear division of roles between agents.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If multiple agents are used to present information, then the user understanding of essential points improves, but the device complexity increases

Engineering Contradiction:
Improveclarity of essential pointsVSAvoidnumber of agents
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments information presentation into two distinct agents: a main agent for primary content and a sub-agent for support information. This segmentation ensures that essential points are not lost while keeping the number of agents minimal (only two), thus balancing information clarity with system simplicity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The sub-agent serves multiple functions simultaneously: it presents support information, emphasizes essential points, and conveys emotional context. This multi-functionality allows the system to achieve comprehensive information presentation with minimal additional agents, reducing the complexity increase that would otherwise result from having separate specialized agents for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If emotional expressions are added to voice output, then user engagement improves, but the monotony of standard reading is reduced

Engineering Contradiction:
Improveuser engagementVSAvoidemotion synthesis complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

Emotional expression is segmented and assigned specifically to the sub-agent rather than being implemented across the entire system. The sub-agent uses emotional words and expressions to present support information, while the main agent maintains standard reading tone. This segmentation increases user engagement through emotional content while limiting emotion synthesis complexity to only the sub-agent.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10878799B2Information presenting apparatus and information presenting method
Publication Date: 2020.12.29 SONY GROUP CORP
  • US10878799B2 patent drawing
  • US10878799B2 patent drawing
  • US10878799B2 patent drawing

AI summary

There is provided an information presenting apparatus to make it easy for the user to understand presentation information that is output as a voice. The information presenting apparatus includes an acquirer that acquires presentation information to be presented to a user, a generator that generates support information to be presented, together with the presentation information, to the user, based on the acquired presentation information, and synthesizes a voice that corresponds to the generated support information, and a presenter that presents the synthesized voice that corresponds to the support information, as an utterance of a first agent. The present technology is applicable to a robot, a signage apparatus, a car navigation apparatus, a watch-over system, a moving-image reproduction apparatus, or the like, for example.