Spoken Utterance Control for Background Sound Affinity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies fail to effectively control the affinity of spoken utterances for background sounds, particularly when the importance of information notifications is high, leading to notifications being drowned out by music or other background sounds.
Innovation Solution
An information processing apparatus and method that control the output mode of spoken utterances based on the importance degree of notification information and affinity for background sounds, adjusting parameters such as voice quality, prosody, and timing to ensure notifications are clearly heard.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the speech format is selected to match the genre of music being reproduced, then the natural experience during music reproduction is improved, but important information notifications may be drowned out by the music
Solution Approach 1:
The system dynamically changes parameters of the spoken utterance (volume level, pitch, speed, timbre) based on the detected background sound characteristics and notification importance level. When a notification is important, the system adjusts parameters to increase prominence; when less important, it adapts parameters to match the background sound for natural integration.
Solution Approach 2:
The speech output characteristics are made dynamic rather than static. The system continuously monitors the background sound environment and adjusts the speech format in real-time based on changing conditions, including the detected music genre, volume level, and the importance level of the notification being delivered.
2Reliability
If the spoken utterance volume is increased to ensure notification is heard, then the notification noticeability is improved, but the natural experience during background sound reproduction deteriorates
Solution Approach 1:
Different segments of the speech output are given different quality characteristics based on their function. Critical notification content receives enhanced volume and prominence, while non-critical information is blended more seamlessly with the background sound. This creates local variations in speech quality that optimize both noticeability and naturalness.
Solution Approach 2:
The system applies partial enhancement only to the extent necessary for notification delivery. Instead of uniformly increasing all speech parameters, it selectively adjusts specific parameters (volume, pitch) only for critical notification portions, leaving other portions to blend naturally with the background sound.
Data Source
AI summary
[Object] To more flexibly control the affinity of a spoken utterance for a background sound in accordance with the importance degree of an information notification. [Solution] There is provided an information processing apparatus including an utterance control unit that controls an output of a spoken utterance corresponding to notification information. The utterance control unit controls an output mode of the spoken utterance on the basis of an importance degree of the notification information and affinity for a background sound. In addition, there is provided an information processing method including controlling, by a processor, an output of a spoken utterance corresponding to notification information. The controlling further includes controlling an output mode of the spoken utterance on the basis of an importance degree of the notification information and affinity for a background sound.


