Multilingual Voice Mixing for Dialog Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Monolingual voice dialog systems inadequately serve multilingual users, requiring sequential announcements in additional languages, which lengthens interaction time and reduces user acceptance.
Innovation Solution
A method for multilingual voice output that simultaneously outputs a primary language sequence and secondary language sequences with differing signal properties, allowing users to recognize and respond in their native language without sequential explanations, utilizing a mixing system to combine and transmit voice sequences in real-time or offline, mimicking the 'cocktail party effect' to highlight the primary language.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If sequential announcements in additional languages are provided, then foreign-language users are informed of multilingual options, but interaction time is extended and user acceptance is reduced
Solution Approach 1:
The patent merges multiple language announcements into a single simultaneous output instead of sequential delivery. The mixing system combines voice sequences in different languages (e.g., German and English) that are output at the same time, allowing foreign-language users to receive information about multilingual options without extending the interaction time for primary language users.
Solution Approach 2:
The patent applies different signal properties to different language components within the same output. The mixing system assigns varying volume levels, frequencies, or spatial positions to different language sequences, allowing users to selectively attend to their native language while hearing other languages in the background, thus providing localized information quality for different user groups simultaneously.
2Adaptability or versatility
If sequential announcements in additional languages are provided, then foreign-language users are informed of multilingual options, but user acceptance in primary language is reduced
Solution Approach 1:
The patent merges multiple language announcements into a single simultaneous output instead of sequential delivery. The mixing system combines voice sequences in different languages (e.g., German and English) that are output at the same time, allowing foreign-language users to receive information about multilingual options without extending the interaction time for primary language users.
Solution Approach 2:
The patent applies different signal properties to different language components within the same output. The mixing system assigns varying volume levels, frequencies, or spatial positions to different language sequences, allowing users to selectively attend to their native language while hearing other languages in the background, thus providing localized information quality for different user groups simultaneously.
3Adaptability or versatility
If multiple access points for different languages are created, then multilingual access is enabled, but system complexity increases
Solution Approach 1:
The patent creates a single access point that serves multiple language groups simultaneously. The mixing system is designed to handle multiple language sequences through one unified interface, eliminating the need for separate access points or phone numbers for different languages. This multi-functional system reduces complexity by consolidating what would otherwise require multiple independent entry points.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
In order to provide improved multilingual speech output in an automated spoken dialog system, the invention provides a method in which a connection is set up between a telecommunications terminal device (100) and a spoken dialog system (301, 302, 303) and, in response to the setup of the connection, a multilingual speech output is provided which comprises the output of a first speech sequence (410) in a first language and the output of at least a second speech sequence (411-41N) in at least one second language different from the first, wherein the first and the at least one second speech sequence are output at least partially isochronously. The invention further provides a spoken dialog system (301, 302, 303) and a telecommunications terminal device (100) constructed respectively for implementing the method.