Multilingual Voice Mixing for Dialog Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Monolingual voice dialog systems inadequately serve multilingual users, requiring sequential announcements in additional languages, which lengthens interaction time and reduces user acceptance.

Innovation Solution

A method for multilingual voice output that simultaneously outputs a primary language sequence and secondary language sequences with differing signal properties, allowing users to recognize and respond in their native language without sequential explanations, utilizing a mixing system to combine and transmit voice sequences in real-time or offline, mimicking the 'cocktail party effect' to highlight the primary language.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If sequential announcements in additional languages are provided, then foreign-language users are informed of multilingual options, but interaction time is extended and user acceptance is reduced

Engineering Contradiction:
Improvemultilingual supportVSAvoidinteraction time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent merges multiple language announcements into a single simultaneous output instead of sequential delivery. The mixing system combines voice sequences in different languages (e.g., German and English) that are output at the same time, allowing foreign-language users to receive information about multilingual options without extending the interaction time for primary language users.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent applies different signal properties to different language components within the same output. The mixing system assigns varying volume levels, frequencies, or spatial positions to different language sequences, allowing users to selectively attend to their native language while hearing other languages in the background, thus providing localized information quality for different user groups simultaneously.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If sequential announcements in additional languages are provided, then foreign-language users are informed of multilingual options, but user acceptance in primary language is reduced

Engineering Contradiction:
Improvemultilingual supportVSAvoiduser acceptance
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent merges multiple language announcements into a single simultaneous output instead of sequential delivery. The mixing system combines voice sequences in different languages (e.g., German and English) that are output at the same time, allowing foreign-language users to receive information about multilingual options without extending the interaction time for primary language users.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent applies different signal properties to different language components within the same output. The mixing system assigns varying volume levels, frequencies, or spatial positions to different language sequences, allowing users to selectively attend to their native language while hearing other languages in the background, thus providing localized information quality for different user groups simultaneously.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If multiple access points for different languages are created, then multilingual access is enabled, but system complexity increases

Engineering Contradiction:
Improvemultilingual accessVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a single access point that serves multiple language groups simultaneously. The mixing system is designed to handle multiple language sequences through one unified interface, eliminating the need for separate access points or phone numbers for different languages. This multi-functional system reduces complexity by consolidating what would otherwise require multiple independent entry points.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP2047668B1Method, spoken dialog system, and telecommunications terminal device for multilingual speech output
Publication Date: 2017.12.27 DEUTSCHE TELEKOM AG
  • EP2047668B1 patent drawingFigure 1
  • EP2047668B1 patent drawingFigure 2
  • EP2047668B1 patent drawingFigure 3

AI summary

In order to provide improved multilingual speech output in an automated spoken dialog system, the invention provides a method in which a connection is set up between a telecommunications terminal device (100) and a spoken dialog system (301, 302, 303) and, in response to the setup of the connection, a multilingual speech output is provided which comprises the output of a first speech sequence (410) in a first language and the output of at least a second speech sequence (411-41N) in at least one second language different from the first, wherein the first and the at least one second speech sequence are output at least partially isochronously. The invention further provides a spoken dialog system (301, 302, 303) and a telecommunications terminal device (100) constructed respectively for implementing the method.