Spoken Dialog Device Multi-Language Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional cross-lingual spoken dialog systems cannot adapt responses based on the input language and struggle to facilitate person-to-person communication in multi-language settings, such as video conferences, where users speak different languages.

Innovation Solution

A spoken dialog device that includes a receiving unit, language identifier acquisition unit, speech recognition unit, dialog control unit, and speech synthesizing unit to detect and respond to input speech in multiple languages, maintaining dialog history and adapting responses accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a conventional cross-lingual spoken dialog system uses a weighted finite-state transducer framework to take over dialog history when language switches, then the system can maintain dialog context across language changes, but the system cannot change responses according to the input language and cannot facilitate person-to-person communication in multi-language settings

Engineering Contradiction:
Improvelanguage adaptabilityVSAvoiddialog control complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The dialog control unit is segmented into multiple independent language-specific control units (e.g., Japanese control unit, English control unit). Each control unit independently manages dialog history and generates responses for its respective language, allowing the system to adapt to different languages without increasing overall system complexity through centralized control.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system achieves multi-functionality by enabling a single spoken dialog device to serve multiple languages simultaneously. Each language-specific control unit can be activated based on the detected input language, allowing the device to function as a universal dialog system that adapts to different linguistic contexts while maintaining separate dialog histories for each language.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If the system maintains separate dialog histories for different languages, then the system can appropriately respond to each language, but the system cannot facilitate natural person-to-person conversation flow where users speak different languages

Engineering Contradiction:
Improvelanguage-specific response accuracyVSAvoidconversation naturalness
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The language identifier acquisition unit continuously detects the input language and provides feedback to the dialog control unit. Based on this feedback, the system dynamically switches between language-specific control units, ensuring that the dialog response matches the input language while maintaining natural conversation flow through real-time language adaptation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The language identifier acquisition unit acts as an intermediary between the user's speech input and the dialog control units. It detects the input language and mediates the selection of the appropriate control unit, enabling seamless transitions between languages and facilitating natural person-to-person conversation even when speakers use different languages.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11049493B2Spoken dialog device, spoken dialog method, and recording medium
Publication Date: 2021.06.29 NAT INST OF INFORMATION & COMM TECH
  • US11049493B2 patent drawing
  • US11049493B2 patent drawing
  • US11049493B2 patent drawing

AI summary

[Problem] With conventional technology, it is impossible to appropriately support spoken dialog that is carried out in multiple languages. [Solution] A spoken dialog device includes: a receiving unit that detects a voice section from a start point to an end point of an input speech that is spoken in any of two or more different languages, and acquires speech data corresponding to the voice section; a language identifier acquisition unit that acquires a language identifier that identifies a language in which the input speech was spoken; a speech recognition unit that generates a text resulting from speech recognition, based on the input speech and the language identifier; a dialog control unit to which a text resulting from speech recognition and a language identifier are input, and that generates a different output sentence depending on a language identifier, while maintaining dialog history even when the language identifier is different from the previous language identifier; a speech synthesizing unit that generates a speech waveform based on the output sentence and the language identifier; and a speech output unit that outputs a speech that is based on a speech waveform generated by the speech synthesizing unit. With such a spoken dialog device, it is possible to appropriately support spoken dialog that is carried out in multiple languages.