Multilingual TTS Segmentation and Symbol Display

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech (TTS) devices struggle to process text inputs containing characters from multiple languages and symbols, leading to information loss and impaired communication, as they cannot convey the emotional feel of text inputs effectively.

Innovation Solution

A method and device that receive text inputs including characters from multiple languages and symbols, segment them by language and symbol, generate sound segments using multiple TTS engines, and merge these segments to produce an output sound while displaying symbols to convey emotional content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single TTS engine is used for speech synthesis, then the device complexity is reduced, but the ability to process multiple languages and symbols is lost

Engineering Contradiction:
Improvelanguage support capabilityVSAvoidTTS engine configuration
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The text input is segmented into character groups by language and symbol groups, with each segment processed by appropriate TTS engines or display mechanisms. This allows multiple languages to be handled by dedicated engines while keeping the overall system manageable through modular processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The TTS device is configured with multiple TTS engines that can handle different languages, making the device universal in its language processing capability. The system can adaptively select which engine to use based on the language detected in the input text.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of information

If symbols are converted to corresponding words by TTS engine, then speech output is generated, but the emotional feel of the text is lost

Engineering Contradiction:
Improveemotional content preservationVSAvoidspeech output quality
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

Symbols are extracted from the text input and separated into a dedicated symbol group, which is then displayed visually rather than being converted to speech. This extraction preserves the emotional information contained in symbols while allowing the remaining text to be processed for speech output.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of converting symbols to audio dimension, the system displays them in the visual dimension on a display device. This dimensional change allows emotional content to be preserved through visual presentation while speech output handles the linguistic content.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Loss of information

If text input with multiple languages and symbols is processed, then comprehensive speech output is achieved, but information loss occurs

Engineering Contradiction:
Improvetext information completenessVSAvoidprocessing system structure
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The text input is segmented into character groups by language and symbol groups, with each segment processed appropriately. This segmentation prevents information loss by ensuring each element is handled by the most suitable processing mechanism while keeping the system structure organized and manageable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses an intermediary processing structure that separates text processing from symbol processing. Text segments are routed to TTS engines while symbol segments are routed to display output, allowing comprehensive information preservation without requiring a single complex processing path.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250166599A1Method and device for speech synthesis
Publication Date: 2025.05.22 HYUNDAI MOTOR CO LTD
  • US20250166599A1 patent drawing
  • US20250166599A1 patent drawing
  • US20250166599A1 patent drawing

AI summary

In a method and device for speech synthesis, the method of outputting text input as sound from an electronic device, includes receiving a text input including characters from at least two languages and at least one symbol, generating a plurality of ordered groups by sequencing, wherein the plurality of ordered groups includes character groups including characters from a common language, and a symbol group including symbols, by segmenting the text input sequentially by language and by symbol, generating sound segments corresponding to each of the character groups by use of a plurality of text-to-speech (TTS) engines supporting different languages, generating an output sound by merging the sound segments according to the sequencing of the character groups within the plurality of ordered groups, and displaying the symbol of the symbol group on a display while outputting the output sound by use of a speaker.