Multilingual TTS Segmentation and Symbol Display
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-to-speech (TTS) devices struggle to process text inputs containing characters from multiple languages and symbols, leading to information loss and impaired communication, as they cannot convey the emotional feel of text inputs effectively.
Innovation Solution
A method and device that receive text inputs including characters from multiple languages and symbols, segment them by language and symbol, generate sound segments using multiple TTS engines, and merge these segments to produce an output sound while displaying symbols to convey emotional content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single TTS engine is used for speech synthesis, then the device complexity is reduced, but the ability to process multiple languages and symbols is lost
Solution Approach 1:
The text input is segmented into character groups by language and symbol groups, with each segment processed by appropriate TTS engines or display mechanisms. This allows multiple languages to be handled by dedicated engines while keeping the overall system manageable through modular processing.
Solution Approach 2:
The TTS device is configured with multiple TTS engines that can handle different languages, making the device universal in its language processing capability. The system can adaptively select which engine to use based on the language detected in the input text.
2Loss of information
If symbols are converted to corresponding words by TTS engine, then speech output is generated, but the emotional feel of the text is lost
Solution Approach 1:
Symbols are extracted from the text input and separated into a dedicated symbol group, which is then displayed visually rather than being converted to speech. This extraction preserves the emotional information contained in symbols while allowing the remaining text to be processed for speech output.
Solution Approach 2:
Instead of converting symbols to audio dimension, the system displays them in the visual dimension on a display device. This dimensional change allows emotional content to be preserved through visual presentation while speech output handles the linguistic content.
3Loss of information
If text input with multiple languages and symbols is processed, then comprehensive speech output is achieved, but information loss occurs
Solution Approach 1:
The text input is segmented into character groups by language and symbol groups, with each segment processed appropriately. This segmentation prevents information loss by ensuring each element is handled by the most suitable processing mechanism while keeping the system structure organized and manageable.
Solution Approach 2:
The system uses an intermediary processing structure that separates text processing from symbol processing. Text segments are routed to TTS engines while symbol segments are routed to display output, allowing comprehensive information preservation without requiring a single complex processing path.
Data Source
AI summary
In a method and device for speech synthesis, the method of outputting text input as sound from an electronic device, includes receiving a text input including characters from at least two languages and at least one symbol, generating a plurality of ordered groups by sequencing, wherein the plurality of ordered groups includes character groups including characters from a common language, and a symbol group including symbols, by segmenting the text input sequentially by language and by symbol, generating sound segments corresponding to each of the character groups by use of a plurality of text-to-speech (TTS) engines supporting different languages, generating an output sound by merging the sound segments according to the sequencing of the character groups within the plurality of ordered groups, and displaying the symbol of the symbol group on a display while outputting the output sound by use of a speaker.


