Multilingual Text-to-Speech Conversion via Language Processor

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in accessing and understanding text data in electronic devices when the language is unknown, as current technologies lack efficient multilingual text-to-speech conversion capabilities, especially for various applications beyond text messages.

Innovation Solution

A system and method that enables text-to-speech conversion in multiple languages, utilizing an event monitor, event manager, text-to-speech engine, user interface controller, language processor, and audio output unit to allow users to select and listen to text data in their preferred language, supporting regional languages like Hindi, Kannada, and Telugu, and converting text data from any language to English or a chosen language.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text-to-speech conversion is implemented for multiple languages, then language accessibility and user convenience are improved, but system complexity and processing requirements increase

Engineering Contradiction:
Improvelanguage accessibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The text-to-speech system is designed to handle multiple languages through a unified architecture. The language processor can identify and process text in different languages, and the speech synthesis engine can generate speech in the corresponding language, making the system universally applicable across multiple languages without requiring separate systems for each language

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

A language processor acts as an intermediary between the text input and speech synthesis engine. This mediator component identifies the language of the input text and routes it to the appropriate speech synthesis model, thereby managing system complexity while maintaining multi-language capability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If real-time text-to-speech conversion is provided across multiple applications, then user convenience and accessibility are improved, but processing time and computational resources increase

Engineering Contradiction:
Improveuser convenienceVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs language identification and speech synthesis preparation in advance when text is copied or selected. By detecting the language early in the process and pre-loading appropriate speech synthesis resources, the system reduces the actual conversion time when the user requests text-to-speech conversion

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces traditional mechanical text-to-speech conversion with neural network-based speech synthesis models. This substitution enables faster, more natural-sounding speech generation while reducing processing time compared to conventional synthesis methods

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11328706B2System and method for multilingual conversion of text data to speech data
Publication Date: 2022.05.10 INDUS APPSTORE PRIVATE LIMITED
  • US11328706B2 patent drawing
  • US11328706B2 patent drawing
  • US11328706B2 patent drawing

AI summary

The present invention provides a system and method for converting text data into speech data. Initially, the system enables a user to select a language from a plurality of languages supported by the operating system (OS) of a computing device. Further, on selecting and copying any text data, the system provides the user with options to listen to an audio output of the text data. The user is provided with options to listen to text data in either English or the selected language, when the language of the text data is one among the plurality of languages supported by the OS. Further, the user is provided with options to listen to text data in English, for the text data in any language. Once the user selects the option, the system converts the text data to speech data. The speech data is provided as the audio output to the user.