Text-to-Speech Language Selection for Multi-Language Content

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-speech systems automatically select a default language synthesizer, which can result in undesirable speech output when the text contains multiple languages, leading to inaccurate pronunciation and accent issues.

Innovation Solution

The system allows users to select a language for text-to-speech conversion from multiple eligible languages, using analysis criteria such as keyboard settings, linguistic tags, user preferences, and location-based information to determine applicable languages and prompt the user to choose the appropriate language for accurate conversion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a default language synthesizer is automatically selected for text-to-speech conversion, then the system operation is simplified and processing speed is improved, but the pronunciation accuracy deteriorates when the text contains multiple languages

Engineering Contradiction:
Improveautomatic language selectionVSAvoidpronunciation accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system performs preliminary language detection and analysis before text-to-speech conversion. It identifies the language of the input text using linguistic tags, keyboard settings, and location information, then pre-selects the appropriate language synthesizer before conversion occurs, ensuring both automatic operation and accurate pronunciation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary language detection and analysis module between the text input and synthesizer selection. This intermediary analyzes linguistic tags, keyboard settings, and location data to determine the correct language, acting as a mediator that bridges the gap between automatic operation and pronunciation accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple language synthesizers are provided to support multiple languages, then the system versatility is improved, but the device complexity increases

Engineering Contradiction:
Improvemulti-language supportVSAvoidnumber of synthesizers
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements a universal language detection mechanism that works across multiple languages and synthesizers. The language detection module analyzes linguistic tags, keyboard settings, and location information to identify the appropriate language, allowing a single unified system to handle multiple languages without requiring separate detection logic for each language

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses feedback from linguistic tags, keyboard settings, and location information to dynamically select the appropriate synthesizer. This feedback mechanism allows the system to adapt to the user's language context and select the correct synthesizer automatically, reducing the perceived complexity while maintaining multi-language support

Inventive Principle:
Principle #23Feedback

3Measurement precision

If the system analyzes text using multiple criteria to identify applicable languages, then the language selection accuracy is improved, but the processing time increases

Engineering Contradiction:
Improvelanguage identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of linguistic tags, keyboard settings, and location information before text-to-speech conversion. By pre-identifying the language using these criteria, the system reduces the processing time during actual conversion while maintaining high language identification accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses multiple criteria (linguistic tags, keyboard settings, location information) for language identification, but applies them in a hierarchical manner. It checks the most reliable criteria first and uses additional criteria only when needed, avoiding unnecessary processing while maintaining high accuracy

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9483461B2Handling speech synthesis of content for multiple languages
Publication Date: 2016.11.01 APPLE INC
  • US9483461B2 patent drawing
  • US9483461B2 patent drawing
  • US9483461B2 patent drawing

AI summary

Techniques that enable a user to select, from among multiple languages, a language to be used for performing text-to-speech conversion. In some embodiments, upon determining that multiple languages may be used to perform text-to-speech conversion for a portion of text, the multiple languages may be displayed to the user. The user may then select a particular language to be used from the multiple languages. The portion of text may then be converted to speech in the user-selected language.