Text-to-Speech Voice Selection for Multi-Language Media Assets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Portable electronic devices with small or no graphical user interfaces face challenges in providing adequate identification of audio content to users, especially when devices are not easily viewable, necessitating alternative mechanisms for content recognition.
Innovation Solution
The implementation of a text-to-speech system that synthesizes speech content in a single, familiar language for media assets, using a voice selection method based on the language of the text strings and device settings, ensuring coherent and understandable audio output even when multiple languages are involved.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a text-to-speech system provides speech content in multiple languages for different text strings, then the system can accommodate diverse media content languages, but the user experience becomes confusing and undesirable
Solution Approach 1:
The TTS system is designed to handle multiple languages through a single unified interface. The voice selection mechanism automatically adapts to the language of each text string while maintaining consistent speech output characteristics, allowing the system to serve multiple language requirements through one versatile component rather than requiring separate TTS systems for each language
Solution Approach 2:
A language detection and voice selection intermediary layer is introduced between the media player and the TTS system. This intermediary automatically identifies the language of each text string and selects the appropriate voice from the available TTS voices, thereby mediating between the diverse language content and the user's listening experience without requiring user intervention
2Ease of operation
If the electronic device uses a default language for all text strings, then the speech output remains consistent and simple, but it may not match the language of the media content
Solution Approach 1:
The voice selection mechanism transitions from a static default language approach to a dynamic language adaptation system. The system automatically adjusts the speech language based on the detected language of each text string, allowing the speech output to dynamically match the media content language while maintaining consistency within each individual speech segment
Solution Approach 2:
The system performs preliminary language detection on each text string before generating speech output. By detecting the language in advance and pre-selecting the appropriate voice from the TTS system, the ensure that the speech language accurately matches the media content language without requiring post-processing or user correction
3Loss of information
If the device provides graphical user interface for media identification, then users can easily identify content visually, but the device size increases and portability decreases
Solution Approach 1:
The patent replaces the mechanical/visual display system with an acoustic speech output system for media content identification. Instead of using a graphical display to show media information, the system uses text-to-speech conversion to audibly present media titles, artist names, and other identifying information, thereby substituting a visual interface with an audio interface that requires no additional display hardware
Solution Approach 2:
The TTS system serves the dual purpose of both playing media content and providing media identification through speech. The same speech synthesis engine that generates the media playback also generates the identifying speech output, eliminating the need for separate identification mechanisms and allowing the device to provide comprehensive functionality without increasing size
Data Source
AI summary
Algorithms for synthesizing speech used to identify media assets are provided. Speech may be selectively synthesized from text strings associated with media assets, where each text string can be associated with a native string language (e.g., the language of the string). When several text strings are associated with at least two distinct languages, a series of rules can be applied to the strings to identify a single voice language to use for synthesizing the speech content from the text strings. In some embodiments, a prioritization scheme can be applied to the text strings to identify the more important text strings. The rules can include, for example, selecting a voice language based on the prioritization scheme, a default language associated with an electronic device, the ability of a voice language to speak text in a different language, or any other suitable rule.


