Server-Side Text-to-Speech Synthesis for Portable Media
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Portable electronic devices face challenges in providing adequate user interfaces, particularly for users who cannot or prefer not to use graphical displays, and existing speech synthesis technologies struggle to produce high-quality, natural-sounding speech efficiently and accurately across various languages and accents.
Innovation Solution
Implementing sophisticated text-to-speech algorithms on a server farm system that normalizes text strings, determines native and target languages, and converts phonemes to produce high-quality, human-sounding speech that can be synthesized and combined with media content, allowing for efficient distribution without modifying the devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If text-to-speech conversion is performed on portable electronic devices, then speech content can be generated, but device size and power consumption increase
Solution Approach 1:
The patent extracts the computationally intensive text-to-speech conversion process from the portable electronic device and relocates it to a remote server. The device only retains lightweight functions for sending text data and receiving synthesized speech, while the heavy processing burden is removed to an external system, thereby reducing device power consumption and size requirements.
Solution Approach 2:
The patent introduces a remote server as an intermediary between the user's text input and the synthesized speech output. This intermediary handles the complex text-to-speech conversion process, allowing the portable device to maintain simple architecture while still providing sophisticated speech synthesis capabilities through the mediating server system.
2Extent of automation
If text-to-speech conversion is performed on portable electronic devices, then speech content can be generated, but device complexity increases
Solution Approach 1:
The patent extracts the complex text-to-speech conversion functionality from the portable electronic device architecture. By removing the need for complex synthesis engines, phoneme databases, and processing circuits from the device, the system achieves speech synthesis capability without increasing device complexity or requiring additional circuitry.
Solution Approach 2:
The patent implements a centralized speech synthesis system on the remote server that can serve multiple devices. Instead of each device having its own complex synthesis capabilities, the server creates copies of synthesized speech content and distributes them to multiple clients, reducing the complexity burden on individual devices.
3Productivity
If speech synthesis is performed remotely, then processing resources are efficiently utilized, but speech quality and naturalness may be compromised
Solution Approach 1:
The patent employs sophisticated text normalization and phoneme generation techniques on the remote server to create high-quality speech outputs. By using advanced algorithms for text preprocessing, phoneme selection, and articulation modeling, the system produces natural-sounding speech that maintains high quality despite being generated remotely rather than on the device itself.
Data Source
AI summary
Algorithms for synthesizing speech used to identify media assets are provided. Speech may be selectively synthesized form text strings associated with media assets. A text string may be normalized and its native language determined for obtaining a target phoneme for providing human-sounding speech in a language (e.g., dialect or accent) that is familiar to a user. The algorithms may be implemented on a system including several dedicated render engines. The system may be part of a back end coupled to a front end including storage for media assets and associated synthesized speech, and a request processor for receiving and processing requests that result in providing the synthesized speech. The front end may communicate media assets and associated synthesized speech content over a network to host devices coupled to portable electronic devices on which the media assets and synthesized speech are played back.


