Incremental Speech Database Download for Mobile Terminals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis systems on mobile terminals with limited memory face challenges in achieving high-quality speech synthesis due to restricted database size and reliance on active network connections, limiting their effectiveness, especially for devices like PDAs.
Innovation Solution
A network system architecture that dynamically downloads incremental databases of speech waveforms and related indexing information, enhancing a reduced database on the terminal to achieve high-quality speech synthesis without increasing memory requirements, using a virtual huge database by downloading context-specific incremental databases as needed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a large speech database is used to improve speech synthesis quality, then speech quality increases, but memory requirements increase beyond the capacity of mobile terminals
Solution Approach 1:
The patent divides the large speech database into a base database stored on the terminal and multiple context-specific incremental databases stored remotely. The base database contains common speech units, while incremental databases contain context-specific speech units. This segmentation allows the terminal to access high-quality speech synthesis without storing the entire large database locally.
Solution Approach 2:
The patent introduces a server as an intermediary between the terminal and the large speech database. The server stores the complete speech database and provides context-specific incremental databases to the terminal upon request. This intermediary enables the terminal to access high-quality speech resources without requiring large local storage capacity.
2Manufacturing precision
If a distributed speech synthesis system is used to achieve high quality synthesis, then speech quality improves, but the system requires an active network connection which limits effectiveness
Solution Approach 1:
The patent performs preliminary action by pre-storing a base database of speech units on the terminal and pre-defining context categories. The terminal can immediately use the base database for general speech synthesis without network connection. When network connection is available, context-specific incremental databases are downloaded in advance or on demand, enabling the terminal to operate independently for basic functions while accessing enhanced capabilities when needed.
3Quantity of substance
If the speech database is statically configured to meet maximum memory capacity, then memory requirements are satisfied, but speech synthesis quality is limited
Solution Approach 1:
The patent transforms the static database configuration into a dynamic system where the effective database size adapts based on context requirements. The terminal dynamically loads different combinations of base and incremental databases depending on the speech context, allowing the system to optimize between memory usage and synthesis quality for different applications.
Solution Approach 2:
The patent applies local quality by providing different levels of database content for different contexts. The base database provides universal coverage, while incremental databases provide enhanced, context-specific speech units. This allows the system to optimize speech quality for specific contexts without requiring the entire large database to be stored locally.
Data Source
AI summary
Service architecture for providing to a user terminal of a communications network textual information and relative speech synthesis, the user terminal being provided with a speech synthesis engine and a basic database of speech waveforms includes: a content server for downloading textual information requested by means of a browser application on the user terminal; a context manager for extracting context information from the textual information requested by the user terminal; a context selector for selecting an incremental database of speech waveforms associated with extracted context information and for downloading the incremental database into the user terminal; a database manager on the user terminal for managing the composition of an enlarged database of speech waveforms for the speech synthesis engine including the basic and the incremental databases of speech waveforms.


