Speech Synthesis Dictionary Delivery Device for Network-Constrained Terminals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies face challenges in efficiently delivering and managing large numbers of speech synthesis dictionaries on terminals with limited network connectivity and storage capacity, making it difficult to use multiple speakers for applications like reading messages from social networking services.
Innovation Solution
A speech synthesis dictionary delivery system that dynamically switches between first dictionaries, which provide high speaker reproducibility but require large data sizes, and second dictionaries, which offer smaller data sizes but lower reproducibility, based on network communication state and user importance, allowing for efficient delivery and synthesis of speech across multiple speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If first dictionaries (acoustic models of individual speakers) are delivered to terminals, then speaker reproducibility is improved, but data transmission volume and storage requirements increase
Solution Approach 1:
The patent segments the speech synthesis data into two distinct types: first dictionaries (individual speaker acoustic models with high reproducibility but large size) and second dictionaries (generic acoustic models with lower reproducibility but small size). This segmentation allows the system to selectively deliver only the necessary data portion based on requirements, resolving the contradiction between speaker reproducibility and data transmission volume.
Solution Approach 2:
The patent changes the parameter of dictionary type selection based on communication state and usage requirements. By dynamically switching between first dictionary mode (high reproducibility, large data) and second dictionary mode (lower reproducibility, small data), the system optimizes the balance between speaker reproducibility and data transmission volume according to current network conditions and application needs.
2Adaptability or versatility
If multiple speech synthesis dictionaries are delivered to terminals, then speaker versatility is improved, but storage capacity requirements increase
Solution Approach 1:
The patent makes the second dictionary universal by designing it to work with multiple speakers through parameter substitution rather than requiring separate first dictionaries for each speaker. This multi-functionality allows a single second dictionary to serve multiple speakers by combining it with different speaker parameter sets, thereby improving speaker versatility while minimizing storage capacity requirements.
Solution Approach 2:
The patent uses parameter sets as lightweight copies that reference the generic second dictionary structure. Instead of storing complete first dictionaries for each speaker, the system stores compact parameter sets that can be combined with the shared second dictionary to generate speaker-specific speech synthesis, significantly reducing storage requirements while maintaining speaker versatility.
3Reliability
If speech synthesis dictionaries are delivered without network connection, then system reliability is improved, but dictionary size must be reduced
Solution Approach 1:
The patent implements a dynamic dictionary selection mechanism that adapts to network availability. When network connection is available, the system can deliver and use first dictionaries for high reproducibility. When network connection is unavailable, the system automatically switches to using second dictionaries with parameter sets, maintaining offline capability with significantly reduced dictionary size requirements.
Data Source
AI summary
A speech synthesis dictionary delivery device that delivers a dictionary for performing speech synthesis to terminals, comprises a storage device for speech synthesis dictionary database that stores a first dictionary which includes an acoustic model of a speaker and is associated with identification information of the speaker, that stores a second dictionary which includes an acoustic model generated using voice data of a plurality of speakers, and that stores parameter sets of the speakers to be used with the second dictionary and which are associated with identification information of the speakers, a processor that determines one of the first dictionary and the second dictionary, which should be used in the terminal for a specified speaker, and an input output interface (I/F) that receives the identification information of a speaker transmitted from the terminal and then delivers at least one of a first dictionary, the second dictionary, and a parameter set of the second dictionary, on the basis of the received identification information of the speaker and a result of the determination by the processor.


