Personalized Audio Program Generation via Dynamic Voice Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-speech (TTS) systems lack the ability to dynamically generate audio programs from diverse textual content sources, such as emails and social network messages, with personalized voice selection based on content source, subject, and user preferences, resulting in monotonous and unengaging audio presentations.
Innovation Solution
An audio program server that processes user-selected content, employs a TTS system with multiple voices and languages, automatically selects voices based on content characteristics and user preferences, and integrates segues and summaries to create engaging audio programs, allowing users to customize voice, speed, and tone for different content types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a TTS system uses a single voice module for all content, then the system complexity is reduced, but the user engagement and listening experience deteriorate due to monotonous audio presentations
Solution Approach 1:
The TTS system is segmented into multiple independent voice modules, each capable of generating audio with distinct characteristics. The system divides content into different types (news, weather, sports, etc.) and assigns appropriate voice modules to each segment, enabling diversified audio presentations while maintaining manageable system architecture through modular design
Solution Approach 2:
The TTS system is designed with universal voice modules that can handle multiple content types and languages. Each voice module is configured to work across different content categories, and the system can automatically select and switch between modules based on content characteristics, providing versatile audio generation without requiring separate dedicated systems for each content type
2Adaptability or versatility
If the TTS system automatically selects voices based on content characteristics, then user engagement is improved, but the processing time and computational resources increase
Solution Approach 1:
Voice modules are pre-configured with specific characteristics suitable for different content types (e.g., formal voices for news, energetic voices for sports). The system maintains a pre-established mapping between content categories and appropriate voice modules, eliminating the need for complex real-time analysis and enabling rapid automatic selection during audio generation
Solution Approach 2:
The system adjusts voice parameters such as pitch, speed, and tone based on content characteristics rather than generating voices from scratch. By modifying parameters of existing voice modules to match content requirements, the system achieves personalized audio presentations with reduced computational overhead and faster generation times
3Ease of operation
If the TTS system processes multiple content sources with different voices, then the listening experience is enhanced, but the device complexity increases due to multiple voice modules
Solution Approach 1:
Different voice modules with distinct characteristics are assigned to specific content types based on their local requirements. For example, formal voices are allocated to news content while more casual voices are assigned to entertainment content. This localized optimization enhances user experience for each content category without requiring all voice modules to be active simultaneously, managing overall system complexity
Data Source
AI summary
Features are disclosed for generating text-to-speech (TTS) audio programs from textual content received from multiple sources. A TTS system may assemble an audio program from several individual audio presentations of user-selected network-accessible content. Users may configure the TTS system to retrieve personal content as well as publically accessible content. The audio program may include segues, introductions, summaries, and the like. Voices may be selected for individual content items based on user selections or on characteristics of the content or content source.


