Audio Content Customization Service for Multi-Voice Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio content, such as audiobooks, often lacks customization options for listeners who prefer specific voice actors for characters, leading to unsatisfactory listening experiences due to the limitations of single-voice recordings, which are impractical and inflexible for multiple voice actor coordination.
Innovation Solution
A computer-implemented content customization service that maps textual content to characters and synchronizes audio portions from different voice actors, allowing for the combination of multiple voices in a single audio content, using techniques like named entity extraction, stochastic prediction, and user input to ensure accurate voice assignments based on attributes like gender and accent.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple voice actors are coordinated to generate a new recording, then voice customization for characters is improved, but production complexity and cost increase significantly
Solution Approach 1:
The patent segments the audio content by separating different character voices and their corresponding audio portions. Each voice actor's performance is extracted as an independent segment, which can then be selectively combined. This segmentation allows the system to handle multiple voice actors without requiring complete re-recording, reducing production complexity while maintaining voice customization capability.
Solution Approach 2:
The patent performs preliminary action by pre-separating and tagging audio portions with metadata indicating character identity and voice actor assignment during the original recording process. This preliminary structuring of the audio content enables flexible recombination later without requiring complex real-time coordination of multiple voice actors, thus improving adaptability while controlling production complexity.
2Manufacturing precision
If a new recording is created for each listener's preferences, then voice assignment accuracy is improved, but production time and resource consumption increase
Solution Approach 1:
The patent uses copying by creating customized audio versions through digital manipulation and recombination of existing recorded segments rather than creating entirely new recordings. The system copies and reassembles pre-recorded character portions with different voice actors based on listener preferences, achieving high voice assignment accuracy without the time and resource costs of new productions.
Solution Approach 2:
The patent applies parameter changes by modifying the configuration parameters of existing audio content (such as voice actor assignment, character mapping, and audio segment selection) rather than changing the fundamental audio recordings themselves. This allows rapid generation of customized versions with different voice assignments, maintaining accuracy while dramatically improving production efficiency compared to creating new recordings.
3Ease of operation
If audio content is recorded with only one voice actor, then production simplicity is maintained, but listener satisfaction decreases due to lack of voice customization
Solution Approach 1:
The patent applies dynamics by transforming static single-voice audio recordings into dynamic, reconfigurable audio compositions. The system enables the same audio content to be dynamically reassembled with different voice actor combinations based on listener preferences, maintaining the simplicity of the original single-actor recording process while adding post-production flexibility for voice customization.
Data Source
AI summary
A content customization service is disclosed. The content customization service may identify one or more speakers in an item of content, and map one or more portions of the item of content to a speaker. A speaker may also be mapped to a voice. In one embodiment, the content customization service obtains portions of audio content synchronized to the mapped portions of the item of content. Each portion of audio content may be associated with a voice to which the speaker of the portion of the item of content is mapped. These portions of audio content may be combined to produce a combined item of audio content with multiple voices.


